The people building Somali AI
Every dataset, standard, bootcamp, and open-source project behind Goobo Labs was shipped by contributors working in the open — 275 of them.
Counted across 13 public GitHub repositories. Snapshot taken 2026-09-02.
Recognising the work that moved a project forward
Shipped the maanka2 Somali Web Corpus downloader into SomNLP-Corpus on 1 September — a new source feeding the text pipeline the rest of the corpus work builds on.
Maintainers of the open stack
Commit access, review duty, and release ownership across the corpora, the bootcamps, and the standards repos.
Ranked by commits, all repos
Commit counts are one signal among many — review load and annotation hours do not show up here.
Everyone who has shipped something
Code, data, docs, and bootcamp work — every merged contribution counts the same on this page.
Where the work happens
How to get on this page
Built with the community, not just for it — annotation, transcription, and translation count as much as code.
Every repo keeps an open queue. Start with one sized for a first pull request — no prior lab context needed.
Draft pull requests get reviewed quickly, and questions in Somali on the issue thread are welcome.
Merged work adds you to this page, to the repo contributor list, and to the dataset acknowledgements.
Build Somali AI in the open
Use our datasets and models, contribute to the research, or partner with the lab. Everything we can open, we do.