Why we build Somali AI
The story, mission, and people behind Goobo Labs.
Every language deserves AI infrastructure. We build it for Somali first — openly, and to a standard the field can trust.
Open by default
Datasets, model weights, and evaluation code are released openly so anyone can build, verify, and improve them.
Community-owned
The people who speak the language help build the technology — and own its direction.
Rigorous & reproducible
Clear benchmarks, documented methods, and results others can reproduce and extend.
From raw data to open models in three steps
Collect & annotate
Community contributors, radio archives, and web crawls feed raw Somali text and speech into our pipeline. Every source is documented and licensed.
Clean, align & train
Data is deduplicated, filtered for quality, and used to train tokenizers, ASR models, and language models — benchmarked against SomBench.
Release in the open
Datasets, model checkpoints, evaluation code, and research notes are published under permissive licenses for anyone to use, study, and improve.
Somali is ~0.005% of language-identified web data — less than 0.01% of English's share
More than 20 million people speak Somali — yet it barely registers in the datasets that train the world's AI. That gap is why Goobo Labs exists.
Source: Common Crawl primary-language page distribution (CLD2, aggregated crawls through CC-MAIN-2026). Percentages are share of language-identified pages. Somali vs English ratio ≈ 0.01% of English's page share.
Open Somali text we've released — not shown on the same scale as global crawl percentages above.
Sharafdin Yusuf
Lead Engineer & Researcher
Somali people should not only use AI tools made elsewhere, but also build, shape, and own the systems that reflect their language and culture.
The people building it
Build Somali AI in the open
Use our datasets and models, contribute to the research, or partner with the lab. Everything we can open, we do.