Goobo Labs
Research
OverviewPublicationsBenchmarks (SomBench)Open SourceRoadmap
DatasetsModelsBootcampsBlogCommunity
Sign inMentorship
Goobo Labs

An open research lab building the foundations of Somali AI.

info@goobolabs.so
Mogadishu, Somalia

Research

  • Overview
  • Publications
  • Benchmarks
  • Roadmap

Datasets

  • SomNLP Corpus
  • SomNLP STT
  • Wikipedia Corpus
  • Citation Guide

Bootcamps

  • Data Science & ML Bootcamp
  • Python for Everyone
  • Git & GitHub Bootcamp

Resources

  • AI & Data Terms
  • Docs
  • Models & Tools
  • Open Source
  • Blog
  • Soplang

Company

  • About
  • Careers
  • Impact
  • Press Kit
  • Verify Certificate
  • Contact

© 2026 Goobo Labs. Open research for Somali AI.

Privacy PolicyTerms of Service
Datasets

Open Somali datasets

Documented, versioned text and speech resources.

Open datasets

Projects you can build on today

Permissively licensed, documented, and versioned. Browse, download, and cite.

View all datasets
Open full dataset page
Open full dataset page
Open full dataset page
Open full dataset page
Try it live

Somali NLP, in your hands

A glimpse of the models we train. Pick a sentence and watch it tokenize, classify, and translate — entirely in Af-Soomaali.

Waxaan jeclahay barashada cilmi-nafsiga iyo AI-ga.

Tokenization · 6 tokens

Waxaanjeclahaybarashadacilmi-nafsigaiyoAI-ga

Model output

LanguageSomali · af
ScriptLatin
SentimentPositive
Confidence94%
Illustrative demo — confidence scores are not evaluation results.
Citation

How to cite our datasets

If you use our datasets in your research, please cite the appropriate entry below.

SomNLP-Corpus v2
@dataset{somnlp-corpus-2026,
  title   = {SomNLP-Corpus: An Open Somali Text Corpus},
  author  = {Osman, Mohamud and Ahmed, Faadumo and Hassan, Abdirahman},
  year    = {2026},
  version = {2.0},
  url     = {https://huggingface.co/datasets/goobolabs/somnlp-corpus},
  license = {CC-BY-4.0}
}
SomNLP-STT-Corpus
@dataset{somnlp-stt-2026,
  title   = {SomNLP-STT-Corpus: Somali Automatic Speech Recognition Dataset},
  author  = {Mohamed, Haawa and Osman, Mohamud},
  year    = {2026},
  url     = {https://huggingface.co/datasets/goobolabs/somnlp-stt},
  license = {CC-BY-4.0}
}
Corpus statistics

Inside the data

Where our text comes from and how the corpus has grown since launch.

Domain breakdown — SomNLP-Corpus

News
38%
Web
28%
Literature
16%
Wikipedia
12%
Social
6%
NewsWebLiteratureWikipediaSocial

Corpus growth — tokens (millions)

0M1M2M3M4MQ1 '25Q2 '25Q3 '25Q4 '25Q1 '26
Q2 '26887M ↑
Build with us

Build Somali AI in the open

Use our datasets and models, contribute to the research, or partner with the lab. Everything we can open, we do.

Contribute on GitHub Browse Datasets
GooboLabs
976+
GitHub stars
490+
Forks
269+
Contributors