Documented, versioned text and speech resources.
Permissively licensed, documented, and versioned. Browse, download, and cite.
A glimpse of the models we train. Pick a sentence and watch it tokenize, classify, and translate — entirely in Af-Soomaali.
Waxaan jeclahay barashada cilmi-nafsiga iyo AI-ga.
Tokenization · 6 tokens
Model output
If you use our datasets in your research, please cite the appropriate entry below.
@dataset{somnlp-corpus-2026,
title = {SomNLP-Corpus: An Open Somali Text Corpus},
author = {Osman, Mohamud and Ahmed, Faadumo and Hassan, Abdirahman},
year = {2026},
version = {2.0},
url = {https://huggingface.co/datasets/goobolabs/somnlp-corpus},
license = {CC-BY-4.0}
}@dataset{somnlp-stt-2026,
title = {SomNLP-STT-Corpus: Somali Automatic Speech Recognition Dataset},
author = {Mohamed, Haawa and Osman, Mohamud},
year = {2026},
url = {https://huggingface.co/datasets/goobolabs/somnlp-stt},
license = {CC-BY-4.0}
}Where our text comes from and how the corpus has grown since launch.
Use our datasets and models, contribute to the research, or partner with the lab. Everything we can open, we do.