Somali people should not only use AI tools made elsewhere, but also build, shape, and own the systems that reflect their language and culture.
Sharafdin Yusuf is the Lead Engineer and Researcher at Goobo Labs, leading the design and development of the company's technology platforms and research initiatives. His work spans software engineering, artificial intelligence, cybersecurity, developer infrastructure, and distributed systems. He focuses on building reliable, scalable, and secure technologies while guiding projects from research and architecture through production and long-term maintenance. He also leads the engineering direction behind Goobo Labs' open-source projects, AI infrastructure, and developer tools.
We Gave a Frontier Model Somali Language Standards, and the Result Surprised Us
A demanding mathematics translation experiment revealed how the Somali Language Standard can help frontier models produce Somali that is more accurate, natural, consistent, and explainable.
Aug 7, 2026 · 11 min
Somali Has 20 Million Speakers and No Source of Truth. SLS Fixes That.
Somali is spoken by more than 20 million people, yet almost nothing in software or AI can point to a single authoritative, machine-readable source for its spelling, grammar, or terminology. Here is what the Somali Language Standard is, why it exists, what it's built on, and what changes once it's complete.
Jul 18, 2026 · 10 min
Introducing the Somali Language Standard (SLS) – Phase 1 Complete
We've formally published the first two standards of the Somali Language Standard (SLS) — an open, machine-readable, CI/CD-validated framework for the Somali language, starting with the alphabet and the standards process itself.
Jul 9, 2026 · 6 min
Soplang v2.0.0: a compiler for the first Somali programming language
Our team shipped a Rust-based compiler for Soplang — Cranelift JIT and ahead-of-time builds, with syntax written entirely in Somali keywords.
Jun 30, 2026 · 6 min
Why Somali needs its own tokenizer
Foreign multilingual tokenizers over-segment Somali — BERT-base spends 2.69 tokens per word. Our native BPE tokenizer, trained on 529M words of SomNLP-Corpus, gets that down to 1.53: a 1.75× improvement.
Nov 11, 2025 · 7 min