Somali Has 20 Million Speakers and No Source of Truth. SLS Fixes That.
ResearchJul 18, 2026 · 10 min read · Updated Aug 10, 2026
Somali Has 20 Million Speakers and No Source of Truth. SLS Fixes That.
Somali is spoken by more than 20 million people, yet almost nothing in software or AI can point to a single authoritative, machine-readable source for its spelling, grammar, or terminology. Here is what the Somali Language Standard is, why it exists, what it's built on, and what changes once it's complete.
Somali is spoken by more than 20 million people, and almost none of the software, AI models, or publishers that serve them have a single authoritative source to check a spelling, a grammar rule, or a technical term against. The Somali Language Standard (SLS) is our attempt to build that source: an open, machine-readable, versioned standard for the Somali language, developed in the open and released under a license that lets anyone — including commercial AI labs — build on it.
This post is a plain-language walkthrough of what SLS is, why we think Somali needs it, what it's built on, who it's for, and what the ecosystem looks like once it's done.
What Is the Somali Language Standard?
Think of SLS the way you'd think of an RFC series — the numbered, versioned documents that define how the internet works — or a Unicode Technical Report, except applied to the Somali language. SLS is not a dictionary you look words up in. It's a specification that systems and humans implement against.
When a software company builds a Somali spell-checker, a hospital builds a Somali patient portal, or an AI lab trains a Somali language model, they need an authoritative reference to rely on. That's the role SLS is designed to play.
Every fact SLS asserts — a word, a grammar rule, a technical term — is numbered with a permanent identifier (e.g. sls:lex:0042), versioned so adopters can cite exactly which revision they implement, sourced with full provenance back to an original Somali publication, and validated by an automated CI pipeline that rejects incorrect or incomplete records.
Why Somali Needs This
Somali is the official language of Somalia and one of the most widely spoken Cushitic languages on earth, with speakers concentrated across Somalia, Djibouti, Ethiopia, and Kenya, plus large diaspora communities worldwide.
Despite that reach, Somali is one of the most under-resourced major languages in computing. Its Latin-script writing system was only officially standardized in 1972, which means it has a far shorter formal written tradition than most other major world languages. When modern tools — spell-checkers, translation engines, AI assistants — need to handle Somali, they have almost no canonical, machine-readable reference to build on.
Dataset releases, new bootcamp cohorts, research papers, and lab updates — straight to your inbox. No spam, unsubscribe anytime.
20M+
Somali speakers worldwide
1972
Year the alphabet was standardized
4
Countries with major Somali-speaking populations
0
Prior machine-readable authoritative standards
The Problem SLS Solves
Today, anyone building software or AI systems that handle Somali runs into the same wall: there is no single authoritative source of truth for the language in machine-readable form. The consequences show up everywhere:
Spell-checkers are unreliable because there is no agreed, canonical word list.
AI language models trained on Somali produce inconsistent spelling, grammatical errors, and invented technical terminology.
Translation systems produce awkward or inconsistent results because there are no normative translation guidelines.
Technical communication in Somali is nearly impossible, because most modern domains — AI, medicine, law, engineering — have no standardized terminology.
Educators and publishers use inconsistent spelling and grammar conventions because no modern standard has the institutional authority to settle disputes.
Researchers who want to evaluate AI systems on Somali have no standard benchmarks to work with.
SLS is designed to close all of these gaps in a single, coherent, open-source project — rather than leaving every company and researcher to invent their own conventions in isolation.
What SLS Actually Standardizes
SLS is organized as a numbered catalog of standards, similar in spirit to Python PEPs or W3C Web Standards. Each standard addresses a specific layer of the language:
Orthography — official spelling rules, the Somali alphabet, punctuation, and capitalization (SLS-0001 Alphabet, SLS-0002 Spelling Rules).
Grammar — parts of speech, verb conjugation, noun morphology, and sentence structure (SLS-0003 Core Grammar).
Lexicon — a curated, schema-validated Somali dictionary with definitions, provenance, and unique IDs for every entry.
Terminology — standardized Somali vocabulary across 20 modern domains, from AI and medicine to law and engineering.
Translation — normative guidance for translating between Somali and English, including technical and idiomatic translation (SLS-0300, SLS-0301).
AI resources — datasets, system prompts, fine-tuning corpora, and RAG-ready knowledge chunks built for AI consumption.
Benchmarks — evaluation suites for testing AI systems on Somali grammar, spelling, and translation accuracy.
Built on Evidence, Not Invention
SLS follows one foundational rule: we do not invent Somali. We collect, preserve, study, and analyze existing authoritative Somali publications and scholarly works, and those form the empirical evidence base from which every standard, dataset, and specification is derived.
The project's evidence library holds original materials including:
Monolingual Somali dictionaries, including a digitized dictionary spanning 31 letter-based chapter files.
Authoritative grammar books, such as Barashada Naxwaha Af Soomaaliga by Puglielli & Mansuur.
Literature, poetry, and proverbs — gabay, maanso, maahmaahyo.
School and university textbooks.
Academic linguistic studies on Somali morphology, phonology, and syntax.
Domain-specific references in medicine, law, and science.
Historical texts from the post-1972 standardization period.
Every standard SLS publishes is traceable back to this evidence base. No rule, word, or grammatical claim is made without a citation to an authoritative source.
Comparable Projects in Other Languages
SLS draws design inspiration from well-established language infrastructure projects in other languages — none of them a perfect match, but each solving a piece of the same problem:
Unicode CLDR — locale-specific data (numbers, dates, plurals) for every major language, used by nearly all software.
The RFC Series / IETF — the model for SLS's numbered, versioned, lifecycle-governed standards process.
Real Academia Española and Académie française — normative authorities for Spanish and French spelling, grammar, and vocabulary.
KNAB (Latvia) — an official terminology authority protecting Latvian technical vocabulary from English dominance.
W3C Internationalization (i18n) — standards for representing languages in web technologies.
Masader (Arabic NLP) and IndicNLP / AI4Bharat — open repositories of NLP datasets and benchmarks, analogous to SLS's data layer.
AfricanNLP / Masakhane — community-driven NLP and translation research for African languages.
SLS is the first project of this kind for Somali at this level of formalism, governance, and machine-readability. Each project above solves part of the problem; SLS is attempting to build the full infrastructure stack for a single language.
Who Will Use SLS, and How
SLS is designed to serve several audiences, each consuming it differently:
AI and NLP researchers use the datasets for fine-tuning, the benchmarks for evaluation, and the RAG chunks to ground AI assistants in normative Somali knowledge — with a citable standard to claim conformance against.
Language model developers (Google, Meta, Anthropic, OpenAI, and others) use the structured lexicon, grammar specs, terminology datasets, and translation pairs as a canonical reference, instead of guessing at Somali conventions.
Software engineers and product teams use the validated wordlist for spell-checkers, the lexicon schema for dictionary features, and the system prompts for Somali AI assistants — implementing against a known, versioned standard rather than reverse-engineering scattered resources.
Translators and translation agencies consult the normative translation guidance and seed translation-memory tools with SLS-standard pairs for terminological consistency.
Educators, universities, and publishers cite SLS standards for spelling and grammar decisions, resolving disputes with reference to an evidence-backed, publicly governed standard rather than personal opinion.
Government and institutional bodies adopt SLS as the official machine-readable reference for Somali in digital services, official publications, and regulatory documents.
Linguists and language researchers cite SLS data in academic work and contribute new findings back through the formal proposal process.
The Somali-speaking public benefits indirectly, through every tool and AI system that implements SLS — better spell-checkers, more accurate AI assistants, and software that speaks correct Somali.
What the Ecosystem Looks Like After Completion
When SLS reaches v1.0 and the core standards are ratified Stable, a number of things become possible that aren't possible today.
A language model will be able to state: "This system implements SLS-0001, SLS-0002, and SLS-0003 v1.0" — a precise, verifiable claim that its Somali spelling, grammar, and core vocabulary conform to the standard. Third parties can then test that claim against SLS's own benchmark suites.
For software, a Somali spell-checker can be built by loading the lexicon and validated against the spelling benchmarks — and any update to the lexicon automatically flows downstream to every tool built on it. For education, publishers can cite a specific SLS section when explaining why a word is spelled a particular way, backed by the same evidence base linguists consulted when drafting the rule.
For AI training, fine-tuning datasets let any lab or researcher improve their model's Somali capability against a known, contamination-audited benchmark. For terminology, a question like "what's the Somali word for artificial intelligence" gets a standard, citable, permanently versioned answer instead of five competing guesses.
And for the global NLP community, SLS becomes a reference in academic papers, dataset cards, and model cards — establishing Somali as a language with proper, citable infrastructure, comparable in rigor, if smaller in scope, to what Unicode and CLDR provide for writing systems generally.
What SLS Is Not
To avoid confusion, it's worth being explicit about what SLS is not:
Not a dictionary app or website — SLS is the data and specifications that power such applications, not the application itself.
Not a corpus dump — every record has provenance, is schema-validated, and is reviewed; volume without quality is explicitly rejected.
Not a prescriptive authority that invents language — SLS standardizes what already exists in authoritative Somali sources, and only proposes new vocabulary through a governed terminology program requiring community consensus.
Not a finished product — SLS is a living standard that will keep growing as new domains are covered, new research emerges, and the language itself evolves.
Not a replacement for native speaker judgment — every standard is authored or reviewed by native speakers and linguists; SLS provides the infrastructure, human expertise provides the authority.
A Note on Governance and Trust
Language standards are only as trustworthy as the process that creates them. SLS is governed by a Language Council — a small body of named, accountable linguists and community representatives, modeled on how W3C working groups and Python's PEP editors operate. No standard reaches Stable status without Council review, and every ratified standard is permanently archived and versioned.
The standard is open-source under a Creative Commons BY 4.0 license for all linguistic content. That means anyone — including commercial AI companies — can freely use and build on SLS, as long as they credit the source. Fencing off Somali language infrastructure behind proprietary walls is explicitly rejected by the project's design.
SLS is maintained by Goobo Labs and the Somali Language Standard contributor community. The repository is open on GitHub — explore the specifications, review the evidence library, or contribute linguistic, technical, or domain expertise.