Releases, bootcamp news, research notes, and announcements.
One week after we uploaded a 13-hour, fully Somali-language Data Science & Machine Learning course to YouTube, it passed 12,000 views — the biggest video and course we've published on the channel so far. Here's what it was, and what we're doing about what it told us.
A demanding mathematics translation experiment revealed how the Somali Language Standard can help frontier models produce Somali that is more accurate, natural, consistent, and explainable.
Our fifth program in this series, and the first not built around data science at all. Git & GitHub closes a gap every earlier cohort surfaced: people who could write working code but weren't yet confident collaborating on it the way professional teams and open source projects do.
In June 2026 we ran our third Data Science & Machine Learning Bootcamp, and it's the biggest curriculum we've built for the program — eight core lessons on the full ML workflow, plus seven bonus sessions covering deep learning, generative AI, ethics, and career paths.
Somali is spoken by more than 20 million people, yet almost nothing in software or AI can point to a single authoritative, machine-readable source for its spelling, grammar, or terminology. Here is what the Somali Language Standard is, why it exists, what it's built on, and what changes once it's complete.
We've formally published the first two standards of the Somali Language Standard (SLS) — an open, machine-readable, CI/CD-validated framework for the Somali language, starting with the alphabet and the standards process itself.
After two Data Science & Machine Learning cohorts, we kept seeing the same bottleneck: people excited about ML who got stuck on plain Python before they ever reached a model. Python for Everyone is the bootcamp we built to close that gap on its own.
Our team shipped a Rust-based compiler for Soplang — Cranelift JIT and ahead-of-time builds, with syntax written entirely in Somali keywords.
Our largest Somali text release yet: 1.77M clean documents, 591M words, and ~887M tokens from HPLT, CC100, mC4, OPUS, MADLAD, and MT560 — filtered through a six-stage pipeline.
Five months after our first cohort, we ran a second Data Science & Machine Learning Bootcamp in February 2026. The core format held — one month, the same seven-stage ML workflow — but the curriculum itself changed in specific, deliberate ways.
Foreign multilingual tokenizers over-segment Somali — BERT-base spends 2.69 tokens per word. Our native BPE tokenizer, trained on 529M words of SomNLP-Corpus, gets that down to 1.53: a 1.75× improvement.
In September 2025 we ran our first Data Science & Machine Learning Bootcamp — a one-month, hands-on program built to take a complete beginner through the full ML workflow to a deployed project. Here's what we taught, why we built it, and what it set in motion.
Dataset releases, new bootcamp cohorts, research papers, and lab updates — straight to your inbox. No spam, unsubscribe anytime.