What generic benchmarks miss
Generic translation benchmarks miss local phrasing, code-switching, tone, intent, and cultural context.
DataYetu gives AI teams deploying in Kenya and East Africa continuously refreshed, rights-cleared data to measure and reduce Swahili and Sheng meaning failures.
For AI labs, voice AI companies, translation teams, conversational AI products, and other model builders.
The meaning gap
What generic benchmarks miss
Generic translation benchmarks miss local phrasing, code-switching, tone, intent, and cultural context.
How Swahili & Sheng failures show up
Observed Swahili and Sheng meaning failures often show up as wrong urgency, flattened tone, missed negation, or intent that flips when speakers mix languages mid-sentence — failures English-centric evaluation sets rarely surface.
What we collect and validate
We collect and validate real Kenyan language usage so teams can test whether their models actually work for the people they serve.
Failure signatures English-centric sets rarely surface
Product
What design partners receive from the language pilot — not a multi-industry platform claim.
Swahili and Sheng examples collected through consent-based contributors and approved data partners.
Independent native reviewers validate meaning, intent, tone, and translation quality.
Code-switching, slang, negation, urgency, tone, and intent labels attached to each example.
Transcripts or audio included only where permissions allow, with personal information removed.
Documented cases where models miss local meaning — so teams can reproduce and debug.
Evaluation results showing improvement or regression as models and language usage change.
Consent, licensing context, and permitted uses attached to each dataset release.
Sample evaluation record
Illustrative schema only — field structure for a pilot evaluation pack. No invented source text, translations, or model results.
| Field | Contents |
|---|---|
| source_text | Consent-based Swahili / Sheng / code-switched example |
| language_variety | sw | sheng | code-switched (+ context notes) |
| meaning_annotation | Human-reviewed translation and intended meaning |
| labels | code-switching · slang · negation · urgency · tone · intent |
| model_output | Customer model response under test |
| failure_mode | Where meaning, tone, or intent breaks (when observed) |
| provenance | Consent, permitted uses, release version |
A repeatable, consent-based pipeline — designed to be auditable at each step.
Collect language examples through consent-based contributors and approved data partners.
Record provenance, language variety, context, and permitted uses.
De-identify personal information and apply access controls.
Independent native reviewers validate meaning, intent, tone, and translation quality.
Test customer models against the dataset.
Deliver evaluation packs, training-ready data where licensed, failure reports, and benchmark results.
Refresh the dataset as language usage and model behaviour change.
Early stage
We're building the first pilot around a focused question: can rights-cleared, context-rich Kenyan language data help AI teams detect and reduce failures that generic benchmarks miss?
The numbers below are pilot targets being validated — not results we claim to have achieved.
Target
Human-reviewed cases where models miss Swahili or Sheng meaning.
Target
Errors a design partner can re-run and inspect.
Target
Measure how consistently native reviewers agree on meaning and labels.
Target
Structured conversations with AI teams deploying in Kenya or East Africa.
Target
One partner runs their model against the pilot evaluation set.
Not because we claim a proven multi-industry platform — because the pilot is narrow, local, and designed to produce countable evidence.
Kenyan Swahili and Sheng — including code-switching — not a generic “African data” abstraction.
We start from where models break on local phrasing, tone, and intent that translation benchmarks miss.
Each release is designed to carry contributor consent, language variety, context, and permitted uses.
Meaning, intent, tone, and translation quality are validated by independent native reviewers before delivery.
Success is defined in countable outputs: reviewed failures, reproducible errors, reviewer agreement, and a first design-partner evaluation.
We're early. If you're deploying models in Kenya or East Africa and care about meaning failures generic benchmarks miss, join the pilot.
Language and Linguistics is the active category. Other domains are future applications — not current proof points.
Continuously refreshed Swahili and Sheng evaluation and training data for AI teams deploying in Kenya and East Africa.
Later applications
Later application — not part of the current pilot.
Later application — not part of the current pilot.
Later application — not part of the current pilot.
Later application — not part of the current pilot.
Later application — not part of the current pilot.