DataYetu gives AI teams deploying in Kenya and East Africa continuously refreshed, rights-cleared data to measure and reduce Swahili and Sheng meaning failures.

For AI labs, voice AI companies, translation teams, conversational AI products, and other model builders.

Request a pilot briefing

The meaning gap

Generic benchmarks miss how Kenyans actually speak

local phrasing
code-switching
tone
intent
cultural context

What generic benchmarks miss

Generic translation benchmarks miss local phrasing, code-switching, tone, intent, and cultural context.

How Swahili & Sheng failures show up

Observed Swahili and Sheng meaning failures often show up as wrong urgency, flattened tone, missed negation, or intent that flips when speakers mix languages mid-sentence — failures English-centric evaluation sets rarely surface.

What we collect and validate

We collect and validate real Kenyan language usage so teams can test whether their models actually work for the people they serve.

Failure signatures English-centric sets rarely surface

wrong urgency
flattened tone
missed negation
flipped intent

Product

Continuously refreshed Swahili and Sheng evaluation and training data

What design partners receive from the language pilot — not a multi-industry platform claim.

Consent-based language examples

Swahili and Sheng examples collected through consent-based contributors and approved data partners.

Human-reviewed translations and meaning annotations

Independent native reviewers validate meaning, intent, tone, and translation quality.

Linguistic and pragmatic labels

Code-switching, slang, negation, urgency, tone, and intent labels attached to each example.

De-identified transcripts or audio

Transcripts or audio included only where permissions allow, with personal information removed.

Reproducible model failure cases

Documented cases where models miss local meaning — so teams can reproduce and debug.

Benchmark results over time

Evaluation results showing improvement or regression as models and language usage change.

Provenance and usage permissions

Consent, licensing context, and permitted uses attached to each dataset release.

Sample evaluation record

Illustrative schema only — field structure for a pilot evaluation pack. No invented source text, translations, or model results.

FieldContents
source_textConsent-based Swahili / Sheng / code-switched example
language_varietysw | sheng | code-switched (+ context notes)
meaning_annotationHuman-reviewed translation and intended meaning
labelscode-switching · slang · negation · urgency · tone · intent
model_outputCustomer model response under test
failure_modeWhere meaning, tone, or intent breaks (when observed)
provenanceConsent, permitted uses, release version

From collection to delivery

A repeatable, consent-based pipeline — designed to be auditable at each step.

  1. 1

    Collect

    Collect language examples through consent-based contributors and approved data partners.

  2. 2

    Record provenance

    Record provenance, language variety, context, and permitted uses.

  3. 3

    De-identify

    De-identify personal information and apply access controls.

  4. 4

    Validate

    Independent native reviewers validate meaning, intent, tone, and translation quality.

  5. 5

    Test models

    Test customer models against the dataset.

  6. 6

    Deliver

    Deliver evaluation packs, training-ready data where licensed, failure reports, and benchmark results.

  7. 7

    Refresh

    Refresh the dataset as language usage and model behaviour change.

Early stage

What we're validating

We're building the first pilot around a focused question: can rights-cleared, context-rich Kenyan language data help AI teams detect and reduce failures that generic benchmarks miss?

The numbers below are pilot targets being validated — not results we claim to have achieved.

  • Target

    30 to 50 reviewed failure examples

    Human-reviewed cases where models miss Swahili or Sheng meaning.

  • Target

    5 to 10 reproducible model errors

    Errors a design partner can re-run and inspect.

  • Target

    Independent reviewer agreement

    Measure how consistently native reviewers agree on meaning and labels.

  • Target

    3 to 5 buyer discovery conversations

    Structured conversations with AI teams deploying in Kenya or East Africa.

  • Target

    First design-partner evaluation

    One partner runs their model against the pilot evaluation set.

Why DataYetu is positioned to do this

Not because we claim a proven multi-industry platform — because the pilot is narrow, local, and designed to produce countable evidence.

Narrow language focus

Kenyan Swahili and Sheng — including code-switching — not a generic “African data” abstraction.

Meaning-failure lens

We start from where models break on local phrasing, tone, and intent that translation benchmarks miss.

Consent and provenance pipeline

Each release is designed to carry contributor consent, language variety, context, and permitted uses.

Independent native review

Meaning, intent, tone, and translation quality are validated by independent native reviewers before delivery.

Measurable pilot targets

Success is defined in countable outputs: reviewed failures, reproducible errors, reviewer agreement, and a first design-partner evaluation.

Help us build the first Swahili and Sheng evaluation set

We're early. If you're deploying models in Kenya or East Africa and care about meaning failures generic benchmarks miss, join the pilot.

Pilot focus

Language and Linguistics is the active category. Other domains are future applications — not current proof points.

Active pilot

Language and Linguistics

Continuously refreshed Swahili and Sheng evaluation and training data for AI teams deploying in Kenya and East Africa.

Later applications

Later

Healthcare

Later application — not part of the current pilot.

Later

Finance

Later application — not part of the current pilot.

Later

Agriculture

Later application — not part of the current pilot.

Later

Customer support

Later application — not part of the current pilot.

Later

IoT & wearables

Later application — not part of the current pilot.