One recording, six datasets.
A single consented Bantu→English code-switch verb take is not one labelled example. Its switch point is logged as switch_ms, so it self-segments with no manual annotation — and becomes the ground truth for six things frontier labs otherwise buy from six vendors.
Code-switch ASR
Natural Bantu→English switching with the exact switch point pre-labelled — training + evaluation data for code-switch recognition and language ID.
Bantu-accented English
The English half, spoken by a native Bantu speaker — accent-robustness and fairness data for Bantu-accented speech.
FSI alignment target
The verb decomposes into its Full Syllable Inventory — the tone-bearing units a forced aligner needs.
Pronunciation & TTS
A consented native pronunciation keyed to meaning and its syllabary — a clean, tone-aware basis for synthesis.
Self-grading benchmark
The segmentation is already structured, so the take grades a model against itself — no labelling vendor, no drift.
Domain & homograph truth
The verb is anchored to a specific meaning (concept + VT cell) — the meaning flat text can't encode.
| Language | Recorded takes | Verb types recorded | Curated source rows |
|---|---|---|---|
| Luganda lug | 1863 | 1 | 934 |
| Bemba bem | 1646 | 1 | 4174 |
| Tonga toi | 121 | 1 | 1547 |
| Lozi loz | 0 | 0 | 52 |
| Nyanja nya | 0 | 0 | 519 |
| Swahili swh | 0 | 0 | 121 |
The corpus is early and climbing — infinitives first, then up the paradigm. The audio itself is licensed; get evaluation access.
One consented Bantu→English take, shown two ways at once — the timed switch you hear (Bantu, then Bantu-accented English at switch_ms) and the exact record a model ingests. Tab between the code-switch takes and the Bantu-only takes.
▸ Record schema — the exact fields the API returns (public)
Never present in any record: — speaker identity & storage paths are stripped.