Voices that sound like us: Nigerian and African speech for Cast
Global TTS defaults to US/UK English. We are building toward Cast-bound voices for Nigerian English, Pidgin, and indigenous languages — consent-first data, adapters per language pack, and continuity across Motion shots — not a fake accent slider.
Why we care
If you generate dialogue with a stock global TTS voice, African stories often get flattened: Nigerian English pulled toward US/UK rhythm, Pidgin mangled, Yoruba / Igbo / Hausa missing or wrong, and the “same” character sounding like a different person every shot.
Qweek’s bar is higher than “Nigerian accent preset.” A Cast member should own a stable voice across the board — same speaker, honest language — especially once Motion locks lip sync to dialogue.
This note is our research direction: how we plan to adapt voice models for Nigerian and broader African speech and bind them to Cast via voice_model_id.
Languages we care about first
| Code | Variety | Notes |
|------|---------|-------|
| en-NG | Nigerian Standard English | Not British, not American |
| pcm | Nigerian Pidgin (Naijá) | Everywhere in dialogue; spelling varies |
| yo | Yoruba | Tone matters |
| ig | Igbo | Tone matters |
| ha | Hausa | Different phonotactics |
v1 focus: en-NG + Pidgin with real code-switching, then indigenous packs.
People do not speak one language per sentence. Mixed English–Pidgin–indigenous lines are the product, not an edge case. Models that assume one language tag for the whole utterance will fail here.
What “good” means for us
Three things at once:
- Same person — enroll a Cast voice; later lines still sound like them
- Honest variety — listeners from that variety do not hear “generic African” or US default
- Understandable — accent authenticity cannot destroy intelligibility
Commercial clone APIs are useful for design-partner trials. They are not the long-term moat. Open / region-aware adapters we control are.
How we will adapt models
We will evaluate open TTS families and zero-shot speaker models, plus commercial baselines.
Preference: LoRA / adapters per language pack, and optional per-character speaker adapters on top — so we do not smash multilingual ability every time we add Yoruba.
For tone languages, grapheme-to-phoneme is not a side quest. Wrong tone marks wreck naturalness. Pidgin needs a light normalization pass before synthesis because orthography is not standardized — and we should still respect talent-preferred spellings where we can.
Data and consent (non-negotiable)
Rough scale we are planning around:
| Goal | Audio | Notes | |------|-------|-------| | Instant clone trial | minutes | Accent fidelity varies hard | | Solid character voice | ~0.5–2 hours | Read + conversational | | Language pack | tens–hundreds of hours, many speakers | Diversity over one celebrity |
Every clip needs consent scope (previs / commercial / transferable), language tags, and workspace isolation — tenant A’s voices never train tenant B’s models.
We will compensate talent, prefer commissioned recordings over scraped YouTube, and stay away from celebrity cloning without a license.
Coverage matters: gender, age, region (Lagos, Abuja, Port Harcourt, Kano, …). A pack that only fits one coastal studio accent just encodes a new bias.
How we will measure
Automatic checks: speaker similarity to enrollment, accent/variety classification, ASR intelligibility, prosody vs reference, code-switch mistakes.
Human ratings with Nigerian and diaspora listeners — naturalness, accent authenticity, character consistency across a sequence. US-only raters under-penalize accent erasure. We will not use that as a ship gate.
Product gate for Motion: same Cast voice still feels like the same person across a short sequence of shots.
Hooking into Cast
Cast.voice_model_id → voice registry entry
entry → provider / adapter / language pack / enrollment profile
At Motion time the worker looks up the voice, synthesizes dialogue, then conditions lip sync / motion on that audio plus the locked still.
Fallback chain if something is missing:
- Workspace adapter for that character + language
- Shared
en-NG/ Pidgin pack - Stock provider voice with accent tag (degraded)
- Motion without dialogue track
Latency still matters — adapters have to stay small enough for workers.
Phased plan
| Phase | What |
|-------|------|
| A | Zero-shot / commercial baselines on en-NG scripts |
| B | Multi-speaker LoRA language pack for en-NG + Pidgin |
| C | Per-talent adapters bound to Cast; sequence consistency |
| D | Yoruba / Igbo G2P + packs; harder code-switch suites |
We will publish Phase A/B numbers as design-partner trials land — not before we have listener panels that can actually hear the difference.
Risks we are not waving away
- Voice cloning abuse → verified consent; watch watermark / detection work
- Pidgin orthography politics → document norms; allow preferred spellings
- Public-figure extraction → out of scope without license
- Extractive data culture → pay people; do not scrape first and apologize later
Bottom line
This is not a cosmetic accent slider. It is Cast-native voice continuity for stories that sound like Nigeria and Africa — with consent, measurement, and a clear path from language packs to per-character adapters.
If Motion looks right but sounds like nowhere, we still failed.