Most text-to-speech sounds like a polite robot reading a phone bill. That's fine for a recipe. It's terrible for a story. When every sentence has the same pitch, pace and emotional weight, dialogue blurs into description and a reader who is already working hard to decode loses the thread of who is speaking.
what wobble does differently
When you scan a page, Wobble sends the recognised text to a language model with one job: split it into narrator, named characters and unattributed speech. Each character keeps a consistent voice across chapters, the wolf is always gravelly, the small sister is always small. Sound effects ("BANG!", "meow") get their own short, punchy treatment.
The result sounds closer to a radio play than a screen reader. And that turns out to matter for comprehension.
why it lifts comprehension
Two effects show up consistently in classrooms we've tested with:
- Characters become tracking handles. "Wait, was that the wolf?", yes, because the wolf has a voice. Children stop losing the through-line.
- Emotional cues come back. Sarcasm, fear, jokes, all the things that make a story worth reading, survive the trip from page to ear.
what we got wrong first
Our first version used twenty different voices and let the model assign them freely. It was chaos. Children couldn't track who was who because the same character sometimes got a new voice in chapter three. Now the model gets the cast list before it speaks, and voices are sticky for the whole book.
We also learned that narrator voice matters more than character voice. A warm, mid-paced narrator is the spine of the experience; you can be more playful with the cast around it.