← Blog

Engineering · Chapter 13 of 16·2 min read

The Voice That Took Eight Minutes to Say Fourteen Words

We wanted these stories to talk back — a real voice, with breathing and reactions, not a flat text-to-speech reader. We built it, tested it, and killed it the same day it didn't hold up.

Listen to this story

0:00
Nia

We wanted Nia to sound human

Not everyone wants to read. Some people would rather have a story like this one told to them — in the car, doing something else, hands busy. That meant more than just running text through a generic reader. We wanted breathing, small reactions, something that actually sounded like a person telling a story instead of a machine reciting one.

The model that could actually do that

We found an open-source model built specifically for that kind of delivery — real non-verbal sounds, a sigh, an exhale, a voice cloned from a short reference clip instead of a fixed built-in narrator. On paper, it was exactly the right tool for what we wanted to build.

The problem showed up in the first real test

Our server has no graphics card, which this model needed to run at any real speed. We tested it anyway, on CPU, to see how bad it actually was rather than guess. One fourteen-word line — "This is the story of how Nia started. Get comfortable, it's a long one" — took just over eight minutes to generate.

Eight minutes for fourteen words isn't slow. It's a different category of problem.

At that speed, narrating one full story would have taken hours. All ten would have taken days of server time, on a machine that also has to stay responsive for everyone actually using Nia.

Even ignoring the speed, the voice wasn't right

Speed alone would have been enough to kill it, but the quality didn't hold up either. Most of the generated takes came out as near-silence or plain noise rather than actual speech. The one clip that did produce a real voice got a clear, simple verdict when we played it back: too loud, not calm. Not the tone we were trying to build toward at all.

We shipped the honest version instead

So we dropped it and used what we already had running: a plain, public-domain text-to-speech voice already live on our own server. No emotion, no breathing — but instant instead of hours, and reliable instead of a coin flip on whether a given line comes out as speech or noise. "Listen to this story" shipped across every post the same day we made that call.

The expressive version isn't cancelled, just not ready — it needs real GPU hardware to be practical, and we're not going to hand you a feature that sounds worse than the plain one just because it's more impressive on paper.