Engineering · Chapter 14 of 16·2 min read
The Fifteen Hours a Model Just Stopped Answering
One of the AI providers we route through ran out of account balance. For fifteen hours, that looked exactly like our software being broken — because nothing told anyone, including us, what was actually going on.
Listen to this story
A model that ran out of money
Nia doesn't run its own models for every request — it routes through several AI providers, and one of them ran its account balance down to zero. That's not a bug in anything we wrote. It's a billing event. Our system just had no idea that's what it was looking at.
What people actually saw
A live chat question failed outright, showing a raw internal billing string that meant nothing to whoever was reading it. Separately, a handful of scheduled automated runs failed the exact same way, with nobody told any of it had happened — they just silently didn't run.
The error wasn't wrong. It just looked exactly like every other kind of broken.
Fifteen hours before anyone noticed
The account sat empty for more than fifteen hours before anyone realized what was actually going on. Not because nobody was paying attention — because a drained account and a genuine software failure produced the identical symptom: a request that just didn't work, with no label telling you which one you were looking at.
What we built instead
The moment a provider account runs dry mid-request, the system now retries automatically on a different provider — and says so out loud, right in the conversation: which model it switched to and why, instead of either failing silently or swapping providers without telling you. Every place that kind of error can surface — a live chat, a scheduled run, the command line — shows the same plain, specific sentence instead of a raw provider string nobody outside an engineering team could read.
And if it happens again, we get told directly instead of waiting for someone to notice fifteen hours later.
The actual fix wasn't "watch the balance more closely"
An account running low is going to happen again eventually — that's just true of running a service that depends on other services. The fix that matters isn't promising to top it up sooner. It's making sure that when it does happen, it can never again look identical to the software just being broken.