Engineering · Chapter 6 of 16·2 min read
The Day One Conversation Spent 2.4 Million Tokens Searching for a Single Tool
A real chat spiraled into calling the same lookup tool 41 times in a row before it stopped on its own. Here's how we found it, and how fast we had to fix it.

Listen to this story
It happened in a real conversation
The morning of August 10th, a live chat — a real conversation, not a test — spiraled into calling the same tool-lookup function 41 times in a row, hunting for a tool it had already been told about. By the time it stopped, that single conversation had burned 2.4 million tokens.
Not a benchmark. Not a stress test. One real chat, on a real day.
We already had a system that should have caught this
The frustrating part of the finding wasn't that no protection existed — it's that a very similar one already did, for a related failure mode: a model re-searching for information instead of using what it had already found. That existing budget just had never been extended to cover this specific lookup tool. The gap wasn't a missing idea. It was one tool that never got added to a list.
A debugging near-miss
While scoping the fix, an early search across the core loop file came back completely empty — as if none of the existing protection logic was even in there. For a moment, that looked like the real explanation: the whole guard system had apparently never been wired in.
It hadn't been. The search tool itself was silently treating that file as binary and skipping it, because of one legitimate low-level separator character sitting in a single line of otherwise ordinary code. Switching tools turned up the real picture: the guard system existed and mostly worked. It just had this one gap.
The empty result almost became the conclusion, instead of a clue that something else was wrong.
Same day: scoped, built, and shipped
The finding, the design doc, the fix, and a second fix extending the same protection to sub-agents — every one of those landed the same day. Scoped first, deliberately, in writing, before any code changed. Then implemented. Then mirrored into the separate code path that runs when Nia spawns its own sub-agents, because a guard that only protects the main conversation and not the agents it can spawn isn't actually a fix.
Every limit like this has a real story behind it
None of the budgets, caps, or nudges scattered through Nia's core loop were designed in from a whiteboard on day one. Each one traces back to something that actually happened — an actual conversation, an actual token count, an actual moment where someone had to explain why a single chat cost what it cost, and then made sure it couldn't happen the same way twice.
