← Blog

Engineering · Chapter 16 of 16·2 min read

The Panel That Said Zero When the Answer Was Sixteen

A real 16-agent run came back mislabeled as zero, one result got stuck showing a spinner forever, and fifteen of the sixteen agents' actual answers were silently cut to fit a size limit. Here's what we found and fixed.

Listen to this story

0:00
Nia

Watching agents work, or supposed to

The whole point of the live parallel-agents panel is trust: you should be able to see what's actually happening — how many agents are running, what each one is doing — instead of staring at one spinner and hoping. One real run showed exactly why that matters, by failing at it in three different ways at once.

Three things wrong in the same run

A real fan-out had spawned sixteen agents at once. The panel's own label said "Spawned 0 parallel agents." One saved result was left in a permanently pending state — a spinner with no answer ever coming, indefinitely. And the combined answer that sixteen agents had actually produced came back as a large block of text that ran into an internal size limit on how much a single result can carry.

The part that actually cost something

That size limit didn't fail loudly — it just quietly cut the block down to fit, from the end. Fifteen of the sixteen agents' real answers were discarded in that cut, with nothing on screen indicating any of it had been trimmed at all. Close to 450 seconds of genuine compute, gone, meaning the fan-out had to be run again from scratch to get the answer back.

A spinner that never resolves and a result that quietly loses fifteen out of sixteen answers are the same kind of bug.

What we fixed

The count shown now reflects what actually ran, not a stale read taken before the agents had spawned. A call that's genuinely finished — succeeded, failed, or interrupted — can no longer be left displaying as still running; there's now an explicit step that closes those out instead of leaving them open indefinitely. And a large combined result now shares its space fairly across every agent that contributed to it, instead of the earliest ones eating the whole budget and the rest simply vanishing — short answers are kept whole, long ones keep the part that actually matters.

The one rule this comes back to

A feature built specifically so you can watch and trust what's happening can't ever quietly claim something is fine when it isn't — not the count, not the spinner, not the result. That's the one thing it exists to get right, and this run showed us three separate ways it hadn't.

That’s every post so far.

Back to the Blog