Engineering · Chapter 8 of 16·2 min read
The Messages That Disappeared
A successful "sent an email" would vanish from your history while a failed attempt right next to it stayed visible. Real conversations, silently corrupted — and the debug log meant to catch it had been quietly broken too.

Listen to this story
A message that should have existed didn't
The report was strange in a specific way: a successful action — an email that genuinely got sent — was missing entirely from a conversation's history, while a failed attempt sitting right next to it in the same chat displayed just fine. Not slow. Not wrong. Gone, as if it had never happened.
The JSON was getting cut in half
The actual mechanism: occasionally, a result block would land in storage with its content cut off mid-value, and whatever came right after it in the response got fused directly onto the broken end — no proper close, no valid structure. Every single place downstream that tried to read that block failed to parse it and silently dropped it rather than showing anything broken. The record wasn't corrupted-looking. It was just absent.
About 40 real conversations, out of roughly 1,150
Checked against the real, live data: a small but real fraction of stored action records were affected — the ones with unusually large results were disproportionately likely to hit it. Rare. Also real, and real enough that the people it happened to had already noticed something felt off.
The tool meant to catch this had been broken too
The debugging log that should have made this easy to chase down had itself been silently failing in production, for an unrelated reason — the environment it ran in didn't allow it to create the folder it expected to write into, and it had been designed to fail quietly rather than crash anything. So there was no trail. The very thing built to explain a mystery like this one had nothing to say about it.
The safety net meant to catch this had a hole in it too, and nobody knew until they reached for it.
Repairing, not just preventing
The fix wasn't only about stopping it from happening again. It included going back and actually repairing the messages that were already broken — structurally closing the cut-off content rather than discarding it, recovering every real field that was still intact. Tested directly against the real corrupted data: every affected message came back clean, and nothing that was already fine got touched.
