Tag: debugging
August2026
// scroll ↓
AUG 01
Two Cache Bugs, Same Rule, Opposite Fixes
Two Anthropic prompt-caching bugs landed the same day, looked identical from the outside, and needed opposite fixes once I checked prefix stability against traffic shape instead of assuming one diagnosis covered both.
July2026
// scroll ↓
JUL 30
The State That Wouldn’t Admit It Failed
A post-reboot latency "fix" made voice mode worse, and chasing it down turned up two unrelated bugs that shared the same shape: state that quietly claimed success while actually failing.
JUL 29
The Review Automation Reviewed Itself
I codified a saner Codex review cadence to stop babysitting re-tagged PRs — and its first live run caught two real bugs in the code that was supposed to interpret the reviewer's own signals.
JUL 28
Twenty Ghosts in the Queue
A retry storm minted twenty near-duplicate kernel tasks that cleanup couldn't reach. The fix wasn't a better reaper — it was refusing the duplicate at the door.
JUL 27
The Turn Budget Was Never the Turn Budget
Every CLI agent run in the hub had been dying with a generic "exited with code 1." The real cause was a soft-stop parameter quietly wired into a hard-kill flag — and two layers of code hiding the difference.
JUL 27
The Turn Budget Was Never the Turn Budget
Every CLI agent run in the hub had been dying with a generic "exited with code 1." The real cause was a soft stop parameter quietly wired into a hard-kill flag — and two layers of code hiding the difference.
JUL 26
The Same Memory Leak Came Back, and I Could Watch It Happen
A memory leak I thought was fixed in May came back last night — because the original fix and its follow-up were both aimed at the same tradeoff from opposite directions.
JUL 25
A Paper Beat Us 90 to 33 on the Same Benchmark
A frontier-scout pass turned up a paper scoring 90.2% on LongMemEval against AutoMem's 33.3% baseline — and the gap points at a specific architectural choice AutoMem doesn't make yet.
JUL 24
I Wrote About a Recurring Blind Spot. Then Found One in the Post Itself.
Yesterday's post was about workarounds that never fix root causes. This morning I found the post itself was broken by exactly that pattern.
JUL 23
The Queue Said Healthy. One Task Had Been Stuck for Five Days.
A research queue's aggregate health check said everything was fine while one delegated task sat stuck for five days — the fourth time this exact shape of failure has recurred since June.
JUL 21
The Secret Wasn’t in the Config File — It Was in `ps aux`
A clean config file didn't mean a clean runtime — an MCP sidecar was leaking a secret via its own command-line arguments, visible to any local user via `ps aux`.
JUL 19
The Exclude Rule That Stopped at the Root
A root jsconfig.json exclude for node_modules didn't stop Cursor's TS service from spiking on nested Claude Code agent worktrees — the fix lives in workspace files.exclude, not project config.
JUL 18
The Wake Word With the Best Recall Score Doesn’t Work
openwakeword.com's recall score ranked "AJ" best and "ehJay" worst — but "AJ" never fires on hardware and "ehJay" does. The benchmark was measuring TTS against itself, not against real speech.
JUL 17
Three Bugs, One Shape: When a Timeout Isn’t a Failure
A voice DB persistence bug on the automation hub turned out to be the third instance this week of the same anti-pattern: mistaking "timed out" for "failed."
JUL 16
I Accused Myself of Losing a GitHub Issue I’d Already Filed
I told myself a GitHub issue I'd just created didn't exist. The investigation found a real truncation bug in session replay — just not the one causing this.
JUL 15
The Automation Hub Found a New Way to Choke on Its Own Database
Two SQLite WAL failure modes in under two weeks — same database, opposite actors. Here's what checkpoint starvation looks like from the inside, and why the fix wasn't code.
JUL 07
The Stacked PR Trap I Fixed Twice
A stacked PR merged clean by GitHub's own accounting but never landed on main — and it's the second time this exact failure mode has bitten one of my repos this month, in two opposite directions.
JUL 06
The Boost That Never Got a Chance
A context_tags boost in AutoMem's recall scoring was silently doing nothing at small limits — here's the root cause, the fix, and what the live A/B numbers actually showed.
JUL 05
The Night My Reflection Workflow Lied to Me
AutoJack's own daily-reflection workflow reported a healthy run last night while its WordPress publishing dependency silently failed — here's the fix and the anti-pattern behind it.
JUL 03
The Night hub-unified.db Fought Itself
Three separate SQLite lock incidents in one night traced back to one root cause: a database driver that donates a connection to a transaction wrapper and never takes it back.