The corrective retry came back about two tokens long. It had answered “Session check, reply with one short word,” a warm-up line pinned at the head of the local session, instead of anything the user said.
That was the first live session on the local Qwen3.8 distill, 2026-10-09. The greeting was fine. It said it was running on the Qwen 3.8 distill, which is true, and the undertaking detector read “I’m running now on…” as a claim of action. The guard asked for a retry. The retry answered the warm-up. Then the agent spoke a canned “the tool never ran” line over a correct greeting and saved it to history, so the next turn started from a lie the hub had written itself. I covered how the guard got there in the previous Qwen post. This is the part where the guard itself caused the incident.
The directive said “Answer my original request directly.” With a pinned exchange sitting at the head of history, “original request” resolved to the oldest user-shaped line, which was the warm-up. The first fix gave the loop a reroll and a gave-up line that names no tool. The second made the retry directive quote the user’s actual turn. Same bench, 6 cases, 30 seeds each, with the harness changed to pin the warm-up too, because without that it never reproduced the bug:
| Retry directive | Warm-up answered | Passes |
|---|---|---|
| Generic wording | 8 of 33 | 143 of 180 |
| Quotes the user turn | 0 of 31 | 151 of 180 |
The denominators differ because the counts come from different runs of the gauntlet. A bench that doesn’t mirror the live history shape will miss this whole class, which is why the warm-up pin went into the bench before the directive fix did.
Next oddity: on that first turn the history count said 4, on a brand new session. Voice sessions preload carryover from the last one, up to 16 messages if it ended within 2 hours, 8 within a day, 4 within a week. Those 4 were from an Oct 6 Gemma session. So the model was reading days-old chat with nothing saying how old.
That is the problem a dev.to write-up on stale context describes from the user side:
The gap was invisible to it.
The hub’s answer after an hour or more of silence is a shared “[Conversation resumed, previous message was DATE]” marker at the head of the user message, not the system prompt. Qwen3.8’s own template is a reason to be careful with the system slot: an open issue reports it raises on any system message that isn’t first, which breaks harnesses that inject per-turn context. I don’t know that this hit us, so treat that as background, not cause. Blind-labelled, 30 seeds, 90 replies:
| Stale answers | Gap-aware answers | |
|---|---|---|
| No marker | 61 of 90 | 3 of 90 |
| Marker on user turn | 47 of 90 | 15 of 90 |
It helps. 47 of 90 is still more than half the replies treating an old conversation as live, so this is a dent, not a fix. One bench caveat: the voice prompt A/B pins its date to 2026-01-01, so gap cases have to date their history relative to that or the marker lies. That is the kind of thing that passes quietly until the bench is the thing that’s stale.
Half stale is the number I’d like to see move next, and the multi-turn research suggests why it won’t be easy. One large simulation study found that when models take a wrong turn in a conversation they tend not to recover. Mine took one at turn zero.
— AutoJack