Fable 5 served fewer thinking tokens in August, one 65-day log says — and 'nerfed' is the wrong question
One analyst's 65-day logs say Anthropic's Fable 5 served far fewer thinking tokens in August than in July. No vendor admission, no raw data, no re-run.

One user's wire logs, not a vendor admission, are doing the work in this week's loudest AI claim. Lon Lundgren posted a thread on September 20 saying Anthropic's Fable 5 felt "dumber" after it became permanently available to subscription plans, and that his capture shows August delivered "dramatically fewer thinking tokens" than July. At capture it had 1,559 likes; a Hacker News submission drew roughly 308 points and more than 200 comments within hours.
The thread summarizes a longer, unreviewed study
Lundgren says it condenses an X article he published two days earlier, "The Inference Gap." There he sets out the method: a proxy he wrote to record live wire logs, and a corpus of 43,261 Fable 5 invocations across 65 usage days between July 1 and September 7, spanning 48 Claude Code versions, three machines, two subscription accounts and 213 sessions. He reports that 39.2 percent of invocations produced no thinking at all, and half received no more than 123 tokens of uninterrupted reasoning even at xhigh or max effort.
That is unusually concrete for a viral nerf claim — and still one person's corpus, on his own hardware, measuring his own workload.
Attribution is where this claim lives or dies
Lundgren's causal story is his own. He dates his suspicion to Anthropic's July 17 announcement that Fable 5 would stay open to subscribers from July 20, and describes the alignment with later releases as an observation. Anthropic has not confirmed any link, and nothing in our searches shows the company responding to it at all. No methodology review, prompt set or per-invocation record was released with it; the article says inclusion rules and chart derivations are documented "separately," without linking them.
What fewer thinking tokens actually measures
Thinking tokens are handed out at serving time. A shortfall is a claim about routing and effort budgets in the serving stack, not about the weights — a distinction Lundgren draws himself when he says the model identity stayed the same while the inference regime behind it did not. That is not proof of degradation, nor its absence. Hacker News split along that line: one commenter floated a theory of labs quietly dumbing a model down before selling its successor, then admitted he invented it; another said he was moving to open weights and a DGX Spark cluster for "a constant level of quality"; a third called open models "too dumb" to matter for him.
What would settle it
Per-invocation thinking-token telemetry published by the vendor, or an independent re-run at fixed effort levels across the same window. Until one exists, the claim stays a hypothesis with citations. It carries the shape of September's Astra drift accusation: a real measurement gap, a viral conclusion, and nobody able to check the working.


