DeepSeek V4.1 Flash tops AA's new agentic benchmark by 0.4 points. The index under it moved twice in a week
An open-weights model took the top row of Artificial Analysis's new private agentic benchmark by 0.4 points, on an index revised twice in just three days.

An open-weights Chinese model is sitting on the top row of the newest agentic benchmark from Artificial Analysis. The chart reads DeepSeek V4.1 Flash at maximum reasoning effort at 68.9% on AutomationBench-AA, ahead of GPT-6 Astra at 68.5%. That is a 0.4-point lead, and it landed on an index the same operator had already rebuilt twice inside one week. The margin matters precisely because it is small: it is not the size of a capability gap, it is the size of a rounding artefact on a private test set nobody outside Artificial Analysis can inspect.
First place, described by its publisher as a tie
Artificial Analysis posted the row on September 10 at 20:36 UTC, and the wording of its own announcement is the most useful sentence in the episode:
"DeepSeek V4.1 Flash takes first place on AutomationBench-AA with 69%, equal to GPT-6 Astra (69%) and slightly above Grok 4.6 (67%)."
Equal. The live chart separates the models anyway, at 68.9% against 68.5%, and the gap stays under half a point down the rest of the table. On a benchmark that awards credit per completed objective across 657 tasks, 0.4 points is roughly two or three objectives out of thousands — the kind of distance that disappears when a serving configuration changes or one task flips from a guardrail violation to a pass. AA's post spends its words on the size of the jump instead: "DeepSeek V4.1 Flash gains 15 percentage points over DeepSeek V4 Flash 0731 (54%), sitting 12 points above DeepSeek V4 Pro 0813 (57%) and 7 points above GLM-5.3 (62%)."
That 15-point gain is the harder number, and Reddit supplied the framing. A post in r/LocalLLaMA on September 14, titled "DeepSeek V4.1 Flash beats Astra on AA's new benchmark", drew roughly 630 points and more than 120 comments — asserting in its title a win the chart supports by 0.4 points and the operator declined to say.
Two index revisions in three days, and what they do not disclose
Artificial Analysis published Intelligence Index v4.2 on September 4, billed as "more complex and realistic tasks, and more private test sets to prevent gaming," then v4.3 three days later, on September 7. The second revision adds AutomationBench-AA, upgrades Terminal-Bench from v2.1 to v4.0, and replaces the tau-3-Banking eval outright. It also raises the weight assigned to evaluations with private tasks or answers from 40% to 45%; category weights stay where v4.2 left them, at Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%.
Two things follow, and only one of them is a criticism. The disclosure is real: each revision ships with a rationale naming difficulty, private sets, and saturation as the drivers, and private sets are a defensible answer to contamination. The gap is real too. Nothing public explains how the score-to-task mapping behaves run to run, no variance or seeds are published, and no note tells a reader how a v4.2 score compares with the same model's v4.3 score. With headline gaps at 0.4 points and rubric versions shipping days apart, a capability change and a scoring change look identical. That is the measurement problem we documented in the Astra ARC-AGI-3 harness gap, one layer up: there the scaffold moved the number, here the ruler does.
The thread goes further and supplies a motive — that Astra was "farming a ton of points" on the new private eval, and that AA "changed the index twice in three days to make Astra look not-quite-worse than Fable." Treat that as the poster's claim. The verifiable parts are narrower: two revisions in three days, tau-3-Banking replaced, and AA's stated aim of adding private test sets to prevent gaming. AA's own rationale points the other way, since it made the index harder rather than easier, and the thread offers no evidence that anyone gamed anything.
The benchmark underneath the row
AutomationBench is Zapier's: 657 tasks across Finance, HR, Marketing, Operations, Sales, and Support, run against simulated Gmail, Google Sheets, Slack, Salesforce, Zendesk, Jira, and HubSpot environments, described in arXiv 2604.18934. Artificial Analysis runs its own implementation on a private held-out test set, with Zapier's involvement. The scoring rule is what lifts the numbers into the sixties — a model earns credit for the share of each task's objectives it completes, and a single guardrail violation zeroes that task.
Scores from AutomationBench-AA, top rows as published (max reasoning effort unless noted):
| Model (configuration) | Score |
|---|---|
| DeepSeek V4.1 Flash (max) | 68.9% |
| GPT-6 Astra (max) | 68.5% |
| Grok 4.6 (xhigh) | 67.0% |
| GLM-5.3 (max) | 62.2% |
| DeepSeek V4 Pro 0813 (max) | 57% |
| DeepSeek V4 Flash 0731 | 54% |
Partial credit is also why the AA figures should never be placed beside Zapier's own leaderboard. The paper's abstract states plainly that "even the best frontier models currently score below 10%" — because the vendor counts only fully completed tasks. Same tasks, different rules, incomparable numbers, and a first place that reads as a rout on one board and a floor on the other.
The win that still costs the most
The other figure AA published alongside the row is a bill. V4.1 Flash consumes roughly 89k output tokens per Intelligence Index task, which AA calls one of the highest it has measured: 25% more than GLM-5.3 at 71k, 62% more than DeepSeek V4 Pro 0813 at 55k, and more output than either frontier model in the comparison, Fable 5.1 at 78k and Claude Opus 5 at 73k. At $0.30 per million input tokens and $1.20 per million output tokens, with cached input at $0.006, that works out to $0.27 per Intelligence Index task — cheap per task, expensive per token, and a reminder that a row near the top of a capability chart is not the same as a row on the Pareto frontier.
The metadata has one unresolved seam. The headline of AA's post puts the model at 552B parameters; its own details section lists it as "763B total parameters, 8B active parameters input and 16B active parameters output." Our launch coverage of the open weights reported the Hugging Face model card's 552B figure. AA has not reconciled the two, and there is no basis here for picking a winner between them — it is an inconsistency in a publisher's own page, and it should be reported as one.
What would settle it
A leaderboard row is one benchmark, one run, one configuration, and this one is maximum reasoning effort — a setting that costs more than whatever DeepSeek ships as the default. Nothing in the chart shows that V4.1 Flash is better at agentic work than Astra. It shows that at max effort, on a partial-credit rubric, on a private test set, in that run, it completed a fraction more objectives than the model AA described as its equal in the same breath.
Three things would settle it, and all are cheap to publish. Repeat runs of the same models on the same private set with a variance figure attached, so a 0.4-point gap can be read against measurement noise instead of argued about. Third-party replication, by Zapier or an academic group with access to the task set, rather than one operator scoring everyone. And per-effort cost printed next to every row, so readers can see what a point at max effort buys. Until then the honest reading of the top of that chart is the one its own publisher used, in a post almost nobody quoted: equal. The structural risk is the larger one. An index re-cut twice in a week, replacing one eval and upgrading another, makes every score published before the change a slightly different measurement from every score after it — and these benchmarks now move faster than the models they rank.


