Scale AI's HSS benchmark finds models fail at seeing, not thinking
Scale AI's Humanity's Sixth Sense benchmark puts humans at 93.1% and the best model at 53.6%. Its own error data says the gap is perception, not reasoning.
Scale AI released Humanity's Sixth Sense this week, a visual-reasoning benchmark built with the human-intelligence firm Elorian, and the headline gap is wide: people score 93.1% on it, the best model reaches 53.6%. What makes the release worth reading past its top line is the error analysis underneath. Of 8,573 recorded failures, 94% trace to perception or latent inference and only 5% to faulty logic. The shortfall is not mainly a reasoning deficit. It is a seeing deficit.
A benchmark for seeing at a glance
Scale's framing is that today's multimodal models have high visual IQ: they read charts, pass exams, pull text off a page. What they lack, on this account, is the intuition a person uses to read a room or judge whether a car will fit between two parked ones. HSS is built to isolate that inference, which cognitive science says people perform within roughly 150 milliseconds. The set is deliberately small: 522 open-ended tasks drawn from 288 images and 234 video clips, 17.6 hours of footage. Each sits in one of four domains, from temporal and causal dynamics to social understanding, across eleven subdomains.
Two of Scale's own examples show the flavour. Asked why a woman in a green coat slowed after running, models missed that she slowed because the bus had stopped to wait for her. Asked whether two more books would fit on a shelf, they said no, even though the gaps were plainly visible in the frame.
| System | HSS accuracy |
|---|---|
| Human participants (n=20) | 93.1% |
| Best agentic setup (388-task subset) | 59.3% |
| GPT-6-astra, maximum effort | 53.6% |
| Median model | 30.9% |
| GPT-6-astra, media removed | 6.6% |
The last row is the control. Strip the image or video and leave only the question, and the strongest model collapses to 6.6% — evidence, Scale argues, that the tasks cannot be answered from language priors alone.
Only 15.1 percent of tasks survived review
The construction of the set is the part that carries the most weight. Trained annotators picked a clip, assigned a subdomain, then wrote a question, a reference answer and a rubric of atomic criteria a correct answer must meet. Three rules governed acceptance: the task had to require inference beyond what is depicted, so a high-resolution crop that makes the answer obvious was grounds for rejection; it had to need no specialised expertise; and its answer had to command unanimous human agreement. Of 3,466 authored tasks, 522 survived — a 15.1% acceptance rate. The first review round alone removed two in three, mostly because the answer was visible in the scene or reviewers could not agree on it. Three independent rounds followed.
That process is the benchmark's real claim: not that it is hard, but that it is clean.
The gap is perception, not logic
The failure taxonomy is where HSS stops being another leaderboard. Reviewers traced 8,573 failures to their source, and the distribution is lopsided: perception and latent inference account for 94%, deliberate reasoning for 5%. The two largest single causes are missing the decisive visual cue, at 21%, and misidentifying an object, person or role, at 20%. These failures are systematic rather than noisy — on tasks that five or more models miss, a median of 80% miss for the same reason. The weakest domain is social understanding, the lowest-scoring area for 21 of 25 models, averaging 24.4% against 34.1% for the other three. Video is harder than stills for 23 of the 25, by 7.3 points on average.
The taxonomy hints at a mechanism. Models tend to identify the objects in a scene and still fail to bind them into a coherent whole, which is why a two-dimensional overlap gets read as three-dimensional alignment. The parts are seen; the arrangement is not inferred.
Why more thinking does not close it
Models burn an average of 4,046 reasoning tokens per task on questions people answer at a glance, and the extra compute does not reliably help. Raising effort lifts the overall score but hurts specific subdomains: GPT-6-astra drops 14 points on retrodiction moving from high to xhigh effort, and peak accuracy often arrives at an intermediate setting. Agentic tooling helps a little more. Inside Claude Code and Codex, where a model can crop, zoom, search the web and re-sample the media, the best setup reaches 59.3% — but on a 388-task subset, and the gains are shallow. Errors from reading a 2D overlap as 3D alignment barely moved, from 102 to 101.
Scale graded its own benchmark
None of this is independent. Scale and Elorian built HSS, and Scale graded it. Models answer in free form with no multiple choice, an LLM judge scores each answer against the rubric, and a task counts as solved only when every criterion is met. To test the judge, Scale re-graded five models using judges from three different vendors and got identical rankings, a Spearman rho of 1.0, with agreement above 95%. That is a sensible robustness check, but it is still the vendor auditing its own subjectivity. The all-or-nothing rubric scoring hides a second blind spot: a task counts only when every criterion is met, so a model that gets most of an answer right and misses one atomic detail scores the same as one that misses everything. The human baseline rests on 20 participants. No outside group has reproduced the numbers, and the tasks have only just become public, on Hugging Face, with the paper posted to arXiv on 6 October.
What would actually settle the question
Scale's own conclusion is that frontier models fall well short of humans on HSS, and that the gap holds even with more test-time compute or agentic tooling — which makes intuitive visual reasoning, in its words, "a measurable axis of multimodal intelligence." The claim is plausible, and the failure data is unusually detailed. What it still needs is a second party: an independent evaluation on the released tasks, a human baseline larger than 20, and a judge audit run by someone who did not write the rubric. If the perception-versus-logic split survives that, HSS will have earned its place. Until then it is a well-argued claim from an interested party.

