Anthropic releases Claude Sonnet 5.5 (28 Sep): GDPval-AA 1844 Elo vs 1449 for Sonnet 5, same $2/$10 token price
2 items, from 28 Sep 2026 to 7 Oct 2026, oldest first.
AI and technology · New data ·
Anthropic releases Claude Sonnet 5.5 (28 Sep): GDPval-AA 1844 Elo vs 1449 for Sonnet 5, same $2/$10 token price
Anthropic released Claude Sonnet 5.5 on 28 Sep. On GDPval-AA v2.1 (economically valuable professional tasks) it scores 1844 Elo against Sonnet 5's 1449, at unchanged token prices ($2 in, $10 out per million) and, Anthropic says, up to 30% lower cost per task.
Why it matters. A cheaper, mid-priced AI model now scores far higher on a benchmark built from real professional work, which lowers the cost of handing office tasks to machines.
Next. Independent benchmark results and customer pricing over the coming weeks.
Anthropic's Claude Haiku 5.5 (7 Oct): 1620 on GDPval-AA vs 735 for Haiku 4.5, at a tenth of the price
Anthropic released Claude Haiku 5.5 on 7 October at $0.10 input and $0.50 output per million tokens (up to 100k tokens), against $1/$5 for Haiku 4.5. It scores 1620 Elo on GDPval-AA, a test of professional work across 44 occupations, against 735 for Haiku 4.5, 1437 for OpenAI's GPT-6 Luna and 1840 for Sonnet 5.5; on OSWorld computer use it scores 72.4% against 15.7%.
Why it matters. Anthropic's new small model scores more than twice its predecessor on a benchmark of professional work across 44 occupations, at a tenth of the price. If scores of this kind translate into practice, the cost of automating routine office tasks might be falling faster than firms can reorganise around it.
Next. Independent benchmark results and any price responses from OpenAI and Google in the coming weeks.