← Back to Hello, AI
2 min

Opus 5.5's Own Text Score Takes the Lead

The September 30 Arena board scores Opus 5.5 as itself, at 1504, and helloai's lead moves with it. Grok 4.7's own score lands lower than the predecessor it replaced.

The September 30, 2026 Arena text board is the first snapshot that scores Claude Opus 5.5 as itself. claude-opus-5.5-high is 1504, with an interval of 10 points either side, on 3,932 votes, rank 4. helloai had been printing 1493 on that card. That figure was claude-opus-5-high on the September 13 snapshot, which is Opus 5, not 5.5. The borrowed number is retired. Opus 5.5 is now the highest Elo in the tracked set, still at $4 per million input tokens and $20 per million output.

Claude Fable 5.1-max sits one step behind at 1501, interval 7, on 11,241 votes. Those intervals overlap, so the three-point gap is not a settled ranking by itself. Muse Spark 1.3-max moved from 1493 to 1495. Gemini 3.1 Pro Preview is still 1487, now on 121,225 votes, and Qwen3.8-Max is still 1481. The large move is Grok. grok-4.7-xhigh is 1442, interval 8, on 5,675 votes. The card had been showing 1456 from grok-4.6-high. On this same board that predecessor is 1453 on 23,012 votes, still above 4.7. A replacement can score lower than the model it replaced, and the directory now shows the newer measurement.

Code Arena WebDev, also dated September 30, moves the coding lead with the text board. claude-opus-5.5-max is 1818 on 1,976 votes, rank 1. gpt-6-astra-max is 1789 on 5,918 votes. claude-fable-5.1-max is 1751 on 6,137 votes, so Fable no longer leads the models this site tracks there. grok-4.7-xhigh is 1636 on 3,062 votes. Overall Preference and Coding & Engineering both name Opus 5.5, because those labels follow the highest rated model that claims the category.

Two untracked models clear the text gate and still stay off the six. gpt-6-astra-max is 1476, interval 7, on 8,565 votes. The September 13 snapshot this site recorded had it at 1480, interval 12, on 2,693 votes, so September 30 is a second published point, and both sit inside 25 points of the floor. glm-5.3-max is 1479, interval 6, on 17,268 votes. The set is already full. The written rule would drop Grok 4.7, the lowest model with an Elo of its own, and that drop is the owner's confirmation, not an automatic edit. GLM was held on August 30 rather than pull Grok days after it was admitted.

Gemini 4 Argon is rank 1 on the same text board, at 1525, interval 9, marked Preliminary, on 4,942 votes, and it does not replace Gemini 3.1 Pro. Google's September 30 post says Argon is rolling out through the Fairwind program to trusted cyber defenders, with a developer API still ahead, at an introductory $2 and $10 per million tokens. One preliminary snapshot and no public API fails admission. GPT-6.1 Sol is on the API at the same $2 and $10 and is third on WebDev at 1759 on 1,264 votes, but the text board does not list it. The next change to the set is a second Argon snapshot developers can call, a text row for Sol, or an explicit decision to drop Grok.

← More from Hello, AI