Your unbiased guide to the world's smartest AIs
One-click access to today's frontier leaders. Ranked by capability, updated weekly.
Latest Mythos-class model for long-horizon coding and knowledge work. Same $10/$50 list as Fable 5 with cache reads at $0.25 — Anthropic estimates typical workloads ~25% cheaper. Code Arena WebDev lists claude-fable-5.1-max at 1765.
Meta's Muse Spark 1.3 — agentic and coding upgrade over 1.2, using ~20% fewer tool calls and ~25% fewer tokens in Meta's own comparisons. Cheapest tracked API at $1.25/$4.25 with 1M context.
Anthropic's July 2026 successor to Opus 4.8: 96% on SWE-bench Verified, doubles Frontier-Bench performance, and scores 3x the next-best model on ARC-AGI-3 — at the same $5/$25 price.
Dominates reasoning and multimodal tasks. Scored 77% on ARC-AGI-2 — double its predecessor — and leads on graduate-level science benchmarks.
Alibaba's 2.4T MoE flagship (95B active): long-horizon coding and professional work at $2/$6 with 1M context. Arena ~1480 Elo — launch sample cooled below Gemini 3.1 Pro.
xAI's current flagship for chat and code at $2/$6 with 500K context. Arena grok-4.6-high ~1461 Elo overall; Code Arena WebDev 1629 now sits well behind Claude Fable 5.1.
Elo ratings from Chatbot Arena blind votes. These shift weekly — here's the current snapshot.
No API key, no usage limits, no data leaving your machine. The best open-weight models ranked by Elo.
Google's Gemma 4 31B ranks #3 among all open-weight models on LMArena with an Elo of 1451 — beating models 20x its size. Fits a single RTX 4090 at Q4_K_M and ships Apache 2.0.
Alibaba's Qwen3.8 generation dense flagship, replacing the Qwen3 32B card at a smaller footprint: 27B parameters, 262K native context, and hybrid tool-use. Fits a single RTX 4090 at Q4 (~15.3 GiB weights); community-reported 65 tok/s decode. Apache 2.0.
The strongest model that comfortably fits a single 24 GB GPU. Adds vision and tool-use over 3.1, tightly instruction-tuned, and Apache 2.0 with no usage restrictions — a dependable all-rounder for local daily use.
A 30B mixture-of-experts with only ~3B parameters active per token. We measured 44.7 tok/s on a two-GPU budget cluster — faster than the dense 8B while packing 4x the capacity. Apache 2.0.
OpenAI's Apache 2.0 MoE (~21B total / ~3.6B active). Native MXFP4 fits ~16 GB; we measured 53.73 tok/s on the budget two-GPU cluster — the fastest open-weight card we track. Arena ~1317 Elo.
The model that makes pooling two budget GPUs worth it: ~10.5 GB of weights fit neither of our 8 GB cards alone, but the cluster ran it at a measured 19.3 tok/s — 12x faster than single-card CPU offload. Apache 2.0.
No hype. Where each model actually leads, based on benchmarks and real-world usage as of today.
Highest Elo in the tracked set at 1504 on LMArena text overall. Mythos-class capability with cheaper cache reads than Fable 5 and more precise cyber/biology safeguards.
Leads Code Arena WebDev at 1765 — well above Qwen3.8-Max and Grok 4.6. Same $10/$50 list as Fable 5, with cache reads cut to $0.25.
Leads on PhD-level benchmarks like GPQA and ARC-AGI subsets. Claude and Qwen3.8-Max are strong contenders.
Cheapest frontier API we track at $1.25/$4.25 with 1M context — the default when token volume dominates and Mythos-class quality is not required.
Weekly analysis, honest takes, and hidden gems. No engagement bait.
Claude Fable 5.1 keeps the $10/$50 list but cuts cache reads 75%. Code Arena WebDev lists it at 1765, well ahead of Grok 4.6.
OpenAI slashed GPT-5.6 Sol's price 20% the same week it topped Terminal-Bench 2.1 against Opus 5 — an invoice argument for a model helloai's Elo-gated table still leaves off the board.
Grok Bot, launched August 11, is a cloud teammate with its own VM. The local grok CLI is what can actually operate a production checkout like helloai.
Qwen3.8-Max's launch-week arena lead is gone. Two weeks of votes drop it to 1481 Elo, five points under Gemini 3.1 Pro. The $2/$6 card did not change.
DeepSeek V4 Pro's hike is live today: $0.66/$1.98 off-peak, up from $0.435/$0.87. It remains helloai's cheapest tracked frontier API.
xAI ships Grok 4.6 today at the same $2/$6 as 4.5, five days after helloai dropped Grok. Arena votes are still too thin to reopen the slot.
helloai admits Qwen3.8-Max and DeepSeek V4 Pro, drops Grok 4.5 and Kimi K3, and adds gpt-oss-20b to the open-weight shelf under the six-model caps.
Meta ships Muse Spark 1.2 with Muse Code at the same $1.25/$4.25 rates. helloai bumps 1.1 to 1.2 at 1498 Elo — still the cheapest agentic pick.
Anthropic's $5/$25 Opus 5 beats its own $10/$50 Fable 5 on Frontier-Bench and ARC-AGI-3, raising an uncomfortable question about who really leads the Claude lineup.
Google GA'd Gemini 3.6 Flash on July 21 while 3.5 Pro stays partner-only. helloai still tracks Gemini 3.1 Pro — and arena already scores the new Flash near frontier.