Your unbiased guide to the world's smartest AIs
Six APIs, ranked by capability. Filter by task or price. Updated weekly.
Latest Mythos-class model for long-horizon coding and knowledge work. Same $10/$50 list as Fable 5 with cache reads at $0.25 — Anthropic estimates typical workloads ~25% cheaper. Code Arena WebDev lists claude-fable-5.1-max at 1758, behind gpt-6-astra-max at 1800.
Meta's Muse Spark 1.3 — agentic and coding upgrade over 1.2, using ~20% fewer tool calls and ~25% fewer tokens in Meta's own comparisons. Max reasoning is live (arena muse-spark-1.3-max). Cheapest tracked API at $1.25/$4.25 with 1M context.
Anthropic's July 2026 successor to Opus 4.8: 96% on SWE-bench Verified, doubles Frontier-Bench performance, and scores 3x the next-best model on ARC-AGI-3 — at the same $5/$25 price.
Dominates reasoning and multimodal tasks. Scored 77% on ARC-AGI-2 — double its predecessor — and leads on graduate-level science benchmarks.
Alibaba's 2.4T MoE flagship (95B active): Sep 1 0902 snapshot further post-trained on coding and cowork, still $2/$6 with 1M context. Arena ~1480 Elo — launch sample cooled below Gemini 3.1 Pro.
xAI's current flagship for chat and code at $2/$6 with 500K context. Arena grok-4.6-high ~1456 Elo overall on a 15k-vote sample, no longer Preliminary; Code Arena WebDev 1618 sits well behind Claude Fable 5.1 and GPT-6 Astra.
LMArena text-overall Elo with list price and context. Updated weekly.
No API key, no usage limits, no data leaving your machine. The best open-weight models ranked by Elo.
Google's Gemma 4 31B ranks #3 among all open-weight models on LMArena with an Elo of 1451 — beating models 20x its size. Fits a single RTX 4090 at Q4_K_M and ships Apache 2.0.
Alibaba's Qwen3.8 generation dense flagship, replacing the Qwen3 32B card at a smaller footprint: 27B parameters, 262K native context, and hybrid tool-use. Fits a single RTX 4090 at Q4 (~15.3 GiB weights); community-reported 65 tok/s decode. Apache 2.0.
The strongest model that comfortably fits a single 24 GB GPU. Adds vision and tool-use over 3.1, tightly instruction-tuned, and Apache 2.0 with no usage restrictions — a dependable all-rounder for local daily use.
A 30B mixture-of-experts with only ~3B parameters active per token. We measured 44.7 tok/s on a two-GPU budget cluster — faster than the dense 8B while packing 4x the capacity. Apache 2.0.
OpenAI's Apache 2.0 MoE (~21B total / ~3.6B active). Native MXFP4 fits ~16 GB; we measured 53.73 tok/s on the budget two-GPU cluster — the fastest open-weight card we track. Arena ~1317 Elo.
The model that makes pooling two budget GPUs worth it: ~10.5 GB of weights fit neither of our 8 GB cards alone, but the cluster ran it at a measured 19.3 tok/s — 12x faster than single-card CPU offload. Apache 2.0.
Category leaders from the tracked set, as of this week's snapshot.
Weekly analysis, honest takes, and hidden gems. No engagement bait.
Astra is on the Sep 13 text board at 1480 Elo. helloai still waits on two weeks and a thicker sample before it can take a slot.
OpenAI shipped GPT-6 Astra at Fable's $10/$50. Code Arena WebDev lists it at 1797. helloai still waits on two weeks of text Elo before it can take a slot.
Claude Fable 5.1 keeps the $10/$50 list but cuts cache reads 75%. Code Arena WebDev lists it at 1765, well ahead of Grok 4.6.