Updated September 30, 2026

Hello, Ai

Your unbiased guide to the world's smartest AIs

Great ideas start
with a conversation.

Find the right AI. Introduce it to another.
See what you can make together.

Or catch up on the latest AI dispatches →

Different minds. Shared possibilities. Start a conversation in our app ↗

This week's frontier

6 frontier models we track, chosen by us. Filter by task or price. Updated weekly.

1
40% Cheaper
Anthropic

Claude Opus 5.5

Anthropic's September 22 successor to Opus 5 at $4/$20 with 1M context and cache reads at $0.20. Anthropic says it matches Fable 5.1 on most work and costs about 40% less to run than Opus 5. Text Arena now lists claude-opus-5.5-high, so this card's Elo is the model's own score. Code Arena WebDev lists claude-opus-5.5-max first in the tracked set.

$4/M in$20/M out1M ctx
2
Mythos Class
Anthropic

Claude Fable 5.1

Latest Mythos-class model for long-horizon coding and knowledge work. Same $10/$50 list as Fable 5 with cache reads at $0.25 — Anthropic estimates typical workloads ~25% cheaper. Code Arena WebDev lists claude-fable-5.1-max at 1751, behind gpt-6-astra-max at 1789 and claude-opus-5.5-max at 1818.

$10/M in$50/M out1M ctx
3
Budget Agentic
Meta

Muse Spark 1.3

Meta's Muse Spark 1.3 — agentic and coding upgrade over 1.2, using ~20% fewer tool calls and ~25% fewer tokens in Meta's own comparisons. Max reasoning is live (arena muse-spark-1.3-max). Cheapest tracked API at $1.25/$4.25 with 1M context.

$1.25/M in$4.25/M out1M ctx
4
Multimodal Leader
Google

Gemini 3.1 Pro

Dominates reasoning and multimodal tasks. Scored 77% on ARC-AGI-2 — double its predecessor — and leads on graduate-level science benchmarks.

$2/M in$12/M out1M ctx
5
2.4T Max
Alibaba

Qwen3.8-Max

Alibaba's 2.4T MoE flagship (95B active): Sep 1 0902 snapshot further post-trained on coding and cowork, still $2/$6 with 1M context. Arena ~1480 Elo — launch sample cooled below Gemini 3.1 Pro.

$2/M in$6/M out1M ctx
6
Agentic Code
xAI

Grok 4.7

xAI's recommended model for chat and code. Same $2/$6 and 500K window as Grok 4.6, with a $4/$12 surcharge once a prompt crosses 200K tokens. Text Arena lists grok-4.7-xhigh at 1442 on 5,675 votes. Code Arena WebDev lists it at 1636 on 3,062 votes.

$2/M in$6/M out500K ctx

This week's ranking

LMArena text Elo from the arena.ai board of Sep 30, 2026, with list price and context. Models without a score of their own are listed below the ranking, unranked.

Run it yourself

No API key, no usage limits, no data leaving your machine. The best open-weight models ranked by Elo.

1
Efficiency Champion
Google DeepMind

Gemma 4 31B

Google's Gemma 4 31B ranks among the top open-weight models on LMArena with an Elo of 1453 — beating models 20x its size. Fits a single RTX 4090 at Q4_K_M and ships Apache 2.0.

18 GB VRAM50 t/sApache 2.0
RTX 4090 (24 GB)
2
Coding Champion
Qwen Team

Qwen3.8-27B

Alibaba's Qwen3.8 generation dense flagship, replacing the Qwen3 32B card at a smaller footprint: 27B parameters, 262K native context, and hybrid tool-use. Fits a single RTX 4090 at Q4 (~15.3 GiB weights); community-reported 65 tok/s decode. Apache 2.0.

17 GB VRAM65 t/sApache 2.0
RTX 4090 (24 GB)
3
Single-GPU Pick
Mistral AI

Mistral Small 3.2 24B

The strongest model that comfortably fits a single 24 GB GPU. Adds vision and tool-use over 3.1, tightly instruction-tuned, and Apache 2.0 with no usage restrictions — a dependable all-rounder for local daily use.

14 GB VRAM14.8 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)
4
MoE Champion
Qwen Team

Qwen3 30B-A3B

A 30B mixture-of-experts with only ~3B parameters active per token. We measured 44.7 tok/s on a two-GPU budget cluster — faster than the dense 8B while packing 4x the capacity. Apache 2.0.

14 GB VRAM44.7 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)
5
OpenAI Local
OpenAI

gpt-oss-20b

OpenAI's Apache 2.0 MoE (~21B total / ~3.6B active). Native MXFP4 fits ~16 GB; we measured 53.73 tok/s on the budget two-GPU cluster — the fastest open-weight card we track. Arena ~1317 Elo.

16 GB VRAM53.73 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)
6
Two-GPU Pick
Qwen Team

Qwen3 14B

The model that makes pooling two budget GPUs worth it: ~10.5 GB of weights fit neither of our 8 GB cards alone, but the cluster ran it at a measured 19.3 tok/s — 12x faster than single-card CPU offload. Apache 2.0.

11 GB VRAM19.3 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)

Where each one leads

Category leaders from the tracked set, as of this week's snapshot.

Dispatches from the frontier

Weekly analysis, honest takes, and hidden gems. No engagement bait.