Updated August 7, 2026

Hello, Ai

Your unbiased guide to the world's smartest AIs

Pick your companion

One-click access to today's frontier leaders. Ranked by capability, updated weekly.

1
Mythos Class
Anthropic

Claude Fable 5

First Mythos-class model cleared for general use. Leads LMArena text overall at 1507 Elo with 95% on SWE-bench Verified and top GDPval-AA knowledge-work scores — gains widen on long-horizon agentic runs.

2
Budget Agentic
Meta

Muse Spark 1.2

Meta's Muse Spark 1.2 — agentic coding upgrade co-trained with Muse Code. #2 in helloai's tracked set at 1498 Elo — cheapest agentic option we track at $1.25/$4.25 with 1M context.

3
2.4T Max
Alibaba

Qwen3.8-Max

Alibaba's 2.4T MoE flagship (95B active): long-horizon coding and professional work at $2/$6 with 1M context. Arena ~1497 Elo — re-admits Alibaba after the July Muse swap.

4
Coding King
Anthropic

Claude Opus 5

Anthropic's July 2026 successor to Opus 4.8: 96% on SWE-bench Verified, doubles Frontier-Bench performance, and scores 3x the next-best model on ARC-AGI-3 — at the same $5/$25 price.

5
Multimodal Leader
Google

Gemini 3.1 Pro

Dominates reasoning and multimodal tasks. Scored 77% on ARC-AGI-2 — double its predecessor — and leads on graduate-level science benchmarks.

6
Price Leader
DeepSeek

DeepSeek V4 Pro

DeepSeek's 1.6T MoE flagship (49B active): MIT-licensed API at $0.435/$0.87 with 1M context — an order of magnitude under Western mid-tier rates. Arena deepseek-v4-pro ~1457 Elo.

Who's actually winning

Elo ratings from Chatbot Arena blind votes. These shift weekly — here's the current snapshot.

1
Claude Fable 5Anthropic
1507
2
Muse Spark 1.2Meta
1498
3
Qwen3.8-MaxAlibaba
1497
4
Claude Opus 5Anthropic
1493
5
Gemini 3.1 ProGoogle
1487
6
DeepSeek V4 ProDeepSeek
1457

Run it yourself

No API key, no usage limits, no data leaving your machine. The best open-weight models ranked by Elo.

1
Efficiency Champion
Google DeepMind

Gemma 4 31B

Google's Gemma 4 31B ranks #3 among all open-weight models on LMArena with an Elo of 1452 — beating models 20x its size. Fits a single RTX 4090 at Q4_K_M and ships Apache 2.0.

18 GB VRAM50 t/sApache 2.0
RTX 4090 (24 GB)
2
Single-GPU Pick
Mistral AI

Mistral Small 3.2 24B

The strongest model that comfortably fits a single 24 GB GPU. Adds vision and tool-use over 3.1, tightly instruction-tuned, and Apache 2.0 with no usage restrictions — a dependable all-rounder for local daily use.

14 GB VRAM14.8 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)
3
Coding Champion
Qwen Team

Qwen3 32B

Alibaba's Qwen3 flagship dense model. Matches Qwen2.5-72B performance in a 32B package that fits a single RTX 4090, with hybrid thinking/non-thinking modes and Apache 2.0.

20 GB VRAM48 t/sApache 2.0
RTX 4090 (24 GB)
4
MoE Champion
Qwen Team

Qwen3 30B-A3B

A 30B mixture-of-experts with only ~3B parameters active per token. We measured 44.7 tok/s on a two-GPU budget cluster — faster than the dense 8B while packing 4x the capacity. Apache 2.0.

14 GB VRAM44.7 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)
5
OpenAI Local
OpenAI

gpt-oss-20b

OpenAI's Apache 2.0 MoE (~21B total / ~3.6B active). Native MXFP4 fits ~16 GB; we measured 53.73 tok/s on the budget two-GPU cluster — the fastest open-weight card we track. Arena ~1317 Elo.

16 GB VRAM53.73 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)
6
Two-GPU Pick
Qwen Team

Qwen3 14B

The model that makes pooling two budget GPUs worth it: ~10.5 GB of weights fit neither of our 8 GB cards alone, but the cluster ran it at a measured 19.3 tok/s — 12x faster than single-card CPU offload. Apache 2.0.

11 GB VRAM19.3 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)

The real picture

No hype. Where each model actually leads, based on benchmarks and real-world usage as of today.

Overall Preference

Leader: Claude Fable 5

Highest Elo in the tracked set at 1507 on LMArena text overall. Mythos-class capability with conservative safety fallbacks to Opus on flagged queries.

Coding & Engineering

Leader: Claude Fable 5

95.0% on SWE-bench Verified and 80% on SWE-bench Pro — still the long-horizon agentic coding pick when the 2× price premium over Opus 5 is justified.

Hard Reasoning & Science

Leader: Gemini 3.1 Pro

Leads on PhD-level benchmarks like GPQA and ARC-AGI subsets. Claude and Qwen3.8-Max are strong contenders.

Honest Daily Use

Leader: DeepSeek V4 Pro

Cheapest frontier API we track at $0.435/$0.87 with 1M context — the default when token volume dominates and Mythos-class quality is not required.

Dispatches from the frontier

Weekly analysis, honest takes, and hidden gems. No engagement bait.

Qwen3.8-Max and DeepSeek Join; Grok and Kimi Exit

helloai admits Qwen3.8-Max and DeepSeek V4 Pro, drops Grok 4.5 and Kimi K3, and adds gpt-oss-20b to the open-weight shelf under the six-model caps.

Muse Spark 1.2 Ships With Muse Code

Meta ships Muse Spark 1.2 with Muse Code at the same $1.25/$4.25 rates. helloai bumps 1.1 to 1.2 at 1498 Elo — still the cheapest agentic pick.

Claude Opus 5 Just Beat Anthropic's Own Flagship

Anthropic's $5/$25 Opus 5 beats its own $10/$50 Fable 5 on Frontier-Bench and ARC-AGI-3, raising an uncomfortable question about who really leads the Claude lineup.

Gemini 3.6 Flash Ships — 3.5 Pro Still Waiting

Google GA'd Gemini 3.6 Flash on July 21 while 3.5 Pro stays partner-only. helloai still tracks Gemini 3.1 Pro — and arena already scores the new Flash near frontier.

Kimi K3 Replaces GPT-5.6 Sol in helloai's Frontier Set

Moonshot's 2.8T Kimi K3 lands on the public API at $3/$15 with 1M context. helloai drops GPT-5.6 Sol and adds Moonshot under the six-model cap.

ChatGPT Work vs Claude Cowork: The Agentic Seat War

OpenAI and Anthropic are no longer fighting only on LMArena — ChatGPT Work and Claude Cowork compete for the surface where finished work ships.

Muse Spark 1.1 Replaces Qwen3.7-Max in helloai's Frontier Set

Meta ships Muse Spark 1.1 with a public API at $1.25/$4.25. helloai drops Qwen3.7-Max and adds Meta as its budget agentic pick at 1487 Elo.

Grok 4.5 Ships as xAI's Default Chat Model

xAI launched Grok 4.5 on July 8 — 83.3% on Terminal-Bench 2.1, 4.2× token efficiency, and $2/$6 pricing. helloai's tracked Grok entry moves from 4.3 to 4.5.

GPT-5.6 Sol Reaches General Availability

OpenAI ends the partner-only preview on July 9. GPT-5.6 Sol replaces GPT-5.5 in helloai's tracked set at the same $5/$30 rate — the upgrade path teams waited on since June 26.

How Caching and Batching Cut Frontier API Costs by 90%

helloai's leaderboard shows nominal per-token rates. Prompt caching and batch APIs stack underneath — turning a $5/MTok flagship into $0.25 on repeated context.