Updated August 12, 2026

Hello, Ai

Your unbiased guide to the world's smartest AIs

Pick your companion

One-click access to today's frontier leaders. Ranked by capability, updated weekly.

1
Mythos Class
Anthropic

Claude Fable 5

First Mythos-class model cleared for general use. Leads LMArena text overall at 1507 Elo with 95% on SWE-bench Verified and top GDPval-AA knowledge-work scores — gains widen on long-horizon agentic runs.

2
Budget Agentic
Meta

Muse Spark 1.2

Meta's Muse Spark 1.2 — agentic coding upgrade co-trained with Muse Code. #2 in helloai's tracked set at 1499 Elo — cheapest agentic option we track at $1.25/$4.25 with 1M context.

3
Coding King
Anthropic

Claude Opus 5

Anthropic's July 2026 successor to Opus 4.8: 96% on SWE-bench Verified, doubles Frontier-Bench performance, and scores 3x the next-best model on ARC-AGI-3 — at the same $5/$25 price.

4
2.4T Max
Alibaba

Qwen3.8-Max

Alibaba's 2.4T MoE flagship (95B active): long-horizon coding and professional work at $2/$6 with 1M context. Arena ~1491 Elo — re-admits Alibaba after the July Muse swap.

5
Multimodal Leader
Google

Gemini 3.1 Pro

Dominates reasoning and multimodal tasks. Scored 77% on ARC-AGI-2 — double its predecessor — and leads on graduate-level science benchmarks.

6
Price Leader
DeepSeek

DeepSeek V4 Pro

DeepSeek's 1.6T MoE flagship (49B active): MIT-licensed API at $0.435/$0.87 with 1M context — an order of magnitude under Western mid-tier rates. Arena deepseek-v4-pro ~1458 Elo.

Who's actually winning

Elo ratings from Chatbot Arena blind votes. These shift weekly — here's the current snapshot.

1
Claude Fable 5Anthropic
1507
2
Muse Spark 1.2Meta
1499
3
Claude Opus 5Anthropic
1494
4
Qwen3.8-MaxAlibaba
1491
5
Gemini 3.1 ProGoogle
1486
6
DeepSeek V4 ProDeepSeek
1458

Run it yourself

No API key, no usage limits, no data leaving your machine. The best open-weight models ranked by Elo.

1
Efficiency Champion
Google DeepMind

Gemma 4 31B

Google's Gemma 4 31B ranks #3 among all open-weight models on LMArena with an Elo of 1451 — beating models 20x its size. Fits a single RTX 4090 at Q4_K_M and ships Apache 2.0.

18 GB VRAM50 t/sApache 2.0
RTX 4090 (24 GB)
2
Single-GPU Pick
Mistral AI

Mistral Small 3.2 24B

The strongest model that comfortably fits a single 24 GB GPU. Adds vision and tool-use over 3.1, tightly instruction-tuned, and Apache 2.0 with no usage restrictions — a dependable all-rounder for local daily use.

14 GB VRAM14.8 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)
3
Coding Champion
Qwen Team

Qwen3 32B

Alibaba's Qwen3 flagship dense model. Matches Qwen2.5-72B performance in a 32B package that fits a single RTX 4090, with hybrid thinking/non-thinking modes and Apache 2.0.

20 GB VRAM48 t/sApache 2.0
RTX 4090 (24 GB)
4
MoE Champion
Qwen Team

Qwen3 30B-A3B

A 30B mixture-of-experts with only ~3B parameters active per token. We measured 44.7 tok/s on a two-GPU budget cluster — faster than the dense 8B while packing 4x the capacity. Apache 2.0.

14 GB VRAM44.7 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)
5
OpenAI Local
OpenAI

gpt-oss-20b

OpenAI's Apache 2.0 MoE (~21B total / ~3.6B active). Native MXFP4 fits ~16 GB; we measured 53.73 tok/s on the budget two-GPU cluster — the fastest open-weight card we track. Arena ~1318 Elo.

16 GB VRAM53.73 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)
6
Two-GPU Pick
Qwen Team

Qwen3 14B

The model that makes pooling two budget GPUs worth it: ~10.5 GB of weights fit neither of our 8 GB cards alone, but the cluster ran it at a measured 19.3 tok/s — 12x faster than single-card CPU offload. Apache 2.0.

11 GB VRAM19.3 t/sApache 2.0⚡ Independently measured
GTX 1070 8GB + RTX 5060 8GB (llama.cpp RPC split, 1GbE)

The real picture

No hype. Where each model actually leads, based on benchmarks and real-world usage as of today.

Overall Preference

Leader: Claude Fable 5

Highest Elo in the tracked set at 1507 on LMArena text overall. Mythos-class capability with conservative safety fallbacks to Opus on flagged queries.

Coding & Engineering

Leader: Claude Fable 5

95.0% on SWE-bench Verified and 80% on SWE-bench Pro — still the long-horizon agentic coding pick when the 2× price premium over Opus 5 is justified.

Hard Reasoning & Science

Leader: Gemini 3.1 Pro

Leads on PhD-level benchmarks like GPQA and ARC-AGI subsets. Claude and Qwen3.8-Max are strong contenders.

Honest Daily Use

Leader: DeepSeek V4 Pro

Cheapest frontier API we track at $0.435/$0.87 with 1M context — the default when token volume dominates and Mythos-class quality is not required.

Dispatches from the frontier

Weekly analysis, honest takes, and hidden gems. No engagement bait.

Grok 4.6 Ships the Week After helloai Dropped It

xAI ships Grok 4.6 today at the same $2/$6 as 4.5, five days after helloai dropped Grok. Arena votes are still too thin to reopen the slot.

Qwen3.8-Max and DeepSeek Join; Grok and Kimi Exit

helloai admits Qwen3.8-Max and DeepSeek V4 Pro, drops Grok 4.5 and Kimi K3, and adds gpt-oss-20b to the open-weight shelf under the six-model caps.

Muse Spark 1.2 Ships With Muse Code

Meta ships Muse Spark 1.2 with Muse Code at the same $1.25/$4.25 rates. helloai bumps 1.1 to 1.2 at 1498 Elo — still the cheapest agentic pick.

Claude Opus 5 Just Beat Anthropic's Own Flagship

Anthropic's $5/$25 Opus 5 beats its own $10/$50 Fable 5 on Frontier-Bench and ARC-AGI-3, raising an uncomfortable question about who really leads the Claude lineup.

Gemini 3.6 Flash Ships — 3.5 Pro Still Waiting

Google GA'd Gemini 3.6 Flash on July 21 while 3.5 Pro stays partner-only. helloai still tracks Gemini 3.1 Pro — and arena already scores the new Flash near frontier.

Kimi K3 Replaces GPT-5.6 Sol in helloai's Frontier Set

Moonshot's 2.8T Kimi K3 lands on the public API at $3/$15 with 1M context. helloai drops GPT-5.6 Sol and adds Moonshot under the six-model cap.

ChatGPT Work vs Claude Cowork: The Agentic Seat War

OpenAI and Anthropic are no longer fighting only on LMArena — ChatGPT Work and Claude Cowork compete for the surface where finished work ships.

Muse Spark 1.1 Replaces Qwen3.7-Max in helloai's Frontier Set

Meta ships Muse Spark 1.1 with a public API at $1.25/$4.25. helloai drops Qwen3.7-Max and adds Meta as its budget agentic pick at 1487 Elo.

Grok 4.5 Ships as xAI's Default Chat Model

xAI launched Grok 4.5 on July 8 — 83.3% on Terminal-Bench 2.1, 4.2× token efficiency, and $2/$6 pricing. helloai's tracked Grok entry moves from 4.3 to 4.5.

GPT-5.6 Sol Reaches General Availability

OpenAI ends the partner-only preview on July 9. GPT-5.6 Sol replaces GPT-5.5 in helloai's tracked set at the same $5/$30 rate — the upgrade path teams waited on since June 26.