Your unbiased guide to the world's smartest AIs
One-click access to today's frontier leaders. Ranked by capability, updated weekly.
First Mythos-class model cleared for general use. Leads LMArena text overall at 1507 Elo with 95% on SWE-bench Verified and top GDPval-AA knowledge-work scores — gains widen on long-horizon agentic runs.
Meta's Muse Spark 1.2 — agentic coding upgrade co-trained with Muse Code. #2 in helloai's tracked set at 1498 Elo — cheapest agentic option we track at $1.25/$4.25 with 1M context.
Alibaba's 2.4T MoE flagship (95B active): long-horizon coding and professional work at $2/$6 with 1M context. Arena ~1497 Elo — re-admits Alibaba after the July Muse swap.
Anthropic's July 2026 successor to Opus 4.8: 96% on SWE-bench Verified, doubles Frontier-Bench performance, and scores 3x the next-best model on ARC-AGI-3 — at the same $5/$25 price.
Dominates reasoning and multimodal tasks. Scored 77% on ARC-AGI-2 — double its predecessor — and leads on graduate-level science benchmarks.
DeepSeek's 1.6T MoE flagship (49B active): MIT-licensed API at $0.435/$0.87 with 1M context — an order of magnitude under Western mid-tier rates. Arena deepseek-v4-pro ~1457 Elo.
Elo ratings from Chatbot Arena blind votes. These shift weekly — here's the current snapshot.
No API key, no usage limits, no data leaving your machine. The best open-weight models ranked by Elo.
Google's Gemma 4 31B ranks #3 among all open-weight models on LMArena with an Elo of 1452 — beating models 20x its size. Fits a single RTX 4090 at Q4_K_M and ships Apache 2.0.
The strongest model that comfortably fits a single 24 GB GPU. Adds vision and tool-use over 3.1, tightly instruction-tuned, and Apache 2.0 with no usage restrictions — a dependable all-rounder for local daily use.
Alibaba's Qwen3 flagship dense model. Matches Qwen2.5-72B performance in a 32B package that fits a single RTX 4090, with hybrid thinking/non-thinking modes and Apache 2.0.
A 30B mixture-of-experts with only ~3B parameters active per token. We measured 44.7 tok/s on a two-GPU budget cluster — faster than the dense 8B while packing 4x the capacity. Apache 2.0.
OpenAI's Apache 2.0 MoE (~21B total / ~3.6B active). Native MXFP4 fits ~16 GB; we measured 53.73 tok/s on the budget two-GPU cluster — the fastest open-weight card we track. Arena ~1317 Elo.
The model that makes pooling two budget GPUs worth it: ~10.5 GB of weights fit neither of our 8 GB cards alone, but the cluster ran it at a measured 19.3 tok/s — 12x faster than single-card CPU offload. Apache 2.0.
No hype. Where each model actually leads, based on benchmarks and real-world usage as of today.
Highest Elo in the tracked set at 1507 on LMArena text overall. Mythos-class capability with conservative safety fallbacks to Opus on flagged queries.
95.0% on SWE-bench Verified and 80% on SWE-bench Pro — still the long-horizon agentic coding pick when the 2× price premium over Opus 5 is justified.
Leads on PhD-level benchmarks like GPQA and ARC-AGI subsets. Claude and Qwen3.8-Max are strong contenders.
Cheapest frontier API we track at $0.435/$0.87 with 1M context — the default when token volume dominates and Mythos-class quality is not required.
Weekly analysis, honest takes, and hidden gems. No engagement bait.
helloai admits Qwen3.8-Max and DeepSeek V4 Pro, drops Grok 4.5 and Kimi K3, and adds gpt-oss-20b to the open-weight shelf under the six-model caps.
Meta ships Muse Spark 1.2 with Muse Code at the same $1.25/$4.25 rates. helloai bumps 1.1 to 1.2 at 1498 Elo — still the cheapest agentic pick.
Anthropic's $5/$25 Opus 5 beats its own $10/$50 Fable 5 on Frontier-Bench and ARC-AGI-3, raising an uncomfortable question about who really leads the Claude lineup.
Google GA'd Gemini 3.6 Flash on July 21 while 3.5 Pro stays partner-only. helloai still tracks Gemini 3.1 Pro — and arena already scores the new Flash near frontier.
Moonshot's 2.8T Kimi K3 lands on the public API at $3/$15 with 1M context. helloai drops GPT-5.6 Sol and adds Moonshot under the six-model cap.
OpenAI and Anthropic are no longer fighting only on LMArena — ChatGPT Work and Claude Cowork compete for the surface where finished work ships.
Meta ships Muse Spark 1.1 with a public API at $1.25/$4.25. helloai drops Qwen3.7-Max and adds Meta as its budget agentic pick at 1487 Elo.
xAI launched Grok 4.5 on July 8 — 83.3% on Terminal-Bench 2.1, 4.2× token efficiency, and $2/$6 pricing. helloai's tracked Grok entry moves from 4.3 to 4.5.
OpenAI ends the partner-only preview on July 9. GPT-5.6 Sol replaces GPT-5.5 in helloai's tracked set at the same $5/$30 rate — the upgrade path teams waited on since June 26.
helloai's leaderboard shows nominal per-token rates. Prompt caching and batch APIs stack underneath — turning a $5/MTok flagship into $0.25 on repeated context.