How we rank
How the Best LLM Models ladders are ranked: Arena Agent IPS on home, Text ratings on the Text page, and what we do not claim.
Core Scoring Standards & Baselines
Two distinct scoring systems. Agent uses LMSYS Arena IPS (net improvement percentage across complex multi-step tool-use sessions). Text, Coding, Open source, Local, Writing, Image, and Video use Arena Elo human preference ratings. Open source and Local track the Text overall scores; Local is an open-weights alias, not a VRAM ranking. Image uses text-to-image. Video uses image-to-video. Do not mix IPS with Elo ratings.
Visual Scaling & Spec Alignment
Bar lengths are normalized from the lowest-ranked model on that wall to #1. Each bar displays the official Arena score โ IPS on Agent, Elo rating on other ladders. Rows also display independent battle sessions or votes, alongside API pricing (input/output per 1M tokens) and context window capacity verified by models.dev.
Platform Principles & Neutrality
- Third-party independence: We do not run closed internal evaluations or organize commercial battles.
- No unofficial composite scores: We never synthesize proprietary weighted averages.
- Strict dimensional separation: We never mix Agent IPS with general text or image ratings.
- Rigorous technical specifications: All token pricing and context window metrics are aligned with verified industry benchmarks.
- Complete data traceability: Direct links and snapshots to official LMSYS dataset repositories.
Open-Weights Standard
We adhere strictly to open-weights definitions: A model is classified as open on this platform if models.dev marks open_weights, or if Arenaโs official license is non-proprietary. Permissive and community licenses count as open.