How we rank

How the Best LLM Models ladders are ranked: Arena Agent IPS on home, Text ratings on the Text page, and what we do not claim.

Core Scoring Standards & Baselines

Two distinct scoring systems. Agent uses LMSYS Arena IPS (net improvement percentage across complex multi-step tool-use sessions). Text, Coding, Open source, Local, Writing, Image, and Video use Arena Elo human preference ratings. Open source and Local track the Text overall scores; Local is an open-weights alias, not a VRAM ranking. Image uses text-to-image. Video uses image-to-video. Do not mix IPS with Elo ratings.

Authoritative data source: lmarena-ai/leaderboard-dataset (CC BY 4.0). We do not scrape arena.ai. Automated ingestion pulls each subsetโ€™s latest snapshot directly from the official dataset.lmarena-ai/leaderboard-dataset

Visual Scaling & Spec Alignment

Bar lengths are normalized from the lowest-ranked model on that wall to #1. Each bar displays the official Arena score โ€” IPS on Agent, Elo rating on other ladders. Rows also display independent battle sessions or votes, alongside API pricing (input/output per 1M tokens) and context window capacity verified by models.dev.

API pricing & context window specification: models.dev

Platform Principles & Neutrality

  • Third-party independence: We do not run closed internal evaluations or organize commercial battles.
  • No unofficial composite scores: We never synthesize proprietary weighted averages.
  • Strict dimensional separation: We never mix Agent IPS with general text or image ratings.
  • Rigorous technical specifications: All token pricing and context window metrics are aligned with verified industry benchmarks.
  • Complete data traceability: Direct links and snapshots to official LMSYS dataset repositories.

Open-Weights Standard

We adhere strictly to open-weights definitions: A model is classified as open on this platform if models.dev marks open_weights, or if Arenaโ€™s official license is non-proprietary. Permissive and community licenses count as open.