LM Arena gives everyday users a simple way to judge AI models without reading a technical benchmark report. You type a prompt, two anonymous models answer, and you pick the better response. Millions of these votes build a live leaderboard. The tool suits developers picking a model, curious users comparing chatbots, and researchers studying human preference. The main takeaway: Arena AI is a trustworthy, free signal of real-world quality, but it measures what people liked, not pure correctness.
Quick Answer
Arena AI, formerly LMArena and originally LMSYS Chatbot Arena, ranks AI models using anonymous human votes instead of self-reported lab benchmarks. It suits anyone deciding which chatbot, coding model, or image generator to use. Its strongest benefit is a free, bias-resistant leaderboard built from millions of blind comparisons across text, code, image, video, and agent tasks. The main limitation: rankings reflect subjective preference and can shift quickly as new models launch, so they work best as a starting point, not a final verdict.
In battle mode, you’ll be served 2 anonymous models. Dig into the responses and decide which answer best fits your needs.
What Is Arena AI?
Arena AI is a public platform that ranks AI models through anonymous, crowdsourced comparisons. Researchers at UC Berkeley’s LMSYS group launched the project in 2023 as Chatbot Arena. It later spun out as an independent company called LMArena, then completed a rebrand to Arena on January 28, 2026, moving to the arena.ai domain.
The platform’s main purpose is simple: let real people decide which AI model gives the better answer, then turn those votes into a public leaderboard. Its target users include developers, product teams, AI researchers, and anyone comparing chatbots before choosing one to rely on.
How Does Arena AI Work?
The core workflow, called Battle Mode, follows a set pattern. A user enters a prompt. The system samples two models from its active pool and shows both answers side by side, labeled only as “Model A” and “Model B.” The user votes for the better response, a tie, or “both bad.” Model identities appear only after the vote is cast.
Each vote becomes one data point in a large preference dataset. Arena AI’s classic leaderboards calculate scores with a Bradley-Terry statistical model, a method similar in spirit to chess Elo ratings. A separate leaderboard, Agent Arena, uses a different method called causal tracing. Instead of head-to-head votes, it studies long single-model sessions and measures signals such as task completion, user corrections, and how often a model invents tools it does not have.
LM Arena: Current Top 10
The Text Arena leaderboard is Arena AI’s flagship board, built from millions of blind votes across coding, math, creative writing, and instruction-following prompts. Anthropic currently holds six of the top ten spots. Scores below are Bradley-Terry ratings pulled from the live leaderboard; models with overlapping ranges are statistically tied rather than strictly ordered.
| Rank | Model | Lab | Score |
|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 1506 ± 5 |
| 2 | Claude Opus 4.6 (High) | Anthropic | 1505 ± 4 |
| 3 | Claude Opus 4.7 (High) | Anthropic | 1502 ± 4 |
| 4 | Muse Spark 1.2 (xHigh) | Meta | 1498 ± 10 |
| 5 | Claude Opus 4.6 | Anthropic | 1497 ± 3 |
| 6 | Claude Opus 4.7 | Anthropic | 1494 ± 4 |
| 7 | Claude Opus 5 (High) | Anthropic | 1493 ± 5 |
| 8 | Qwen3.8 Max | Alibaba | 1491 ± 8 |
| 9 | Gemini 3.7 Flash (High) | 1490 ± 8 | |
| 10 | Claude Opus 5 (Max) | Anthropic | 1489 ± 7 |
Key Features
- Blind, anonymous Battle Mode voting across dozens of commercial and open-source models.
- Separate leaderboards for Text, Agent, Code and WebDev, Image, Video, Vision, Search, and Document tasks.
- Agent Mode for testing models on real multi-step tasks such as building apps, dashboards, or games.
- Direct Chat and Side-by-Side comparison modes for testing a specific named model.
- Open leaderboard methodology, published through the Arena-Rank open-source package.
- Editorial-style signal breakdowns in Agent Arena, including steerability and bash recovery scores.
Performance and Experience
Battle Mode feels fast. Prompts return streaming answers within seconds, and voting takes one click. The interface stays simple: a text box, two response panes, and vote buttons. New users need no tutorial to start comparing models.
Accuracy is a different question. Arena AI measures preference, not correctness. A confident but wrong answer can beat a hedged but accurate one. The site itself notes that top-ranked models often sit within a small margin of each other, so a “#1” spot does not always mean a clearly better model. The Agent Arena leaderboard adds more objective signals, such as confirmed task success and tool hallucination rate, which partly offsets this limitation for agentic use cases.
Integrations and Compatibility
Arena AI runs entirely in a web browser and needs no installation. It does not currently offer a public API for developers to embed its voting system into other products. Instead, it links directly to each model provider’s own documentation and pricing pages from the leaderboard table. There is no dedicated mobile app; the responsive web layout covers phone and tablet use.
Who Should Use It?
Best for:
- Developers deciding which LLM API to build on
- Product teams benchmarking chatbot or coding assistant options
- Researchers studying human preference in AI evaluation
- Curious users who want to compare top chatbots without technical benchmarks
Not ideal for:
- Teams that need a certified accuracy benchmark rather than a preference score
- Developers who want to call Arena AI’s ranking system through their own API
Is LM Arena Worth It? Our Verdict
Arena AI earns its popularity honestly. The blind voting format removes a lot of the marketing spin that surrounds AI model launches, and the leaderboard updates fast enough to stay relevant as new models ship almost weekly. Coverage across text, image, video, code, and now agentic tasks makes it one of the few places to compare models across so many formats in one view. The tradeoff is that popularity is not the same as correctness. Prompt selection skews toward what visitors choose to ask, and close leaderboard gaps can look more decisive than they really are. Used as a starting shortlist rather than a final verdict, Arena AI is a genuinely useful, free resource.
Suggestions For LM Arena Improvements
- Offer a public API so developers can pull leaderboard data programmatically
- Add clearer confidence-interval visuals on the main leaderboard view, not just in Agent Arena
- Publish a lightweight mobile app for on-the-go voting
- Expand documented use-case guides for non-technical visitors
- Add topic-specific leaderboard filters, such as legal or medical writing
- Provide more transparency on how often the active model pool rotates























