Minimalist black-and-white logo of an Ionic column capital on a cream background, representing LMArena

LM Arena

LLM, image, video, code rankings
LM Arena lets anyone compare AI models by voting on anonymous, head-to-head answers. It turns millions of real votes into a live leaderboard, so you can see which model people actually prefer instead of relying on a lab's own benchmark claims.
Popularity Score
83%
95/100
Easy To Use
76/100
AI Quality
64/100
Speed
95/100
Integrations
55/100
Value for Money
74/100
Customer Support
In This Guide

LM Arena gives everyday users a simple way to judge AI models without reading a technical benchmark report. You type a prompt, two anonymous models answer, and you pick the better response. Millions of these votes build a live leaderboard. The tool suits developers picking a model, curious users comparing chatbots, and researchers studying human preference. The main takeaway: Arena AI is a trustworthy, free signal of real-world quality, but it measures what people liked, not pure correctness.

Quick Answer
Arena AI, formerly LMArena and originally LMSYS Chatbot Arena, ranks AI models using anonymous human votes instead of self-reported lab benchmarks. It suits anyone deciding which chatbot, coding model, or image generator to use. Its strongest benefit is a free, bias-resistant leaderboard built from millions of blind comparisons across text, code, image, video, and agent tasks. The main limitation: rankings reflect subjective preference and can shift quickly as new models launch, so they work best as a starting point, not a final verdict.

In battle mode, you’ll be served 2 anonymous models. Dig into the responses and decide which answer best fits your needs.
Arena AI, How Arena Works, arena.ai/how-it-works

What Is Arena AI?

Arena AI is a public platform that ranks AI models through anonymous, crowdsourced comparisons. Researchers at UC Berkeley’s LMSYS group launched the project in 2023 as Chatbot Arena. It later spun out as an independent company called LMArena, then completed a rebrand to Arena on January 28, 2026, moving to the arena.ai domain.

The platform’s main purpose is simple: let real people decide which AI model gives the better answer, then turn those votes into a public leaderboard. Its target users include developers, product teams, AI researchers, and anyone comparing chatbots before choosing one to rely on.

How Does Arena AI Work?

The core workflow, called Battle Mode, follows a set pattern. A user enters a prompt. The system samples two models from its active pool and shows both answers side by side, labeled only as “Model A” and “Model B.” The user votes for the better response, a tie, or “both bad.” Model identities appear only after the vote is cast.

Each vote becomes one data point in a large preference dataset. Arena AI’s classic leaderboards calculate scores with a Bradley-Terry statistical model, a method similar in spirit to chess Elo ratings. A separate leaderboard, Agent Arena, uses a different method called causal tracing. Instead of head-to-head votes, it studies long single-model sessions and measures signals such as task completion, user corrections, and how often a model invents tools it does not have.

LM Arena: Current Top 10

The Text Arena leaderboard is Arena AI’s flagship board, built from millions of blind votes across coding, math, creative writing, and instruction-following prompts. Anthropic currently holds six of the top ten spots. Scores below are Bradley-Terry ratings pulled from the live leaderboard; models with overlapping ranges are statistically tied rather than strictly ordered.

RankModelLabScore
1Claude Fable 5Anthropic1506 ± 5
2Claude Opus 4.6 (High)Anthropic1505 ± 4
3Claude Opus 4.7 (High)Anthropic1502 ± 4
4Muse Spark 1.2 (xHigh)Meta1498 ± 10
5Claude Opus 4.6Anthropic1497 ± 3
6Claude Opus 4.7Anthropic1494 ± 4
7Claude Opus 5 (High)Anthropic1493 ± 5
8Qwen3.8 MaxAlibaba1491 ± 8
9Gemini 3.7 Flash (High)Google1490 ± 8
10Claude Opus 5 (Max)Anthropic1489 ± 7

Key Features

  • Blind, anonymous Battle Mode voting across dozens of commercial and open-source models.
  • Separate leaderboards for Text, Agent, Code and WebDev, Image, Video, Vision, Search, and Document tasks.
  • Agent Mode for testing models on real multi-step tasks such as building apps, dashboards, or games.
  • Direct Chat and Side-by-Side comparison modes for testing a specific named model.
  • Open leaderboard methodology, published through the Arena-Rank open-source package.
  • Editorial-style signal breakdowns in Agent Arena, including steerability and bash recovery scores.

Performance and Experience

Battle Mode feels fast. Prompts return streaming answers within seconds, and voting takes one click. The interface stays simple: a text box, two response panes, and vote buttons. New users need no tutorial to start comparing models.

Accuracy is a different question. Arena AI measures preference, not correctness. A confident but wrong answer can beat a hedged but accurate one. The site itself notes that top-ranked models often sit within a small margin of each other, so a “#1” spot does not always mean a clearly better model. The Agent Arena leaderboard adds more objective signals, such as confirmed task success and tool hallucination rate, which partly offsets this limitation for agentic use cases.

Integrations and Compatibility

Arena AI runs entirely in a web browser and needs no installation. It does not currently offer a public API for developers to embed its voting system into other products. Instead, it links directly to each model provider’s own documentation and pricing pages from the leaderboard table. There is no dedicated mobile app; the responsive web layout covers phone and tablet use.

Who Should Use It?

Best for:

  • Developers deciding which LLM API to build on
  • Product teams benchmarking chatbot or coding assistant options
  • Researchers studying human preference in AI evaluation
  • Curious users who want to compare top chatbots without technical benchmarks

Not ideal for:

  • Teams that need a certified accuracy benchmark rather than a preference score
  • Developers who want to call Arena AI’s ranking system through their own API

Is LM Arena Worth It? Our Verdict

Arena AI earns its popularity honestly. The blind voting format removes a lot of the marketing spin that surrounds AI model launches, and the leaderboard updates fast enough to stay relevant as new models ship almost weekly. Coverage across text, image, video, code, and now agentic tasks makes it one of the few places to compare models across so many formats in one view. The tradeoff is that popularity is not the same as correctness. Prompt selection skews toward what visitors choose to ask, and close leaderboard gaps can look more decisive than they really are. Used as a starting shortlist rather than a final verdict, Arena AI is a genuinely useful, free resource.

Suggestions For LM Arena Improvements

  • Offer a public API so developers can pull leaderboard data programmatically
  • Add clearer confidence-interval visuals on the main leaderboard view, not just in Agent Arena
  • Publish a lightweight mobile app for on-the-go voting
  • Expand documented use-case guides for non-technical visitors
  • Add topic-specific leaderboard filters, such as legal or medical writing
  • Provide more transparency on how often the active model pool rotates

Capabilities

Arena AI Capabilities

Six core capabilities that define what Arena AI actually does.

Battle Mode Voting

Anonymous head-to-head model comparisons decided by a single click.

Live Leaderboards

Rankings update continuously as new votes and models arrive.

Agent Mode

Tests models on real multi-step tasks like building apps or dashboards.

Multimodal Coverage

Separate boards for text, image, video, code, and vision models.

Direct Chat Access

Chat with a specific named model outside of blind Battle Mode.

Open Methodology

Ranking code is published openly through the Arena-Rank package.

Use cases

Arena AI Use Cases

Practical ways people put the leaderboard to work.

Choosing an LLM API

Compare candidate models before committing to one for a product.

Evaluating Coding Assistants

Check WebDev and Agent leaderboards before picking a coding model.

Picking an Image Generator

Use the Text to Image leaderboard to shortlist top-rated tools.

Comparing Video Models

Review Text to Video and Image to Video boards for output quality.

Research and Teaching

Use open preference data to study or teach AI evaluation methods.

Casual Model Comparison

Settle "which AI is smarter" questions with real blind votes.

The honest verdict

Arena AI Pros and Cons

A balanced look at where Arena AI helps and where it falls short.

The good

Pros

Fully Free

Voting and leaderboard access cost nothing, no account needed.

Bias Resistant

Blind format hides model identity until after a vote is cast.

Broad Coverage

Ranks text, code, image, video, and agent models in one place.

Frequently Updated

New models are added quickly after public release.

Open Methodology

Ranking code and statistical approach are published openly.

The not-so-good

Cons

No Public API

Developers cannot pull leaderboard data into their own apps.

Subjective Voting

Preference is not the same as factual correctness.

Prompt Skew

Votes reflect what visitors ask, not every real-world task.

Fast-Moving Ranks

Leaderboard order can shift within days of a new launch.

Close Margins

Top models often sit within each other's confidence interval.

FAQ

Questions everyone eventually asks.

Clear answers to the common questions people ask before choosing an AI tool.

Is Arena AI (LMArena) free to use?
Yes. Browsing the leaderboard and voting in Battle Mode is free and does not require an account. Arena AI generates revenue by working with AI labs and through its Agent Mode product rather than by charging voters.
How does Arena AI rank models?
Battle Mode leaderboards use a Bradley-Terry statistical model built from anonymous head-to-head votes. Agent Mode leaderboards use a separate method called causal tracing, which scores models on implicit behavioral signals from real agent sessions.
What is the difference between LMArena and Arena AI?
They are the same project. It launched in 2023 as LMSYS Chatbot Arena, became an independent company called LMArena, and completed a rebrand to Arena (arena.ai) on January 28, 2026.
Does Arena AI require a login to vote?
No login is required to browse the leaderboard or cast votes in Battle Mode. An account may unlock extra features such as saved history, but core voting stays open to anonymous visitors.
Keep reading

Explore More AI Tools.

Editor-tested tools that pair well with this one.

Leonardo.Ai Review – AI Image Generation Platform

Leonardo AI

Leonardo AI is a complete image generation studio, not just a prompt box. Owned by Canva, it pairs its own...
ImageToSTL Review showing AI-powered image to 3D model conversion tool on a desktop monitor

ImageToSTL AI

ImageToSTL turns a single PNG or JPG photo into a full 3D model, and bundles free browser-based STL, GLB, OBJ...
Black Forest Labs Review feature image showing FLUX AI image generation platform on a desktop monitor

Black Forest Labs

Black Forest Labs is the German-American research lab behind FLUX, one of the world's most used AI image models. Built...
Get Weekly AI Tools & Expert Prompts.

10,000+ readers · Spam-free since 2026

Scroll to Top