Anthropic's claude-opus-5-max leads the Code Arena WebDev leaderboard with a top AutoEval score of 1691, well ahead of Moonshot's kimi-k3-max and Alibaba's qwen3.8-max. This edge stems from its superior agentic coding performance on multi-step front-end and full-stack web development tasks, including image-to-site generation, iterative refactoring, and tool-use workflows that mirror real developer needs. Recent August updates from competitors have narrowed gaps in cost-efficiency and open-weight categories but have not displaced the top proprietary model. With resolution days away, a late leaderboard surge from a new release or vote influx could theoretically shift outcomes, though current momentum strongly favors sustained Anthropic dominance in this benchmark.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · UpdatedAnthropic 96.4%
OpenAI 2.0%
Alibaba <1%
Moonshot <1%
$53,554 Vol.
$53,554 Vol.

Anthropic
96%

OpenAI
2%

Alibaba
1%

Moonshot
1%

DeepSeek
<1%

Z.ai
<1%

<1%

SpaceXAI
<1%

Meta
<1%

ByteDance
<1%

MiniMax
<1%

Xiaomi
<1%

Thinky
<1%

Tencent
<1%

Poolside
<1%

Mistral
<1%
Anthropic 96.4%
OpenAI 2.0%
Alibaba <1%
Moonshot <1%
$53,554 Vol.
$53,554 Vol.

Anthropic
96%

OpenAI
2%

Alibaba
1%

Moonshot
1%

DeepSeek
<1%

Z.ai
<1%

<1%

SpaceXAI
<1%

Meta
<1%

ByteDance
<1%

MiniMax
<1%

Xiaomi
<1%

Thinky
<1%

Tencent
<1%

Poolside
<1%

Mistral
<1%
Results from the "Rank" column under the "Code Arena | WebDev" Leaderboard tab at https://arena.ai/leaderboard/code/webdev (Overall) filtered for "Models" will be used to resolve this market.
Models will be ordered primarily by their leaderboard rank at the market’s check time. If two or more models are tied on rank, they will be ordered by which model is listed higher on the leaderboard. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the arena.ai Code Arena | WebDev Leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to "Other".
Market Opened: Jul 16, 2026, 10:56 PM ET
Resolution Source
https://arena.ai/leaderboard/code/webdevResolver
0x69c47De9D...Results from the "Rank" column under the "Code Arena | WebDev" Leaderboard tab at https://arena.ai/leaderboard/code/webdev (Overall) filtered for "Models" will be used to resolve this market.
Models will be ordered primarily by their leaderboard rank at the market’s check time. If two or more models are tied on rank, they will be ordered by which model is listed higher on the leaderboard. If a tie still remains, alphabetical order of company names as listed in this market group will be used as a final tiebreaker (e.g., “Google” would be ranked ahead of “SpaceXAI”). This market will resolve based on the company that occupies first place under this ranking.
The resolution source for this market is the arena.ai Code Arena | WebDev Leaderboard. If this resolution source is unavailable at check time, this market will remain open until the leaderboard comes back online and will resolve based on the first check after it becomes available. If it becomes permanently unavailable, this market will resolve to "Other".
Resolution Source
https://arena.ai/leaderboard/code/webdevResolver
0x69c47De9D...Anthropic's claude-opus-5-max leads the Code Arena WebDev leaderboard with a top AutoEval score of 1691, well ahead of Moonshot's kimi-k3-max and Alibaba's qwen3.8-max. This edge stems from its superior agentic coding performance on multi-step front-end and full-stack web development tasks, including image-to-site generation, iterative refactoring, and tool-use workflows that mirror real developer needs. Recent August updates from competitors have narrowed gaps in cost-efficiency and open-weight categories but have not displaced the top proprietary model. With resolution days away, a late leaderboard surge from a new release or vote influx could theoretically shift outcomes, though current momentum strongly favors sustained Anthropic dominance in this benchmark.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated
Beware of external links.
Beware of external links.
Frequently Asked Questions