Meta’s Muse Spark 1.1 currently leads its lineup with scores reaching 62.1% on select HLE leaderboards (such as BenchLM’s August 2026 update) and 46.2% on standard Artificial Analysis text-only evaluations, placing it within a few points of Anthropic’s Claude Opus 5 and Mythos 5. Recent gains stem from Meta’s focus on advanced reasoning modes, chain-of-thought scaling, and closed-weight optimizations that have accelerated progress beyond earlier Llama releases. Traders weigh the likelihood of further refinements or a new Muse iteration before December 31 against Anthropic and OpenAI’s continued frontier pushes, noting HLE’s design as a hard-to-saturate benchmark of graduate-level questions across math, science, and humanities. Key near-term catalysts include Meta’s next model announcements, developer conference updates, or documented benchmark submissions that could shift the highest Meta score toward or past the 60% threshold.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated$24,154 Vol.
55%+
49%
60%+
32%
65%+
14%
70%+
7%
$24,154 Vol.
55%+
49%
60%+
32%
65%+
14%
70%+
7%
For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Market Opened: Jul 23, 2026, 6:48 PM ET
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...For resolution, “accuracy” refers solely to the value labeled “HLE Accuracy” or a clear equivalent metric if the data’s presentation or terminology is restructured, regardless of the model’s Calibration Error or any other displayed metric.
The resolution source will be the official Humanity’s Last Exam leaderboard at https://agi.safe.ai/. If the resolution source becomes unavailable during the listed timeframe, this market will remain open to allow the relevant data to become available again. If the source remains unavailable after the end of the listed timeframe or is otherwise confirmed to be permanently unavailable, official Humanity’s Last Exam results published elsewhere may be used. If no official alternative source is available, this market will resolve to "No".
Resolution Source
https://agi.safe.ai/Resolver
0x65070BE91...Meta’s Muse Spark 1.1 currently leads its lineup with scores reaching 62.1% on select HLE leaderboards (such as BenchLM’s August 2026 update) and 46.2% on standard Artificial Analysis text-only evaluations, placing it within a few points of Anthropic’s Claude Opus 5 and Mythos 5. Recent gains stem from Meta’s focus on advanced reasoning modes, chain-of-thought scaling, and closed-weight optimizations that have accelerated progress beyond earlier Llama releases. Traders weigh the likelihood of further refinements or a new Muse iteration before December 31 against Anthropic and OpenAI’s continued frontier pushes, noting HLE’s design as a hard-to-saturate benchmark of graduate-level questions across math, science, and humanities. Key near-term catalysts include Meta’s next model announcements, developer conference updates, or documented benchmark submissions that could shift the highest Meta score toward or past the 60% threshold.
Experimental AI-generated summary referencing Polymarket data. This is not trading advice and plays no role in how this market resolves. · Updated



Beware of external links.
Beware of external links.
Frequently Asked Questions