Skip to main content
icon for Next Grok Model (4.6+): Humanity’s Last Exam Debut?

Next Grok Model (4.6+): Humanity’s Last Exam Debut?

icon for Next Grok Model (4.6+): Humanity’s Last Exam Debut?

Next Grok Model (4.6+): Humanity’s Last Exam Debut?

NEW
Dec 31, 2026
Polymarket

$8,572 Vol.

Polymarket

35%+

$920 Vol.

91%

40%+

$2,393 Vol.

59%

45%+

$2,921 Vol.

20%

50%+

$1,662 Vol.

8%

55%+

$676 Vol.

9%

This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No". If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results. A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify. The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere. If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered. A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market. The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".**Trader sentiment around whether the next Grok model (4.6 or higher) will mark xAI’s notable debut on Humanity’s Last Exam (HLE) reflects xAI’s steady but still trailing progress on the expert-level 2,500-question benchmark.** Created by the Center for AI Safety and Scale AI, HLE tests graduate-level reasoning across math, physics, biology, and other fields, with current no-tools leaderboards showing Claude Opus 5 and related Anthropic models at 55–65%, ahead of Grok 4.6 at 42.9%. Grok 4 variants demonstrated meaningful gains in 2025, including tool-augmented or multi-agent scores near 44–50% in internal testing, yet verified public results have not closed the gap to frontrunners. xAI’s smaller engineering team emphasizes rapid iteration and long-horizon reasoning agents over benchmark chasing, while competitors continue releasing frequent updates. Key upcoming catalysts include further 4.x iterations or a Grok 5-scale release expected to target extended agentic workflows and higher HLE performance before year-end. Market odds hinge on whether these advances produce a verified leap sufficient for a competitive debut relative to Anthropic, OpenAI, and Meta entries.

This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No".

If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results.

A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify.

The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere.

If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered.

A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market.

The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".
Volume
$8,572
End Date
Dec 31, 2026
Market Opened
Aug 10, 2026, 6:25 PM ET

Resolution Source

https://agi.safe.ai/
This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No". If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results. A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify. The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere. If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered. A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market. The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".
This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No". If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results. A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify. The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere. If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered. A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market. The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".**Trader sentiment around whether the next Grok model (4.6 or higher) will mark xAI’s notable debut on Humanity’s Last Exam (HLE) reflects xAI’s steady but still trailing progress on the expert-level 2,500-question benchmark.** Created by the Center for AI Safety and Scale AI, HLE tests graduate-level reasoning across math, physics, biology, and other fields, with current no-tools leaderboards showing Claude Opus 5 and related Anthropic models at 55–65%, ahead of Grok 4.6 at 42.9%. Grok 4 variants demonstrated meaningful gains in 2025, including tool-augmented or multi-agent scores near 44–50% in internal testing, yet verified public results have not closed the gap to frontrunners. xAI’s smaller engineering team emphasizes rapid iteration and long-horizon reasoning agents over benchmark chasing, while competitors continue releasing frequent updates. Key upcoming catalysts include further 4.x iterations or a Grok 5-scale release expected to target extended agentic workflows and higher HLE performance before year-end. Market odds hinge on whether these advances produce a verified leap sufficient for a competitive debut relative to Anthropic, OpenAI, and Meta entries.

This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No".

If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results.

A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify.

The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere.

If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered.

A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market.

The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".
Volume
$8,572
End Date
Dec 31, 2026
Market Opened
Aug 10, 2026, 6:25 PM ET

Resolution Source

https://agi.safe.ai/
This market will resolve to "Yes" if the next Grok Model (4.6+) model added to the Humanity’s Last Exam results at https://agi.safe.ai/ has an HLE Accuracy of at least the specified percentage at 12:00 PM ET on the calendar date following the date on which it first appears on the site. Otherwise, this market will resolve to "No". If a model first appears on the site but is removed before 12:00 PM ET on the following calendar date, its appearance will not qualify as added to the Humanity’s Last Exam results. A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a version designation below 4.6, such as grok-4.1-expert, grok-4.3-heavy, grok-4.5-GA will not qualify. The percentage displayed as “HLE Accuracy” for the model in its result card on the “AI Progress on Humanity’s Last Exam” chart at https://agi.safe.ai/ will be used to resolve this market. This market will resolve solely based on the displayed HLE Accuracy, regardless of the model’s Calibration Error, chart position, other configuration scores, or any underlying granular or unrounded data presented elsewhere. If multiple qualifying models are added to the Humanity’s Last Exam results on the same calendar date (ET), the model with the highest HLE Accuracy will be used for resolution. Models added to the results on the calendar date following the initial qualifying model’s first appearance will not be considered. A qualifying model must be newly added to the Humanity’s Last Exam results at https://agi.safe.ai/. Whether the model was previously released, publicly accessible, in beta, or otherwise available before appearing on the site is irrelevant for this market. The resolution source for this market is the official Humanity’s Last Exam website found at https://agi.safe.ai/. If this resolution source is unavailable at 12:00 PM ET on the calendar date following the date on which the qualifying model first appears on the site, this market will resolve based on the first subsequent instance at which the model’s HLE Accuracy becomes available on the site. If it remains unavailable through the end of the seventh day after the qualifying model first appears on the site or if no qualifying model is added by December 31, 2026, 11:59 PM ET, this market will resolve to "No".

Beware of external links.

Frequently Asked Questions

"Next Grok Model (4.6+): Humanity’s Last Exam Debut?" is a prediction market on Polymarket with 5 possible outcomes where traders buy and sell shares based on what they believe will happen. The current leading outcome is "35%+" at 91%, followed by "40%+" at 59%. Prices reflect real-time crowd-sourced probabilities. For example, a share priced at 91¢ implies that the market collectively assigns a 91% chance to that outcome. These odds shift continuously as traders react to new developments and information. Shares in the correct outcome are redeemable for $1 each upon market resolution.

"Next Grok Model (4.6+): Humanity’s Last Exam Debut?" is a newly created market on Polymarket, launched on Aug 11, 2026. As an early market, this is your opportunity to be among the first traders to set the odds and establish the market's initial price signals. You can also bookmark this page to track volume and trading activity as the market gains traction over time.

To trade on "Next Grok Model (4.6+): Humanity’s Last Exam Debut?," browse the 5 available outcomes listed on this page. Each outcome displays a current price representing the market's implied probability. To take a position, select the outcome you believe is most likely, choose "Yes" to trade in favor of it or "No" to trade against it, enter your amount, and click "Trade." If your chosen outcome is correct when the market resolves, your "Yes" shares pay out $1 each. If it's incorrect, they pay out $0. You can also sell your shares at any time before resolution if you want to lock in a profit or cut a loss.

The current frontrunner for "Next Grok Model (4.6+): Humanity’s Last Exam Debut?" is "35%+" at 91%, meaning the market assigns a 91% chance to that outcome. The next closest outcome is "40%+" at 59%. These odds update in real-time as traders buy and sell shares, so they reflect the latest collective view of what's most likely to happen. Check back frequently or bookmark this page to follow how the odds shift as new information emerges.

The resolution rules for "Next Grok Model (4.6+): Humanity’s Last Exam Debut?" define exactly what needs to happen for each outcome to be declared a winner — including the official data sources used to determine the result. You can review the complete resolution criteria in the "Rules" section on this page above the comments. We recommend reading the rules carefully before trading, as they specify the precise conditions, edge cases, and sources that govern how this market is settled.