Skip to main content
Events
GroupKALSHI

Highest score on Humanity's Last Exam before Dec 31, 2026?

Highest score on Humanity's Last Exam before Dec 31, 2026?
Vol

$0.00

|
Events

1

|
Markets

9

AI Analysis

Trader mode: Actionable analysis for identifying opportunities and edge

90%
Top Probability
$0.00
Volume
9
Markets
1
Platforms

About This Event

Before Dec 31, 2026 If any language model achieves an accuracy of at least X on Humanity's Last Exam before Dec 31, 2026, then the market resolves to Yes. Early close condition: This market will close and expire early if the event occurs. This market will close and expire early if the event occurs.

Current Market Outlook

Kalshi traders are pricing a 90% probability that some language model will hit at least 50% accuracy on Humanity's Last Exam before the end of 2026. That is a high conviction bet. A 90% chance means the market views this outcome as nearly inevitable, with only a 10% chance that no LLM reaches the threshold.

Humanity's Last Exam is a benchmark designed to test the absolute frontier of AI capability. It contains questions that are intended to be extremely difficult, requiring multi-step reasoning, specialized knowledge, and novel problem-solving. The test was created by AI researchers who wanted a metric that would not be saturated quickly. A 50% score would represent a major milestone.

Key Factors Driving the Odds

The market is pricing this high for two concrete reasons.

First, the rapid pace of capability gains. GPT-4 scored around 10% on similar hard benchmarks in early 2023. By late 2024, Claude 3.5 Opus and Gemini Ultra were pushing past 30% on comparable tests. The trend line suggests a 50% score is within reach within 18 to 24 months, not by 2026. The market is betting that progress accelerates, not slows.

Second, the sheer number of competitors. OpenAI, Google DeepMind, Anthropic, Meta, and several Chinese labs are all racing. Even if one lab hits a wall, another may not. The market is betting on the collective output of the field, not any single model. The 90% number reflects that redundancy.

What Could Change These Odds

The biggest risk is that the benchmark is genuinely harder than the market assumes. The exam was designed by experts who specifically tried to create questions that would resist scaling. If the test contains long-tail reasoning chains that current architectures cannot handle, 50% could be a ceiling, not a floor.

A second risk is regulatory intervention. If governments impose safety testing requirements that slow deployment, some labs might delay releases. That could push the date past December 2026.

The main catalyst to watch is the release of GPT-5 or Gemini 2.0. If either scores above 40% on a similar hard benchmark in 2025, the market will likely move to 99%. If both score under 20%, expect a sharp correction.

AI-generated analysis based on market data. Not financial advice.

Overview

Humanity's Last Exam is a benchmark designed to test the capabilities of artificial intelligence systems at the frontier of knowledge. It was created by a coalition of researchers, including Dan Hendrycks of the Center for AI Safety and others, as a successor to earlier tests like the Massive Multitask Language Understanding (MMLU) benchmark. The exam consists of hundreds of extremely difficult questions across mathematics, physics, biology, law, and other specialized domains, intended to measure whether AI can solve problems that require deep reasoning and expert-level knowledge. The questions were crowdsourced from experts worldwide and vetted to ensure they are not easily answerable by current models through memorization or pattern matching. The prediction market question asks whether any language model will achieve a specific accuracy threshold, set at 90% or higher, on this exam before December 31, 2026. This benchmark is distinct from earlier tests because it was specifically designed to be harder than any existing AI evaluation, with questions that often require multi-step reasoning, novel problem-solving, and integration of knowledge from multiple fields. The interest in this market reflects broader debates about the pace of AI progress, the reliability of benchmarks, and whether current scaling laws will continue to produce rapid gains in capability. Some researchers argue that achieving 90% on Humanity's Last Exam would represent a significant milestone, potentially indicating that AI systems have reached a level of competence that approaches or exceeds human experts in many domains. Others caution that benchmark scores can be misleading due to data contamination, overfitting, or the narrowness of the test itself. The market's resolution depends on publicly verifiable results from the official benchmark, which is maintained by the Center for AI Safety and associated researchers. As of early 2025, no model has publicly reported scores above 50% on this exam, though some frontier labs like OpenAI, Google DeepMind, and Anthropic are likely testing their systems internally.

Historical Context

The development of AI benchmarks has evolved significantly over the past decade. Early benchmarks like the Stanford Question Answering Dataset (SQuAD) and the General Language Understanding Evaluation (GLUE) focused on narrow tasks such as reading comprehension and sentiment analysis. These were superseded by more comprehensive tests like MMLU, introduced in 2020 by Dan Hendrycks and colleagues, which covered 57 subjects from elementary mathematics to professional law. MMLU quickly became a standard for evaluating large language models, with GPT-4 achieving approximately 86% accuracy in 2023, a result that surprised many researchers. That same year, the concept of 'Humanity's Last Exam' was proposed as a response to concerns that MMLU and other benchmarks were becoming saturated. The idea was to create a test that would remain challenging for several years, similar to how the ImageNet benchmark for computer vision was eventually surpassed. The exam was officially released in early 2024, with questions curated from over 500 experts across 100 fields. Each question required at least 30 minutes for a human expert to solve, and the authors explicitly designed them to resist memorization by using novel problem formulations. The benchmark's name reflects its intended role as the final evaluation needed before AI systems could be considered broadly superhuman. In 2024, OpenAI's GPT-4o scored approximately 44% on the exam, while Google's Gemini Ultra scored around 42%, far below the 90% threshold. These results set the baseline for the prediction market. The historical pattern of AI progress on benchmarks suggests that scores often improve rapidly after initial releases, but the difficulty of Humanity's Last Exam may slow this trend. For example, MMLU saw a 20 percentage point improvement in its first three years, but the harder questions on the new exam may require breakthroughs in reasoning rather than simple scaling.

Why It Matters

The outcome of this prediction market has implications for understanding the trajectory of AI development. If a language model achieves 90% on Humanity's Last Exam before 2027, it would suggest that current scaling approaches and architectural innovations are sufficient to reach expert-level performance across many domains. This could accelerate investment in AI, influence regulatory discussions, and affect public perception of AI capabilities. Companies like OpenAI, Google, and Anthropic would face increased pressure to deploy such models safely and to address risks associated with misuse. On the other hand, if no model reaches the threshold, it may indicate that fundamental limitations exist in current approaches, potentially redirecting research toward new architectures or training methods. The benchmark's difficulty also matters for AI safety research. Tests like this are used to evaluate whether models can cause harm, for example by generating bioweapons or disinformation. A model that scores highly on the exam might possess dangerous capabilities, making it essential to develop alignment techniques before deployment. The market thus reflects not only technical curiosity but also practical concerns about risk management. For policymakers, the results could inform decisions about AI regulation, export controls, and research funding. For the public, it offers a concrete measure of progress toward artificial general intelligence, a topic that has become central to discussions about the future of work, education, and society.

Current Status

As of early 2025, no language model has publicly achieved a score above 44% on Humanity's Last Exam. The most recent public results came from OpenAI's GPT-4o and Google's Gemini Ultra, both scoring in the low 40s. Anthropic's Claude 3.5 Sonnet has not released official scores but is believed to be in a similar range. Frontier labs are likely conducting internal evaluations but have not published results. The prediction market opened in late 2024 with initial odds suggesting a low probability of success, around 10-15%. However, given the rapid pace of AI development, some analysts expect that new models like GPT-5 or Gemini 2.0 could show significant improvements. The market will close early if any model achieves the threshold, and results will be verified by the Center for AI Safety. There is ongoing debate about whether the exam's difficulty is appropriate, with some researchers arguing that it may be too hard for current architectures and others claiming that it is only a matter of time before it is surpassed.

Frequently Asked Questions

What is Humanity's Last Exam?

Humanity's Last Exam is a benchmark of 500 extremely difficult questions across many fields, designed to test the limits of AI capabilities. It was created by the Center for AI Safety and over 500 experts to be harder than any existing AI evaluation.

How is the score on Humanity's Last Exam calculated?

The score is the percentage of questions answered correctly by the AI model. Each question has a single correct answer, and models are evaluated under standard conditions without external tools or internet access.

Why is the threshold set at 90%?

The 90% threshold was chosen by the benchmark creators to represent a level of performance that would indicate expert-level competence across all tested domains. It is significantly higher than current scores, making it a meaningful milestone.

Was this helpful?
Updated Jul 27, 2026

Educational content is AI-generated and sourced from Wikipedia. It should not be considered financial advice.

Market Insights

Average Yes Price
36¢
Kalshi
Arbitrage Opps
0
Cross-Platform
0

Trade This Market