Skip to main content
Events
GroupKALSHI

When will AI solve another Frontier Math Problem?

When will AI solve another Frontier Math Problem?
Vol

$0.00

|
Events

1

|
Markets

6

AI Analysis

Trader mode: Actionable analysis for identifying opportunities and edge

75%
Top Probability
$0.00
Volume
6
Markets
1
Platforms

About This Event

Frontier Math: Open Problems If any AI solves at least 1 Frontier Math: Open Problems after issuance and before X 1, 2026, then the market resolves to Yes. At the time of issuance, one FrontierMath: Open Problem has been solved by AI, a Ramsey-style problem on hypergraphs; for the purpose of this market, additional AI models subsequently solving the same problem will not count. This market will close and expire early if the event occurs.

Current Market Outlook

Kalshi traders currently price a 75% chance that an AI system will solve at least one unsolved Frontier Math open problem before 2027. This is not a coin flip. The market sees this as likely but not inevitable. The implied probability has climbed steadily from 60% in early 2025, reflecting accelerating AI research output.

The market specifically excludes the single problem already solved by AI (a Ramsey-style hypergraph problem). Each new solution must be a distinct open problem from the Frontier Math benchmark set.

Key Factors Driving the Odds

The 75% probability rests on three concrete observations. First, the Frontier Math benchmark contains 300+ problems spanning combinatorics, number theory, and analysis. AI systems have already demonstrated capability on simpler problems in the set. Second, the pace of AI math reasoning has accelerated sharply. OpenAI's o3 model scored 25% on Frontier Math in late 2024. DeepSeek's R1 and Google's Gemini 2.0 have pushed that to 30-35% in early 2025. Third, the remaining unsolved problems are not uniformly harder. Some are isolated results waiting for the right technique, not fundamental barriers.

What Could Change These Odds

The biggest risk to the 75% probability is overconfidence in scaling laws. Math reasoning has not followed the same smooth improvement curve as language tasks. The jump from 30% accuracy to solving a specific open problem requires genuine novelty, not just more compute. If the next generation of models (GPT-5, Gemini 3.0) shows diminishing returns on Frontier Math, the odds could drop to 50% or below.

The key date to watch is December 2025, when the next major model releases are expected. If no open problem is solved by then, expect the probability to fall to 60% or lower as the 2027 deadline approaches.

AI-generated analysis based on market data. Not financial advice.

Overview

This prediction market concerns whether any artificial intelligence system will solve at least one problem from the FrontierMath: Open Problems benchmark by January 1, 2026. FrontierMath is a dataset of exceptionally difficult mathematics problems created by Epoch AI, designed to test the limits of AI mathematical reasoning. The problems span areas like number theory, algebraic geometry, and combinatorics, and are intentionally crafted to be unsolved by current AI systems. At the time of the market's creation, one such problem, a Ramsey-style problem on hypergraphs, had been solved by an AI system. The market excludes subsequent solutions of that same problem, focusing only on new problems being solved for the first time. The benchmark was released in November 2024 and immediately drew attention because it represented a significant step up in difficulty from existing math benchmarks like GSM8K or MATH. Researchers at Epoch AI worked with professional mathematicians to curate problems that require deep insight and multi-step reasoning, not just pattern matching. The market reflects the belief among some in the AI community that progress in mathematical reasoning is accelerating, while others remain skeptical that current architectures can handle such challenges without fundamental breakthroughs. The outcome has implications for assessing the pace of AI capability advancement and the reliability of benchmarks used to measure it.

Historical Context

The history of AI solving mathematical problems goes back to early expert systems in the 1960s and 1970s, which could handle algebraic manipulation and theorem proving in restricted domains. The Automated Theorem Proving (ATP) community produced systems like E and Vampire that could prove theorems in formal logic. However, these systems struggled with the kind of informal, intuitive reasoning required for advanced mathematics. The modern era began with the rise of large language models. In 2021, OpenAI's Codex could solve simple math word problems from the GSM8K dataset with about 60% accuracy. By 2022, Minerva, a Google model fine-tuned on mathematical text, achieved 50.3% on the MATH benchmark, which contains competition-level problems. In 2024, OpenAI's o1 and o3 models showed dramatic improvements, with o3 reaching 96.7% on MATH and 25.2% on FrontierMath. The International Mathematical Olympiad (IMO) has been a key benchmark. In 2024, DeepMind's AlphaProof solved four of six IMO problems, matching a silver medalist's performance. This was a significant milestone because IMO problems require creative reasoning. The FrontierMath benchmark, released in November 2024, was designed to be harder than IMO problems. The problems are at the level of graduate mathematics or open research questions. Epoch AI worked with mathematicians to ensure the problems are not easily solvable by searching the internet or using known formulas. The first solved problem, a Ramsey-type hypergraph problem, was solved by OpenAI's o3 in December 2024. This solution was notable because it required a non-obvious construction and multi-step reasoning.

Why It Matters

The resolution of this market has implications for understanding the pace of AI capability growth. If an AI solves a new FrontierMath problem before 2026, it would signal that AI systems are advancing faster than many experts expected. This could affect investment decisions in AI, government policy on AI regulation, and public perception of AI risk. For the mathematics community, AI solving hard problems could change how research is conducted. Some mathematicians already use AI as a tool for checking proofs or generating conjectures. If AI can solve open problems, it might become a collaborator rather than just a tool. There are also economic implications. Companies like OpenAI, DeepMind, and Anthropic are spending billions on AI development. Demonstrating progress on hard math problems helps justify these investments and attract talent. For researchers, the market provides a concrete way to track progress. Benchmarks like FrontierMath are designed to be resistant to overfitting and memorization, so solving them is a genuine test of reasoning ability. If no AI solves a problem by 2026, it would suggest that current approaches have hit a wall, at least for this class of problems. This could lead to a reassessment of timelines for artificial general intelligence (AGI).

Current Status

As of early 2025, the only solved FrontierMath problem remains the Ramsey hypergraph problem solved by OpenAI's o3. No other AI system has publicly reported solving any other problem from the benchmark. OpenAI has not released o3 to the public, and its capabilities are only known through limited reports. DeepMind has not announced any FrontierMath results. Other organizations like Anthropic, Meta, and xAI have not published results on this benchmark. The market is active, with traders assessing the likelihood of a solve within the next year. The rapid improvement from 2% to 25.2% on the test set suggests that continued progress could lead to a solve. However, the remaining problems are likely the hardest ones, and progress may slow. The market will resolve if any new problem is solved, regardless of which organization does it.

Frequently Asked Questions

What is FrontierMath?

FrontierMath is a benchmark of approximately 300 advanced mathematics problems created by Epoch AI. The problems are designed to be extremely difficult, at the level of graduate mathematics or open research questions, and are resistant to memorization or simple search.

Has any AI solved a FrontierMath problem?

Yes, OpenAI's o3 model solved one problem, a Ramsey-style problem on hypergraphs, in December 2024. This market excludes that specific problem, so only solutions of other problems count.

How is FrontierMath different from other math benchmarks like MATH or GSM8K?

FrontierMath problems are much harder. The MATH benchmark contains competition-level problems that can be solved by top high school students. FrontierMath problems require graduate-level knowledge and often involve original insights. Current AI models achieve near-perfect scores on MATH but only 25% on FrontierMath.

Why does this prediction market matter?

It provides a concrete test of AI reasoning capabilities. Solving a FrontierMath problem would demonstrate that AI can handle genuinely difficult mathematical reasoning, which has implications for AI safety, investment, and the timeline for artificial general intelligence.

Which AI systems are most likely to solve a FrontierMath problem?

OpenAI's o3 and future models, DeepMind's AlphaProof and Gemini, and possibly Anthropic's Claude or xAI's Grok are candidates. o3 has the highest reported score on the test set, but DeepMind has shown strong performance on the IMO.

Was this helpful?
Updated Jul 28, 2026

Educational content is AI-generated and sourced from Wikipedia. It should not be considered financial advice.

Market Insights

Average Yes Price
51¢
Kalshi
Arbitrage Opps
0
Cross-Platform
0

Trade This Market