{"id": 1306024, "name": "Share of FrontierMath problems solved correctly by AI models", "unit": "%", "createdAt": "2026-09-07T17:19:33.000Z", "updatedAt": "2026-09-07T17:19:33.000Z", "coverage": "", "timespan": "", "datasetId": 8150, "shortUnit": "%", "columnOrder": 0, "shortName": "mean_score", "catalogPath": "grapher/artificial_intelligence/2026-09-07/frontiermath/epoch_benchmark_data#mean_score", "descriptionShort": "FrontierMath evaluates models on 295 difficult, research-level problems in advanced mathematics (Tiers 1\u20133), which can take expert mathematicians hours or days to work through.", "descriptionFromProducer": "[FrontierMath](https://epoch.ai/frontiermath) is a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians. The questions cover most major branches of modern mathematics \u2013 from computationally intensive problems in number theory and real analysis to abstract questions in algebraic geometry and category theory. Solving a typical problem requires multiple hours of effort from a researcher in the relevant branch of mathematics, and for the upper end questions, multiple days.\n\nOn 2026-06-12, we released a major update, addressing errors in 42% of problems. Following this update, the full FrontierMath dataset consists of 338 problems. This is split into a base set of 295 problems, which we call Tiers 1-3, and an expansion set of 43 exceptionally difficult problems, which we call Tier 4. We have made twelve problems public: ten from Tiers 1-3 and two from Tier 4. Unless stated otherwise, all the numbers on this hub correspond to evaluations on the private sets. You can find the public problems [here](https://epoch.ai/frontiermath/tiers-1-4/benchmark-problems).\n\nFrontierMath was developed with funding from OpenAI, who has exclusive access to a subset of the benchmark.", "type": "float", "datasetName": "Epoch AI Benchmark Data", "updatePeriodDays": 31, "datasetVersion": "2026-09-07", "nonRedistributable": false, "display": {"unit": "%", "zeroDay": "2024-01-25", "shortUnit": "%", "timeInterval": "day", "numDecimalPlaces": 1}, "schemaVersion": 2, "processingLevel": "minor", "presentation": {"topicTagsLinks": ["Artificial Intelligence"]}, "descriptionKey": "- This indicator shows the share of FrontierMath problems that AI models solve correctly, based on Epoch AI's evaluation.\n- FrontierMath is a set of 338 original math problems written by experts, covering many areas of advanced mathematics. Many problems are difficult enough that human specialists might need hours or days to solve them.\n- The benchmark has four difficulty tiers. This indicator shows accuracy on Tiers 1\u20133 (295 problems). Tier 4 contains 43 exceptionally difficult problems and is not included here.\n- In June 2026, Epoch AI reissued FrontierMath after addressing errors in 42% of its problems. This indicator uses the corrected problem set. Scores on it are not comparable with scores on the original set, where the flawed problems held every model well below 100%; models evaluated on both scored around 12 percentage points higher on the corrected set.\n- Scoring is all-or-nothing: models get 1 point for a correct final answer and 0 for anything else, with no partial credit. Models submit their answers as Python code and can use Python while working on problems. This means scores reflect math ability with access to computational tools, not just pen-and-paper reasoning.\n- Only 12 of the problems are publicly available: 10 from Tiers 1\u20133 and 2 from Tier 4. They are published so researchers can inspect how evaluations work, not to report scores.\n- FrontierMath was developed by Epoch AI with funding from OpenAI, whose GPT models are among those evaluated on this benchmark. OpenAI has exclusive access to a subset of the problems.", "dimensions": {"years": {"values": [{"id": 0}, {"id": 75}, {"id": 175}, {"id": 194}, {"id": 327}, {"id": 372}, {"id": 445}, {"id": 447}, {"id": 509}, {"id": 558}, {"id": 560}, {"id": 608}, {"id": 613}, {"id": 621}, {"id": 669}, {"id": 686}, {"id": 692}, {"id": 742}, {"id": 750}, {"id": 754}, {"id": 756}, {"id": 762}, {"id": 768}, {"id": 770}, {"id": 782}, {"id": 796}, {"id": 803}, {"id": 810}, {"id": 812}, {"id": 813}, {"id": 816}, {"id": 818}, {"id": 819}, {"id": 820}, {"id": 823}, {"id": 831}, {"id": 845}, {"id": 854}, {"id": 859}, {"id": 866}, {"id": 869}, {"id": 873}, {"id": 887}, {"id": 895}, {"id": 896}, {"id": 902}, {"id": 903}, {"id": 908}, {"id": 911}, {"id": 914}, {"id": 918}, {"id": 920}, {"id": 930}, {"id": 931}, {"id": 932}, {"id": 938}, {"id": 950}, {"id": 952}]}, "entities": {"values": [{"id": 369965, "name": "GPT-3.5 Turbo", "code": null}, {"id": 372306, "name": "GPT-4 Turbo (Apr 2024)", "code": null}, {"id": 372907, "name": "GPT-4o mini (Jul 2024)", "code": null}, {"id": 372336, "name": "GPT-4o (Aug 2024)", "code": null}, {"id": 372567, "name": "o1 (Dec 2024), high", "code": null}, {"id": 372891, "name": "o1 (Dec 2024), low", "code": null}, {"id": 372932, "name": "o1 (Dec 2024), medium", "code": null}, {"id": 372804, "name": "o3-mini (Jan 2025), high", "code": null}, {"id": 372879, "name": "o3-mini (Jan 2025), low", "code": null}, {"id": 372826, "name": "o3-mini (Jan 2025), medium", "code": null}, {"id": 372814, "name": "GPT-4.1 (Apr 2025)", "code": null}, {"id": 372830, "name": "GPT-4.1 mini (Apr 2025)", "code": null}, {"id": 372578, "name": "o3 (Apr 2025), high", "code": null}, {"id": 372548, "name": "o3 (Apr 2025), low", "code": null}, {"id": 372558, "name": "o3 (Apr 2025), medium", "code": null}, {"id": 372768, "name": "o4-mini (Apr 2025), high", "code": null}, {"id": 372803, "name": "o4-mini (Apr 2025), low", "code": null}, {"id": 372816, "name": "o4-mini (Apr 2025), medium", "code": null}, {"id": 372345, "name": "Gemini 2.5 Pro (Jun 2025)", "code": null}, {"id": 372915, "name": "Claude Opus 4.1 (Aug 2025), 32K", "code": null}, {"id": 372819, "name": "GPT-5 (Aug 2025), high", "code": null}, {"id": 372933, "name": "GPT-5 (Aug 2025), low", "code": null}, {"id": 372924, "name": "GPT-5 (Aug 2025), minimal", "code": null}, {"id": 372769, "name": "GPT-5 mini (Aug 2025), high", "code": null}, {"id": 372889, "name": "GPT-5 mini (Aug 2025), low", "code": null}, {"id": 372934, "name": "GPT-5 mini (Aug 2025), minimal", "code": null}, {"id": 372795, "name": "GPT-5 nano (Aug 2025), high", "code": null}, {"id": 372916, "name": "GPT-5 nano (Aug 2025), low", "code": null}, {"id": 372886, "name": "GPT-5 nano (Aug 2025), minimal", "code": null}, {"id": 372942, "name": "Qwen3-Max (Sep 2025)", "code": null}, {"id": 372779, "name": "Claude Sonnet 4.5 (Sep 2025), 32K", "code": null}, {"id": 372892, "name": "GPT-5 Pro (Oct 2025), high", "code": null}, {"id": 372788, "name": "Claude Opus 4.5 (Nov 2025), 32K", "code": null}, {"id": 372783, "name": "GPT-5.2 (Dec 2025), xhigh", "code": null}, {"id": 372919, "name": "GPT-5.2 Pro (Dec 2025), xhigh", "code": null}, {"id": 372774, "name": "Gemini 3 Flash", "code": null}, {"id": 372775, "name": "Claude Opus 4.6, max", "code": null}, {"id": 372692, "name": "Qwen3.5 397B-A17B", "code": null}, {"id": 372909, "name": "Qwen3.5 397B-A17B, none", "code": null}, {"id": 372723, "name": "Grok 4.20", "code": null}, {"id": 372701, "name": "Gemini 3.1 Pro", "code": null}, {"id": 372806, "name": "Qwen 3.5 Flash (hosted 35B-A3B)", "code": null}, {"id": 372885, "name": "Qwen 3.5 Flash (hosted 35B-A3B), none", "code": null}, {"id": 372898, "name": "Gemini 3.1 Flash-Lite, high", "code": null}, {"id": 372903, "name": "Gemini 3.1 Flash-Lite, low", "code": null}, {"id": 372883, "name": "Gemini 3.1 Flash-Lite, minimal", "code": null}, {"id": 372801, "name": "GPT-5.4 (Mar 2026), xhigh", "code": null}, {"id": 372785, "name": "GPT-5.4 Pro (Mar 2026), xhigh", "code": null}, {"id": 372930, "name": "GPT-5.4 Mini (Mar 2026), low", "code": null}, {"id": 372923, "name": "GPT-5.4 Mini (Mar 2026), none", "code": null}, {"id": 372887, "name": "GPT-5.4 Mini (Mar 2026), xhigh", "code": null}, {"id": 372820, "name": "GPT-5.4 Nano (Mar 2026), high", "code": null}, {"id": 372904, "name": "GPT-5.4 Nano (Mar 2026), low", "code": null}, {"id": 372925, "name": "GPT-5.4 Nano (Mar 2026), none", "code": null}, {"id": 372736, "name": "Qwen 3.6 Plus", "code": null}, {"id": 372935, "name": "Qwen 3.6 Plus, none", "code": null}, {"id": 372724, "name": "GLM-5.1", "code": null}, {"id": 372896, "name": "GLM-5.1, none", "code": null}, {"id": 372893, "name": "Qwen 3.6 35B-A3B", "code": null}, {"id": 372941, "name": "Qwen 3.6 35B-A3B, none", "code": null}, {"id": 372926, "name": "Claude Opus 4.7, max", "code": null}, {"id": 372910, "name": "Grok 4.3 Beta, high", "code": null}, {"id": 372717, "name": "Kimi K2.6", "code": null}, {"id": 372884, "name": "Qwen3.6 27B", "code": null}, {"id": 372880, "name": "Qwen3.6 27B, none", "code": null}, {"id": 372802, "name": "GPT-5.5 Pro, xhigh", "code": null}, {"id": 372807, "name": "GPT-5.5, xhigh", "code": null}, {"id": 372920, "name": "DeepSeek-V4-Pro, max", "code": null}, {"id": 372787, "name": "Qwen 3.6 Flash", "code": null}, {"id": 372913, "name": "Qwen 3.6 Flash, none", "code": null}, {"id": 372838, "name": "GPT-5.5 Instant", "code": null}, {"id": 372733, "name": "Gemini 3.5 Flash, high", "code": null}, {"id": 372931, "name": "Qwen3.7-Max", "code": null}, {"id": 372831, "name": "Claude Opus 4.8, max", "code": null}, {"id": 372917, "name": "Qwen3.7-Plus, none", "code": null}, {"id": 372905, "name": "Claude Fable 5, max", "code": null}, {"id": 372936, "name": "Kimi K2.7 Code", "code": null}, {"id": 372940, "name": "GLM-5.2, low", "code": null}, {"id": 372927, "name": "GLM-5.2, max", "code": null}, {"id": 372888, "name": "GLM-5.2, none", "code": null}, {"id": 372901, "name": "Claude Sonnet 5, max", "code": null}, {"id": 372897, "name": "Grok 4.5, high", "code": null}, {"id": 372902, "name": "GPT-5.6 Luna, low", "code": null}, {"id": 372911, "name": "GPT-5.6 Luna, max", "code": null}, {"id": 372899, "name": "GPT-5.6 Luna, none", "code": null}, {"id": 372894, "name": "GPT-5.6 Sol, max", "code": null}, {"id": 372918, "name": "GPT-5.6 Terra, max", "code": null}, {"id": 372943, "name": "Inkling, xhigh", "code": null}, {"id": 372939, "name": "Inkling-Small, xhigh", "code": null}, {"id": 372881, "name": "Kimi K3, max", "code": null}, {"id": 372937, "name": "Gemini 3.5 Flash-Lite, high", "code": null}, {"id": 372895, "name": "Gemini 3.6 Flash, high", "code": null}, {"id": 372906, "name": "Claude Opus 5, max", "code": null}, {"id": 372928, "name": "Qwen3.7 Flash", "code": null}, {"id": 372914, "name": "Qwen3.7 Flash, none", "code": null}, {"id": 372890, "name": "DeepSeek V4 Flash 0731, max", "code": null}, {"id": 372921, "name": "Qwen 3.8 Max, xhigh", "code": null}, {"id": 372922, "name": "Grok 4.6, xhigh", "code": null}, {"id": 372929, "name": "DeepSeek V4 Pro 0813, max", "code": null}, {"id": 372882, "name": "Gemini 3.7 Flash, high", "code": null}, {"id": 372900, "name": "GLM-5.3, max", "code": null}, {"id": 372938, "name": "GLM-5.3-Flash, max", "code": null}, {"id": 372908, "name": "Claude Fable 5.1, max", "code": null}, {"id": 372912, "name": "GPT-6 Astra, max", "code": null}]}}, "origins": [{"id": 21255, "title": "Epoch AI Benchmark Data", "description": "Comprehensive collection of AI benchmark datasets from Epoch AI, including FrontierMath and other performance benchmarks.", "producer": "Epoch AI", "citationFull": "Epoch AI, \u2018AI Benchmarking Hub\u2019. Published online at epoch.ai. Retrieved from \u2018https://epoch.ai/benchmarks\u2019 [online resource].", "urlMain": "https://epoch.ai/benchmarks", "urlDownload": "https://epoch.ai/data/benchmark_data.zip", "dateAccessed": "2026-09-04", "datePublished": "2026-04-17", "license": {"url": "https://epoch.ai/about", "name": "CC BY 4.0"}}]}