We ran Torki 2, the model behind Torki, on 22 public benchmarks, the same model our users get, with thinking on and no tools, then set every result next to the numbers other labs publish. Every number links to its source.
GPQA Diamond89.719th of 28
Humanity's Last Exam38.016th of 21
MMLU-Pro88.12nd of 14
AIME 202596.32nd of 7
AIME 202695.02nd of 6
MATH-50099.41st of 4
LiveCodeBench77.16th of 10
IFBench73.34th of 6
MMMU84.72nd of 4
MILU85.14th of 5
0255075100
Each grey mark is another model's published score on the same benchmark; the white marker is Torki 2. Rank counts every model in the three tables below that has a number for that benchmark.
Against Anthropic and OpenAI
The current Claude and GPT models. Their makers publish few of these benchmarks for their newest models, so most numbers here are independent measurements.
Bold marks the best score in each row, whoever holds it. A dash means no comparable number is published.
† Independent measurement, used where the model's maker publishes none: Artificial Analysis, Epoch AI, Vals AI.
Small text under a number is its reasoning setting; Torki 2 runs at its default thinking depth.
Torki 2 samples: Humanity's Last Exam (first 400 text-only questions, graded with the official prompt), MMLU-Pro and MILU. Torki 2's MGSM is 93.4 if Telugu and Bengali spelling variants of 'Answer' are accepted.
Claude numbers, from Anthropic and from independent testers, include Anthropic's safeguard fallback to another Claude model, as Anthropic's system cards describe.
Against Google, Meta and xAI
Gemini, Gemma, Meta's Muse and Llama, and Grok. Google and Meta publish most of their own numbers; Grok's come from independent testing.
Bold marks the best score in each row, whoever holds it. A dash means no comparable number is published.
† Independent measurement, used where the model's maker publishes none: Artificial Analysis.
Small text under a number is its reasoning setting; Torki 2 runs at its default thinking depth.
Torki 2 samples: Humanity's Last Exam (first 400 text-only questions, graded with the official prompt), MMLU-Pro and MILU. Torki 2's MGSM is 93.4 if Telugu and Bengali spelling variants of 'Answer' are accepted.
Llama 4 Maverick has no thinking mode. Gemini 3.1 Flash-Lite's LiveCodeBench window is January to May 2025; Torki 2's is February to April 2025.
Against China and India
The current flagships from DeepSeek, Moonshot (Kimi), Zhipu (GLM), MiniMax, ByteDance (Seed), Tencent (Hunyuan), Baidu (ERNIE) and StepFun, and Sarvam from India. Most numbers are the labs' own.
Bold marks the best score in each row, whoever holds it. A dash means no comparable number is published.
† Independent measurement, used where the model's maker publishes none: Artificial Analysis, Vals AI.
Small text under a number is its reasoning setting; Torki 2 runs at its default thinking depth.
Torki 2 samples: Humanity's Last Exam (first 400 text-only questions, graded with the official prompt), MMLU-Pro and MILU. Torki 2's MGSM is 93.4 if Telugu and Bengali spelling variants of 'Answer' are accepted.
DeepSeek-V4-Pro (preview) numbers use DeepSeek's maximum-effort setup with a special system prompt. Seed 2.0 Pro is ByteDance's latest model with published numbers on these benchmarks; Seed 2.1 reports few of them. Hy4 and Step 5 are previews. Baidu labels ERNIE 5.1's science score "GPQA" without saying it is the Diamond set.
Every result
All 22 benchmarks, thinking on unless marked. Unanswered questions, including answers that ran out of room, count as wrong. A sample is a fixed subset; where subjects were sampled unevenly the score is weighted back to the full subject mix.
Reasoning and knowledge
GPQA Diamondaccuracy, average of 4 runs
89.7
Humanity's Last Exam (text-only)sampleaccuracy, weighted to full subject mix
38.0
MMLU-Prosampleaccuracy, weighted to full subject mix
88.1
Mathematics
AIME 2026accuracy, average of 4 attempts
95.0
AIME 2025accuracy, average of 8 attempts
96.3
AIME 2024accuracy, average of 8 attempts
95.8
MATH-500accuracy
99.4
Coding
LiveCodeBenchpass@1, v6 window 2025-02..2025-04
77.1
LiveCodeBenchpass@1, v5 window 2024-08..2025-01
79.6
HumanEval+pass@1 (base + extended tests)
94.5
MBPP+pass@1 (base + extended tests)
83.9
Instruction following
IFEvalprompt-level strict accuracy
90.9
IFBenchprompt-level strict accuracy
73.3
Vision
MMMUaccuracy
84.7
MathVistaaccuracy
89.1
DocVQAthinking offANLS
94.6
ChartQAthinking offrelaxed accuracy
85.2
Languages
MILU (11 Indian languages)sampleaccuracy
85.1
MGSMaccuracy, reference scorer
86.0
MMMLU (14 languages)samplethinking offaccuracy
75.3
Long context
RULER (32K-128K)samplethinking offaverage of 13 tasks x 3 lengths
94.1
Tool use
BFCL v4accuracy, non-live, official AST average
85.2
BFCL v4accuracy, live (real-world calls)
75.4
BFCL v4accuracy, multi-turn
50.8
BFCL v4accuracy, irrelevance detection
74.6
Built by Torki 2
Each project was written by Torki 2 from the prompt shown, as a single file with no libraries, and not edited by hand. Play with them here or open them full screen.
Chess against a computer opponent
The only complete attempt of two; the other ran out of room. In our test it answered 1.e4 with 1…d5 after searching 12,498 positions.
Read the prompt
Build a complete chess game I can play in the browser against a computer opponent. Implement all the rules: legal move generation, check, checkmate, stalemate, castling, en passant and promotion. The computer should search a few moves ahead with minimax and alpha-beta pruning. Show captured pieces, a move list in algebraic notation, highlight legal moves when I pick a piece, and offer undo and new game.
Return the complete program in ONE html code block. Everything must be in that single file: no external libraries, CDNs, fonts, images or network requests of any kind. It must work by opening the file in a modern desktop browser, and it should also be usable on a phone.
Better of two attempts. Sound is synthesised in the browser; press Start, then move with the arrow keys or drag.
Read the prompt
Build a polished neon arcade space shooter. Waves of enemies with different movement patterns, power-ups, a boss every fifth wave, particle explosions, screen shake, a score and a high score saved in localStorage, sound effects synthesised with the Web Audio API, keyboard and touch controls, pause, and a game-over screen with restart.
Return the complete program in ONE html code block. Everything must be in that single file: no external libraries, CDNs, fonts, images or network requests of any kind. It must work by opening the file in a modern desktop browser, and it should also be usable on a phone.
Better of two attempts. Real relative orbital periods on a compressed distance scale; tap a planet for its facts.
Read the prompt
Build an interactive solar system simulator on a canvas. The eight planets orbit the Sun with their real relative orbital periods and roughly correct relative distances (use a compressed scale and say so on screen). Include a time-speed slider, pause, zoom and pan, orbit trails, and clicking a planet shows a panel with real facts about it (diameter, distance from the Sun, orbital period, number of moons).
Return the complete program in ONE html code block. Everything must be in that single file: no external libraries, CDNs, fonts, images or network requests of any kind. It must work by opening the file in a modern desktop browser, and it should also be usable on a phone.
Better of two attempts. Every sound is synthesised with the Web Audio API; tap Play.
Read the prompt
Build a 16-step drum machine and bass synth using only the Web Audio API, with every sound synthesised (no samples): kick, snare, closed hat, open hat, clap and a bass line with a selectable note per step. Tempo and swing controls, per-track volume and mute, a playhead that lights the current step, three preset patterns, and save and load of patterns in localStorage.
Return the complete program in ONE html code block. Everything must be in that single file: no external libraries, CDNs, fonts, images or network requests of any kind. It must work by opening the file in a modern desktop browser, and it should also be usable on a phone.
The working attempt of two; the other had a script error. Data stays in your browser.
Read the prompt
Build a personal finance tracker app. Add income and expenses with amount, category, date and note; set a monthly budget per category; show this month's totals, budget progress bars, a spending-by-category donut chart and a 6-month income-versus-expense bar chart, all drawn by hand on canvas. Search and filter the transaction list, edit and delete entries, save everything in localStorage, and import and export CSV. Use Indian rupee formatting.
Return the complete program in ONE html code block. Everything must be in that single file: no external libraries, CDNs, fonts, images or network requests of any kind. It must work by opening the file in a modern desktop browser, and it should also be usable on a phone.
Better of two attempts. Drag to pull the cloth; hold Shift, or switch on Tear, to cut it.
Read the prompt
Build an interactive cloth simulation on a canvas using Verlet integration: a grid of points joined by springy constraints, pinned along the top edge, swaying under gravity and a gentle wind. Dragging the mouse or a finger pulls the cloth; holding a modifier key (or a toggle button on touch screens) tears it. Add sliders for gravity, wind and stiffness, and a reset button. Keep it smooth at 60 frames per second.
Return the complete program in ONE html code block. Everything must be in that single file: no external libraries, CDNs, fonts, images or network requests of any kind. It must work by opening the file in a modern desktop browser, and it should also be usable on a phone.
Torki 2 was tested on the same production model that serves Torki, between 30 September and 1 October 2026. Every score was checked by hand from the raw responses before publishing.
Settings
Thinking on at the depth Torki's Stellar and Cosmos modes use, temperature 0.7, answers up to 65,536 tokens. No web search, code execution or other tools on any benchmark.
Prompts and scoring
Each benchmark's public reference prompts and scorers. GPQA Diamond is averaged over 4 runs with reshuffled options; AIME over 8 attempts per problem (4 for AIME 2026).
Samples
Humanity's Last Exam (first 400 text-only questions), MMLU-Pro (788 questions), MILU (50 per language), MMMLU (200 per language) and RULER (50 per task and length) are fixed samples. Humanity's Last Exam and MMLU-Pro are weighted to the full subject mix.
Grading Humanity's Last Exam
The reference grader is a commercial model from another lab, so we ran the official grading prompt on Torki 2; 16 of 16 decisions checked by hand matched the official answers. A strict rule-based reading gives 27.0%, the floor.
Other models
Their makers' published numbers, or an independent measurement (Artificial Analysis, Epoch AI, Vals AI) where the maker publishes none, at the highest stated reasoning setting and without tools.