AIUpdateWatch Intelligence
See what technical benchmark results mean for real tasks. Each system receives an A, B or C grade inside its own verified category, while raw percentages, scope and original sources remain available for technical review.
Technical evidence translated for everyday decisions
The six categories are translated into Reasoning Ability, Advanced Math Ability, Code Repair Ability, Computer & Tool Use, Long-Document Ability and Scanned Document Reading. A grade applies only inside that exact benchmark category.
Plain-English grade key
The grades use position within each category’s verified displayed results. This prevents a difficult 55% benchmark from being treated as automatically worse than an easier 90% benchmark.
Registry health
Registry as of 2026-08-22. A category becomes build-blocking when its explicit verification date passes its review deadline.
Older official source snapshot: General reasoning (Apr 13, 2026) · Terminal agents (Jul 11, 2026) · Long context (Jan 8, 2026). These sources were rechecked on Jul 25, 2026, but no newer directly comparable official result set was available.
Open the machine-readable benchmark registry · Open its SHA-256 manifest
General reasoning
Accuracy across 365 diverse seed problems and 1,095 variants covering complex constraints, branching, spatial and temporal reasoning, recursion, semantic interference, implicit information, strategy and uncertainty. Higher is better.
Google · International
Planning, comparing options, following constraints and solving unfamiliar problems.
This does not prove factual accuracy or professional judgment.
Google · International
Planning, comparing options, following constraints and solving unfamiliar problems.
This does not prove factual accuracy or professional judgment.
Z.ai · China
Planning, comparing options, following constraints and solving unfamiliar problems.
This does not prove factual accuracy or professional judgment.
OpenAI · International
Planning, comparing options, following constraints and solving unfamiliar problems.
This does not prove factual accuracy or professional judgment.
OpenAI · International
Planning, comparing options, following constraints and solving unfamiliar problems.
This does not prove factual accuracy or professional judgment.
Qwen · China
Planning, comparing options, following constraints and solving unfamiliar problems.
This does not prove factual accuracy or professional judgment.
DeepSeek · China
Planning, comparing options, following constraints and solving unfamiliar problems.
This does not prove factual accuracy or professional judgment.
Z.ai · China
Planning, comparing options, following constraints and solving unfamiliar problems.
This does not prove factual accuracy or professional judgment.
Qwen · China
Planning, comparing options, following constraints and solving unfamiliar problems.
This does not prove factual accuracy or professional judgment.
Moonshot AI · China
Planning, comparing options, following constraints and solving unfamiliar problems.
This does not prove factual accuracy or professional judgment.
General365 contains 365 manually designed seed problems expanded into 1,095 variants across eight reasoning categories. The project keeps part of the benchmark held out to monitor contamination and reports a single within-benchmark accuracy score.
Mathematics
Raw percentages from the OpenEvals official benchmark aggregation. The two exams remain separate columns; the displayed mean ranks this mathematics category only.
Moonshot AI · China
Competition-style mathematics and multi-step quantitative reasoning.
This is not a guarantee of reliable everyday arithmetic or financial advice.
StepFun · China
Competition-style mathematics and multi-step quantitative reasoning.
This is not a guarantee of reliable everyday arithmetic or financial advice.
Z.ai · China
Competition-style mathematics and multi-step quantitative reasoning.
This is not a guarantee of reliable everyday arithmetic or financial advice.
Qwen · China
Competition-style mathematics and multi-step quantitative reasoning.
This is not a guarantee of reliable everyday arithmetic or financial advice.
DeepSeek · China
Competition-style mathematics and multi-step quantitative reasoning.
This is not a guarantee of reliable everyday arithmetic or financial advice.
Qwen · China
Competition-style mathematics and multi-step quantitative reasoning.
This is not a guarantee of reliable everyday arithmetic or financial advice.
NVIDIA · International
Competition-style mathematics and multi-step quantitative reasoning.
This is not a guarantee of reliable everyday arithmetic or financial advice.
Qwen · China
Competition-style mathematics and multi-step quantitative reasoning.
This is not a guarantee of reliable everyday arithmetic or financial advice.
AIME 2026 and HMMT 2026 remain visible as separate official-feed scores. Their arithmetic mean is used only to order this mathematics section and is never reused as a universal model score.
Software engineering
Percentage of the 500 human-validated software issues resolved by the listed system result in the official benchmark feed.
Qwen · China
Understanding a real software repository, locating a bug and producing a tested fix.
Language, framework, tools and agent setup still affect real production results.
StepFun · China
Understanding a real software repository, locating a bug and producing a tested fix.
Language, framework, tools and agent setup still affect real production results.
Z.ai · China
Understanding a real software repository, locating a bug and producing a tested fix.
Language, framework, tools and agent setup still affect real production results.
Qwen · China
Understanding a real software repository, locating a bug and producing a tested fix.
Language, framework, tools and agent setup still affect real production results.
Moonshot AI · China
Understanding a real software repository, locating a bug and producing a tested fix.
Language, framework, tools and agent setup still affect real production results.
DeepSeek · China
Understanding a real software repository, locating a bug and producing a tested fix.
Language, framework, tools and agent setup still affect real production results.
Qwen · China
Understanding a real software repository, locating a bug and producing a tested fix.
Language, framework, tools and agent setup still affect real production results.
NVIDIA · International
Understanding a real software repository, locating a bug and producing a tested fix.
Language, framework, tools and agent setup still affect real production results.
SWE-bench Verified evaluates systems on 500 human-validated GitHub issues and checks proposed fixes with repository tests. Model version, agent scaffold, tools and inference settings remain part of a fair comparison.
Terminal agents
Terminal-Bench 2.1 evaluates 89 terminal tasks across software engineering, system administration, data processing, model training and security. Scores are not model-only ratings: the agent harness, effort setting, tool behavior and cost are part of each result. This selected table keeps one official entry per model so repeated harness submissions do not dominate the comparison.
Anthropic · International
Using a terminal, tools and multi-step actions to complete technical tasks.
This is a complete agent-and-model system result, not a model-only grade.
OpenAI · International
Using a terminal, tools and multi-step actions to complete technical tasks.
This is a complete agent-and-model system result, not a model-only grade.
xAI · International
Using a terminal, tools and multi-step actions to complete technical tasks.
This is a complete agent-and-model system result, not a model-only grade.
Anthropic · International
Using a terminal, tools and multi-step actions to complete technical tasks.
This is a complete agent-and-model system result, not a model-only grade.
OpenAI · International
Using a terminal, tools and multi-step actions to complete technical tasks.
This is a complete agent-and-model system result, not a model-only grade.
Meta · International
Using a terminal, tools and multi-step actions to complete technical tasks.
This is a complete agent-and-model system result, not a model-only grade.
Google · International
Using a terminal, tools and multi-step actions to complete technical tasks.
This is a complete agent-and-model system result, not a model-only grade.
Z.ai · China
Using a terminal, tools and multi-step actions to complete technical tasks.
This is a complete agent-and-model system result, not a model-only grade.
At the July 24 cutoff, the official maintained leaderboard included GLM-5.1 as its Chinese-model entry. Kimi K3, GLM-5.2, Qwen3.7 Max, MiniMax-M3 and DeepSeek V4 Pro were not added here because no official Terminal-Bench 2.1 row for those exact models appeared on the maintained leaderboard at verification time.
Long context
Average performance across 1,500 English and Chinese long-context tasks spanning 11 primary task families and input lengths from 8K to 256K tokens. Scores use the official leaderboard evaluation setting shown for each model.
Google · International
Finding, connecting and summarizing information across long documents.
A high grade does not guarantee that every detail in a very long file will be handled correctly.
OpenAI · International
Finding, connecting and summarizing information across long documents.
A high grade does not guarantee that every detail in a very long file will be handled correctly.
Anthropic · International
Finding, connecting and summarizing information across long documents.
A high grade does not guarantee that every detail in a very long file will be handled correctly.
DeepSeek · China
Finding, connecting and summarizing information across long documents.
A high grade does not guarantee that every detail in a very long file will be handled correctly.
Qwen · China
Finding, connecting and summarizing information across long documents.
A high grade does not guarantee that every detail in a very long file will be handled correctly.
Qwen · China
Finding, connecting and summarizing information across long documents.
A high grade does not guarantee that every detail in a very long file will be handled correctly.
Qwen · China
Finding, connecting and summarizing information across long documents.
A high grade does not guarantee that every detail in a very long file will be handled correctly.
Moonshot AI · China
Finding, connecting and summarizing information across long documents.
A high grade does not guarantee that every detail in a very long file will be handled correctly.
MiniMax · China
Finding, connecting and summarizing information across long documents.
A high grade does not guarantee that every detail in a very long file will be handled correctly.
OpenAI · International
Finding, connecting and summarizing information across long documents.
A high grade does not guarantee that every detail in a very long file will be handled correctly.
LongBench Pro uses 1,500 naturally occurring English and Chinese samples across 11 primary task families, six length bands from 8K to 256K tokens and four calibrated difficulty levels. The ranking uses the official leaderboard’s average overall metric for the displayed evaluation mode.
Document and OCR
This category is document/OCR evaluation, not a universal vision score. General chat models are not assigned an OCR result unless the official feed contains one.
Datalab · International
Extracting text, tables and page structure from scans, PDFs and document images.
This does not measure general photography, image generation or conversational quality.
RedNote HiLab · China
Extracting text, tables and page structure from scans, PDFs and document images.
This does not measure general photography, image generation or conversational quality.
LightOn AI · International
Extracting text, tables and page structure from scans, PDFs and document images.
This does not measure general photography, image generation or conversational quality.
Datalab · International
Extracting text, tables and page structure from scans, PDFs and document images.
This does not measure general photography, image generation or conversational quality.
Infiny AI · International
Extracting text, tables and page structure from scans, PDFs and document images.
This does not measure general photography, image generation or conversational quality.
AllenAI · International
Extracting text, tables and page structure from scans, PDFs and document images.
This does not measure general photography, image generation or conversational quality.
PaddlePaddle · China
Extracting text, tables and page structure from scans, PDFs and document images.
This does not measure general photography, image generation or conversational quality.
Baidu · China
Extracting text, tables and page structure from scans, PDFs and document images.
This does not measure general photography, image generation or conversational quality.
DeepSeek · China
Extracting text, tables and page structure from scans, PDFs and document images.
This does not measure general photography, image generation or conversational quality.
Z.ai · China
Extracting text, tables and page structure from scans, PDFs and document images.
This does not measure general photography, image generation or conversational quality.
The olmOCR category focuses on document extraction specialists. General chat models receive no score unless the maintained feed contains a directly comparable result for the exact evaluated system.
This matrix shows which capability areas the daily registry monitors. It does not claim that every model has a verified score in every category.
| Origin | Provider | Model | General | Coding | Image / multimodal | Agents / tools | Long context |
|---|---|---|---|---|---|---|---|
| International | OpenAI | GPT-5.6 Sol | Tracked | Tracked | Tracked | Tracked | Not evaluated |
| International | OpenAI | GPT-5.6 Terra | Tracked | Tracked | Tracked | Tracked | Not evaluated |
| International | OpenAI | GPT-5.6 Luna | Tracked | Not evaluated | Not evaluated | Not evaluated | Not evaluated |
| International | Anthropic | Claude Fable 5 | Tracked | Tracked | Tracked | Tracked | Not evaluated |
| International | xAI | Grok 4.5 | Tracked | Tracked | Not evaluated | Tracked | Tracked |
| International | Gemini 3.5 Flash | Tracked | Tracked | Tracked | Tracked | Not evaluated | |
| China | DeepSeek | V4 Flash | Tracked | Tracked | Not evaluated | Not evaluated | Tracked |
| China | Alibaba Cloud | Qwen3.7-Max | Tracked | Tracked | Not evaluated | Tracked | Not evaluated |
| China | Moonshot AI | Kimi K3 | Tracked | Tracked | Not evaluated | Tracked | Tracked |
| China | Zhipu AI | GLM-5.2 | Tracked | Tracked | Not evaluated | Tracked | Tracked |
| China | Baidu | ERNIE 5.0 | Tracked | Not evaluated | Tracked | Not evaluated | Tracked |
| China | ByteDance | Doubao Seed 2.1 | Tracked | Tracked | Tracked | Tracked | Not evaluated |
| China | MiniMax | MiniMax-M3 | Tracked | Tracked | Tracked | Tracked | Tracked |
| China | StepFun | Step 3.7 Flash | Tracked | Tracked | Tracked | Tracked | Not evaluated |
| China | Tencent | Hunyuan A13B | Tracked | Tracked | Not evaluated | Tracked | Tracked |
Assign A, B and C only inside the exact benchmark category being displayed.
Keep the raw percentage available and never replace the original technical evidence.
Show the official source snapshot date separately from the latest verification date.
Label selected feeds, deduplicated lists and incomplete coverage instead of calling them complete leaderboards.
Distinguish model-only evaluations from agent-and-model system results.
Do not invent grades for capabilities—such as creative writing—that the current registry does not directly test.