AIUpdateWatch Intelligence

AI benchmark results in plain English

See what technical benchmark results mean for real tasks. Each system receives an A, B or C grade inside its own verified category, while raw percentages, scope and original sources remain available for technical review.

Saturday complete daily edition: Saturday, August 22, 2026 · Data cutoff Aug 22, 2026, 9:15 AM (America/New_York)

Technical evidence translated for everyday decisions

Grades explain the use case—not overall intelligence

The six categories are translated into Reasoning Ability, Advanced Math Ability, Code Repair Ability, Computer & Tool Use, Long-Document Ability and Scanned Document Reading. A grade applies only inside that exact benchmark category.

Plain-English grade key

A = Excellent · B = Good · C = Fair

The grades use position within each category’s verified displayed results. This prevents a difficult 55% benchmark from being treated as automatically worse than an easier 90% benchmark.

A
ExcellentTop third of the verified systems displayed in this benchmark category.
B
GoodMiddle third of the verified systems displayed in this benchmark category.
C
FairLower third of the verified systems displayed in this benchmark category.

Registry health

Benchmark review status

Registry as of 2026-08-22. A category becomes build-blocking when its explicit verification date passes its review deadline.

Current reviews
1 of 6
Review due soon
4
Older source snapshots
3
Next mandatory review
2026-08-08

Older official source snapshot: General reasoning (Apr 13, 2026) · Terminal agents (Jul 11, 2026) · Long context (Jan 8, 2026). These sources were rechecked on Jul 25, 2026, but no newer directly comparable official result set was available.

Open the machine-readable benchmark registry · Open its SHA-256 manifest

General reasoning

General365 — 2026 general-reasoning leaderboard

Held-out reasoning benchmarkCurrent reviewChecked Jul 25, 2026Review by Sep 8, 202610 displayed systems

Accuracy across 365 diverse seed problems and 1,095 variants covering complex constraints, branching, spatial and temporal reasoning, recursion, semantic interference, implicit information, strategy and uncertainty. Higher is better.

#1

Gemini 3 Pro

Google · International

Reasoning AbilityExcellent
A

Planning, comparing options, following constraints and solving unfamiliar problems.

Show technical score
Overall accuracy62.8%
Result snapshot
2026-04-13
Mode
Published General365 evaluation

This does not prove factual accuracy or professional judgment.

#2

Gemini 3 Flash

Google · International

Reasoning AbilityExcellent
A

Planning, comparing options, following constraints and solving unfamiliar problems.

Show technical score
Overall accuracy60.8%
Result snapshot
2026-04-13
Mode
Published General365 evaluation

This does not prove factual accuracy or professional judgment.

#3

GLM-5 Thinking

Z.ai · China

Reasoning AbilityExcellent
A

Planning, comparing options, following constraints and solving unfamiliar problems.

Show technical score
Overall accuracy59.9%
Result snapshot
2026-04-13
Mode
Published General365 evaluation

This does not prove factual accuracy or professional judgment.

#4

GPT-5 Thinking

OpenAI · International

Reasoning AbilityExcellent
A

Planning, comparing options, following constraints and solving unfamiliar problems.

Show technical score
Overall accuracy58.6%
Result snapshot
2026-04-13
Mode
Published General365 evaluation

This does not prove factual accuracy or professional judgment.

#5

GPT-5.1 Thinking

OpenAI · International

Reasoning AbilityGood
B

Planning, comparing options, following constraints and solving unfamiliar problems.

Show technical score
Overall accuracy58.2%
Result snapshot
2026-04-13
Mode
Published General365 evaluation

This does not prove factual accuracy or professional judgment.

Show displayed ranks 6–10
#6

Qwen3.5-397B-A17B Thinking

Qwen · China

Reasoning AbilityGood
B

Planning, comparing options, following constraints and solving unfamiliar problems.

Show technical score
Overall accuracy57.7%
Result snapshot
2026-04-13
Mode
Published General365 evaluation

This does not prove factual accuracy or professional judgment.

#7

DeepSeek-V3.2-Speciale

DeepSeek · China

Reasoning AbilityGood
B

Planning, comparing options, following constraints and solving unfamiliar problems.

Show technical score
Overall accuracy57.5%
Result snapshot
2026-04-13
Mode
Published General365 evaluation

This does not prove factual accuracy or professional judgment.

#8

GLM-4.7 Thinking

Z.ai · China

Reasoning AbilityFair
C

Planning, comparing options, following constraints and solving unfamiliar problems.

Show technical score
Overall accuracy57.4%
Result snapshot
2026-04-13
Mode
Published General365 evaluation

This does not prove factual accuracy or professional judgment.

#9

Qwen3-Max Thinking

Qwen · China

Reasoning AbilityFair
C

Planning, comparing options, following constraints and solving unfamiliar problems.

Show technical score
Overall accuracy57.2%
Result snapshot
2026-04-13
Mode
Published General365 evaluation

This does not prove factual accuracy or professional judgment.

#10

Kimi K2.5 Thinking

Moonshot AI · China

Reasoning AbilityFair
C

Planning, comparing options, following constraints and solving unfamiliar problems.

Show technical score
Overall accuracy56.4%
Result snapshot
2026-04-13
Mode
Published General365 evaluation

This does not prove factual accuracy or professional judgment.

Benchmark methodology, limitations and original sources

General365 contains 365 manually designed seed problems expanded into 1,095 variants across eight reasoning categories. The project keeps part of the benchmark held out to monitor contamination and reports a single within-benchmark accuracy score.

Mathematics

AIME 2026 and HMMT 2026 — selected official-feed rows

Two exam scores kept separateReview due soonChecked Jul 25, 2026Review by Aug 24, 20268 displayed systems

Raw percentages from the OpenEvals official benchmark aggregation. The two exams remain separate columns; the displayed mean ranks this mathematics category only.

#1

Kimi K2.5

Moonshot AI · China

Advanced Math AbilityExcellent
A

Competition-style mathematics and multi-step quantitative reasoning.

Show technical score
Category mean91.5%
AIME 2026
95.8%
HMMT 2026
87.1%
Snapshot
2026-07-24

This is not a guarantee of reliable everyday arithmetic or financial advice.

#2

Step-3.5-Flash

StepFun · China

Advanced Math AbilityExcellent
A

Competition-style mathematics and multi-step quantitative reasoning.

Show technical score
Category mean91.5%
AIME 2026
96.7%
HMMT 2026
86.4%
Snapshot
2026-07-24

This is not a guarantee of reliable everyday arithmetic or financial advice.

#3

GLM-5

Z.ai · China

Advanced Math AbilityExcellent
A

Competition-style mathematics and multi-step quantitative reasoning.

Show technical score
Category mean91.1%
AIME 2026
95.8%
HMMT 2026
86.4%
Snapshot
2026-07-24

This is not a guarantee of reliable everyday arithmetic or financial advice.

#4

Qwen3.5-397B-A17B

Qwen · China

Advanced Math AbilityGood
B

Competition-style mathematics and multi-step quantitative reasoning.

Show technical score
Category mean90.6%
AIME 2026
93.3%
HMMT 2026
87.9%
Snapshot
2026-07-24

This is not a guarantee of reliable everyday arithmetic or financial advice.

#5

DeepSeek-V3.2

DeepSeek · China

Advanced Math AbilityGood
B

Competition-style mathematics and multi-step quantitative reasoning.

Show technical score
Category mean89.1%
AIME 2026
94.2%
HMMT 2026
84.1%
Snapshot
2026-07-24

This is not a guarantee of reliable everyday arithmetic or financial advice.

Show displayed ranks 6–8
#6

Qwen3.5-35B-A3B

Qwen · China

Advanced Math AbilityGood
B

Competition-style mathematics and multi-step quantitative reasoning.

Show technical score
Category mean87.6%
AIME 2026
93.3%
HMMT 2026
81.8%
Snapshot
2026-07-24

This is not a guarantee of reliable everyday arithmetic or financial advice.

#7

NVIDIA Nemotron 3 Super 120B-A12B

NVIDIA · International

Advanced Math AbilityFair
C

Competition-style mathematics and multi-step quantitative reasoning.

Show technical score
Category mean87.4%
AIME 2026
90%
HMMT 2026
84.8%
Snapshot
2026-07-24

This is not a guarantee of reliable everyday arithmetic or financial advice.

#8

Qwen3.5-27B

Qwen · China

Advanced Math AbilityFair
C

Competition-style mathematics and multi-step quantitative reasoning.

Show technical score
Category mean85.9%
AIME 2026
90.8%
HMMT 2026
81.1%
Snapshot
2026-07-24

This is not a guarantee of reliable everyday arithmetic or financial advice.

Benchmark methodology, limitations and original sources

AIME 2026 and HMMT 2026 remain visible as separate official-feed scores. Their arithmetic mean is used only to order this mathematics section and is never reused as a universal model score.

Software engineering

SWE-bench Verified — selected official-feed rows

Human-validated repository issuesReview due soonChecked Jul 25, 2026Review by Aug 24, 20268 displayed systems

Percentage of the 500 human-validated software issues resolved by the listed system result in the official benchmark feed.

#1

Qwen3.5-397B-A17B

Qwen · China

Code Repair AbilityExcellent
A

Understanding a real software repository, locating a bug and producing a tested fix.

Show technical score
Issues resolved76.4%
Snapshot
2026-07-24
Comparison scope
Official feed system result; exact harness may vary

Language, framework, tools and agent setup still affect real production results.

#2

Step-3.5-Flash

StepFun · China

Code Repair AbilityExcellent
A

Understanding a real software repository, locating a bug and producing a tested fix.

Show technical score
Issues resolved74.4%
Snapshot
2026-07-24
Comparison scope
Official feed system result; exact harness may vary

Language, framework, tools and agent setup still affect real production results.

#3

GLM-5

Z.ai · China

Code Repair AbilityExcellent
A

Understanding a real software repository, locating a bug and producing a tested fix.

Show technical score
Issues resolved72.8%
Snapshot
2026-07-24
Comparison scope
Official feed system result; exact harness may vary

Language, framework, tools and agent setup still affect real production results.

#4

Qwen3.5-27B

Qwen · China

Code Repair AbilityGood
B

Understanding a real software repository, locating a bug and producing a tested fix.

Show technical score
Issues resolved72.4%
Snapshot
2026-07-24
Comparison scope
Official feed system result; exact harness may vary

Language, framework, tools and agent setup still affect real production results.

#5

Kimi K2.5

Moonshot AI · China

Code Repair AbilityGood
B

Understanding a real software repository, locating a bug and producing a tested fix.

Show technical score
Issues resolved70.8%
Snapshot
2026-07-24
Comparison scope
Official feed system result; exact harness may vary

Language, framework, tools and agent setup still affect real production results.

Show displayed ranks 6–8
#6

DeepSeek-V3.2

DeepSeek · China

Code Repair AbilityGood
B

Understanding a real software repository, locating a bug and producing a tested fix.

Show technical score
Issues resolved70%
Snapshot
2026-07-24
Comparison scope
Official feed system result; exact harness may vary

Language, framework, tools and agent setup still affect real production results.

#7

Qwen3.5-35B-A3B

Qwen · China

Code Repair AbilityFair
C

Understanding a real software repository, locating a bug and producing a tested fix.

Show technical score
Issues resolved69.2%
Snapshot
2026-07-24
Comparison scope
Official feed system result; exact harness may vary

Language, framework, tools and agent setup still affect real production results.

#8

NVIDIA Nemotron 3 Super 120B-A12B

NVIDIA · International

Code Repair AbilityFair
C

Understanding a real software repository, locating a bug and producing a tested fix.

Show technical score
Issues resolved53.7%
Snapshot
2026-07-24
Comparison scope
Official feed system result; exact harness may vary

Language, framework, tools and agent setup still affect real production results.

Benchmark methodology, limitations and original sources

SWE-bench Verified evaluates systems on 500 human-validated GitHub issues and checks proposed fixes with repository tests. Model version, agent scaffold, tools and inference settings remain part of a fair comparison.

Terminal agents

Terminal-Bench 2.1 — verified international and Chinese agent–model results

Live system benchmarkReview overdueChecked Jul 25, 2026Review by Aug 8, 20268 displayed systems

Terminal-Bench 2.1 evaluates 89 terminal tasks across software engineering, system administration, data processing, model training and security. Scores are not model-only ratings: the agent harness, effort setting, tool behavior and cost are part of each result. This selected table keeps one official entry per model so repeated harness submissions do not dominate the comparison.

#1

Fable 5 + Claude Code

Anthropic · International

Computer & Tool UseExcellent
A

Using a terminal, tools and multi-step actions to complete technical tasks.

Show technical score
Accuracy83.8%
Uncertainty
± 1.2%
Agent
Claude Code
Effort
xhigh
Run cost
$552.67
Submitted
2026-06-07

This is a complete agent-and-model system result, not a model-only grade.

#2

GPT-5.5 + Codex

OpenAI · International

Computer & Tool UseExcellent
A

Using a terminal, tools and multi-step actions to complete technical tasks.

Show technical score
Accuracy83.1%
Uncertainty
± 1.1%
Agent
Codex
Effort
xhigh
Run cost
$2,059.19
Submitted
2026-05-01

This is a complete agent-and-model system result, not a model-only grade.

#3

Grok 4.5 + Cursor CLI

xAI · International

Computer & Tool UseExcellent
A

Using a terminal, tools and multi-step actions to complete technical tasks.

Show technical score
Accuracy79.3%
Uncertainty
± 1.5%
Agent
Cursor CLI
Effort
high
Run cost
$134.09
Submitted
2026-07-09

This is a complete agent-and-model system result, not a model-only grade.

#4

Opus 4.8 + Claude Code

Anthropic · International

Computer & Tool UseGood
B

Using a terminal, tools and multi-step actions to complete technical tasks.

Show technical score
Accuracy78.9%
Uncertainty
± 1.3%
Agent
Claude Code
Effort
high
Run cost
$286.94
Submitted
2026-07-09

This is a complete agent-and-model system result, not a model-only grade.

#5

GPT-5.6 Terra + Codex

OpenAI · International

Computer & Tool UseGood
B

Using a terminal, tools and multi-step actions to complete technical tasks.

Show technical score
Accuracy78.4%
Uncertainty
± 1.3%
Agent
Codex
Effort
max
Run cost
$421.15
Submitted
2026-07-11

This is a complete agent-and-model system result, not a model-only grade.

Show displayed ranks 6–8
#6

Muse Spark 1.1 + mini-SWE-agent

Meta · International

Computer & Tool UseGood
B

Using a terminal, tools and multi-step actions to complete technical tasks.

Show technical score
Accuracy76.2%
Uncertainty
± 1.2%
Agent
mini-SWE-agent
Effort
xhigh
Run cost
$198.05
Submitted
2026-07-09

This is a complete agent-and-model system result, not a model-only grade.

#7

Gemini 3 Pro + Terminus 2

Google · International

Computer & Tool UseFair
C

Using a terminal, tools and multi-step actions to complete technical tasks.

Show technical score
Accuracy73.9%
Uncertainty
± 1.3%
Agent
Terminus 2
Effort
high
Run cost
$224.44
Submitted
2026-05-01

This is a complete agent-and-model system result, not a model-only grade.

#8

GLM-5.1 + Claude Code

Z.ai · China

Computer & Tool UseFair
C

Using a terminal, tools and multi-step actions to complete technical tasks.

Show technical score
Accuracy58.7%
Uncertainty
± 1.2%
Agent
Claude Code
Effort
max
Run cost
$277.14
Submitted
2026-05-01

This is a complete agent-and-model system result, not a model-only grade.

Benchmark methodology, limitations and original sources

At the July 24 cutoff, the official maintained leaderboard included GLM-5.1 as its Chinese-model entry. Kimi K3, GLM-5.2, Qwen3.7 Max, MiniMax-M3 and DeepSeek V4 Pro were not added here because no official Terminal-Bench 2.1 row for those exact models appeared on the maintained leaderboard at verification time.

Long context

LongBench Pro — 2026 realistic long-context leaderboard

1,500 bilingual long-context tasksReview due soonChecked Jul 25, 2026Review by Aug 24, 202610 displayed systems

Average performance across 1,500 English and Chinese long-context tasks spanning 11 primary task families and input lengths from 8K to 256K tokens. Scores use the official leaderboard evaluation setting shown for each model.

#1

Gemini 2.5 Pro

Google · International

Long-Document AbilityExcellent
A

Finding, connecting and summarizing information across long documents.

Show technical score
Overall score73.4%
Evaluation mode
Thinking
Model context cap
1M
Result snapshot
2026-01-08

A high grade does not guarantee that every detail in a very long file will be handled correctly.

#2

GPT-5

OpenAI · International

Long-Document AbilityExcellent
A

Finding, connecting and summarizing information across long documents.

Show technical score
Overall score72.6%
Evaluation mode
Thinking
Model context cap
272K
Result snapshot
2026-01-08

A high grade does not guarantee that every detail in a very long file will be handled correctly.

#3

Claude 4 Sonnet

Anthropic · International

Long-Document AbilityExcellent
A

Finding, connecting and summarizing information across long documents.

Show technical score
Overall score69.9%
Evaluation mode
Thinking
Model context cap
1M
Result snapshot
2026-01-08

A high grade does not guarantee that every detail in a very long file will be handled correctly.

#4

DeepSeek-V3.2

DeepSeek · China

Long-Document AbilityExcellent
A

Finding, connecting and summarizing information across long documents.

Show technical score
Overall score67.8%
Evaluation mode
Thinking
Model context cap
120K
Result snapshot
2026-01-08

A high grade does not guarantee that every detail in a very long file will be handled correctly.

#5

Qwen3-235B-A22B Thinking

Qwen · China

Long-Document AbilityGood
B

Finding, connecting and summarizing information across long documents.

Show technical score
Overall score67%
Evaluation mode
Thinking
Model context cap
224K
Result snapshot
2026-01-08

A high grade does not guarantee that every detail in a very long file will be handled correctly.

Show displayed ranks 6–10
#6

Qwen3-Next-80B-A3B Thinking

Qwen · China

Long-Document AbilityGood
B

Finding, connecting and summarizing information across long documents.

Show technical score
Overall score64%
Evaluation mode
Thinking
Model context cap
224K
Result snapshot
2026-01-08

A high grade does not guarantee that every detail in a very long file will be handled correctly.

#7

Qwen3-Next-80B-A3B Instruct

Qwen · China

Long-Document AbilityGood
B

Finding, connecting and summarizing information across long documents.

Show technical score
Overall score60.8%
Evaluation mode
Thinking prompt
Model context cap
224K
Result snapshot
2026-01-08

A high grade does not guarantee that every detail in a very long file will be handled correctly.

#8

Kimi K2 Instruct 0905

Moonshot AI · China

Long-Document AbilityFair
C

Finding, connecting and summarizing information across long documents.

Show technical score
Overall score55.5%
Evaluation mode
Thinking prompt
Model context cap
224K
Result snapshot
2026-01-08

A high grade does not guarantee that every detail in a very long file will be handled correctly.

#9

MiniMax M2

MiniMax · China

Long-Document AbilityFair
C

Finding, connecting and summarizing information across long documents.

Show technical score
Overall score53.2%
Evaluation mode
Thinking
Model context cap
1M
Result snapshot
2026-01-08

A high grade does not guarantee that every detail in a very long file will be handled correctly.

#10

GPT-OSS-120B

OpenAI · International

Long-Document AbilityFair
C

Finding, connecting and summarizing information across long documents.

Show technical score
Overall score52.6%
Evaluation mode
Thinking
Model context cap
120K
Result snapshot
2026-01-08

A high grade does not guarantee that every detail in a very long file will be handled correctly.

Benchmark methodology, limitations and original sources

LongBench Pro uses 1,500 naturally occurring English and Chinese samples across 11 primary task families, six length bands from 8K to 256K tokens and four calibrated difficulty levels. The ranking uses the official leaderboard’s average overall metric for the displayed evaluation mode.

Document and OCR

olmOCR benchmark — document-understanding specialists

Specialist document benchmarkReview due soonChecked Jul 25, 2026Review by Aug 24, 202610 displayed systems

This category is document/OCR evaluation, not a universal vision score. General chat models are not assigned an OCR result unless the official feed contains one.

#1

chandra-ocr-2

Datalab · International

Scanned Document ReadingExcellent
A

Extracting text, tables and page structure from scans, PDFs and document images.

Show technical score
Document score85.9%
Snapshot
2026-07-24
Mode
Maintained specialist benchmark feed

This does not measure general photography, image generation or conversational quality.

#2

dots.mocr

RedNote HiLab · China

Scanned Document ReadingExcellent
A

Extracting text, tables and page structure from scans, PDFs and document images.

Show technical score
Document score83.9%
Snapshot
2026-07-24
Mode
Maintained specialist benchmark feed

This does not measure general photography, image generation or conversational quality.

#3

LightOnOCR-2-1B

LightOn AI · International

Scanned Document ReadingExcellent
A

Extracting text, tables and page structure from scans, PDFs and document images.

Show technical score
Document score83.2%
Snapshot
2026-07-24
Mode
Maintained specialist benchmark feed

This does not measure general photography, image generation or conversational quality.

#4

chandra

Datalab · International

Scanned Document ReadingExcellent
A

Extracting text, tables and page structure from scans, PDFs and document images.

Show technical score
Document score83.1%
Snapshot
2026-07-24
Mode
Maintained specialist benchmark feed

This does not measure general photography, image generation or conversational quality.

#5

Infinity-Parser-7B

Infiny AI · International

Scanned Document ReadingGood
B

Extracting text, tables and page structure from scans, PDFs and document images.

Show technical score
Document score82.5%
Snapshot
2026-07-24
Mode
Maintained specialist benchmark feed

This does not measure general photography, image generation or conversational quality.

Show displayed ranks 6–10
#6

olmOCR-2-7B FP8

AllenAI · International

Scanned Document ReadingGood
B

Extracting text, tables and page structure from scans, PDFs and document images.

Show technical score
Document score82.4%
Snapshot
2026-07-24
Mode
Maintained specialist benchmark feed

This does not measure general photography, image generation or conversational quality.

#7

PaddleOCR-VL

PaddlePaddle · China

Scanned Document ReadingGood
B

Extracting text, tables and page structure from scans, PDFs and document images.

Show technical score
Document score80%
Snapshot
2026-07-24
Mode
Maintained specialist benchmark feed

This does not measure general photography, image generation or conversational quality.

#8

Qianfan-OCR

Baidu · China

Scanned Document ReadingFair
C

Extracting text, tables and page structure from scans, PDFs and document images.

Show technical score
Document score79.8%
Snapshot
2026-07-24
Mode
Maintained specialist benchmark feed

This does not measure general photography, image generation or conversational quality.

#9

DeepSeek-OCR-2

DeepSeek · China

Scanned Document ReadingFair
C

Extracting text, tables and page structure from scans, PDFs and document images.

Show technical score
Document score76.3%
Snapshot
2026-07-24
Mode
Maintained specialist benchmark feed

This does not measure general photography, image generation or conversational quality.

#10

GLM-OCR

Z.ai · China

Scanned Document ReadingFair
C

Extracting text, tables and page structure from scans, PDFs and document images.

Show technical score
Document score75.2%
Snapshot
2026-07-24
Mode
Maintained specialist benchmark feed

This does not measure general photography, image generation or conversational quality.

Benchmark methodology, limitations and original sources

The olmOCR category focuses on document extraction specialists. General chat models receive no score unless the maintained feed contains a directly comparable result for the exact evaluated system.

Technical appendix: evidence coverage and benchmark-governance rules

Global model evidence coverage

This matrix shows which capability areas the daily registry monitors. It does not claim that every model has a verified score in every category.

OriginProviderModelGeneralCodingImage / multimodalAgents / toolsLong context
InternationalOpenAIGPT-5.6 SolTrackedTrackedTrackedTrackedNot evaluated
InternationalOpenAIGPT-5.6 TerraTrackedTrackedTrackedTrackedNot evaluated
InternationalOpenAIGPT-5.6 LunaTrackedNot evaluatedNot evaluatedNot evaluatedNot evaluated
InternationalAnthropicClaude Fable 5TrackedTrackedTrackedTrackedNot evaluated
InternationalxAIGrok 4.5TrackedTrackedNot evaluatedTrackedTracked
InternationalGoogleGemini 3.5 FlashTrackedTrackedTrackedTrackedNot evaluated
ChinaDeepSeekV4 FlashTrackedTrackedNot evaluatedNot evaluatedTracked
ChinaAlibaba CloudQwen3.7-MaxTrackedTrackedNot evaluatedTrackedNot evaluated
ChinaMoonshot AIKimi K3TrackedTrackedNot evaluatedTrackedTracked
ChinaZhipu AIGLM-5.2TrackedTrackedNot evaluatedTrackedTracked
ChinaBaiduERNIE 5.0TrackedNot evaluatedTrackedNot evaluatedTracked
ChinaByteDanceDoubao Seed 2.1TrackedTrackedTrackedTrackedNot evaluated
ChinaMiniMaxMiniMax-M3TrackedTrackedTrackedTrackedTracked
ChinaStepFunStep 3.7 FlashTrackedTrackedTrackedTrackedNot evaluated
ChinaTencentHunyuan A13BTrackedTrackedNot evaluatedTrackedTracked

Responsible benchmark and grade rules

  1. 1

    Assign A, B and C only inside the exact benchmark category being displayed.

  2. 2

    Keep the raw percentage available and never replace the original technical evidence.

  3. 3

    Show the official source snapshot date separately from the latest verification date.

  4. 4

    Label selected feeds, deduplicated lists and incomplete coverage instead of calling them complete leaderboards.

  5. 5

    Distinguish model-only evaluations from agent-and-model system results.

  6. 6

    Do not invent grades for capabilities—such as creative writing—that the current registry does not directly test.