V4 Pro
Best for long-context reasoning, coding and agent workflows.
Value for moneyPerformance evidence pendingThe model has a price, but no exact or recent same-family result in the six-category benchmark registry.
Where it is available: DeepSeek API, app and web
Important limitation: Official August 13 production revision DeepSeek-V4-Pro-0813. Prices shown are the current pre-August 17 tariff; announced peak/off-peak rates are future-dated. No independent 0813 benchmark package was verified by the edition cutoff.
Check official model information↗ (opens in a new tab)Claude Opus 5
Best for long-horizon reasoning, interactive agent work and complex professional tasks.
Value for moneyPerformance evidence pendingThe model has a price, but no exact or recent same-family result in the six-category benchmark registry.
Where it is available: Anthropic API, Amazon Bedrock, Google Cloud and Microsoft Foundry
Important limitation: ARC-AGI-3 performance is benchmark-specific. Anthropic effort labels are provider-specific and should not be treated as standardized units of compute across vendors.
Check official model information↗ (opens in a new tab)GPT-5.6 Sol
Best for difficult planning, research and multi-step professional work.
Value for moneyProvisional premium-priced16/100 value indexCapability may be strong, but the comparable API price is high relative to other plotted models. The performance side uses a recent same-family proxy rather than the exact current model.
Where it is available: API and eligible ChatGPT plans
Important limitation: Launch comparisons are provider-published; confirm task-specific performance independently.
Check official model information↗ (opens in a new tab)GPT-5.6 Terra
Best for everyday writing, email drafting, analysis and balanced cost.
Value for moneyPremium-priced20/100 value indexCapability may be strong, but the comparable API price is high relative to other plotted models.
Where it is available: API and eligible ChatGPT plans
Important limitation: Use current API documentation for limits and regional availability.
Check official model information↗ (opens in a new tab)GPT-5.6 Luna
Best for quick drafts, summaries and routine high-volume tasks.
Value for moneyProvisional premium-priced36/100 value indexCapability may be strong, but the comparable API price is high relative to other plotted models. The performance side uses a recent same-family proxy rather than the exact current model.
Where it is available: API and eligible ChatGPT plans
Important limitation: Lower price does not imply best performance for every workload.
Check official model information↗ (opens in a new tab)Claude Fable 5
Best for long documents, careful editing and sustained project work.
Value for moneyPremium-priced13/100 value indexCapability may be strong, but the comparable API price is high relative to other plotted models.
Where it is available: Claude and API where available
Important limitation: Anthropic documents additional safeguards and retention behavior for Mythos-class traffic.
Check official model information↗ (opens in a new tab)Grok 4.5
Best for coding, agent-style tasks and workflows that use current information.
Value for moneyPremium-priced35/100 value indexCapability may be strong, but the comparable API price is high relative to other plotted models.
Where it is available: xAI API and Grok products
Important limitation: Requests exceeding 200K context use higher prices: $4 input and $12 output per 1M tokens.
Check official model information↗ (opens in a new tab)Gemini 3.5 Flash
Best for fast multimodal work, Google users and quick everyday assistance.
Value for moneyPrice verification neededNo stable comparable input-and-output API price is available in the current registry.
Where it is available: Gemini ecosystem and developer services
Important limitation: This edition did not capture a stable region-neutral official API price; verify before purchase.
Check official model information↗ (opens in a new tab)V4 Flash
Best for lower-cost long-document and technical workflows where regional access works.
Value for moneyProvisional outstanding value100/100 value indexHigh evidence-backed capability relative to the blended USD API price. The performance side uses a recent same-family proxy rather than the exact current model.
Where it is available: DeepSeek API through the explicit deepseek-v4-flash and deepseek-v4-pro model names; legacy aliases are past their retirement deadline.
Important limitation: DeepSeek primary documentation identifies April 24 as the V4 Preview API availability date. Later third-party score snapshots vary and should be read with their evaluation date and methodology.
Check official model information↗ (opens in a new tab)Qwen3.7-Max
Best for complex reasoning, coding and work inside the Alibaba Cloud ecosystem.
Value for moneyCompare in source currencyThe tracker does not convert CNY pricing into USD because exchange rates and regional billing change.
Where it is available: Alibaba Cloud Model Studio; regional endpoints vary
Important limitation: Pricing, endpoint names and availability differ between China and international Model Studio regions.
Check official model information↗ (opens in a new tab)Kimi K3
Best for very long documents, software work and extended agent-style tasks.
Value for moneyCompare in source currencyThe tracker does not convert CNY pricing into USD because exchange rates and regional billing change.
Where it is available: Kimi API platform; region and account eligibility vary
Important limitation: Prices are official CNY rates per million tokens and must not be displayed as USD.
Check official model information↗ (opens in a new tab)GLM-5.2
Best for long-running coding and autonomous workflow experiments.
Value for moneyCompare in source currencyThe tracker does not convert CNY pricing into USD because exchange rates and regional billing change.
Where it is available: Zhipu BigModel platform; regional access varies
Important limitation: Provider capability statements require independent task-specific reproduction.
Check official model information↗ (opens in a new tab)ERNIE 5.0
Best for Chinese-language and multimodal tasks in Baidu services.
Value for moneyPerformance evidence pendingThe model has a price, but no exact or recent same-family result in the six-category benchmark registry.
Where it is available: Baidu Qianfan and ERNIE services; international catalog availability varies
Important limitation: The displayed API price is from Baidu’s international Qianfan catalog; China-region billing differs.
Check official model information↗ (opens in a new tab)Doubao Seed 2.1
Best for production-oriented coding, agent and multimodal tasks in Volcengine.
Value for moneyCompare in source currencyThe tracker does not convert CNY pricing into USD because exchange rates and regional billing change.
Where it is available: Volcengine Ark; regional availability varies
Important limitation: Volcengine pricing uses tiered regional tables; verify the exact model ID and input-length tier.
Check official model information↗ (opens in a new tab)MiniMax-M3
Best for coding, reasoning and long-context work through compatible endpoints.
Value for moneyPrice verification neededNo stable comparable input-and-output API price is available in the current registry.
Where it is available: MiniMax API platform and compatible endpoints
Important limitation: Use the current regional pricing table and distinguish API pay-as-you-go from Token Plan subscriptions.
Check official model information↗ (opens in a new tab)Step 3.7 Flash
Best for lower-cost multimodal reasoning and computer-use-style tasks.
Value for moneyCompare in source currencyThe tracker does not convert CNY pricing into USD because exchange rates and regional billing change.
Where it is available: StepFun open platform; regional access varies
Important limitation: Official prices are CNY per million tokens. Confirm current quotas and model lifecycle notices.
Check official model information↗ (opens in a new tab)Hunyuan A13B
Best for efficient reasoning in Tencent Cloud and Chinese enterprise workflows.
Value for moneyCompare in source currencyThe tracker does not convert CNY pricing into USD because exchange rates and regional billing change.
Where it is available: Tencent Cloud Hunyuan / migration path to TokenHub
Important limitation: Tencent is migrating newer model access toward TokenHub; verify the active platform before integration.
Check official model information↗ (opens in a new tab)Qwen-Image-3.0
Best for images containing text, interfaces, diagrams and dense layouts.
Value for moneyPrice verification neededNo stable comparable input-and-output API price is available in the current registry.
Where it is available: Qwen and Alibaba Cloud ecosystem; verify regional endpoint availability
Important limitation: Capabilities and examples are provider-published; independent benchmark reproduction is pending.
Check official model information↗ (opens in a new tab)Qwen-Audio-3.0-TTS Flash / Plus
Best for real-time or higher-fidelity text-to-speech generation.
Value for moneyPrice verification neededNo stable comparable input-and-output API price is available in the current registry.
Where it is available: Alibaba Cloud Model Studio; rollout and regional availability may vary
Important limitation: The 300ms-level first-packet latency and multilingual results are provider-reported; verify independently for production workloads.
Check official model information↗ (opens in a new tab)Llama 4 Scout
Best for very long context and multimodal local or private-cloud deployment.
Value for moneyLocal test requiredThe weights are free to obtain, but hardware, power, hosting and operations determine total cost.
Local deployment: Meta says an Int4 version fits on one NVIDIA H100 GPU.
Where it is available: Downloadable open weights through Meta and ecosystem partners
Important limitation: The weights have no per-token license fee, but a practical deployment still needs GPU capacity, storage, power, security and operations.
Check official model information↗ (opens in a new tab)Llama 4 Maverick
Best for higher-capability multimodal workloads with open-weight control.
Value for moneyLocal test requiredThe weights are free to obtain, but hardware, power, hosting and operations determine total cost.
Local deployment: Meta describes deployment on a single H100 DGX host or distributed infrastructure.
Where it is available: Downloadable open weights through Meta and ecosystem partners
Important limitation: Maverick stores 400B total parameters even though 17B are active per token, so hosting is materially heavier than the active-parameter figure alone suggests.
Check official model information↗ (opens in a new tab)gpt-oss-20b
Best for local reasoning on constrained hardware and rapid private experimentation.
Value for moneyLocal test requiredThe weights are free to obtain, but hardware, power, hosting and operations determine total cost.
Local deployment: OpenAI says the model can run with about 16 GB of memory.
Where it is available: Free downloadable weights; self-hosted or third-party hosted
Important limitation: OpenAI does not serve this model through ChatGPT or the OpenAI API. The operator pays for compute, storage, monitoring and maintenance.
Check official model information↗ (opens in a new tab)gpt-oss-120b
Best for high-capacity private reasoning and agent workflows.
Value for moneyPotentially strong locallyThe weights are free to obtain, but hardware, power, hosting and operations determine total cost.
Local deployment: OpenAI says the model fits on one 80 GB GPU.
Where it is available: Free downloadable weights; self-hosted or third-party hosted
Important limitation: The weights are free, but OpenAI says the model needs about 80 GB of memory; total deployment cost depends on utilization and operations.
Check official model information↗ (opens in a new tab)Ministral 3 14B
Best for smaller local enterprise assistant with multimodal support.
Value for moneyLocal test requiredThe weights are free to obtain, but hardware, power, hosting and operations determine total cost.
Local deployment: Dense 14B model; actual memory depends on precision, quantization and runtime.
Where it is available: Downloadable weights and hosted ecosystem endpoints
Important limitation: A free model license does not make production inference free. Quantization, hardware, throughput and support requirements determine cost.
Check official model information↗ (opens in a new tab)Mistral Large 3
Best for frontier-class open-weight enterprise customization.
Value for moneyLocal test requiredThe weights are free to obtain, but hardware, power, hosting and operations determine total cost.
Local deployment: Data-center-class sparse model; deployment normally requires substantial accelerator infrastructure.
Where it is available: Downloadable weights and hosted ecosystem endpoints
Important limitation: The model has 675B total parameters and 41B active parameters; storage and serving requirements remain substantial despite sparse activation.
Check official model information↗ (opens in a new tab)Gemma 3 27B
Best for multilingual local assistant, document and image workflows.
Value for moneyLocal test requiredThe weights are free to obtain, but hardware, power, hosting and operations determine total cost.
Local deployment: Dense 27B model; actual memory depends on precision, quantization and runtime.
Where it is available: Downloadable open model through Google and ecosystem partners
Important limitation: The model is free to obtain under Google’s terms, but businesses must still budget for hardware, inference software, monitoring and updates.
Check official model information↗ (opens in a new tab)Nemotron 3 Super 120B-A12B
Best for efficient long-running multi-agent and code workflows.
Value for moneyPotentially strong locallyThe weights are free to obtain, but hardware, power, hosting and operations determine total cost.
Local deployment: 120B total / 12B active sparse model; optimized for NVIDIA Blackwell and data-center deployment.
Where it is available: Open weights, datasets and recipes; self-hosting and NVIDIA ecosystem deployment
Important limitation: NVIDIA optimizes the model for its accelerator stack. Cost and throughput can differ substantially on other hardware and runtimes.
Check official model information↗ (opens in a new tab)