Key Specifications
Parameters
1.0T
Context
256.0K
Release Date
February 13, 2026
Average Score
64.2%
Timeline
Key dates in the model's history
Announcement
February 13, 2026
Last Update
September 10, 2026
Today
September 20, 2026
Technical Specifications
Parameters
1.0T
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Pricing & Availability
Input (per 1M tokens)
$1.20
Output (per 1M tokens)
$6.00
Max Input Tokens
256.0K
Max Output Tokens
256.0K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning
Benchmark Results
Model performance metrics across various tests and benchmarks
Programming
Programming skills tests
SWE-Bench Verified
Publisher comparison-table evaluation; scaffolding and agent setup not specified. • Self-reported
Reasoning
Logical reasoning and analysis
GPQA
Publisher comparison-table evaluation; setup not further specified. • Self-reported
Other Tests
Specialized benchmarks
AA-LCR
Publisher comparison-table evaluation; setup not further specified. • Self-reported
AIME 2026
Publisher comparison-table evaluation on AIME26; setup not further specified. • Self-reported
BFCL-V4
Publisher comparison-table evaluation; setup not further specified. • Self-reported
BrowseComp
Simple context-folding with a 256k context window. • Self-reported
BrowseComp-zh
Publisher comparison-table evaluation; setup not further specified. • Self-reported
C-Eval
Publisher comparison-table evaluation; setup not further specified. • Self-reported
DeepPlanning
Publisher comparison-table evaluation; setup not further specified. • Self-reported
Global PIQA
Publisher comparison-table evaluation; setup not further specified. • Self-reported
HLE-Verified
Verified and revised HLE subset with component-wise verification and a fine-grained error taxonomy. • Self-reported
Humanity's Last Exam (no tools, text-only)
Text-only HLE subset without tools. • Self-reported
Humanity's Last Exam (with tools, text-only)
Text-only HLE subset with tool use enabled. • Self-reported
IFBench
Publisher comparison-table evaluation; setup not further specified. • Self-reported
IFEval
Publisher comparison-table evaluation; setup not further specified. • Self-reported
IMO-AnswerBench
Publisher comparison-table evaluation; setup not further specified. • Self-reported
LiveCodeBench v6
Publisher comparison-table evaluation; setup not further specified. • Self-reported
LongBench v2
Publisher comparison-table evaluation; setup not further specified. • Self-reported
MAXIFE
Accuracy on English plus multilingual original prompts across 23 settings. • Self-reported
MCP-Mark
GitHub MCP server v0.30.3 from api.githubcopilot.com; Playwright responses truncated at 32k tokens. • Self-reported
MMLU-Pro
Publisher comparison-table evaluation; setup not further specified. • Self-reported
MMLU-Redux
Publisher comparison-table evaluation; setup not further specified. • Self-reported
MMMLU
Publisher comparison-table evaluation; setup not further specified. • Self-reported
Multi-Challenge
Publisher comparison-table evaluation; setup not further specified. • Self-reported
NOVA-63
Publisher comparison-table evaluation; setup not further specified. • Self-reported
Seal-0
Publisher comparison-table evaluation; setup not further specified. • Self-reported
SecCodeBench
Publisher comparison-table evaluation; setup not further specified. • Self-reported
SuperGPQA
Publisher comparison-table evaluation; setup not further specified. • Self-reported
SWE-bench Multilingual
Publisher comparison-table evaluation; scaffolding and agent setup not specified. • Self-reported
Terminal-Bench 2.0
Publisher comparison-table evaluation; agent harness and setup not specified. • Self-reported
Toolathlon
Publisher comparison-table evaluation; setup not further specified. • Self-reported
VITA-Bench
Publisher comparison-table evaluation; setup not further specified. • Self-reported
WideSearch
256k context window without context management. • Self-reported
WMT24++
Average across 55 languages using XCOMET-XXL on the harder rebalanced WMT24++ subset. • Self-reported
License & Metadata
License
proprietary
Announcement Date
February 13, 2026
Last Updated
September 10, 2026
Similar Models
All ModelsQwen3.5 122B A10B
Alibaba
122.0B
Best score:0.9 (GPQA)
Released:Mar 2026
Qwen3-235B-A22B-Thinking-2507
Alibaba
235.0B
Best score:0.8 (GPQA)
Released:Jul 2025
Price:$0.30/1M tokens
Qwen3 235B A22B
Alibaba
235.0B
Best score:0.9 (MMLU)
Released:Apr 2025
Price:$0.20/1M tokens
Qwen3-235B-A22B-Instruct-2507
Alibaba
235.0B
Best score:0.8 (GPQA)
Released:Jul 2025
Price:$0.15/1M tokens
Qwen3-Coder 480B A35B Instruct
Alibaba
480.0B
Best score:0.8 (TAU)
Released:Jan 2025
Qwen3.7 Max
Alibaba
Best score:0.9 (GPQA)
Released:May 2026
Price:$2.50/1M tokens
Qwen3.5 9B
Alibaba
9.0B
Best score:0.8 (GPQA)
Released:Mar 2026
Qwen3.5 35B A3B
Alibaba
35.0B
Best score:0.8 (GPQA)
Released:Mar 2026
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.