InclusionAI logo

Ling 3.0 Flash

InclusionAI

The model prioritizes token efficiency and agentic inference at production scale, stretching what developers can achieve within limited token, latency, and serving-cost budgets.

Key Specifications

Parameters
124.0B
Context
131.1K
Release Date
August 4, 2026
Average Score
68.5%

Timeline

Key dates in the model's history
Announcement
August 4, 2026
Last Update
September 10, 2026
Today
September 20, 2026

Technical Specifications

Parameters
124.0B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Pricing & Availability

Input (per 1M tokens)
$0.06
Output (per 1M tokens)
$0.18
Max Input Tokens
131.1K
Max Output Tokens
131.1K
Supported Features
Function CallingStructured OutputCode ExecutionWeb SearchBatch InferenceFine-tuning

Benchmark Results

Model performance metrics across various tests and benchmarks

Other Tests

Specialized benchmarks
AA-LCR
Official model-card evaluation table; detailed setup not specified.Self-reported
65.1%
AIME 2026
Official model-card evaluation table; detailed setup not specified.Self-reported
93.2%
Artifacts Bench
Official model-card evaluation table; detailed setup not specified.Self-reported
77.0%
BFCL-V4
Official model-card evaluation table; detailed setup not specified.Self-reported
73.0%
BrowseComp
Internal single-agent search harness; context resume at 64K via trajectory summarization and history discard; average pass@1.Self-reported
72.2%
DRACO
Internal search-agent harness; official per-question rubrics averaged across questions; Claude Opus 4.6 scoring model.Self-reported
70.4%
GDPval-AA
GDPval v2-AA Elo on the public 220-task set using the official Stirrup harness; 250-turn limit; 5-hour timeout.Self-reported
36.9%
HMMT Feb 26
Official model-card evaluation table; detailed setup not specified.Self-reported
87.0%
IFBench
Official model-card evaluation table; detailed setup not specified.Self-reported
74.5%
IMO-AnswerBench
Official model-card evaluation table; detailed setup not specified.Self-reported
83.7%
LiveCodeBench v6
Official model-card evaluation table; detailed setup not specified.Self-reported
82.8%
MCP Atlas
500-task public set; official v1 harness; 20-turn limit; Gemini-2.5-Pro claim-coverage judge.Self-reported
65.5%
Multi-IF
Official model-card evaluation table; detailed setup not specified.Self-reported
87.7%
SkillsBench
Kilo Code harness on 87 default runnable tasks, excluding external-API-dependent tasks; average of 3 runs.Self-reported
44.8%
SWE-bench Multilingual
OpenHands agent harness with tailored prompts; temperature=0.6, top_p=0.95, max_new_tokens=32K; 256K context.Self-reported
72.4%
SWE-Bench Pro
OpenHands agent harness with tailored prompts; temperature=0.6, top_p=0.95, max_new_tokens=32K; 256K context.Self-reported
56.6%
Tau3 Banking
Artificial Analysis-aligned protocol; GPT-5.4-mini with medium reasoning used for user simulation and natural-language assertion judging.Self-reported
28.0%
Terminal-Bench 2.1
Artificial Analysis protocol; default Terminus 2 harness; 2-hour timeout; preserve-thinking JSON parser; 3-run task mean; temperature=0.6, top_p=1.0, max_new_tokens=32K; 256K context.Self-reported
57.0%
WideSearch
Internal basic-ReAct single-agent search harness; official prompt and GPT-4.1 judge on the corrected dataset; average pass@1.Self-reported
73.6%

License & Metadata

License
unknown
Announcement Date
August 4, 2026
Last Updated
September 10, 2026

Compare Ling 3.0 Flash

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.