DeepSeek-V4.1-Flash vs Inkling-Small: Specs & Benchmark Comparison
DeepSeek-V4.1-Flash is developed by DeepSeek, while Inkling-Small comes from Thinking Machines Lab. Inkling-Small was released in July 2026, and DeepSeek-V4.1-Flash followed 2 months later in September 2026. DeepSeek-V4.1-Flash is the larger model at roughly 552 billion parameters, against 276 billion for Inkling-Small.
The two models share 2 published benchmarks. DeepSeek-V4.1-Flash leads on 2 of them. The widest gaps are on Terminal-Bench 2.1, where DeepSeek-V4.1-Flash scores 90.6% against 64.7%; GPQA Diamond, where DeepSeek-V4.1-Flash scores 90.9% against 89.5%. Averaged across everything we track, DeepSeek-V4.1-Flash sits at 60.9% and Inkling-Small at 60.3%.
On price the two are close: DeepSeek-V4.1-Flash costs $0.30 per million input tokens and $1.20 per million output tokens, Inkling-Small $0.30 and $1.20. DeepSeek-V4.1-Flash takes the larger context window at 1M tokens, compared with 256K for Inkling-Small. DeepSeek-V4.1-Flash accepts text and images as input, while Inkling-Small accepts text, images, and audio. On tooling, only DeepSeek-V4.1-Flash supports function calling and only DeepSeek-V4.1-Flash offers structured output.
| Characteristic | DeepSeek-V4.1-Flash | Inkling-Small |
|---|---|---|
| Company | DeepSeek | Thinking Machines Lab |
| Release Date | September 10, 2026 | July 30, 2026 |
| Parameters | 552B | 276B |
| Multimodal | Yes | Yes |
| Context (input) | 1.0M | 256K |
| Context (output) | 393K | 256K |
| Input Price / 1M | $0.30 | $0.30 |
| Output Price / 1M | $1.20 | $1.20 |
| Average Score | 60.9% | 60.3% |
| Benchmarks | ||
| Terminal-Bench 2.1 | 90.6% | 64.7% |
| GPQA Diamond | 90.9% | 89.5% |
Visual Benchmark Comparison
Verdict
DeepSeek-V4.1-Flash leads in 2 out of 4 comparison categories.
Both models show comparable average scores: DeepSeek-V4.1-Flash — 0.6, Inkling-Small — 0.6.
API cost is identical: input $0.30/1M, output $1.20/1M tokens.
DeepSeek-V4.1-Flash supports a larger context: 1M vs 256K tokens.
DeepSeek-V4.1-Flash is newer: released 9/10/2026 vs 7/30/2026.
More About These Models
Related Comparisons
Frequently Asked Questions
Which is better for coding — DeepSeek-V4.1-Flash or Inkling-Small?
Which model is cheaper — DeepSeek-V4.1-Flash or Inkling-Small?
Which has a larger context window — DeepSeek-V4.1-Flash or Inkling-Small?
The DeepSeek-V4.1-Flash and Inkling-Small comparison is updated for 2026. Data includes benchmark results, API pricing, context window size and other specifications. For more detailed information, visit the DeepSeek-V4.1-Flash or Inkling-Small page. See also the complete list of AI model comparisons.