Gemini 3.7 Flash vs Inkling-Small: Specs & Benchmark Comparison

Gemini 3.7 Flash is developed by Google, while Inkling-Small comes from Thinking Machines Lab. Inkling-Small was released in July 2026, and Gemini 3.7 Flash followed a month later in August 2026. Inkling-Small has a published size of about 276 billion parameters; Google has not disclosed the parameter count of Gemini 3.7 Flash.

The two models share 4 published benchmarks. Gemini 3.7 Flash leads on 4 of them. The widest gaps are on Terminal-Bench 2.1, where Gemini 3.7 Flash scores 85.8% against 64.7%; Artificial Analysis, where Gemini 3.7 Flash scores 56.0% against 40.0%. Averaged across everything we track, Gemini 3.7 Flash sits at 59.8% and Inkling-Small at 60.3%.

Inkling-Small is the cheaper API at $0.30 per million input tokens and $1.20 per million output tokens, roughly 3 times cheaper than Gemini 3.7 Flash at $0.75 and $3.75. Gemini 3.7 Flash takes the larger context window at 1M tokens, compared with 256K for Inkling-Small. Gemini 3.7 Flash accepts text and images as input, while Inkling-Small accepts text, images, and audio. On tooling, only Gemini 3.7 Flash supports function calling and only Gemini 3.7 Flash offers structured output.

CharacteristicGemini 3.7 FlashInkling-Small
CompanyGoogleThinking Machines Lab
Release DateAugust 13, 2026July 30, 2026
Parameters276B
MultimodalYesYes
Context (input)1.0M256K
Context (output)66K256K
Input Price / 1M$0.75$0.30
Output Price / 1M$3.75$1.20
Average Score59.8%60.3%
Benchmarks
Terminal-Bench 2.185.8%64.7%
Artificial Analysis56.0%40.0%
CharXiv-R88.7%77.4%
GDPval-AA50.8%42.3%

Visual Benchmark Comparison

Gemini 3.7 Flash
Inkling-Small
Terminal-Bench 2.10.9 vs 0.6
0.9
0.6
Artificial Analysis0.6 vs 0.4
0.6
0.4
CharXiv-R0.9 vs 0.8
0.9
0.8
GDPval-AA0.5 vs 0.4
0.5
0.4

Verdict

Gemini 3.7 Flash leads in 2 out of 4 comparison categories.

Overall Performance

Both models show comparable average scores: Gemini 3.7 Flash — 0.6, Inkling-Small — 0.6.

API Cost

Inkling-Small is 3.0x cheaper: input $0.30/1M vs $0.75/1M tokens.

Context Window

Gemini 3.7 Flash supports a larger context: 1M vs 256K tokens.

Recency

Gemini 3.7 Flash is newer: released 8/13/2026 vs 7/30/2026.

More About These Models

Related Comparisons

Frequently Asked Questions

Which is better for coding — Gemini 3.7 Flash or Inkling-Small?
Direct comparison on the SWE-Bench benchmark is not available. We recommend reviewing other metrics on the comparison page.
Which model is cheaper — Gemini 3.7 Flash or Inkling-Small?
Inkling-Small is cheaper for input: $0.30 per 1M tokens vs $0.75.
Which has a larger context window — Gemini 3.7 Flash or Inkling-Small?
Gemini 3.7 Flash supports a larger context: 1,048,576 tokens vs 256,000.

The Gemini 3.7 Flash and Inkling-Small comparison is updated for 2026. Data includes benchmark results, API pricing, context window size and other specifications. For more detailed information, visit the Gemini 3.7 Flash or Inkling-Small page. See also the complete list of AI model comparisons.