Inkling-Small vs Nemotron 3.5 Lightning (30B A3B): Specs & Benchmark Comparison

Inkling-Small is developed by Thinking Machines Lab, while Nemotron 3.5 Lightning (30B A3B) comes from NVIDIA. Inkling-Small was released in July 2026, and Nemotron 3.5 Lightning (30B A3B) followed a month later in August 2026. Inkling-Small is the larger model at roughly 276 billion parameters, against 30 billion for Nemotron 3.5 Lightning (30B A3B).

The two models share 8 published benchmarks. Inkling-Small leads on 8 of them. The widest gaps are on BrowseComp, where Inkling-Small scores 77.4% against 37.0%; Terminal-Bench 2.1, where Inkling-Small scores 64.7% against 24.6%. Averaged across everything we track, Inkling-Small sits at 60.3% and Nemotron 3.5 Lightning (30B A3B) at 45.3%.

Nemotron 3.5 Lightning (30B A3B) is the cheaper API at $0.05 per million input tokens and $0.20 per million output tokens, roughly 6 times cheaper than Inkling-Small at $0.30 and $1.20. Both accept a context window of about 256K tokens. Inkling-Small accepts text, images, and audio as input, while Nemotron 3.5 Lightning (30B A3B) accepts text. On tooling, only Nemotron 3.5 Lightning (30B A3B) supports function calling and only Nemotron 3.5 Lightning (30B A3B) offers structured output.

CharacteristicInkling-SmallNemotron 3.5 Lightning (30B A3B)
CompanyThinking Machines LabNVIDIA
Release DateJuly 30, 2026August 11, 2026
Parameters276B30B
MultimodalYesNo
Context (input)256K262K
Context (output)256K262K
Input Price / 1M$0.30$0.05
Output Price / 1M$1.20$0.20
Average Score60.3%45.3%
Benchmarks
BrowseComp77.4%37.0%
Terminal-Bench 2.164.7%24.6%
SWE-Bench Verified80.2%51.6%
Humanity's Last Exam31.6%11.7%
SciCode48.7%32.6%
GPQA89.5%75.0%
IFBench82.2%72.0%
Tau3 Banking15.5%9.3%

Visual Benchmark Comparison

Inkling-Small
Nemotron 3.5 Lightning (30B A3B)
BrowseComp0.8 vs 0.4
0.8
0.4
Terminal-Bench 2.10.6 vs 0.2
0.6
0.2
SWE-Bench Verified0.8 vs 0.5
0.8
0.5
Humanity's Last Exam0.3 vs 0.1
0.3
0.1
SciCode0.5 vs 0.3
0.5
0.3
GPQA0.9 vs 0.8
0.9
0.8
IFBench0.8 vs 0.7
0.8
0.7
Tau3 Banking0.2 vs 0.1
0.2
0.1

Verdict

Nemotron 3.5 Lightning (30B A3B) leads in 2 out of 5 comparison categories.

Overall Performance

Both models show comparable average scores: Inkling-Small — 0.6, Nemotron 3.5 Lightning (30B A3B) — 0.5.

Programming

On SWE-Bench, both models are nearly equal: Inkling-Small — 0.8, Nemotron 3.5 Lightning (30B A3B) — 0.5.

API Cost

Nemotron 3.5 Lightning (30B A3B) is 6.0x cheaper: input $0.05/1M vs $0.30/1M tokens.

Context Window

Nemotron 3.5 Lightning (30B A3B) supports a larger context: 262K vs 256K tokens.

Recency

Both models were released around the same time: 7/30/2026 and 8/11/2026.

More About These Models

Related Comparisons

Frequently Asked Questions

Which is better for coding — Inkling-Small or Nemotron 3.5 Lightning (30B A3B)?
On the SWE-Bench benchmark, Inkling-Small shows a better result: 80.2% vs 51.6%.
Which model is cheaper — Inkling-Small or Nemotron 3.5 Lightning (30B A3B)?
Nemotron 3.5 Lightning (30B A3B) is cheaper for input: $0.05 per 1M tokens vs $0.30.
Which has a larger context window — Inkling-Small or Nemotron 3.5 Lightning (30B A3B)?
Nemotron 3.5 Lightning (30B A3B) supports a larger context: 262,100 tokens vs 256,000.

The Inkling-Small and Nemotron 3.5 Lightning (30B A3B) comparison is updated for 2026. Data includes benchmark results, API pricing, context window size and other specifications. For more detailed information, visit the Inkling-Small or Nemotron 3.5 Lightning (30B A3B) page. See also the complete list of AI model comparisons.