Inkling-Small vs Qwen3.8 Max: Specs & Benchmark Comparison

Inkling-Small is developed by Thinking Machines Lab, while Qwen3.8 Max comes from Alibaba. Inkling-Small was released in July 2026, and Qwen3.8 Max followed a month later in August 2026. Qwen3.8 Max is the larger model at roughly 2.4 trillion parameters, against 276 billion for Inkling-Small.

We have 24 benchmark results for Inkling-Small and 2 benchmark results for Qwen3.8 Max, but they were measured on different benchmarks, so there is no like-for-like scoreboard.

Inkling-Small is the cheaper API at $0.30 per million input tokens and $1.20 per million output tokens, roughly 8 times cheaper than Qwen3.8 Max at $2.50 and $6.25. Qwen3.8 Max takes the larger context window at 1M tokens, compared with 256K for Inkling-Small. Inkling-Small accepts text, images, and audio as input, while Qwen3.8 Max accepts text and images. On tooling, only Qwen3.8 Max supports function calling and only Qwen3.8 Max offers structured output.

CharacteristicInkling-SmallQwen3.8 Max
CompanyThinking Machines LabAlibaba
Release DateJuly 30, 2026August 2, 2026
Parameters276B2.4T
MultimodalYesYes
Context (input)256K1.0M
Context (output)256K131K
Input Price / 1M$0.30$2.50
Output Price / 1M$1.20$6.25
Average Score60.3%93.0%

Verdict

Both models show equal results — the choice depends on your specific use case.

Overall Performance

Both models show comparable average scores: Inkling-Small — 0.6, Qwen3.8 Max — 0.9.

API Cost

Inkling-Small is 5.8x cheaper: input $0.30/1M vs $2.50/1M tokens.

Context Window

Qwen3.8 Max supports a larger context: 1M vs 256K tokens.

Recency

Both models were released around the same time: 7/30/2026 and 8/2/2026.

More About These Models

Related Comparisons

Frequently Asked Questions

Which is better for coding — Inkling-Small or Qwen3.8 Max?
Direct comparison on the SWE-Bench benchmark is not available. We recommend reviewing other metrics on the comparison page.
Which model is cheaper — Inkling-Small or Qwen3.8 Max?
Inkling-Small is cheaper for input: $0.30 per 1M tokens vs $2.50.
Which has a larger context window — Inkling-Small or Qwen3.8 Max?
Qwen3.8 Max supports a larger context: 1,000,000 tokens vs 256,000.

The Inkling-Small and Qwen3.8 Max comparison is updated for 2026. Data includes benchmark results, API pricing, context window size and other specifications. For more detailed information, visit the Inkling-Small or Qwen3.8 Max page. See also the complete list of AI model comparisons.