North Micro Vision Instruct
MultimodalCompact open-weight vision-language model (2.4B) with native-resolution image support for VQA, captioning, grounding, OCR, charts, and documents. Custom 400M SigLIP 2-based vision encoder + 2B North Micro LLM backbone (Command A+ style). Multilingual and multi-image. LM context 128K; multimodal training validated to 8K. Not a reasoning/tool-calling model. Intended for prototyping and fine-tuning.
Key Specifications
Parameters
2.4B
Context
-
Release Date
August 12, 2026
Average Score
61.1%
Timeline
Key dates in the model's history
Announcement
August 12, 2026
Last Update
September 12, 2026
Today
September 20, 2026
Technical Specifications
Parameters
2.4B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval
Benchmark Results
Model performance metrics across various tests and benchmarks
General Knowledge
Tests on general knowledge and understanding
MMLU
HF model card table (MMLU test via VLMEvalKit) • Self-reported
Multimodal
Working with images and visual data
AI2D
HF model card table (AI2D_TEST via VLMEvalKit) • Self-reported
ChartQA
HF model card table (ChartQA Test via VLMEvalKit) • Self-reported
DocVQA
HF model card table (DocVQA VAL via VLMEvalKit) • Self-reported
Other Tests
Specialized benchmarks
BLINK
HF model card table (BLINK via VLMEvalKit) • Self-reported
CharXiv-D
HF model card table (CharXiv DQ via VLMEvalKit) • Self-reported
CountBench
HF model card table (CountBench via VLMEvalKit) • Self-reported
Hallusion Bench
HF model card table (HallusionBench via VLMEvalKit) • Self-reported
IFEval
HF model card table (IFEval via VLMEvalKit) • Self-reported
InfoVQA
HF model card table (InfoVQA VAL via VLMEvalKit) • Self-reported
MMBench-V1.1
HF model card table (MMBenchDEV_EN_V11 via VLMEvalKit) • Self-reported
MMLU-Pro
HF model card table (MMLU-Pro test via VLMEvalKit) • Self-reported
MMMU (val)
HF model card table (MMMU DEV_VAL via VLMEvalKit) • Self-reported
MMStar
HF model card table (MMStar via VLMEvalKit) • Self-reported
Multi-IF
HF model card table (Multi-If via VLMEvalKit) • Self-reported
OCRBench
HF model card table (OCRBench via VLMEvalKit) • Self-reported
OCRBench-V2 (en)
HF model card table (OCRBench v2_en via VLMEvalKit) • Self-reported
RealWorldQA
HF model card table (RealWorldQA via VLMEvalKit) • Self-reported
RefCOCO-avg
HF model card table (RefCOCO avg P@1 over 8 splits via VLMEvalKit) • Self-reported
License & Metadata
License
apache_2_0
Announcement Date
August 12, 2026
Last Updated
September 12, 2026
Compare North Micro Vision Instruct
All comparisonsSimilar Models
All ModelsGemma 3n E2B
MM8.0B
Best score:0.5 (ARC)
Released:Jun 2025
Gemma 3 4B
MM4.0B
Best score:0.7 (HumanEval)
Released:Mar 2025
Price:$0.02/1M tokens
Gemma 3n E2B Instructed LiteRT (Preview)
MM1.9B
Best score:0.7 (HumanEval)
Released:May 2025
Gemma 3n E4B Instructed
MM8.0B
Best score:0.8 (HumanEval)
Released:Jun 2025
Price:$20.00/1M tokens
Gemma 3n E2B Instructed
MM8.0B
Best score:0.7 (HumanEval)
Released:Jun 2025
Qwen2.5-Omni-7B
Alibaba
MM7.0B
Best score:0.8 (HumanEval)
Released:Mar 2025
Qwen3.5-2B
Alibaba
MM2.0B
Best score:0.5 (GPQA)
Released:Mar 2026
Command A+
Cohere
MM218.0B
Best score:0.8 (TAU)
Released:May 2026
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.