Cohere logo

North Micro Vision Instruct

Multimodal
Cohere

Compact open-weight vision-language model (2.4B) with native-resolution image support for VQA, captioning, grounding, OCR, charts, and documents. Custom 400M SigLIP 2-based vision encoder + 2B North Micro LLM backbone (Command A+ style). Multilingual and multi-image. LM context 128K; multimodal training validated to 8K. Not a reasoning/tool-calling model. Intended for prototyping and fine-tuning.

Key Specifications

Parameters
2.4B
Context
-
Release Date
August 12, 2026
Average Score
61.1%

Timeline

Key dates in the model's history
Announcement
August 12, 2026
Last Update
September 12, 2026
Today
September 20, 2026

Technical Specifications

Parameters
2.4B
Training Tokens
-
Knowledge Cutoff
-
Family
-
Capabilities
MultimodalZeroEval

Benchmark Results

Model performance metrics across various tests and benchmarks

General Knowledge

Tests on general knowledge and understanding
MMLU
HF model card table (MMLU test via VLMEvalKit)Self-reported
50.4%

Multimodal

Working with images and visual data
AI2D
HF model card table (AI2D_TEST via VLMEvalKit)Self-reported
77.5%
ChartQA
HF model card table (ChartQA Test via VLMEvalKit)Self-reported
80.8%
DocVQA
HF model card table (DocVQA VAL via VLMEvalKit)Self-reported
92.1%

Other Tests

Specialized benchmarks
BLINK
HF model card table (BLINK via VLMEvalKit)Self-reported
52.7%
CharXiv-D
HF model card table (CharXiv DQ via VLMEvalKit)Self-reported
60.0%
CountBench
HF model card table (CountBench via VLMEvalKit)Self-reported
72.5%
Hallusion Bench
HF model card table (HallusionBench via VLMEvalKit)Self-reported
61.5%
IFEval
HF model card table (IFEval via VLMEvalKit)Self-reported
74.9%
InfoVQA
HF model card table (InfoVQA VAL via VLMEvalKit)Self-reported
65.2%
MMBench-V1.1
HF model card table (MMBenchDEV_EN_V11 via VLMEvalKit)Self-reported
68.7%
MMLU-Pro
HF model card table (MMLU-Pro test via VLMEvalKit)Self-reported
30.7%
MMMU (val)
HF model card table (MMMU DEV_VAL via VLMEvalKit)Self-reported
32.9%
MMStar
HF model card table (MMStar via VLMEvalKit)Self-reported
51.8%
Multi-IF
HF model card table (Multi-If via VLMEvalKit)Self-reported
37.3%
OCRBench
HF model card table (OCRBench via VLMEvalKit)Self-reported
79.2%
OCRBench-V2 (en)
HF model card table (OCRBench v2_en via VLMEvalKit)Self-reported
36.7%
RealWorldQA
HF model card table (RealWorldQA via VLMEvalKit)Self-reported
62.2%
RefCOCO-avg
HF model card table (RefCOCO avg P@1 over 8 splits via VLMEvalKit)Self-reported
73.2%

License & Metadata

License
apache_2_0
Announcement Date
August 12, 2026
Last Updated
September 12, 2026

Compare North Micro Vision Instruct

All comparisons

Similar Models

All Models

Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.