LFM2.5-VL-3B
MultimodalLFM2.5-VL-3B is Liquid AI's open-weight vision-language model (~3.12B parameters; HF safetensors total 3,123,483,888) for on-device and edge image-text→text. Non-reasoning: answers directly for low latency. Builds on the LFM2.5-2.6B language backbone with a SigLIP2 400M NaFlex vision encoder; improves screen/UI understanding, grounding, function calling, and multi-image input over LFM2-VL-3B. Pre-trained on ~34T tokens; 128K vocabulary; 32,768-token context. Day-one support for llama.cpp, MLX, vLLM, SGLang, and ONNX. License: LFM Open License v1.0 (lfm1.0).
Key Specifications
Timeline
Technical Specifications
Benchmark Results
Model performance metrics across various tests and benchmarks
Multimodal
Other Tests
License & Metadata
Compare LFM2.5-VL-3B
All comparisonsSimilar Models
All ModelsLFM2.5-2.6B
Liquid AI
Granite 3.3 8B Base
IBM
North Micro Vision Instruct
Cohere
Gemini 1.5 Flash 8B
Gemma 3n E2B
MedGemma 4B IT
Gemma 3 4B
Gemma 3n E2B Instructed LiteRT (Preview)
Recommendations are based on similarity of characteristics: developer organization, multimodality, parameter size, and benchmark performance. Choose a model to compare or go to the full catalog to browse all available AI models.