LFM2.5-VL-3B for Better and Faster Vision Capabilities for the Edge
Covered by 2 sources
Read full postLFM2.5-VL-3B is a new vision-language model designed for on-device use, offering enhanced screen understanding, object grounding, multi-image reasoning, and function calling. It combines a SigLIP2 vision encoder with a large text model backbone, trained on extensive multimodal data and fine-tuned with reinforcement learning. Benchmark tests show it leads its size class in real-world image tasks and digital content comprehension.




