Molmo2-8B is an open vision-language model developed by the Allen Institute for AI (Ai2) as part of the Molmo2 family, supporting image, video, and multi-image understanding and grounding. It is based on Qwen3-8B and uses SigLIP 2 as its vision backbone, outperforming other open-weight, open-data models on short videos, counting, and captioning, while remaining competitive on long-video tasks.
Modalities
Context
37K
Released
Jan 9, 2026
Molmo2-8B is an open vision-language model developed by the Allen Institute for AI (Ai2) as part of the Molmo2 family, supporting image, video, and multi-image understanding and grounding. It is based on Qwen3-8B and uses SigLIP 2 as its vision backbone, outperforming other open-weight, open-data models on short videos, counting, and captioning, while remaining competitive on long-video tasks.
Molmo2 8B has a 36,864 token context window.
Molmo2 8B accepts text, images and video as input and returns text.
Olmo 3 32B Think is another text model from Ai2.
Molmo2 8B was released on January 9, 2026.
Token volume and request traffic to this model over time.