Multimodal Models · Shanghai AI Laboratory
InternVideo3: An 8B Video Agent with Contextual Reasoning
InternVideo3 is an 8B video model that scores 73.8 on Video-MME single-pass, then adds +2.7 from an agentic reasoning loop on top. M2LA attention holds 768K tokens on one H200 where the base OOMs at 512K.