A vision-language model accepts images (and often video) alongside text, reasoning across both: answering questions about screenshots, reading documents with la
A model that accepts images, and often video, alongside text and reasons across both: answering questions about screenshots, reading documents with layout awareness, describing scenes, and grounding agent actions in what is visually on screen.
They collapsed whole categories of bespoke vision pipelines into general capability, underpinning 2026 staples like document AI, visual inspection copilots, computer-use agents, and multimodal assistants.