Title: Understanding vision language models: a deep dive
Authors: Won Kim
Addresses: Gachon University, Seongnam, South Korea
Abstract: Vision language models (VLMs) are multimodal, generative AI models that can understand and process modalities beyond text, including image, video and audio. Over the past several years, VLMs have remarkably advanced and are continuing to advance at a rapid pace. A large number of VLMs have been released, with a wide range of capabilities and architectures. The literature is vast with academic papers, technical blog posts, YouTube videos and news articles. The objective of this article is to provide a comprehensive guide to understanding essential aspects of VLMs, including the capabilities, use cases, architectures, training methods, benchmarks, challenges and R&D directions.
Keywords: vision language model; VLM; VLM architectures; VLM training; VLM learning objectives; VLM benchmarks; VLM challenges.
DOI: 10.1504/IJWGS.2026.154469
International Journal of Web and Grid Services, 2026 Vol.22 No.2, pp.225 - 269
Received: 15 Dec 2025
Accepted: 15 Feb 2026
Published online: 29 Jun 2026 *