Title: Understanding vision language models: a deep dive

Authors: Won Kim

Addresses: Gachon University, Seongnam, South Korea

Abstract: Vision language models (VLMs) are multimodal, generative AI models that can understand and process modalities beyond text, including image, video and audio. Over the past several years, VLMs have remarkably advanced and are continuing to advance at a rapid pace. A large number of VLMs have been released, with a wide range of capabilities and architectures. The literature is vast with academic papers, technical blog posts, YouTube videos and news articles. The objective of this article is to provide a comprehensive guide to understanding essential aspects of VLMs, including the capabilities, use cases, architectures, training methods, benchmarks, challenges and R&D directions.

Keywords: vision language model; VLM; VLM architectures; VLM training; VLM learning objectives; VLM benchmarks; VLM challenges.

DOI: 10.1504/IJWGS.2026.154469

International Journal of Web and Grid Services, 2026 Vol.22 No.2, pp.225 - 269

Received: 15 Dec 2025
Accepted: 15 Feb 2026

Published online: 29 Jun 2026 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article