Title: Structured radiology report generation using ViT-B/16 and clinical-T5: a multimodal approach on IU X-Ray dataset

Authors: Nilam Sureshrao Khairnar; Shirish S. Sane

Addresses: MET Bhujbal Knowledge City Institute of Engineering, Adgaon, Nashik, India ' MET Bhujbal Knowledge City Institute of Engineering, Adgaon, Nashik, India; R.H. Sapat COE, Nashik, India

Abstract: This work presents a multimodal framework for generating structured radiology reports directly from the chest X-ray images. The model combines the vision transformer (ViT-B/16) to extract rich visual representations with Clinical-T5, which is a domain-tuned language model that interprets the clinical text using a co-attention module. The visual and textual streams interact to allow the system to link features of images with clinical descriptions that match the image. The decoder then generates well-structured reports containing indication, findings, and impression sections. The framework was trained and evaluated on the IU X-Ray dataset and was found to have significant advances in the BLEU and ROUGE-L metrics, and strong CheXbert F1 results compared with previous methods. The results show that a combination of transformer-based vision and language models can generate coherent, interpretable, and clinically reliable radiology reports, highlighting the importance of multimodal learning for automated radiology report generation.

Keywords: BLEU; clinical T5; NLP; radiology report generation; ROUGE; vision transformer; ViT-B/16.

DOI: 10.1504/IJBET.2026.155801

International Journal of Biomedical Engineering and Technology, 2026 Vol.51 No.1, pp.1 - 17

Received: 08 Aug 2025
Accepted: 28 Dec 2025

Published online: 14 Aug 2026 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article