Title: Lightweighting and end-side deployment of multi-modal large model based on cross-modal attention distillation
Authors: Kaijie Liu; Yun Dong; Zhipeng Meng; Qi Chen; Qiwen Tan; Zonghui Wei; Qi Meng
Addresses: Digital & Intelligent Operation Center, Guangxi Power Grid Co., Ltd., Nanning Guangxi, China ' Digital & Intelligent Operation Center, Guangxi Power Grid Co., Ltd., Nanning Guangxi, China ' Digital & Intelligent Operation Center, Guangxi Power Grid Co., Ltd., Nanning Guangxi, China ' Digital & Intelligent Operation Center, Guangxi Power Grid Co., Ltd., Nanning Guangxi, China ' Digital Department, Guangxi Power Grid Co., Ltd., Nanning Guangxi, China ' Digital Department, Guangxi Power Grid Co., Ltd., Nanning Guangxi, China ' Digital & Intelligent Operation Center, Guangxi Power Grid Co., Ltd., Nanning Guangxi, China
Abstract: To address the issue of high cost incurred by multimodal large models in visual-language tasks, this paper proposes a lightweight model, CAD-LM, based on cross-modal attention distillation. It designs an adaptive modal reparameterisation module that utilises a multi-branch structure to enhance representational power during training and reparameterises it into an efficient single-branch structure during inference. This paper integrates an end-to-end deployment optimisation process that encompasses hardware-aware pruning, mixed-precision quantisation, and efficient parameter fine-tuning. Experiments show that CAD-LM successfully compresses the model parameter count to 118.4 M and reduces computational complexity to 12.1 GFLOPs, achieving approximately 20% and 30% reduction, respectively. Its performance on benchmark tasks such as Flickr30k and VQA v2.0 significantly surpasses baseline models like the original-scale CLIP and ALBEF. Edge deployment verification reveals that the final model occupies only 89.4 MB of memory and boasts millisecond-level inference latency, achieving an excellent balance between model performance, computational efficiency, and engineering practicality. This provides an efficient solution for multi-modal applications in resource-constrained environments.
Keywords: attention distillation; multi-modal; large model; lightweight; end-side deployment.
DOI: 10.1504/IJICT.2026.153990
International Journal of Information and Communication Technology, 2026 Vol.27 No.62, pp.92 - 121
Received: 16 Jan 2026
Accepted: 11 Mar 2026
Published online: 09 Jun 2026 *


