Title: Enhancing computational efficiency in video action recognition via temporal token merging in ViViT

Authors: Takashi Higashi; Lin Meng; Ryuto Ishibashi

Addresses: Graduate School of Science and Engineering, Ritsumeikan University, 1-1-1 – Noji-higashi, Kusatsu, Shiga – 525-8577, Japan ' College of Science and Engineering, Ritsumeikan University, 1-1-1 – Noji-higashi, Kusatsu, Shiga – 525-8577 Japan ' College of Science and Engineering, Ritsumeikan University, 1-1-1 – Noji-higashi, Kusatsu, Shiga – 525-8577 Japan

Abstract: This study proposes a temporal token merging (TTM) method for ViViT-based action recognition to improve computational efficiency while maintaining accuracy. The method merges temporally similar tokens at the same spatial locations across frames, reducing redundancy without discarding information. Experimental results show that the proposed method achieves up to 39.2% reduction in FLOPs while limiting accuracy degradation to less than 1%. Compared with frame pruning, which reduces FLOPs but significantly degrades accuracy, TTM preserves temporal coherence and achieves a better trade-off between efficiency and performance. Furthermore, comparison with semantic-aware temporal accumulation (STA) demonstrates that TTM maintains higher accuracy despite slightly lower computational reduction. These results indicate that TTM is effective for efficient and reliable video recognition.

Keywords: video recognition; action recognition; ViViT; token merging.

DOI: 10.1504/IJHFMS.2026.153498

International Journal of Human Factors Modelling and Simulation, 2026 Vol.8 No.2, pp.169 - 190

Received: 05 Nov 2025
Accepted: 16 Jan 2026

Published online: 11 May 2026 *

Full-text access for editors Full-text access for subscribers Purchase this article Comment on this article