Title: Enhancing computational efficiency in video action recognition via temporal token merging in ViViT
Authors: Takashi Higashi; Lin Meng; Ryuto Ishibashi
Addresses: Graduate School of Science and Engineering, Ritsumeikan University, 1-1-1 – Noji-higashi, Kusatsu, Shiga – 525-8577, Japan ' College of Science and Engineering, Ritsumeikan University, 1-1-1 – Noji-higashi, Kusatsu, Shiga – 525-8577 Japan ' College of Science and Engineering, Ritsumeikan University, 1-1-1 – Noji-higashi, Kusatsu, Shiga – 525-8577 Japan
Abstract: This study proposes a temporal token merging (TTM) method for ViViT-based action recognition to improve computational efficiency while maintaining accuracy. The method merges temporally similar tokens at the same spatial locations across frames, reducing redundancy without discarding information. Experimental results show that the proposed method achieves up to 39.2% reduction in FLOPs while limiting accuracy degradation to less than 1%. Compared with frame pruning, which reduces FLOPs but significantly degrades accuracy, TTM preserves temporal coherence and achieves a better trade-off between efficiency and performance. Furthermore, comparison with semantic-aware temporal accumulation (STA) demonstrates that TTM maintains higher accuracy despite slightly lower computational reduction. These results indicate that TTM is effective for efficient and reliable video recognition.
Keywords: video recognition; action recognition; ViViT; token merging.
DOI: 10.1504/IJHFMS.2026.153498
International Journal of Human Factors Modelling and Simulation, 2026 Vol.8 No.2, pp.169 - 190
Received: 05 Nov 2025
Accepted: 16 Jan 2026
Published online: 11 May 2026 *