UPM Institutional Repository

Self-supervised learning for action recognition: trends, models, and applications


Citation

Al-Qaisieh, Mouwiya S. A. and Mustaffa, Mas Rina (2025) Self-supervised learning for action recognition: trends, models, and applications. Journal of Information Systems Engineering and Management, 10 (48s). pp. 1183-1198. ISSN 2468-4376

Abstract

Recent advances in self-supervised learning (SSL) have reshaped the landscape of human action recognition by reducing dependency on large-scale annotated datasets. This survey provides a comprehensive overview of state-of-the-art SSL techniques developed for understanding human actions in videos. We categorize methods into three primary paradigms: contrastive learning, masked video modeling, and multimodal or sensor-based approaches. Across each category, we discuss key innovations including motion-guided contrastive sampling, transformer-based masked autoencoders, and cross-modal alignment strategies that leverage audio, skeleton, or wearable sensor signals. Models such as VideoMAE, ST-MAE, XDC, and Actionlet-Contrastive represent significant milestones in capturing both spatial and temporal cues without supervision. Beyond model design, we identify major challenges facing current SSL systems, including generalization across domains, modeling long-horizon activities, and real-time deployment constraints. We also highlight underexplored areas such as explainability and unified evaluation protocols. To guide future work, we present a structured taxonomy, a comparative table of representative models, and a discussion of promising research directions including multimodal fusion, modality-agnostic learning, and hardware-aware training. This survey aims to equip researchers with a clear understanding of the evolving trends, persistent gaps, and opportunities that lie ahead in self-supervised action recognition.


Download File

[img] Text
127203.pdf - Published Version
Available under License Creative Commons Attribution.

Download (5MB)

Additional Metadata

Item Type: Article
Subject: Computer Science
Subject: Artificial Intelligence
Subject: Machine Learning
Divisions: Faculty of Computer Science and Information Technology
DOI Number: https://doi.org/10.52783/jisem.v10i48s.9738
Publisher: Science Research Society
Keywords: Component; Self-supervised learning; Human action recognition; Contrastive learning; Multimodal fusion
Sustainable Development Goals (SDGs): SDG 9: Industry, Innovation and Infrastructure, SDG 4: Quality Education, SDG 17: Partnerships for the Goals
Depositing User: Ms. Nur Faseha Mohd Kadim
Date Deposited: 21 Jul 2026 07:37
Last Modified: 21 Jul 2026 07:37
Altmetrics: http://www.altmetric.com/details.php?domain=psasir.upm.edu.my&doi=10.52783/jisem.v10i48s.9738
URI: http://psasir.upm.edu.my/id/eprint/127203
Statistic Details: View Download Statistic

Actions (login required)

View Item View Item