research-article

Progressive Cross-modal Knowledge Distillation for Human Action Recognition

Authors:

Jianyuan Ni,

Anne H.H. Ngu,

Yan YanAuthors Info & Claims

MM '22: Proceedings of the 30th ACM International Conference on Multimedia

Pages 5903 - 5912

https://doi.org/10.1145/3503161.3548238

Published: 10 October 2022 Publication History

Get Access

Abstract

Wearable sensor-based Human Action Recognition (HAR) has achieved remarkable success recently. However, the accuracy performance of wearable sensor-based HAR is still far behind the ones from the visual modalities-based system (i.e., RGB video, skeleton and depth). Diverse input modalities can provide complementary cues and thus improve the accuracy performance of HAR, but how to take advantage of multi-modal data on wearable sensor-based HAR has rarely been explored. Currently, wearable devices, i.e., smartwatches, can only capture limited kinds of non-visual modality data. This hinders the multi-modal HAR association as it is unable to simultaneously use both visual and non-visual modality data. Another major challenge lies in how to efficiently utilize multi-modal data on wearable devices with their limited computation resources. In this work, we propose a novel Progressive Skeleton-to-sensor Knowledge Distillation (PSKD) model which utilizes only time-series data, i.e., accelerometer data, from a smartwatch for solving the wearable sensor-based HAR problem. Specifically, we construct multiple teacher models using data from both teacher (human skeleton sequence) and student (time-series accelerometer data) modalities. In addition, we propose an effective progressive learning scheme to eliminate the performance gap between teacher and student models. We also designed a novel loss function called Adaptive-Confidence Semantic (ACS), to allow the student model to adaptively select either one of the teacher models or the ground-truth label it needs to mimic. To demonstrate the effectiveness of our proposed PSKD method, we conduct extensive experiments on Berkeley-MHAD, UTD-MHAD and MMAct datasets. The results confirm that the proposed PSKD method has competitive performance compared to the previous mono sensor-based HAR methods.

Supplementary Material

MP4 File (MM22-fp2032.mp4)

Short presentation for the paper "Progressive Cross-modal Knowledge Distillation for Human Action Recognition"

Download
17.79 MB

References

[1]

Zeeshan Ahmad and Naimul Khan. 2020. CNN-based multistage gated average fusion (MGAF) for human action recognition using depth and inertial sensors. IEEE Sensors Journal, Vol. 21, 3 (2020), 3623--3634.

Abstract

Supplementary Material

References

Cited By

Index Terms

Recommendations

Recognizing multi-user activities using wearable sensors in a smart home

LightHART: Lightweight Human Activity Recognition Transformer

Codebook approach for sensor-based human activity recognition

Comments

Information

Published In

Sponsors

Publisher

Publication History

Permissions

Check for updates

Author Tags

Qualifiers

Funding Sources

Conference

Acceptance Rates

Contributors

Other Metrics

Bibliometrics

Article Metrics

Other Metrics

Citations

Cited By

Login options

Full Access

View options

PDF

eReader

Share

Share this Publication link

Share on social media

Affiliations