TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition

Abstract

Going beyond few-shot action recognition (FSAR), cross-domain FSAR (CDFSAR) has attracted recent research interests by solving the domain gap lying in source-to-target transfer learning. Existing CDFSAR methods mainly focus on joint training of source and target data to mitigate the side effect of domain gap. However, such kind of methods suffer from two limitations: First, pair-wise joint training requires retraining deep models in case of one source data and multiple target ones, which incurs heavy computation cost, especially for large source and small target data. Second, pre-trained models after joint training are adopted to target domain in a straightforward manner, hardly taking full potential of pre-trained models and then limiting recognition performance. To overcome above limitations, this paper proposes a simple yet effective baseline, namely Temporal-Aware Model Tuning (TAMT) for CDFSAR. Specifically, our TAMT involves a decoupled paradigm by performing pre-training on source data and fine-tuning target data, which avoids retraining for multiple target data with single source. To effectively and efficiently explore the potential of pre-trained models in transferring to target domain, our TAMT proposes a Hierarchical Temporal Tuning Network (HTTN), whose core involves local temporal-aware adapters (TAA) and a global temporal-aware moment tuning (GTMT). Particularly, TAA learns few parameters to recalibrate the intermediate features of frozen pre-trained models, enabling efficient adaptation to target domains. Furthermore, GTMT helps to generate powerful video representations, improving match performance on the target domain. Experiments on several widely used video benchmarks show our TAMT outperforms the recently proposed counterparts by 13%~31%, achieving new state-of-the-art CDFSAR results.

Fig.2. (a) Overview of our TAMT paradigm, which pre-trains the models on source data and fine-tunes them on target data. Specifically, for pre-training stage, the model is first optimized with a reconstruction-based SSL solution, while the encoder $\mathcal{E}$ is post-trained with the SL objective. Subsequently, the pre-trained $\mathcal{E}$ is fine-tuned for few-shot adaptation on $\mathcal{T}_{CD}$ by using our HTTN. (b) HTTN for few-shot adaptation, where a metric-based is used for few-shot adaptation. Particularly, our HTTN consists of local Temporal-Aware Adapters (TAA) and Global Temporal-aware Moment Tuning (GTMT).

Fig.3. Overview of our proposed Hierarchical Temporal Tuning Network (HTTN), where (a) local temporal-aware adapters (TAA) are inserted into the last $L$ transformer blocks to recalibrate the intermediate features of frozen pre-training models in an efficient manner. At the end of HTTN, a Global Temporal-aware Moment Tuning (GTMT) module with efficient long-short temporal covariance (ELSTC) is used to obtain powerful video representations for improving matching performance.

TODO

Release the code.
Release the models.
Release the arxiv preprint.

Citation

If our work is helpful to you, please consider citing us by using the following BibTeX entry:

Pre-training on Source Data

Fine-tuning on Target Data

1.Requirements

Code is tested under Pytorch 1.9.1, python 3.6.10, and CUDA 11.3. Mainly libraries:

Or see the requirements_all.txt for detailed libraries.

2.Dataset

Put the data the same with your filelist:

hmdb51_org
├── brush_hair
└── cartwheel

3.Train and Test

Run following commands to start training or testing:

cd scripts/hmdb51/run_meta_deepbdc
sh run_test.sh    # For test only.

sh run_metatrain.sh    # For train and test, for individual training or testing, please comment out parts of the code yourself.

Pre-trained Model

The following table shows the Pre-trained Model on K-400(364 classes) with 112 × 112 resolution.

Pre-trained Model	Checkpoint
vit_s	Download

Finetuned Model

The following table shows the results of TAMT on CDFSAR setting in terms of 5-way 5-shot accuracy.

Dataset	5-way 5-shot Acc(%)	Checkpoint
HMDB	74.14	Download
SSV2	59.18	Download
Diving	45.18	Download
UCF	95.92	Download
RareAct	67.44	Download

Name		Name	Last commit message	Last commit date
Latest commit History 24 Commits
__pycache__		__pycache__
data		data
filelist		filelist
methods		methods
network		network
pic		pic
scripts		scripts
LICENSE.txt		LICENSE.txt
README.md		README.md
assets		assets
meta_train.py		meta_train.py
pretrain.py		pretrain.py
requirements_all.txt		requirements_all.txt
test.py		test.py
test500.py		test500.py
utils.py		utils.py
weight_loaders.py		weight_loaders.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition

Abstract

TODO

Citation

Pre-training on Source Data

Fine-tuning on Target Data

1.Requirements

2.Dataset

3.Train and Test

Pre-trained Model

Finetuned Model

About

Releases

Packages

Languages

License

TJU-YDragonW/TAMT

Folders and files

Latest commit

History

Repository files navigation

TAMT: Temporal-Aware Model Tuning for Cross-Domain Few-Shot Action Recognition

Abstract

TODO

Citation

Pre-training on Source Data

Fine-tuning on Target Data

1.Requirements

2.Dataset

3.Train and Test

Pre-trained Model

Finetuned Model

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages