More Web Proxy on the site http://driver.im/

Article

SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking

Authors:

Luigi Piccinelli,

Martin Danelljan,

Luc Van GoolAuthors Info & Claims

Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part XXVII

Pages 1 - 18

https://doi.org/10.1007/978-3-031-73383-3_1

Published: 03 November 2024 Publication History

Abstract

Open-vocabulary Multiple Object Tracking (MOT) aims to generalize trackers to novel categories not in the training set. Currently, the best-performing methods are mainly based on pure appearance matching. Due to the complexity of motion patterns in the large-vocabulary scenarios and unstable classification of the novel objects, the motion and semantics cues are either ignored or applied based on heuristics in the final matching steps by existing methods. In this paper, we present a unified framework SLAck that jointly considers semantics location, and appearance priors in the early steps of association and learns how to integrate all valuable information through a lightweight spatial and temporal object graph. Our method eliminates complex post-processing heuristics for fusing different cues and boosts the association performance significantly for large-scale open-vocabulary tracking. Without bells and whistles, we outperform previous state-of-the-art methods for novel classes tracking on the open-vocabulary MOT and TAO TETA benchmarks. Our code is available at github.com/siyuanliii/SLAck.

References

[1]

Bergmann, P., Meinhardt, T., Leal-Taixe, L.: Tracking without bells and whistles. In: ICCV (2019)

[2]

Bewley, A., Ge, Z., Ott, L., Ramos, F., Upcroft, B.: Simple online and realtime tracking. In: ICIP (2016)

[3]

Brasó, G., Leal-Taixé, L.: Learning a neural solver for multiple object tracking. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6247–6257 (2020)

[4]

Caesar, H., et al.: nuScenes: a multimodal dataset for autonomous driving. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11621–11631 (2020)

[5]

Cao, J., Pang, J., Weng, X., Khirodkar, R., Kitani, K.: Observation-centric sort: rethinking sort for robust multi-object tracking. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9686–9696 (2023)

[6]

Cetintas, O., Brasó, G., Leal-Taixé, L.: Unifying short and long-term tracking with graph hierarchies. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 22877–22887, June 2023

[7]

Dave A, Khurana T, Tokmakov P, Schmid C, and Ramanan D Vedaldi A, Bischof H, Brox T, and Frahm J-M TAO: a large-scale benchmark for tracking any object Computer Vision – ECCV 2020 2020 Cham Springer 436-454

Digital Library

[8]

Dendorfer, P., et al.: Mot20: a benchmark for multi object tracking in crowded scenes. arXiv preprint arXiv:2003.09003 (2020)

[9]

Du, F., Xu, B., Tang, J., Zhang, Y., Wang, F., Li, H.: 1st place solution to ECCV-TAO-2020: detect and represent any object for tracking. arXiv preprint arXiv:2101.08040 (2021)

[10]

Du, Y., Wei, F., Zhang, Z., Shi, M., Gao, Y., Li, G.: Learning to prompt for open-vocabulary object detection with vision-language model. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 14084–14093 (2022)

[11]

Du Y et al. StrongSORT: make deepSORT great again IEEE Trans. Multimedia 2023 25 8725-8737

Digital Library

[12]

Geiger A, Lenz P, Stiller C, and Urtasun R Vision meets robotics: the kitti dataset Int. J. Robot. Res. 2013 32 11 1231-1237

Digital Library

[13]

Gu, X., Lin, T.Y., Kuo, W., Cui, Y.: Open-vocabulary object detection via vision and language knowledge distillation. arXiv preprint arXiv:2104.13921 (2021)

[14]

Gupta, A., Dollar, P., Girshick, R.: LVIS: a dataset for large vocabulary instance segmentation. In: CVPR (2019)

[15]

He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)

[16]

Hu, H., Gu, J., Zhang, Z., Dai, J., Wei, Y.: Relation networks for object detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3588–3597 (2018)

[17]

Kim, V., Jung, G., Lee, S.W.: AM-SORT: adaptable motion predictor with historical trajectory embedding for multi-object tracking. arXiv preprint arXiv:2401.13950 (2024)

[18]

Li S, Danelljan M, Ding H, Huang TE, and Yu F Avidan S, Brostow G, Cissé M, Farinella GM, and Hassner T Tracking every thing in the wild ECCV 2022 2022 Cham Springer 498-515

Digital Library

[19]

Li, S., Fischer, T., Ke, L., Ding, H., Danelljan, M., Yu, F.: OVTrack: open-vocabulary multiple object tracking. In: CVPR (2023)

[20]

Li, S., et al.: Matching anything by segmenting anything. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18963–18973 (2024)

[21]

Liu, S., et al.: Grounding DINO: marrying DINO with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499 (2023)

[22]

Liu, Y., et al.: Opening up open world tracking. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 19045–19055 (2022)

[23]

Liu, Z., et al.: Swin transformer: hierarchical vision transformer using shifted windows. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10012–10022 (2021)

[24]

Liu Z, Segu M, and Yu F Köthe U and Rother C COOLer: class-incremental learning for appearance-based multiple object tracking DAGM GCPR 2023 2023 Cham Springer 443-458

Digital Library

[25]

Meinhardt, T., Kirillov, A., Leal-Taixe, L., Feichtenhofer, C.: TrackFormer: multi-object tracking with transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8844–8854 (2022)

[26]

Milan, A., Leal-Taixé, L., Reid, I., Roth, S., Schindler, K.: MOT16: a benchmark for multi-object tracking. arXiv preprint arXiv:1603.00831 (2016)

[27]

Pang, J., et al.: Quasi-dense similarity learning for multiple object tracking. In: CVPR (2021)

[28]

Radford, A., et al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748–8763. PMLR (2021)

[29]

Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695 (2022)

[30]

Sarlin, P.E., DeTone, D., Malisiewicz, T., Rabinovich, A.: SuperGlue: learning feature matching with graph neural networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4938–4947 (2020)

[31]

Segu, M., Piccinelli, L., Li, S., Van Gool, L., Yu, F., Schiele, B.: Walker: self-supervised multiple object tracking by walking on temporal appearance graphs. In: Computer Vision–ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings. Springer (2024)

[32]

Segu, M., Piccinelli, L., Li, S., Yang, Y.H., Schiele, B., Van Gool, L.: Samba: synchronized set-of-sequences modeling for end-to-end multiple object tracking. arXiv preprint (2024)

[33]

Segu, M., Schiele, B., Yu, F.: Darth: holistic test-time adaptation for multiple object tracking. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9717–9727 (2023)

[34]

Sun, P., et al.: Scalability in perception for autonomous driving: Waymo open dataset. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2446–2454 (2020)

[35]

Sun, P., et al.: TransTrack: multiple object tracking with transformer. arXiv preprint arXiv:2012.15460 (2020)

[36]

Wang Z, Zheng L, Liu Y, and Wang S Vedaldi A, Bischof H, Brox T, and Frahm JM Towards real-time multi-object tracking ECCV 2020 2020 Cham Springer 107-122

Digital Library

[37]

Wojke, N., Bewley, A., Paulus, D.: Simple online and realtime tracking with a deep association metric. In: ICIP (2017)

[38]

Wu, J., Jiang, Y., Liu, Q., Yuan, Z., Bai, X., Bai, S.: General object foundation model for images and videos at scale. arXiv preprint arXiv:2312.09158 (2023)

[39]

Wu J, Liu Q, Jiang Y, Bai S, Yuille A, and Bai X Avidan S, Brostow G, Cissé M, Farinella GM, and Hassner T In defense of online models for video instance segmentation ECCV 2022 2022 Cham Springer 588-605

Digital Library

[40]

Yan, B., et al.: Universal instance perception as object discovery and retrieval. In: CVPR (2023)

[41]

Ye, M., et al.: Cascade-DETR: delving into high-quality universal object detection. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6704–6714 (2023)

[42]

Zeng F, Dong B, Zhang Y, Wang T, Zhang X, and Wei Y Avidan S, Brostow G, Cissé M, Farinella GM, and Hassner T MOTR: end-to-end multiple-object tracking with transformer ECCV 2022 2022 Cham Springer 659-675

Digital Library

[43]

Zhang Y et al. Avidan S, Brostow G, Cissé M, Farinella GM, Hassner T, et al. ByteTrack: multi-object tracking by associating every detection box ECCV 2022 2022 Cham Springer 1-21

Digital Library

[44]

Zhang, Y., Wang, C., Wang, X., Zeng, W., Liu, W.: FairMOT: On the fairness of detection and re-identification in multiple object tracking. IJCV (2021)

[45]

Zheng, G., Lin, S., Zuo, H., Fu, C., Pan, J.: NetTrack: tracking highly dynamic objects with a net. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 19145–19155 (2024)

[46]

Zhou X, Girdhar R, Joulin A, Krähenbühl P, and Misra I Avidan S, Brostow G, Cissé M, Farinella GM, and Hassner T Detecting twenty-thousand classes using image-level supervision ECCV 2022 2022 Cham Springer 350-368

Digital Library

Index Terms

SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking
1. Computing methodologies
  1. Artificial intelligence
    1. Computer vision

Index terms have been assigned to the content through auto-classification.

Recommendations

Online multi-object tracking by detection based on generative appearance models

A multi-object tracking approach based on multiple cues is proposed.A hierarchical data association can improve the multi-object tracking performance by handling occlusion problems between targets.A defined state for each target handles varying number ...
Heterogeneous Fusion of Omnidirectional and PTZ Cameras for Multiple Object Tracking

Dual-camera systems have been widely used in surveillance because of the ability to explore the wide field of view (FOV) of the omnidirectional camera and the wide zoom range of the PTZ camera. Most existing algorithms require a priori knowledge of the ...
Real-time object tracking using bounded irregular pyramids

Target representation and localization is a central component in visual object tracking. In this paper a new approach for target representation and localization is presented. This approach tackles two of the most important causes of failure in object ...

Comments

Please enable JavaScript to view thecomments powered by Disqus.

Information & Contributors

Information

Published In

cover image Guide Proceedings

Computer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part XXVII

Sep 2024

568 pages

ISBN:978-3-031-73382-6

DOI:10.1007/978-3-031-73383-3

Editors:
Aleš Leonardis
https://ror.org/03angcq70University of Birmingham, Birmingham, UK
,
Elisa Ricci
https://ror.org/05trd4x28University of Trento, Trento, Italy
,
Stefan Roth
https://ror.org/05n911h24Technical University of Darmstadt, Darmstadt, Germany
,
Olga Russakovsky
https://ror.org/00hx57361Princeton University, Princeton, NJ, USA
,
Torsten Sattler
https://ror.org/03kqpb082Czech Technical University in Prague, Prague, Czech Republic
,
Gül Varol
https://ror.org/02nwvxz07École des Ponts ParisTech, Marne-la-Vallée, France

© The Author(s), under exclusive license to Springer Nature Switzerland AG 2025.

Publisher

Springer-Verlag

Berlin, Heidelberg

Publication History

Published: 03 November 2024

Author Tags

Qualifiers

Article

Contributors

Other Metrics

View Article Metrics

Bibliometrics & Citations

Bibliometrics

Article Metrics

0
Total Citations
0
Total Downloads

Downloads (Last 12 months)0
Downloads (Last 6 weeks)0

Reflects downloads up to 28 Feb 2025

Other Metrics

View Author Metrics

Citations

View Options

View options

Figures

Tables

Media

View Table of Conten