IEEE Transactions on Circuits and Systems for Video Technology · 2023 · 55 citations · 33 references
Artificial IntelligenceEffective Vision-languageEngineeringMachine LearningToken Generation TaskLanguage ProcessingNatural Language ProcessingMultimodal LlmData SciencePattern RecognitionObject TrackingMachine TranslationMachine VisionVision-language TrackingVision Language ModelMoving Object TrackingComputer ScienceComputer VisionToken QueriesLinguisticsTracking System
In this paper, we present a simple, flexible and effective vision-language (VL) tracking pipeline, termed MMTrack, which casts VL tracking as a token generation task. Traditional paradigms address VL tracking task indirectly with sophisticated prior designs, making them over-specialize on the features of specific architectures or mechanisms. In contrast, our proposed framework serializes language description and bounding box into a sequence of discrete tokens. In this new design paradigm, all token queries are required to perceive the desired target and directly predict spatial coordinates of the target in an auto-regressive manner. The design without other prior modules avoids multiple sub-tasks learning and hand-designed loss functions, significantly reducing the complexity of VL tracking modeling and allowing our tracker to use a simple cross-entropy loss as unified optimization objective for VL tracking task. Extensive experiments on TNL2K, LaSOT, LaSOT <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">$_{\mathrm{ext}}$ </tex-math></inline-formula> and OTB99-Lang benchmarks show that our approach achieves promising results, compared to other state-of-the-arts.
33
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia et al. · 2015 · 46.2K citations
Image Classification, Deep Neural Networks, Image Analysis +15
Focal Loss for Dense Object Detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick et al. · 2017 · 24.4K citations
Image Classification, Convolutional Neural Network, Image Analysis +15