Machine Intelligence Research · 2023 · 176 citations · 166 references
Llm Fine-tuningEngineeringMachine LearningMultimodal LearningMultilingual PretrainingLanguage ProcessingSpeech RecognitionNatural Language ProcessingMultimodal LlmPre-trainingData ScienceMachine TranslationLarge Ai ModelComprehensive SurveyGenerative ModelsPre-trained ModelsMultimodal Signal ProcessingComputer ScienceConventional Deep LearningDeep LearningComputer Vision
The rapid rise of generalized deep models such as BERT, ViT, and GPT has spurred growing interest in multi‑modal pre‑trained big models that extend these successes across vision, language, and speech domains. This survey aims to comprehensively review multi‑modal pre‑trained models, offering fresh insights and guidance for researchers tracking the latest developments. The authors review foundational deep‑learning and pre‑training work in NLP, CV, and speech, define key tasks and challenges, analyze data, objectives, architectures, and knowledge‑enhanced strategies, and evaluate models on generative, classification, and regression downstream tasks. They present visualizations and analyses of model parameters and downstream performance, and provide a continuously updated repository of large‑scale multi‑modal pre‑trained models.
Abstract With the urgent demand for generalized deep models, many pre-trained big models are proposed, such as bidirectional encoder representations (BERT), vision transformer (ViT), generative pre-trained transformers (GPT), etc. Inspired by the success of these models in single domains (like computer vision and natural language processing), the multi-modal pre-trained big models have also drawn more and more attention in recent years. In this work, we give a comprehensive survey of these models and hope this paper could provide new insights and helps fresh researchers to track the most cutting-edge works. Specifically, we firstly introduce the background of multi-modal pre-training by reviewing the conventional deep learning, pre-training works in natural language process, computer vision, and speech. Then, we introduce the task definition, key challenges, and advantages of multi-modal pre-training models (MM-PTMs), and discuss the MM-PTMs with a focus on data, objectives, network architectures, and knowledge enhanced pre-training. After that, we introduce the downstream tasks used for the validation of large-scale MM-PTMs, including generative, classification, and regression tasks. We also give visualization and analysis of the model parameters and results on representative downstream tasks. Finally, we point out possible research directions for this topic that may benefit future works. In addition, we maintain a continuously updated paper list for large-scale pre-trained multi-modal big models: https://github.com/wangxiao5791509/MultiModal_BigModels_Survey .
166
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Sepp Hochreiter, Jürgen Schmidhuber · Neural Computation · 1997 · 93.8K citations
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher et al. · 2009 IEEE Conference on Computer Vision and Pattern Recognition · 2009 · 60.2K citations
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio et al. · Proceedings of the IEEE · 1998 · 56.5K citations · Full text
Engineering, Machine Learning, Multilayer Neural Networks +17