2023 · 39 citations · 28 references
Multimodal LlmPre-trained EncoderImage AnalysisMachine LearningData ScienceEngineeringPattern RecognitionText-to-image RetrievalBiometricsCross-modal Image-text RetrievalFeature LearningSecurity IssueMultimodal LearningDownstream-agnostic Adversarial ExamplesComputer ScienceDeep LearningComputer Vision
Multimodal contrastive learning aims to train a general-purpose feature extractor, such as CLIP, on vast amounts of raw, unlabeled paired image-text data. This can greatly benefit various complex downstream tasks, including cross-modal image-text retrieval and image classification. Despite its promising prospect, the security issue of cross-modal pre-trained encoder has not been fully explored yet, especially when the pre-trained encoder is publicly available for commercial use.
28
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su et al. · International Journal of Computer Vision · 2015 · 39.5K citations
Image Classification, Convolutional Neural Network, Machine Vision +7
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza et al. · Communications of the ACM · 2020 · 12.7K citations · Full text
Tat‐Seng Chua, Jinhui Tang, Richang Hong et al. · 2009 · 3K citations
Natural Language Processing, Media Search, Image Analysis +14
Universal Adversarial Perturbations
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, Omar Fawzi et al. · 2017 · 2.7K citations · Full text
Convolutional Neural Network, Engineering, Machine Learning +16