2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) · 2022 · 102 citations · 32 references
EngineeringMachine LearningSignificant Anchor PointsNatural Language ProcessingMultimodal LlmSecond Language AcquisitionImage AnalysisText-to-image RetrievalData SciencePattern RecognitionComputational LinguisticsVisual Question AnsweringLanguage StudiesMachine TranslationMachine VisionFeature LearningAttention MatrixLinguisticsVision Language ModelComputer ScienceDeep LearningMedical Image ComputingMultimodal TranslationComputer VisionNeural Machine TranslationI2i TranslationMutual InformationSpeech Translation
Unpaired image-to-image (I2I) translation often requires to maximize the mutual information between the source and the translated images across different domains, which is critical for the generator to keep the source content and prevent it from unnecessary modifications. The self-supervised contrastive learning has already been successfully applied in the I2I. By constraining features from the same location to be closer than those from different ones, it implicitly ensures the result to take content from the source. However, previous work uses the features from random locations to impose the constraint, which may not be appropriate since some locations contain less information of source domain. Moreover, the feature itself does not reflect the relation with others. This paper deals with these problems by intentionally selecting significant anchor points for contrastive learning. We design a query-selected attention (QS-Attn) module, which compares feature distances in the source domain, giving an attention matrix with a probability distribution in each row. Then we select queries according to their measurement of significance, computed from the distribution. The selected ones are regarded as anchors for contrastive loss. At the same time, the reduced attention matrix is employed to route features in both domains, so that source relations maintain in the synthesis. We validate our proposed method in three different I2I datasets, showing that it increases the image quality with-out adding learnable parameters. Codes are available at https://github.com/sapphire497/query-selected-attention.
32
Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks
Jun-Yan Zhu, Taesung Park, Phillip Isola et al. · 2017 · 21.3K citations · Full text
Engineering, Machine Learning, Image-to-image Translation +17
Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network
Christian Ledig, Lucas Theis, Ferenc Huszár et al. · 2017 · 12K citations
Momentum Contrast for Unsupervised Visual Representation Learning
Kaiming He, Haoqi Fan, Yuxin Wu et al. · 2020 · 11.6K citations
Convolutional Neural Network, Image Analysis, Machine Learning +14