EngineeringMachine LearningSemantic Image SynthesisNatural Language ProcessingSource ImageData ScienceRealistic ImagesComputational ImagingIntelligent Image ManipulationMachine TranslationSynthetic Image GenerationMachine VisionImage SynthesisVision Language ModelGenerative ModelsComputer ScienceHuman Image SynthesisDeep LearningComputer VisionGenerative Adversarial NetworkGenerative Ai
Intelligent image manipulation. The paper proposes synthesizing realistic images directly from natural language descriptions, aiming to generate images that are both realistic and match the target text while preserving unrelated image features. The authors design an end‑to‑end neural architecture that disentangles image and text semantics and uses adversarial learning to automatically learn implicit loss functions, enabling synthesis of realistic images that match the target description while preserving unrelated features. Experiments on Caltech‑200 and Oxford‑102 demonstrate that the model can synthesize realistic images that match descriptions while preserving other image features.
In this paper, we propose a way of synthesizing realistic images directly with natural language description, which has many useful applications, e.g. intelligent image manipulation. We attempt to accomplish such synthesis: given a source image and a target text description, our model synthesizes images to meet two requirements: 1) being realistic while matching the target text description; 2) maintaining other image features that are irrelevant to the text description. The model should be able to disentangle the semantic information from the two modalities (image and text), and generate new images from the combined semantics. To achieve this, we proposed an end-to-end neural architecture that leverages adversarial learning to automatically learn implicit loss functions, which are optimized to fulfill the aforementioned two requirements. We have evaluated our model by conducting experiments on Caltech-200 bird dataset and Oxford-102 flower dataset, and have demonstrated that our model is capable of synthesizing realistic images that match the given descriptions, while still maintain other features of original images.
28
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su et al. · International Journal of Computer Vision · 2015 · 39.5K citations
Image Classification, Convolutional Neural Network, Machine Vision +7