2024 · 15 citations · 24 references
EngineeringMachine LearningShape-guided DiffusionObject MaskAttentionSocial SciencesMultimodal LlmImage AnalysisVision RecognitionSynthetic Image GenerationMachine VisionPrecise Object SilhouetteVision Language ModelComputer ScienceHuman Image SynthesisVisual ProcessingDeep LearningMedical Image ComputingComputer VisionBiomedical ImagingScene UnderstandingDiffusion-based ModelingShape Constraint
We introduce precise object silhouette as a new constraint in text-to-image diffusion models, which we dub Shape-Guided Diffusion. Our training-free method uses an Inside-Outside Attention mechanism during the inversion and generation process to apply a shape constraint to the cross- and self-attention maps. Our mechanism designates which spatial region is the object (inside) vs. background (outside) then associates edits to the correct region. We demonstrate the efficacy of our method on the shape-guided editing task, where the model must replace an object according to a text prompt and object mask. We curate a new ShapePrompts benchmark derived from MS-COCO and achieve SOTA results in shape faithfulness without a degradation in text alignment or image realism according to both automatic metrics and annotator ratings. Our data and code will be made available at https://shape-guided-diffusion.github.io.
24
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe et al. · 2016 · 30.2K citations
Convolutional Neural Network, Engineering, Machine Learning +17
Denoising Diffusion Probabilistic Models
arXiv (Cornell University) · 2020 · 5.6K citations · Full text
Basic objects in natural categories
Eleanor Rosch, Carolyn Β. Mervis, Wayne D. Gray et al. · Cognitive Psychology · 1976 · 5.5K citations
Deep High-Resolution Representation Learning for Human Pose Estimation