Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images

TLDR

Previous deep learning approaches typically output 3D shapes as volumes or point clouds, making conversion to usable mesh models difficult. The study proposes an end‑to‑end deep learning architecture that generates a 3D triangular mesh from a single RGB image. The network uses a graph‑based convolutional neural network to progressively deform an ellipsoid into a mesh, employing a coarse‑to‑fine strategy and mesh‑related losses to ensure accurate geometry. Experiments demonstrate that the method yields meshes with finer details and higher shape‑estimation accuracy than state‑of‑the‑art approaches.

Abstract

We propose an end-to-end deep learning architecture that produces a 3D shape in triangular mesh from a single color image. Limited by the nature of deep neural network, previous methods usually represent a 3D shape in volume or point cloud, and it is non-trivial to convert them to the more ready-to-use mesh model. Unlike the existing methods, our network represents 3D mesh in a graph-based convolutional neural network and produces correct geometry by progressively deforming an ellipsoid, leveraging perceptual features extracted from the input image. We adopt a coarse-to-fine strategy to make the whole deformation procedure stable, and define various of mesh related losses to capture properties of different levels to guarantee visually appealing and physically accurate 3D geometry. Extensive experiments show that our method not only qualitatively produces mesh model with better details, but also achieves higher 3D shape estimation accuracy compared to the state-of-the-art.

References

Page 1

	Year	Citations

Page 1