2017 · 172 citations · 30 references
Convolutional Neural NetworkScene AnalysisEngineeringMachine LearningMultiple Instance LearningImage AnalysisData SciencePattern RecognitionSemantic SegmentationVideo TransformerInstance-based LearningMachine VisionObject DetectionComputer ScienceDeep LearningMedical Image ComputingBoundary-aware Instance SegmentationComputer VisionBinary MaskScene UnderstandingImage SegmentationInstance Segmentation
We address the problem of instance-level semantic segmentation, which aims at jointly detecting, segmenting and classifying every individual object in an image. In this context, existing methods typically propose candidate objects, usually as bounding boxes, and directly predict a binary mask within each such proposal. As a consequence, they cannot recover from errors in the object candidate generation process, such as too small or shifted boxes. In this paper, we introduce a novel object segment representation based on the distance transform of the object masks. We then design an object mask network (OMN) with a new residual-deconvolution architecture that infers such a representation and decodes it into the final binary object mask. This allows us to predict masks that go beyond the scope of the bounding boxes and are thus robust to inaccurate object candidates. We integrate our OMN into a Multitask Network Cascade framework, and learn the resulting boundary-aware instance segmentation (BAIS) network in an end-to-end manner. Our experiments on the PASCAL VOC 2012 and the Cityscapes datasets demonstrate the benefits of our approach, which outperforms the state-of-the-art in both object proposal generation and instance segmentation.
30
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, Trevor Darrell · 2015 · 36.2K citations
Ross Girshick · 2015 · 27.2K citations
Image Classification, Convolutional Neural Network, Image Analysis +11
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick et al. · arXiv (Cornell University) · 2015 · 18.2K citations · Full text