arXiv (Cornell University) · 2019 · 18 citations · 32 references
EngineeringMachine LearningAccurate DetectorArbitrary ShapesNatural Language ProcessingImage AnalysisText-to-image RetrievalData ScienceVisual GroundingPattern RecognitionText RecognitionCharacter RecognitionComputational GeometryGeometric ModelingMachine VisionOptical Character RecognitionObject DetectionVision Language ModelComputer ScienceIterative Refinement ModuleDeep LearningComputer VisionBorder OffsetsNatural SciencesDocument ProcessingIterative Refinement
Previous scene text detection methods have progressed substantially over the past years. However, limited by the receptive field of CNNs and the simple representations like rectangle bounding box or quadrangle adopted to describe text, previous methods may fall short when dealing with more challenging text instances, such as extremely long text and arbitrarily shaped text. To address these two problems, we present a novel text detector namely LOMO, which localizes the text progressively for multiple times (or in other word, LOok More than Once). LOMO consists of a direct regressor (DR), an iterative refinement module (IRM) and a shape expression module (SEM). At first, text proposals in the form of quadrangle are generated by DR branch. Next, IRM progressively perceives the entire long text by iterative refinement based on the extracted feature blocks of preliminary proposals. Finally, a SEM is introduced to reconstruct more precise representation of irregular text by considering the geometry properties of text instance, including text region, text center line and border offsets. The state-of-the-art results on several public benchmarks including ICDAR2017-RCTW, SCUT-CTW1500, Total-Text, ICDAR2015 and ICDAR17-MLT confirm the striking robustness and effectiveness of LOMO.
32
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. · 2016 · 214.9K citations · Full text
Image Classification, Deep Neural Networks, Machine Vision +14
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, Trevor Darrell · 2015 · 36.2K citations
Kaiming He, Georgia Gkioxari, Piotr Dollár et al. · 2017 · 27.9K citations
Object Instance Segmentation, Scene Analysis, Machine Vision +13
Feature Pyramid Networks for Object Detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick et al. · 2017 · 27.7K citations
Feature Pyramid Networks, Convolutional Neural Network, Image Analysis +14
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
Shaoqing Ren, Kaiming He, Ross Girshick et al. · arXiv (Cornell University) · 2015 · 18.2K citations · Full text