ICDAR 2015 Nancy, France  -   August 23-26, 2015

Hervé Dejean: Extracting Structured Data from Unstructured Documents with Incomplete Resources

Abstract: We present a method for extracting structured elements of information, called structured data (sdata), from ocr’ed pages. The method first analyzes the layout of the page, building several concurrent layout structures. Then a tagging step is performed in order to tag textual elements based on their content. Combining the layout structures and the tagged elements, layout models for representing the structured data are inferred for the current page. These models are used to correct or tag some elements missed by the tagging step. The final set of structured data is extracted. An evaluation is presented.Abstract: