Synthetic Data Generation for a Document Parsing AI
US20250103794A1

Description (excerpt)
CROSS-REFERENCE This application is a nonprovisional application of and claims the benefit of priority under 35 U.S.C. § 119 based on U.S. Provisional Patent Application No. 63/441,199 filed on Jan. 26, 2023. The Provisional application and all references cited herein are hereby incorporated by reference into the present disclosure in their entirety. FEDERALLY-SPONSORED RESEARCH AND DEVELOPMENT The United States Government has ownership rights in this invention. Licensing inquiries may be directed to Office of Technology Transfer, US Naval Research Laboratory, Code 1004, Washington, DC 20375, USA; +1.202.767.7230; techtran@nrl.navy.mil, referencing Navy Case #211377. TECHNICAL FIELD The present disclosure is related to parsing a document that is formatted in a first format and generating another file format based on the parsing, and more specifically to, but not limited to, creating training, testing, and validation data for an artificial intelligence (AI) system to use for training and evaluation purposes in order to parse a document. BACKGROUND Currently, in order to get training data for a document-parsing AI, one may either manually parse a document into JSON (JSON, n.d.), use pre-made configuration files to parse a document into JSON, or use an external dataset. Manually parsing a document into JSON is time consuming. Relying on pre-made configuration files is better, but sometimes text cannot be extracted from a PDF (Adobe, 2024), meaning some documents cannot be parsed into JSON correctly. Also, after a document is parsed using a configuration file, it should be manually checked, which is also time consuming. Furthermore, in order to get a specific data variation, such as a unique way of writing a table, one will have to find a pre-existing document with that variation, and then update the configuration file(s) and possibly the parsing code in order to have data that captures that variation. A pre-existing database may parse certain objects into JSON using an undesirable format. Existing methods may include manually parsing documents into JSON, using pre-made configuration files to parse a document into JSON, and relying on pre-existing databases. There are disadvantages to these existing methods. For example, there are many, many different forms for which to manually parse. For example, on an FAA (Federal Aviation Administration) webpage entitled “Forms” (U.S. Department of Transportation, 2024) under “Export All,” there is a download link which downloads an Excel file that lists 1,221 forms as of Jan. 26, 2024.” Therefore, it takes a lot of time and effort to manually parse each one of these forms. For example, even if one were to write a configuration file for all these forms, if a new form is added, one would have to add new configuration file(s), and if an existing form is significantly changed, one would have to edit existing configuration file(s). This can be cumbersome, time-consuming, and expensive. In a paper published in 2021 entitled “DocParser: Hierarchical Document Structure Parsing from Renderings,” when discussing “the effectiveness of DocParser for parsing the complete document structures, the authors state “that both suitable baselines and datasets for this task are hitherto lacking.” (Rausch et al., 2021) Because of these problems, there exists a need for a solution for efficiently parsing a wide-variety of forms and/or documents, which is provided by disclosed embodiments and/or aspects described herein. SUMMARY This summary is intended to introduce, in simplified form, a selection of concepts that are further described in the Detailed Description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. Instead, it is merely presented as a brief overview of the subject matter described and claimed herein. Disclosed aspects provide for creating objects by a synthetic data generation software. The data created by this software will be the training, validation, and testing data for an Artificial Intelligence (AI) system that can, for example, convert a document (e.g., an image of a document) into a JSON representation of that document. Disclosed aspects provide for a system incorporating AI that can parse a wide variety of forms. By generating robust synthetic data, aspects described herein can increase the total amount of data that exists for training and evaluating document-parsing AI systems. Aspects described herein provide for a synthetic data generati
Filing details
- Inventors
- Charles A. Norsworthy
- Assignee
- The Government Of The United States Of America, As Represented By The Secretary …
- Filed
- Dec 17, 2024
- Granted
- Application pending
Bibliographic data and excerpted text sourced from Google Patents (public record) as part of IP TechMatch's current-filings monitor. This filing is not part of the 2019 historical archive. For the authoritative full text, drawings, and legal status, see the source links above or consult USPTO records directly.