ABBYY is a pioneer in optical character recognition technology, actively researching and innovating in this area since 1993, when our first “omnifont OCR system” ABBYY FineReader was launched to the market. Over the years, the technology has evolved from recognizing individual characters, identifying words, and reproducing page structure, to applying adaptive document recognition technology (ADRT®) that understands documents in their entirety, including layout, multi-page structure, and elements such as header, footer, and table of contents.
With the advancements of AI, ABBYY has developed and solidified its end-to-end approach to OCR and ICR in the last several years. This approach uses the same technologies that are the basis of generative AI tools—convolutional neural networks, transformers, and language models.
The convolutional neural network breaks apart an image of handwritten or printed text on a document into its bits and bytes, trying to make sense of what it actually is. All that input from the CNN then goes into a transformer to provide a potential outcome of a word. Then, we introduce our very own LM, which is trained on billions of parameters, with the specific function of being able to take the context of all of the different words in a group and make the best use of that info to come to a conclusion. This technique drastically improves the performance and accuracy of our OCR and ICR capabilities overall, and it is leveraged in combination with our statistical approach. Our AI will automatically decide which approach is best fit for your document use cases to optimize on the fly for consistency, accuracy, and speed, leading to better straight-through-processing rates.