STOCK TITAN

Elastic Introduces jina-ocr-v1: End-to-End Document Processing in a Single Frontier-Grade Model

Elastic releases jina-ocr-v1, a compact frontier-grade OCR model for complex, multilingual documents, available in cloud, API and on-premises options.

(Neutral)
(Neutral)
Tags
See more from StockTitan in Google Search and AI answers. Adds StockTitan as a preferred source · opens Google
Add on Google

New OCR model processes complex layouts, tables, handwriting, and math across 100+ languages with frontier-grade accuracy at a fraction of the size and cost

SAN FRANCISCO--(BUSINESS WIRE)-- Elastic (NYSE: ESTC) today announced the launch of jina-ocr-v1, a new optical character recognition (OCR) model for end-to-end document processing. At 574M active parameters, it delivers frontier-grade accuracy in a model roughly one tenth the size of the benchmark leader. Jina-ocr-v1 accurately converts complex visual documents into structured, machine-readable text, such as Markdown, in a single pass, making it easy to search, train models, and build agentic applications using the data from scanned documents.

While traditional OCR works well on clean text and simple layouts, complex documents with highly visual content often require separate processing steps, such as page segmentation, element classification, text recognition and reassembly. Each step introduces potential for errors that can accumulate through the fragile processing pipeline. When inaccurate or incomplete data is passed downstream, agents can return incomplete facts, and RAG pipelines can return answers that don't accurately reflect the source documents.

jina-ocr-v1 handles the entire process end to end in a single model. It uses a mixture-of-experts architecture with 3.4B total parameters and 574M active at inference, running at the speed and cost of a sub-600M model. jina-ocr-v1 also adds FastMTP technology, which improves multi-token prediction to accelerate inference.

In a single model, jina-ocr-v1 can:

  • Work across a broad range of imaged documents: Processes images of varying quality, including scanned pages, photographed documents, slides and label images from source formats such as PDF, PPTX and XLSX.
  • Preserve document structure: Processes complex layouts and returns structured Markdown that retains headings, sections, lists and reading order.
  • Extract tables: Converts tables into basic HTML format, suitable for further processing and importing into spreadsheets or other applications.
  • Recognize handwriting: Reads handwriting, including block text in a wide array of languages, and English cursive.
  • Read more than 100 languages: Understands a wide array of global languages and scripts, with all major international languages and scripts represented.
  • Convert mathematical notation: Transforms printed formulas into LaTeX math code for use in documents and scientific applications.

At a tenth the size of the olmOCR-bench leader, jina-ocr-v1 scores 83.4 on olmOCR-bench, the highest published score among models with fewer than 600M active parameters. It delivers frontier-grade accuracy on less hardware, outperforming frontier LLMs on character-level accuracy and reading order.

"Customers need an easy way to digitize their information more than ever in the age of AI,” said Han Xiao, vice president of AI, Elastic. "Traditional OCR pipelines break down with complex layouts, tables, handwriting and other highly visual content. Until now, companies either had to accept those limitations or pay a significant premium to use general-purpose LLMs for ingesting documents. We built jina-ocr-v1 to handle that full range of complexity in a single model, while remaining very efficient at scale.”

Availability

jina-ocr-v1 is available now via the Elastic Inference Service, included with Elastic Cloud, with preconfigured model provisioning and GPU acceleration. Developers can access the model through a preconfigured endpoint without hosting the model or provisioning their own GPUs.

  • Get started with the Jina API: Access jina-ocr-v1 on a pay-per-token basis through the Jina API.
  • Deploy jina-ocr-v1 on-premises: Run jina-ocr-v1 locally or on-premises using pre-built containers with commercial licensing from Elastic, or access the model through Hugging Face under CC BY-NC 4.0 for academic and noncommercial use.

Additional Materials

About Elastic

Elastic (NYSE: ESTC) integrates its deep expertise in search technology with artificial intelligence to help everyone transform all of their data into answers, actions, and outcomes. Elasticsearch, which is the foundation for its search, observability, and security solutions, is used by thousands of companies, including more than 75% of the Fortune 100. Learn more at elastic.co.

Elastic and associated marks are trademarks or registered trademarks of elasticsearch B.V. and its subsidiaries. All other company and product names may be trademarks of their respective owners.

Media Contact
Elastic PR
PR-team@elastic.co

Source: Elastic N.V.

Key Terms

optical character recognition technical
Optical character recognition (OCR) is software that reads printed or handwritten text in images or scanned documents and turns those letter shapes into editable, searchable digital text. Investors care because OCR lets companies extract financial data, contracts, invoices and regulatory filings quickly and with fewer errors, cutting manual work and speeding analysis. Think of it as converting a filing cabinet into a searchable spreadsheet so decisions and audits happen faster and cheaper.
mixture-of-experts technical
A mixture-of-experts is a computer model design that uses a group of smaller specialist models, with a controller that picks which specialist(s) handle each task—like sending a question to the right expert on a team. For investors, this matters because it can deliver faster, cheaper, or more accurate AI features without building one giant model, affecting a company’s product performance, costs, competitive edge, and regulatory or safety risks tied to how the system is managed.
latex technical
Latex is a milky fluid naturally tapped from rubber trees or made synthetically, processed into flexible, rubbery materials used in gloves, medical tubing, condoms, balloons and many industrial goods. Investors care because latex is a basic input whose price, supply or safety profile (for example, allergic reactions or regulatory limits) can affect manufacturers’ costs, product availability and recall risk—like how a gasoline shortage raises costs across many businesses.
cc by-nc 4.0 regulatory
A Creative Commons Attribution-NonCommercial 4.0 International license that lets others copy, share, and modify a work as long as they give credit to the creator and do not use it for commercial purposes. It matters to investors because press releases, research, images, or other content under this license can be reused in internal reports, commentary, or educational materials but cannot be used in products or marketing that generate commercial revenue without separate permission.

Keep reading