AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Baidu has released Unlimited-OCR, a state-of-the-art AI-powered OCR model capable of processing entire multi-page documents in one go. This development could significantly improve long-document digitization, but questions remain about its comparative accuracy and adoption.

Baidu has unveiled Unlimited-OCR, a new AI-powered optical character recognition model capable of analyzing entire multi-page documents in a single forward pass. This marks a significant technical milestone, particularly for applications involving lengthy texts, and positions Baidu as a leader in OCR innovation.

On June 22, 2026, Baidu open-sourced Unlimited-OCR, a 3-billion-parameter model designed to parse entire documents within a standard 32K context window. The model, based on an architecture derived from DeepSeek-OCR, introduces a novel Reference Sliding Window Attention (R-SWA) mechanism that replaces traditional linear cache growth with a fixed-size memory, enabling faster and more memory-efficient processing of long texts.

The technical report, published on arXiv the following day, confirms that Unlimited-OCR can process dozens of pages in a single pass, maintaining low error rates and high throughput. It achieves a 35% faster throughput than previous models like DeepSeek-OCR, with a throughput of approximately 5,580 tokens per second on benchmark tests, and demonstrates superior performance on the OmniDocBench evaluation, scoring over 93 points overall.

Contrary to viral claims, Baidu’s own data shows the model has around 8,400 downloads in the last month, not 1.9 million, clarifying misconceptions about its popularity. The model is available under an MIT license and supports various deployment options, including transformers and Docker, making it accessible for self-hosted applications.

At a glance
reportWhen: announced June 2026, with technical det…
The developmentBaidu officially open-sourced Unlimited-OCR in June 2026, introducing a new architecture that enables efficient, single-pass parsing of lengthy documents.

Implications for Long-Document OCR Applications

This development signifies a potential shift in how long documents are digitized and processed, reducing the need for page-by-page OCR and complex stitching algorithms. The constant-memory architecture allows for more accurate, faster, and more reliable extraction of data from lengthy texts, which could impact industries like legal, academic, and governmental document management.

However, the trade-off appears to be a slight decrease in peak accuracy compared to some existing page-by-page models, though the overall efficiency gains may outweigh this for many practical applications. The open-source availability also encourages broader experimentation and adoption, potentially setting a new standard in OCR technology.

Amazon

AI-powered OCR scanner

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Baidu’s OCR Evolution and Industry Benchmarks

Baidu has a longstanding presence in OCR with models like PaddleOCR and PaddleOCR-VL, which have achieved high accuracy scores in benchmark tests. The release of Unlimited-OCR builds on this legacy, introducing architectural innovations aimed at solving the memory and latency issues faced by traditional decoder-based OCR models.

Prior to this, models like Zhipu’s GLM-OCR and PaddleOCR-VL achieved scores above 94 on OmniDocBench, but they process documents page-by-page, limiting their efficiency on long texts. Baidu’s new model, while slightly behind in peak accuracy, offers a unique advantage in processing entire documents in one pass, which is a significant step forward for applications requiring comprehensive document understanding.

The release also counters the narrative that China is “killing” OCR innovation, demonstrating that Baidu continues to push the boundaries with architectural improvements rooted in open research and existing models.

“Baidu’s Unlimited-OCR is less a moonshot and more a surgical architectural fix, making long-document OCR more efficient and reproducible.”

— Thorsten Meyer, AI researcher

Amazon

multi-page document OCR software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Model Performance and Adoption

While the technical results are promising, it is still unclear how Unlimited-OCR performs in real-world, diverse document collections outside controlled benchmarks. Its accuracy relative to existing models in production environments, and the extent of its adoption, especially among smaller organizations, remain to be seen.

Additionally, the long-term stability of the architecture and its scalability across different hardware setups are still under evaluation, and independent benchmarks are awaited to confirm its standing in the OCR landscape.

Amazon

long document digitization tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Developments and Industry Impact Expectations

Baidu is expected to continue refining Unlimited-OCR, potentially releasing updates that improve accuracy and robustness. Industry adoption will likely grow as more organizations experiment with the open-source model, especially in sectors handling extensive textual data.

Further independent evaluations and real-world case studies will clarify its competitive positioning. Additionally, competitors may develop similar architectures, leading to broader innovations in long-document OCR processing.

Amazon

OCR document processing device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How does Unlimited-OCR differ from previous Baidu models?

Unlimited-OCR introduces a fixed-size memory attention mechanism called R-SWA, enabling it to process entire multi-page documents in a single pass, unlike earlier models that processed pages individually.

Can I use Unlimited-OCR for commercial projects?

Yes, the model is released under an MIT license, making it freely available for commercial and research use, with support for various deployment options.

Will Unlimited-OCR replace existing OCR solutions?

It offers significant advantages for long-document processing, but its adoption will depend on accuracy, integration ease, and specific application needs. It is likely to complement rather than fully replace current solutions initially.

What are the limitations of Unlimited-OCR?

While effective for long documents, it may have slightly lower peak accuracy compared to page-by-page models in some cases. Its performance outside benchmark settings is still under evaluation.

Source: ThorstenMeyerAI.com

You May Also Like

MacBook Neo Deep Dive: Benchmarks, Wafer Economics, and the 8GB Gamble

An in-depth analysis of the MacBook Neo, including benchmarks, wafer economics, and the implications of its 8GB RAM configuration, unveiled on May 8, 2026.

What Network Sharing Really Adds to Scanner Productivity

The true value of network sharing in scanners lies in its ability to enhance productivity and collaboration—discover how it can transform your workflow today.

Cybercriminal Twins Caught After They Forgot to Turn Off Microsoft Teams Recording

Twin hackers pleaded guilty after their Microsoft Teams meeting recording captured their plot to destroy government databases post-termination.

Japan’s Nidec to end China JV as it scales back EV drive parts

Japanese motor maker Nidec plans to dissolve its joint venture in China as it scales back its electric vehicle drive component business, citing intense competition.