Baidu’s AI OCR: The Secret To Fast, One-Pass PDF Transcription
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Baidu released Unlimited-OCR, a 3-billion-parameter model that can transcribe entire multi-page PDFs in one forward pass. It introduces a new attention mechanism that maintains constant memory, enabling faster processing of long documents on standard hardware.

Baidu has open-sourced Unlimited-OCR, a large-scale OCR model capable of parsing entire multi-page documents in a single forward pass. This development, announced on June 22, 2026, introduces a novel attention mechanism that maintains constant GPU memory and latency regardless of document length, representing a significant technical achievement in document OCR processing.

The Unlimited-OCR model is based on a 3-billion-parameter architecture, derived from Baidu’s previous DeepSeek-OCR, and employs a new mechanism called Reference Sliding Window Attention (R-SWA). This replaces the traditional linear growth of memory with a fixed-size cache, allowing dozens of pages to be processed in one pass without splitting or external scheduling. The model is open-sourced under an MIT license and supports various deployment formats, including Docker and community quantizations.

According to Baidu, the model achieves a throughput of approximately 5,580 tokens per second on OmniDocBench, surpassing DeepSeek-OCR by about 12.7%. On long documents, it maintains low error rates, with an edit distance below 0.11 after processing 40+ pages, which is a significant improvement over prior page-by-page OCR methods. However, the model does not outperform Baidu’s own PaddleOCR-VL 1.5 or Zhipu’s GLM-OCR on the same benchmark, as those models prioritize single-page accuracy rather than multi-page efficiency.

Contrary to viral claims, the model has around 8,400 downloads on Hugging Face as of late July 2026, not 1.9 million. The release emphasizes architectural innovations over mere accuracy metrics, positioning Unlimited-OCR as a practical solution for long-document OCR tasks on standard hardware.

At a glance
breakingWhen: announced June 22, 2026, with technical…
The developmentBaidu officially released Unlimited-OCR on June 22, 2026, demonstrating a new architecture that handles multi-page document OCR in a single pass with constant memory, marking a significant technical advancement.

Implications for Long-Document OCR Efficiency

This development signifies a major step forward in OCR technology, enabling fast and accurate transcription of multi-page documents in a single pass. This reduces the need for page splitting, stitching, and complex post-processing, which are common bottlenecks in current OCR pipelines. For industries dealing with large volumes of long documents—such as legal, academic, and governmental sectors—this could lead to more efficient workflows and lower hardware requirements.

Moreover, the open-source release provides researchers and developers with a reproducible architecture, encouraging further innovation. While the model does not currently surpass the highest single-page accuracy models, its ability to handle long documents efficiently makes it a valuable tool in real-world applications where speed and memory constraints are critical.

Amazon

document OCR scanner

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Technical Foundations and Prior Developments

The architecture builds upon Baidu’s earlier DeepSeek-OCR, which used a cascaded SAM-ViT and CLIP-ViT encoder to compress page images. The key innovation is R-SWA, which mimics human-like ‘soft forgetting’ by maintaining a fixed memory cache that slides over the input, avoiding the linear growth of attention cache typical in decoder-based models. This approach allows the model to process entire multi-page documents in a single forward pass, a feat previously hindered by memory limitations and latency issues.

Open-sourcing the model follows Baidu’s pattern of releasing advanced models to the community, with prior models like PaddleOCR-VL and Zhipu’s GLM-OCR setting benchmarks for page-by-page OCR. The new architecture is designed explicitly to address the challenges of long-document transcription, a long-standing bottleneck in OCR technology.

“Unlimited-OCR demonstrates that constant memory and single-pass processing are achievable at scale, opening new doors for long-document OCR.”

— Baidu Research Team

Amazon

multi-page PDF OCR software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About Performance and Adoption

While the technical details are promising, it is still unclear how Unlimited-OCR performs in diverse real-world scenarios beyond the benchmark tests. Its accuracy on complex layouts, heavily formatted documents, or noisy scans remains to be evaluated. Additionally, the extent of community adoption and integration into commercial OCR pipelines is uncertain, as is the performance on hardware beyond Baidu’s testing environment.

Amazon

handheld OCR document scanner

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Community Testing and Industry Adoption

Developers and researchers are likely to experiment with the open-source model, testing its capabilities on various long-document datasets. Baidu may release updates or optimized versions based on community feedback. Industry users will evaluate whether the model’s efficiency gains justify integration into their OCR workflows, and further benchmarks may be released to validate its long-term robustness and accuracy.

Amazon

long document OCR tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How does Unlimited-OCR differ from traditional OCR models?

It introduces a Reference Sliding Window Attention (R-SWA) mechanism that maintains a fixed memory cache, enabling single-pass processing of multi-page documents without splitting or stitching, unlike traditional page-by-page OCR models.

Can Unlimited-OCR run on standard hardware?

Yes, the model is designed to operate within a standard 32K context window on typical hardware, making it accessible for many users without specialized GPU setups.

How accurate is Unlimited-OCR for long documents?

According to Baidu’s internal tests, the model maintains an error rate below 0.11 over 40 pages, which is competitive for long-document OCR tasks, though it may not surpass specialized page-by-page models in single-page accuracy.

What are the limitations of this model?

Its performance on complex layouts, noisy scans, or heavily formatted documents has not been fully evaluated, and real-world adoption may reveal additional challenges.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The gigawatt gap. Why China is structurally positioned for AI power and the US is engineering around its grid.

Analysis of how China’s centralized infrastructure and renewable buildout give it a gigawatt-scale edge in AI deployment, contrasting US fragmentation.

How Fintech Is Changing Invoice Collection For Small Businesses

New fintech tools are automating and tone-calibrating invoice follow-ups, helping small businesses reduce overdue payments and improve cash flow.

How The RayNeo GT Max Enhances VR Experience Through Better Signal Monitoring

The RayNeo GT Max smart glasses incorporate an upgraded signal monitoring system aimed at improving VR data stability and accuracy, supporting operators in dynamic environments.

A War Room for Your Next Idea: Inside IdeaClyst

Discover how IdeaClyst provides founders with a local, AI-driven war room for evaluating and validating startup ideas, reducing costly mistakes.