Models & Research

Datalab’s Marker 2 vs MinerU, Docling and LiteParse: 76.0 on olmOCR-bench at 5× MinerU’s Throughput

· July 25, 2026
Datalab’s Marker 2 vs MinerU, Docling and LiteParse: 76.0 on olmOCR-bench at 5× MinerU’s Throughput

What changed

Datalab rewrote its Marker OCR system into a three-mode pipeline now called Marker 2. This update pushed accuracy to 76.0 on the olmOCR-bench, a strong measure of OCR model performance. Marker 2 also sustains processing speeds of 2.9 pages per second on a single B200 GPU. This throughput rate is over five times higher than MinerU’s backend pipeline, which was one of the better performers before.

Why builders should care

Scaling OCR processing usually involves trade-offs between speed and accuracy. Marker 2’s three-mode pipeline reduces this tension by maintaining solid accuracy while massively boosting processing speed. Builders and operators running high-volume document workflows can extract useful text five times faster compared to MinerU with no accuracy penalty. It also outperforms Docling on both fronts, making it an interesting option for teams building or upgrading document extraction pipelines.

The practical takeaway

If throughput matters in your OCR use case—especially for large batches—Marker 2 cuts GPU time drastically. Faster throughput lowers cloud or hardware costs and speeds up downstream automation like text analytics or compliance checks. For many, this makes the economics of large-scale OCR easier to justify. Marker 2’s architecture also signals a design direction where flexible, multi-mode pipelines unlock both speed and accuracy, not forcing a compromise between them.

What to watch next

It will be important to see how Marker 2 performs beyond the olmOCR-bench benchmark in real-world scenarios and with different document types. Comparisons against LiteParse remain less detailed, so their relative positioning in speed or accuracy is still somewhat open. Adoption trends for Marker 2 versus MinerU or Docling could reveal what operators prioritize most—maximum throughput or niche cases where accuracy edges matter more. Updates to pipeline components or support for additional languages will also shape the competitive landscape.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.