More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Datalab’s new Chandra OCR 2 is a four-billion-parameter open-weight model that turns images and PDFs into structured Markdown, HTML or JSON while keeping the original layout intact. On the independent olmOCR benchmark it scored 85.9%—16 points higher than GPT-4o’s 69.9%. In Datalab’s own 90-language test, Chandra hit 72.7% versus Gemini 2.5 Flash at 60.8%, and on the 43 most common languages it jumped to 77.8%, outpacing GPT-5 Mini’s 60.5%. Scripts like Kannada, Malayalam and Telugu saw gains of 40–46 points compared to Chandra 1.
Beyond printed English, Chandra OCR 2 handles complex tables with nested headers, handwritten math on century-old scans, multi-column layouts, forms with checkboxes, images with auto-captions and even exports flowcharts as Mermaid diagrams. They halved model size from 9B to 4B parameters, boosted accuracy across all categories and doubled throughput—about two pages per second on a single NVIDIA H100. Installation is a five-line pip command with no API keys or per-page fees, or you can deploy via Docker and vLLM for production.
Licensing is Apache 2.0 for code, but weights use a modified OpenRAIL-M license: free for research, personal use and startups under $2 million in funding or revenue; larger companies need a commercial license. The olmOCR results come from an independent benchmark, but Datalab’s multilingual scores await third-party validation. GPT-4o’s score also reflects fixed prompts, so fine-tuned prompting could narrow the gap. Handwriting only hits around 90.8% on notes and falls to about 50.4% on complex forms, so heavy handwriting use still needs testing.
Questions about this article
No questions yet.