Click any tag below to further narrow down your results
Links
Mistral OCR 4 extracts text from PDFs, DOCs and more while also returning bounding boxes, block types and per-word confidence. It supports 170 languages, runs in a single container for self-hosted deployments, and outperforms rivals on human and automated benchmarks.
- Mistral OCR 4 doesn't just extract text—it returns bounding boxes, block types (titles, tables, equations, signatures), and per-word confidence scores, and can be self-hosted in a single container for data-sovereignty needs
- Beat every tested competitor in human evaluations across 600+ documents/12+ languages with a 72% average win rate, and scored 85.20 on OlmOCRBench and 93.07 on OmniDocBench
- Roughly 4x faster than some enterprise OCR providers, and 8x cheaper with 17x lower latency than leading agentic document parsers on finance datasets with charts/figures
- Priced at $4 per 1,000 pages via API ($2 with batch discount) or $5 per 1,000 pages through the no-code Document AI interface
Datalab’s 4 billion-parameter Chandra OCR 2 outperforms GPT-4o and Gemini across independent and multilingual benchmarks, handling complex layouts, math notation, flowcharts and 90 languages with state-of-the-art accuracy. It’s available under Apache 2.0 code with a modified OpenRAIL-M license for weights, runs locally via HuggingFace or vLLM, and doubles throughput over its predecessor.
- Chandra OCR 2 (4B params, open-weight) scored 85.9% on olmOCR vs GPT-4o's 69.9%, and beat Gemini/GPT-5 Mini on multilingual benchmarks, with huge gains on South Asian scripts (Kannada +42.6, Malayalam +46.2, Telugu +39.1).
- It processes full pages in one pass rather than splitting into blocks, giving it an edge on tables, nested headers, checkboxes, handwritten math, and flowcharts exported as Mermaid diagrams.
- Despite shrinking from 9B to 4B parameters, throughput doubled to ~2 pages/sec on an H100 while accuracy improved.
- Code is Apache 2.0 and installable via pip/Docker, but weights use a modified OpenRAIL-M license requiring a paid commercial license for larger companies.
Google’s new Gemini 3.5 Live Translate model converts speech to speech in real time across more than 70 languages, preserving speakers’ intonation, pacing and pitch. It streams audio continuously with minimal delay and is available via the Gemini Live API, Google Meet preview, and Google Translate apps. TAGS: live-translation, speech-translation, multilingual, real-time-ai, gemini-models
- Gemini 3.5 Live Translate does real-time speech-to-speech translation across 70+ languages while preserving intonation, pacing and pitch, with only a few seconds of delay
- Google Meet's speech translation jumps from 5 languages to 70+, now covering over 2,000 language pairs instead of just English-based ones
- Grab is testing it across 10 million voice calls a month for drivers and riders
- All translated audio carries an invisible SynthID watermark to keep AI-generated speech traceable
Chandra OCR 2, a 4 billion-parameter model from Datalab, outperforms GPT-4o and Gemini on AllenAI’s olmOCR benchmark and a 90-language test while halving the model size. It preserves layout, reads complex tables and math notation, converts diagrams to Mermaid, and runs at two pages per second on an NVIDIA H100. The code is Apache 2.0 but the model weights use an OpenRAIL-M license with commercial restrictions.
- Chandra OCR 2 scores 85.9% on olmOCR vs GPT-4o's 69.9%, while cutting model size from 9B to 4B parameters and hitting ~2 pages/sec on an H100
- Multilingual performance beats Gemini 2.5 Flash and GPT-5 Mini, with 40-46 point gains on scripts like Kannada, Malayalam and Telugu over Chandra 1
- Weights carry an OpenRAIL-M license requiring a paid commercial license once a company exceeds $2M in funding or revenue, despite Apache 2.0 code
- Handwriting recognition remains weak, dropping to ~50.4% accuracy on complex forms despite strong printed-text and table/math handling