Click any tag below to further narrow down your results
Links
Xberg is a single engine for detecting, reading, OCR’ing, and extracting text, tables, metadata, and structured data from over 100 document formats and 115 file extensions. It offers transcription, embeddings, layout reconstruction, schema-driven JSON extraction, and code intelligence via language bindings in Rust, Python, Go, Java, and more. You can run it as a library, CLI, REST API, or MCP server with configurable Cargo features.
Mistral OCR 4 extracts text from PDFs, DOCs and more while also returning bounding boxes, block types and per-word confidence. It supports 170 languages, runs in a single container for self-hosted deployments, and outperforms rivals on human and automated benchmarks.
Unlimited-OCR is an open-source OCR framework from Baidu that extends Deepseek-OCR to handle single images, multi-page documents, and PDFs. It provides Hugging Face transformer models, PyMuPDF conversion, and an OpenAI-compatible SGLang server for streaming or batch inference. The repo includes setup instructions, example scripts, and configuration for GPU-accelerated parsing.
This daily roundup highlights new AI tools, models, and research—from Mistral’s OCR 4 and ByteDance’s Seedance 2.5 video generator to Anthropic’s Claude Tag and IBM’s CUGA agent harness. It also covers security deep dives on indirect prompt injection, industry moves like OpenAI’s bidirectional voice and US‐Meta AI reviews, plus several open-source releases.
Datalab’s 4 billion-parameter Chandra OCR 2 outperforms GPT-4o and Gemini across independent and multilingual benchmarks, handling complex layouts, math notation, flowcharts and 90 languages with state-of-the-art accuracy. It’s available under Apache 2.0 code with a modified OpenRAIL-M license for weights, runs locally via HuggingFace or vLLM, and doubles throughput over its predecessor.
Chandra OCR 2, a 4 billion-parameter model from Datalab, outperforms GPT-4o and Gemini on AllenAI’s olmOCR benchmark and a 90-language test while halving the model size. It preserves layout, reads complex tables and math notation, converts diagrams to Mermaid, and runs at two pages per second on an NVIDIA H100. The code is Apache 2.0 but the model weights use an OpenRAIL-M license with commercial restrictions.
A new open-source OCR model outperformed all major commercial tools on standard text and handwriting tests. It accurately transcribed a 1913 handwritten letter by Ramanujan, preserving layout, math notation, and faint ink details.
A powerful CLI tool and browser extension that generates fast summaries from URLs, files, and media, including YouTube videos and podcasts. It features a Chrome Side Panel and Firefox Sidebar, supports various media types, and provides advanced functionalities like OCR and transcript extraction. The tool can be installed via npm or Homebrew, with options for local and paid model endpoints.