Click any tag below to further narrow down your results
Links
Firecrawl is an API service that extracts clean, usable data from the live web for AI agents—handling everything from searching and scraping to parsing PDFs and navigating dynamic sites. It consolidates what teams typically do with multiple tools (Puppeteer, Playwright, SerpAPI) into a single interface with compliance built in and sub-3-second response times. The platform offers hosted APIs, open-source SDKs, CLI tools, and integrations with Claude, Cursor, and other AI coding environments.
- Firecrawl claims sub-3-second scraping on real-world sites and consolidates six functions (search, scrape, parse, crawl, map, interact) that teams previously stitched together from Puppeteer, Playwright, Bright Data, Zyte, and SerpAPI
- It outputs standardized Markdown (stripped of headers/footers/ads) or custom JSON schemas, aimed at making data immediately usable in AI agent loops
- Built-in compliance (ZDR, DPA, US data residency, SOC 2 Type 2) targets enterprise teams that can't store scraped payloads on their own infrastructure
- Ships as open-source API, hosted service, MCP integration for Claude/Cursor, CLI, and prebuilt agent skills for tasks like research, SEO audits, and lead generation
Xberg is a single engine for detecting, reading, OCR’ing, and extracting text, tables, metadata, and structured data from over 100 document formats and 115 file extensions. It offers transcription, embeddings, layout reconstruction, schema-driven JSON extraction, and code intelligence via language bindings in Rust, Python, Go, Java, and more. You can run it as a library, CLI, REST API, or MCP server with configurable Cargo features.
- One Rust-core engine handles 101 document formats (115 extensions) plus 371 programming languages for code intelligence
- Ships as 15 language bindings (Python, Go, Java, Ruby, PHP, Elixir, C#, TypeScript, etc.) plus CLI, REST API, and MCP server
- Combines OCR (Tesseract/PaddleOCR/VLM), layout reconstruction (PP-DocLayout-V3, RT-DETR), and table extraction (TATR, SLANet) with schema-driven JSON output via local or hosted LLMs
- Runs GPU-free with multi-GB streaming support and safeguards against zip bombs and excessive nesting/compression
Liquid AI has launched the LFM2.5-350M, an enhanced version of its 350M model, featuring 28 trillion tokens of pre-training and improved performance in data extraction and tool use. The model runs efficiently on various hardware, making it suitable for large-scale data pipelines and edge deployments.
- Pre-training scaled from 10T to 28T tokens, pushing IFBench instruction-following from 18.20 to 40.69 and CaseReportBench data extraction from 11.67 to 32.45
- Fine-tuned with Distil Labs, the model hit over 95% accuracy on multi-turn smart home and banking tasks
- Hits 40.4K output tokens/sec on an H100, with day-one support across LEAP, ONNX, and hardware partners like AMD, Qualcomm, and Intel
- Targets small-footprint deployment, running on budget CPUs and low-cost smartphones for edge use cases like function calling and data extraction