
Chandra
OCR model that turns images and PDFs into Markdown, HTML or JSON while keeping layout. Handles tables, math, handwriting and forms across 90+ languages.
8 entries tagged with "ocr"

OCR model that turns images and PDFs into Markdown, HTML or JSON while keeping layout. Handles tables, math, handwriting and forms across 90+ languages.

Native macOS screenshot and screen recording app in the vein of CleanShot X. Annotation, OCR, webcam recording, a video editor and capture history built in.

Open-source OCR toolkit that turns images and PDFs into structured, LLM-ready data (Markdown, JSON). Handles 100+ languages, tables, formulas and charts.

Python OCR engine for receipts: raw text extraction via Tesseract and structured data via LLMs (OpenAI, Gemini, Groq). Ships with a CLI and FastAPI.

PyTorch inference engine for vision-language models: detection, segmentation and OCR driven by natural language queries, with paged KV cache and CUDA graphs.

Microsoft's Python library and CLI to convert PDF, Word, Excel, images, audio and more into clean Markdown, ready for LLMs and text analysis pipelines.

Self-hosted open-source PDF toolkit with 50+ built-in tools. Merge, split, OCR, sign, convert and compress PDFs with REST API and workflow automation.

JavaScript OCR library that runs in the browser and Node.js. Recognizes text in 100+ languages via WebAssembly with no server-side processing needed.