#ocr
GitHub repositories that have self-applied the topic "ocr" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #ocr (top 17 by stars)
Typed structured extraction from PDFs, photos & text via any LLM — Instructor for Swift, document-first
korovin-aa97/talkthrough-mcpDon't write a bug report — record it. Local MCP: screen recording → transcript, frames, OCR, wall-clock evidence for coding agents.
amitpatole/agent-visionEyes for AI coding agents 👁️ — render → perceive → report → fix → re-render. A machine-graded visual feedback loop (DOM/contrast/OCR-grounded + optional vision LLM) agents consume to self-correct before claiming done.
sfyyy/dsh-vision-bridgeOn-demand vision for text-only DeepSeek Harness (DSH) sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model
Aevorine/SynoriveOffline semantic search for everything on your disk - documents, code, PDFs, scanned images (OCR), video by the second. 100% local AI: hybrid RAG with keyword + vector + rerank, no API key, no cloud. Desktop app + 24 MCP tools for Claude Code, plus a multi-engine web researcher that quotes sources verbatim. 本地离线多模态语义检索
rohitagr06/mediscanIntelligent Medical Report Analyzer — AI-assisted medical document intelligence platform
FirstCastSolutions423/carrelA library desk for your files — and your agents. Local file toolkit + Claude Code plugin marketplace.
bhaskargurram-ai/verifydocTrust layer for document→JSON extraction & AI agents: calibrated per-field confidence + source grounding + accept/review abstention on any OCR/VLM. Ships VerifyDocBench, a novel grounding-conditioned conformal method, and an MCP server.
yygg693/autocomputerAI-driven desktop GUI automation — Rust core + Python SDK + MCP Server + HTML5 Dashboard. Zero-build Obsidian UI with jelly physics, desktop console, RPA recorder, SQLite audit.
zhangzhangwudi2-creator/aiqiuzhizheAI resume–JD matching assistant with DeepSeek, browser OCR, quota protection, structured output validation, and automated tests.
Rajeshdevandla/ai-document-intelligence-platformEnd-to-end document processing — Spring Boot microservices, Python OCR (AWS Textract), OpenAI GPT-4o extraction, Kafka async pipeline, React dashboard.
renyi9044-png/creator-clone-lab抖音对标采集、创作者蒸馏、AI 分身与 Obsidian 知识图谱 Skill
amajorai/shadowShadow — 14-modality on-device capture (screen/audio/input), OCR, and semantic search for AI agents.
paulookino/processoiaLegal document analysis with AI — ASP.NET Core 8, Azure, OCR, LLM integration
shrey315/hallucination-gateConservative grounding gate for RAG and fine-tuned LLMs — pass, rewrite, or abstain.
SoftChaosByKatie/book-scanner-appA hosted web app that photographs a book cover, extracts title/author/metadata using Claude's vision API, and logs it to a Notion library — built as a reusable "photo → structured data" engine designed to extend beyond books (e.g. small business inventory scanning). Replaces an earlier iPhone Shortcut prototype with a branded, always-on service.
Thedan-1/ScreenQA-ocr-quiz-solver屏幕区域 OCR + 大模型自动答题助手 | Screen-region OCR + LLM quiz solver — 框住题目区域,内容一变自动截图、本地 OCR、大模型作答。支持 DeepSeek 等 OpenAI 兼容接口与多模态视觉模型
RELATED Other topics · full topics ranking →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology