#multimodal
GitHub Topic「multimodal」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「multimodal」と分類されているかを可視化します。
REPOS #multimodal のRepo (TOP 17 / Stars降順)
Turn documents into AI-ready Markdown with visual understanding
whitelonng/dsh-plugin-describe-imageDeepSeek Harness plugin: describe_image — give a text-only model vision through an OpenAI-compatible VLM endpoint
lixiuyin/meeting-agentA full-stack multimodal meeting intelligence system with layered RAG, long-term memory, and skill-based generation.
sfyyy/dsh-vision-bridgeOn-demand vision for text-only DeepSeek Harness (DSH) sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model
Aevorine/SynoriveOffline semantic search for everything on your disk - documents, code, PDFs, scanned images (OCR), video by the second. 100% local AI: hybrid RAG with keyword + vector + rerank, no API key, no cloud. Desktop app + 24 MCP tools for Claude Code, plus a multi-engine web researcher that quotes sources verbatim. 本地离线多模态语义检索
2472786266-spec/deepseek-hsrness-devkitDSH DevKit: multimodal gallery + multi-agent supervision console (DeepSeek Harness dynamic Cordis plugin)
54xkeee/dsh-youreyesEyes for text-only DeepSeek on DeepSeek Harness: model-invokable vision tool + wrapper adapters + general VLM channels (OpenAI-compatible / Gemini / local Ollama)
Qcxiaoyuksp/vid2summary🎬 A powerful multimodal AI video summarizer. Supports YouTube/Bilibili/Local videos, smart frame extraction, Whisper transcription, and generates structured summaries via OpenAI/Qwen/Claude. (AI 视频摘要提取神器)
Agastya51/ScholarAIMultimodal RAG + Agentic AI research paper analyser — summarisation, vision figure analysis, grounded Q&A, and 4-source literature comparison. Built with LangGraph, FAISS and Groq.
usagi20/claude-code-vision-skillGive Claude Code image understanding when its primary model cannot see images. Supports UI screenshots, photos, documents, charts, comparisons, and Windows clipboard input through Qwen.
vijay2411/liteserach-multimodal🔍 100%-offline-capable multimodal semantic search for macOS. Find your files by meaning, not name — text, code, PDFs, images, audio. Ollama + Qwen3-VL local, Gemini/OpenAI optional. Power-user tool, macOS-only.
Tatendaz/VerganceGaze + voice as a multimodal input layer — look, speak, and your gaze-resolved intent goes to Claude.
Zhenghong-Liu/LocalSightLocalSight · 198M-A64M 思考型 MoE LLM,2×RTX4090 从零训练(pretrain→SFT→SimPO→RLAIF→Agent RL),GGUF/Ollama 可运行
xsoc1/dsh-image-visionEyes for text-only DeepSeek: view_image tool (local Ollama or any OpenAI-compatible VLM) + chat image-attachment bridge — paste/drop images in the chat and the model can see them.
Thedan-1/ScreenQA-ocr-quiz-solver屏幕区域 OCR + 大模型自动答题助手 | Screen-region OCR + LLM quiz solver — 框住题目区域,内容一变自动截图、本地 OCR、大模型作答。支持 DeepSeek 等 OpenAI 兼容接口与多模态视觉模型
Wenri/TaskSolverProvider-agnostic VLM query flow: one Agent that dispatches to OpenAI / Anthropic / Gemini / vLLM / Claude Code CLI / local HuggingFace model backends, returning parsed answers. Used by 3D-CoT.
Harvey-Will/dsh-vision-analysisImage understanding for the DeepSeek Harness — analyze_image tool with 8 modes, any OpenAI/Anthropic-compatible vision endpoint
RELATED 他のTopicも見る · 全Topicランキング →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法