AI開発影響研究所 EN
← Topicランキング · 2026-08
GitHub TOPIC

#multimodal

GitHub Topic「multimodal」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「multimodal」と分類されているかを可視化します。

17
タグ付Repo
1,121
TOP17合計★
6
AIツール痕跡あり
17
TOP表示数

REPOS #multimodal のRepo (TOP 17 / Stars降順)

magicrew/doc7

Turn documents into AI-ready Markdown with visual understanding

Go 1,107 AI 70 1 sig
whitelonng/dsh-plugin-describe-image

DeepSeek Harness plugin: describe_image — give a text-only model vision through an OpenAI-compatible VLM endpoint

TypeScript 5 AI 60
lixiuyin/meeting-agent

A full-stack multimodal meeting intelligence system with layered RAG, long-term memory, and skill-based generation.

Python 3 AI 100 1 sig
sfyyy/dsh-vision-bridge

On-demand vision for text-only DeepSeek Harness (DSH) sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model

JavaScript 2 AI 60 個人 1 sig 公開済 ↗
Aevorine/Synorive

Offline semantic search for everything on your disk - documents, code, PDFs, scanned images (OCR), video by the second. 100% local AI: hybrid RAG with keyword + vector + rerank, no API key, no cloud. Desktop app + 24 MCP tools for Claude Code, plus a multi-engine web researcher that quotes sources verbatim. 本地离线多模态语义检索

Python 1 AI 70 個人 公開済 ↗
2472786266-spec/deepseek-hsrness-devkit

DSH DevKit: multimodal gallery + multi-agent supervision console (DeepSeek Harness dynamic Cordis plugin)

JavaScript 1 AI 70
54xkeee/dsh-youreyes

Eyes for text-only DeepSeek on DeepSeek Harness: model-invokable vision tool + wrapper adapters + general VLM channels (OpenAI-compatible / Gemini / local Ollama)

TypeScript 1 AI 60
Qcxiaoyuksp/vid2summary

🎬 A powerful multimodal AI video summarizer. Supports YouTube/Bilibili/Local videos, smart frame extraction, Whisper transcription, and generates structured summaries via OpenAI/Qwen/Claude. (AI 视频摘要提取神器)

Python 1 AI 60
Agastya51/ScholarAI

Multimodal RAG + Agentic AI research paper analyser — summarisation, vision figure analysis, grounded Q&A, and 4-source literature comparison. Built with LangGraph, FAISS and Groq.

Python 0 AI 100
usagi20/claude-code-vision-skill

Give Claude Code image understanding when its primary model cannot see images. Supports UI screenshots, photos, documents, charts, comparisons, and Windows clipboard input through Qwen.

JavaScript 0 AI 100 1 sig
vijay2411/liteserach-multimodal

🔍 100%-offline-capable multimodal semantic search for macOS. Find your files by meaning, not name — text, code, PDFs, images, audio. Ollama + Qwen3-VL local, Gemini/OpenAI optional. Power-user tool, macOS-only.

Python 0 AI 90
Tatendaz/Vergance

Gaze + voice as a multimodal input layer — look, speak, and your gaze-resolved intent goes to Claude.

Swift 0 AI 70 個人 1 sig 公開済 ↗
Zhenghong-Liu/LocalSight

LocalSight · 198M-A64M 思考型 MoE LLM,2×RTX4090 从零训练(pretrain→SFT→SimPO→RLAIF→Agent RL),GGUF/Ollama 可运行

Python 0 AI 70
xsoc1/dsh-image-vision

Eyes for text-only DeepSeek: view_image tool (local Ollama or any OpenAI-compatible VLM) + chat image-attachment bridge — paste/drop images in the chat and the model can see them.

TypeScript 0 AI 60 個人 公開済 ↗
Thedan-1/ScreenQA-ocr-quiz-solver

屏幕区域 OCR + 大模型自动答题助手 | Screen-region OCR + LLM quiz solver — 框住题目区域,内容一变自动截图、本地 OCR、大模型作答。支持 DeepSeek 等 OpenAI 兼容接口与多模态视觉模型

Python 0 AI 60
Wenri/TaskSolver

Provider-agnostic VLM query flow: one Agent that dispatches to OpenAI / Anthropic / Gemini / vLLM / Claude Code CLI / local HuggingFace model backends, returning parsed answers. Used by 3D-CoT.

Python 0 AI 50 1 sig
Harvey-Will/dsh-vision-analysis

Image understanding for the DeepSeek Harness — analyze_image tool with 8 modes, any OpenAI/Anthropic-compatible vision endpoint

0 AI 30

RELATED 他のTopicも見る · 全Topicランキング →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法