#speech-to-text
GitHub Topic「speech-to-text」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「speech-to-text」と分類されているかを可視化します。
REPOS #speech-to-text のRepo (TOP 31 / Stars降順)
Don't write a bug report — record it. Local MCP: screen recording → transcript, frames, OCR, wall-clock evidence for coding agents.
bluejacketblackhawk/saysomethingFree, local voice dictation for Windows. Hold a key, speak, your words appear at the cursor. On-device Whisper, no cloud, no account, no telemetry.
NadirAliOfficial/voice-to-text🎙️ Free, local, offline voice-to-text for Mac. Hold Ctrl → speak → text pastes anywhere. Powered by OpenAI Whisper. No API key, no internet, no limits.
wzwys9/MurmurnoteAndroid voice note app for recording, ASR transcription, AI summaries, todos, and ideas.
HYPERVAPOR/dysubAn open-source tool to extract subtitles (.srt/.vtt) from Douyin video links via OpenAI-compatible ASR APIs. Supports batch processing and 1-click Docker deployment.
luckeyfaraday/athena-whisperLocal-first AI dictation widget: faster-whisper speech-to-text, PyQt desktop UI, and system-wide voice input for terminals and coding agents. Supports Linux (X11/Wayland) and Windows.
p8n-ai/pi-listensSpeech-first Pi package powered by Sarvam AI
PavelLizunov/suflyorNative Windows meeting overlay with real-time transcription and local-or-cloud AI answers. Pure Rust + Slint.
fcavalcantirj/agent-facesGive your AI agent a talking, lip-syncing face — 12-emotion particle face, in-browser Whisper (WebGPU), bring-your-own-agent bridge. MIT.
StanTheGorilla/GlyphA free, local, open-source alternative to Wispr Flow. Hold a hotkey, talk, and your words are typed into any Windows app — on-device Whisper/Nemotron speech-to-text with optional local LLM cleanup. No cloud, no account.
handnewb/hermes-voiceVoice-enabled personal assistant for Hermes Agent. Open mic with wake word, fully local — no API keys, no accounts, 35 languages.
salvoclemenza-hub/vokariLocal-first voice → knowledge: faster-whisper transcription + Claude/Ollama analysis → briefing.md + recap + Obsidian notes. Audio never leaves your device.
oscar-slg-8/muesliPrivate, on-device meeting assistant for macOS. It records, transcribes, identifies who spoke, and writes a structured summary, entirely on your machine.
amcgiluma/AIComputerVoice-driven local Codex assistant for CachyOS and Omarchy, with Whisper transcription, Piper TTS, Hyprland hotkey integration, and persistent memory.
vivekarya2026/Support-Bot-LLMWhite-label multi-workspace support chatbot with self-hosted voice: talk to your bot, hear it answer — Next.js 15 + local RAG (sqlite-vec) + faster-whisper STT + Kokoro/Coqui TTS, multi-language
glensk/my-stt-ttsLocal voice assistant for macOS Apple Silicon — wake-word → on-device STT → Claude → natural TTS, with speaker ID and DE/FR/EN.
adityakumarprasad/AI-video-AgentAn AI Meeting Assistant that downloads audio/video streams, transcribes speech, extracts key decisions & action items using Google Gemini, and builds a local RAG vector store for interactive chat. Powered by a FastAPI backend and a Next.js frontend dashboard.
danzgabrielgabuat/careless-whisperLocal audio transcription app (58 languages, Filipino/Tagalog included) using OpenAI Whisper, with AI-powered cleanup via Ollama + Qwen 2.5. Built for transcribing interviews — fully offline, no internet required.
BenyaminMahdavifar/VoiceFlow-AgentAn open-source AI workflow platform for speech transcription, LLM-powered processing, and interactive transcript chat using any OpenAI-compatible API.
aryeo0908/whisper-loopReal-time local voice agent (Jarvis-style) with instant barge-in, streaming STT/LLM/TTS, and pluggable AWS Bedrock/Ollama/OpenAI backends
HAMED-PAYANDA/watson-openai-voice-assistantA full-stack web-based voice assistant built with Python and Flask. Integrates OpenAI's GPT for intelligent conversational responses and IBM Watson Speech Libraries for Embed for seamless speech-to-text (STT) and text-to-speech (TTS) processing.
1nspectorCat/Claudio-CodeTalk to your Claude Code sessions by voice, hands-free, from your phone
justanotherkevin/VOAOffline meeting transcription and AI summaries for macOS. Whisper speech-to-text plus a built-in local LLM for decisions and action items. No cloud, no account, no API keys.
voxzerr/blurtFree, fully local voice dictation for macOS. Hold a key, talk, release — cleaned-up text appears at your cursor. No account, no network, no cost.
renyi9044-png/creator-clone-lab抖音对标采集、创作者蒸馏、AI 分身与 Obsidian 知识图谱 Skill
hsergiu/conversLocal voice agent (eng) with barge-in. Three layer memory, browser as mic over LAN, ~1.9 s to first audio on a 4070 Super
dan-ince-aai/aura-voice-bookingBrowser voice agent that books appointments by talking — built on the AssemblyAI Voice Agent API. Temp-token auth, inline agent config, client-side (simulated) tools.
erensavass/Kaptan-ChatAI-powered cross-platform chatbot application built with Flutter, OpenAI API, speech recognition, and RESTful services.
zakariaf/SublyTelegram bot & CLI that transcribes a video/audio file and returns translated subtitles — .srt or burned-in video. Whisper (local or API) for transcription; OpenAI, DeepSeek, or Gemini for translation. Right-to-left aware; deploys with Docker/Kamal.
xr204/game-audio-action-typescriptRoute game backend transcripts to bounded actions with the OpenAI-compatible Infrai client.
elloloop/llmrouterPolyglot Go client for OpenAI, Anthropic, Bedrock, Vertex, Gemini, Cohere + 18 more. Chat, embeddings, TTS, STT, realtime sessions. One OpenAI-shaped API. Build LLM gateways and voice agents.
RELATED 他のTopicも見る · 全Topicランキング →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法