vLLM
vLLM is tracked as part of LLM providers, measuring its mention/usage across GitHub AI repos.
REPOS Repos using / mentioning vLLM (Top 27 by stars)
The AI-native JupyterLab — open-source Cursor for notebooks. Cmd+K inline edit, multi-step agent with cell-level tools (read/edit/run), chat with @cell/@file context, ghost-text completion, one-click traceback fix. BYO model: Anthropic, OpenAI, Gemini, Ollama, vLLM. Local-first, privacy-first, fully open source.
ding7015869-alt/agent-meeting-studio🏛️ Multi-Agent Debate & Brainstorm Studio | 本地多Agent辩论&头脑风暴 — run Hermes/Codex/Ollama on your machine
GrokBuildMJW/ironcladReliability for LLM agents through enforcement, not model size — Agent-Contract-Kernel + a fail-closed orchestration engine. Model-agnostic. 🇦🇪 Built in the UAE.
awdemos/toks-benchReproducible token-throughput benchmark for OpenAI-compatible LLM servers, tuned for NVIDIA Spark and GB10 inference.
epsilonagentx/intel_arc_gpu_llmDocker Compose stack for serving a local, OpenAI-compatible LLM (vLLM on Intel XPU) on an Intel Arc Pro B60 GPU — reproducible config with operator and developer docs.
OneCielAI/claude-anyClaude Code provider selector for Anthropic, Ollama, Ollama Cloud, vLLM, NVIDIA hosted, and self-hosted NIM
dosmoon/aistack faizan007jr/local-llm-delegateClaude Code plugin that delegates simple, well-scoped tasks to a locally running LLM (Ollama, LM Studio, llama.cpp, vLLM) via MCP
RamazanKara/private-ai-platform-kitLocal-first Kubernetes platform for private LLM and coding-agent workloads.
wpalish/petrel-rag-v5On-Premise RAG (продвинутая версия): vLLM/Ollama + bge-m3 + reranker + Qdrant + hybrid (BM25+RRF) + PDR + Basic Auth + Prometheus/Grafana + локальная оценка. Запускается на Ollama без GPU. Тех-задача Petrel AI (Astana Hub).
manishklach/k3-inference-platformProduction-oriented Kimi K3 inference control plane with checkpoint release gates, MoE capacity planning, OpenAI gateway, NVL72 deployment, benchmarks, and observability.
gapilongo/pentest-copilotSelf-hosted pentest copilot. Substrate-first (technique catalog + playbooks + RAG + deterministic tools) with structural verifier rules and LLM-as-judge quality eval. Apache 2.0.
yourself-q/local-browser-agentLocal-first browser agent — attaches to existing Chrome via CDP, runs fully on local LLMs (LM Studio, Ollama, vLLM). Loop detection, multi-action chaining, data-agent-ref grounding.
Deep-AI-Evo/qwen3.8-27b-q6k-fp8-rtx-pro5000-serving-benchmarkQwen3.8-27B serving benchmark on RTX PRO 5000: llama.cpp Q6_K vs vLLM FP8/NVFP4 (TTFT/prefill/decode/concurrency)
msradam/xk6-llmLoad test LLM inference servers with k6. TTFT, ITL, TPOT, goodput, cost, and energy metrics for any OpenAI-compatible server. Ships to Prometheus and Grafana.
Hert4/LLM-Certainty-ConsistencyBackend-agnostic black-box hallucination & RAG-faithfulness detection for LLMs — Probabilistic Certainty & Consistency (arXiv:2601.02574). Works on MLX / OpenAI / vLLM via token logprobs; no model internals, no training.
eagle-42/askableAgent-Ops is a technical exploration project focused on instrumentation, evaluation, and GPU serving of LLM agents using OpenTelemetry, Langfuse, RAGAS, and vLLM. It targets generative AI Tech Lead roles in banking and aims to build a robust, long-term LLMOps positioning.
indiser/CivSimTurn-based geopolitical simulator where 5 AI civilizations — militarist, mercantile, theocratic, democratic, and authoritarian — reason via LLM (Llama 3.3 70B / Groq) to form alliances, wage wars, conduct espionage, and pursue competing victory conditions. FastAPI game engine · Flask frontend · fantasy SVG world map · injectable event cards.
Akshitha024/multi-tenant-llm-routerFastAPI multi-LoRA router on top of vLLM: per-tenant auth, rate limits, LRU adapter cache, cost accounting
cyberlife-coder/llm-localThin, zero-dependency CLI to run local LLMs on Apple Silicon (vllm-mlx or mlx_lm) with OpenAI- and Anthropic-compatible endpoints — point Claude Code at a local model.
monthop-gmail/llm-gatewayOpenAI-compatible LLM gateway — LiteLLM + Open WebUI ออก API token เองได้ ต่อ HuggingFace / vLLM / Ollama / cloud providers
stpcoder/here-context-recallHere — 끊긴 업무의 시작점과 다음 행동을 복원하는 데스크톱 앱 · OpenAI-compatible/vLLM
kevinbtalbert/Claude-Workbench-with-CAI-InferenceClaude Workbench using CAI Inference Service vllm hosted models
VAKEELRAKESH/agentos-amd-ai-platform luongnv89/dgx-spark-llm-labBenchmark local coding LLMs on an OpenAI-compatible endpoint, then keep the serving config that won. Hidden executable tests, reference-validated tasks, mermaid reports.
Wenri/TaskSolverProvider-agnostic VLM query flow: one Agent that dispatches to OpenAI / Anthropic / Gemini / vLLM / Claude Code CLI / local HuggingFace model backends, returning parsed answers. Used by 3D-CoT.
Bazaarlinkorg/bazaarlink-byoc-agentBazaarLink BYOC Agent — connect your own GPU (Ollama, LM Studio, vLLM, llama.cpp) to your own BazaarLink account. MIT.
A "mention" means the repo description, topics, or README summary contains the string "vLLM".
RELATED Other LLM providers · See full LLM providers ranking →
OpenAI
1,124Gemini
1,014Anthropic
686Ollama
384Groq
371DeepSeek
329Mistral
86Hugging Face
56Bedrock
44EXPLORE Other categories
AI coding tools
Programming languages
GitHub Topics
AI frameworks
Web/App frameworks
Cloud/hosting platforms
Auth services
Vector databases
General databases
LLM models
Embedding models
Agent frameworks
Aggregated from latest content snapshots of AI-relevant repos (AI score ≥ 40). methodology