#vllm
GitHub repositories that have self-applied the topic "vllm" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #vllm (top 16 by stars)
🏛️ Multi-Agent Debate & Brainstorm Studio | 本地多Agent辩论&头脑风暴 — run Hermes/Codex/Ollama on your machine
GrokBuildMJW/ironcladReliability for LLM agents through enforcement, not model size — Agent-Contract-Kernel + a fail-closed orchestration engine. Model-agnostic. 🇦🇪 Built in the UAE.
OneCielAI/claude-anyClaude Code provider selector for Anthropic, Ollama, Ollama Cloud, vLLM, NVIDIA hosted, and self-hosted NIM
epsilonagentx/intel_arc_gpu_llmDocker Compose stack for serving a local, OpenAI-compatible LLM (vLLM on Intel XPU) on an Intel Arc Pro B60 GPU — reproducible config with operator and developer docs.
awdemos/toks-benchReproducible token-throughput benchmark for OpenAI-compatible LLM servers, tuned for NVIDIA Spark and GB10 inference.
gapilongo/pentest-copilotSelf-hosted pentest copilot. Substrate-first (technique catalog + playbooks + RAG + deterministic tools) with structural verifier rules and LLM-as-judge quality eval. Apache 2.0.
RamazanKara/private-ai-platform-kitLocal-first Kubernetes platform for private LLM and coding-agent workloads.
wpalish/petrel-rag-v5On-Premise RAG (продвинутая версия): vLLM/Ollama + bge-m3 + reranker + Qdrant + hybrid (BM25+RRF) + PDR + Basic Auth + Prometheus/Grafana + локальная оценка. Запускается на Ollama без GPU. Тех-задача Petrel AI (Astana Hub).
manishklach/k3-inference-platformProduction-oriented Kimi K3 inference control plane with checkpoint release gates, MoE capacity planning, OpenAI gateway, NVL72 deployment, benchmarks, and observability.
msradam/xk6-llmLoad test LLM inference servers with k6. TTFT, ITL, TPOT, goodput, cost, and energy metrics for any OpenAI-compatible server. Ships to Prometheus and Grafana.
Hert4/LLM-Certainty-ConsistencyBackend-agnostic black-box hallucination & RAG-faithfulness detection for LLMs — Probabilistic Certainty & Consistency (arXiv:2601.02574). Works on MLX / OpenAI / vLLM via token logprobs; no model internals, no training.
luongnv89/dgx-spark-llm-labBenchmark local coding LLMs on an OpenAI-compatible endpoint, then keep the serving config that won. Hidden executable tests, reference-validated tasks, mermaid reports.
cyberlife-coder/llm-localThin, zero-dependency CLI to run local LLMs on Apple Silicon (vllm-mlx or mlx_lm) with OpenAI- and Anthropic-compatible endpoints — point Claude Code at a local model.
stpcoder/here-context-recallHere — 끊긴 업무의 시작점과 다음 행동을 복원하는 데스크톱 앱 · OpenAI-compatible/vLLM
indiser/CivSimTurn-based geopolitical simulator where 5 AI civilizations — militarist, mercantile, theocratic, democratic, and authoritarian — reason via LLM (Llama 3.3 70B / Groq) to form alliances, wage wars, conduct espionage, and pursue competing victory conditions. FastAPI game engine · Flask frontend · fantasy SVG world map · injectable event cards.
Wenri/TaskSolverProvider-agnostic VLM query flow: one Agent that dispatches to OpenAI / Anthropic / Gemini / vLLM / Claude Code CLI / local HuggingFace model backends, returning parsed answers. Used by 3D-CoT.
RELATED Other topics · full topics ranking →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology