AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#vllm

GitHub repositories that have self-applied the topic "vllm" — a creator-tagged metadata that surfaces how AI projects describe themselves.

16
tagged repos
22
top 16 stars
3
with tool sigs
16
shown

REPOS Repos for #vllm (top 16 by stars)

ding7015869-alt/agent-meeting-studio

🏛️ Multi-Agent Debate & Brainstorm Studio | 本地多Agent辩论&头脑风暴 — run Hermes/Codex/Ollama on your machine

JavaScript 8 AI 100
GrokBuildMJW/ironclad

Reliability for LLM agents through enforcement, not model size — Agent-Contract-Kernel + a fail-closed orchestration engine. Model-agnostic. 🇦🇪 Built in the UAE.

Python 4 AI 70 Solo 1 sig live ↗
OneCielAI/claude-any

Claude Code provider selector for Anthropic, Ollama, Ollama Cloud, vLLM, NVIDIA hosted, and self-hosted NIM

Python 2 AI 80
epsilonagentx/intel_arc_gpu_llm

Docker Compose stack for serving a local, OpenAI-compatible LLM (vLLM on Intel XPU) on an Intel Arc Pro B60 GPU — reproducible config with operator and developer docs.

Shell 2 AI 50
awdemos/toks-bench

Reproducible token-throughput benchmark for OpenAI-compatible LLM servers, tuned for NVIDIA Spark and GB10 inference.

Python 2 AI 50
gapilongo/pentest-copilot

Self-hosted pentest copilot. Substrate-first (technique catalog + playbooks + RAG + deterministic tools) with structural verifier rules and LLM-as-judge quality eval. Apache 2.0.

Python 1 AI 100 Solo live ↗
RamazanKara/private-ai-platform-kit

Local-first Kubernetes platform for private LLM and coding-agent workloads.

Python 1 AI 100 Solo live ↗
wpalish/petrel-rag-v5

On-Premise RAG (продвинутая версия): vLLM/Ollama + bge-m3 + reranker + Qdrant + hybrid (BM25+RRF) + PDR + Basic Auth + Prometheus/Grafana + локальная оценка. Запускается на Ollama без GPU. Тех-задача Petrel AI (Astana Hub).

Python 1 AI 100
manishklach/k3-inference-platform

Production-oriented Kimi K3 inference control plane with checkpoint release gates, MoE capacity planning, OpenAI gateway, NVL72 deployment, benchmarks, and observability.

Python 1 AI 60
msradam/xk6-llm

Load test LLM inference servers with k6. TTFT, ITL, TPOT, goodput, cost, and energy metrics for any OpenAI-compatible server. Ships to Prometheus and Grafana.

Go 0 AI 90 live ↗
Hert4/LLM-Certainty-Consistency

Backend-agnostic black-box hallucination & RAG-faithfulness detection for LLMs — Probabilistic Certainty & Consistency (arXiv:2601.02574). Works on MLX / OpenAI / vLLM via token logprobs; no model internals, no training.

Python 0 AI 90
luongnv89/dgx-spark-llm-lab

Benchmark local coding LLMs on an OpenAI-compatible endpoint, then keep the serving config that won. Hidden executable tests, reference-validated tasks, mermaid reports.

Python 0 AI 90
cyberlife-coder/llm-local

Thin, zero-dependency CLI to run local LLMs on Apple Silicon (vllm-mlx or mlx_lm) with OpenAI- and Anthropic-compatible endpoints — point Claude Code at a local model.

Python 0 AI 70 2 sig
stpcoder/here-context-recall

Here — 끊긴 업무의 시작점과 다음 행동을 복원하는 데스크톱 앱 · OpenAI-compatible/vLLM

TypeScript 0 AI 60 Solo live ↗
indiser/CivSim

Turn-based geopolitical simulator where 5 AI civilizations — militarist, mercantile, theocratic, democratic, and authoritarian — reason via LLM (Llama 3.3 70B / Groq) to form alliances, wage wars, conduct espionage, and pursue competing victory conditions. FastAPI game engine · Flask frontend · fantasy SVG world map · injectable event cards.

JavaScript 0 AI 50
Wenri/TaskSolver

Provider-agnostic VLM query flow: one Agent that dispatches to OpenAI / Anthropic / Gemini / vLLM / Claude Code CLI / local HuggingFace model backends, returning parsed answers. Used by 3D-CoT.

Python 0 AI 50 1 sig

RELATED Other topics · full topics ranking →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology