#llm-inference
GitHub repositories that have self-applied the topic "llm-inference" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #llm-inference (top 10 by stars)
A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.
RightNow-AI/AutoMegaKernelAn agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682
anthony-chaudhary/fakfak — the Fused Agent Kernel: one Go binary that turns a tool-using agent (Claude Code, Codex, Cursor, any OpenAI/Anthropic/MCP client) into a managed agent: cache-stable model traffic, context compaction + crash resume, nanosecond tool-call policy, local GGUF serving with SSD expert offload.
RedHillsMediaFL/caixcaix — native Apple Core AI inference server for Apple silicon (beta): OpenAI/Anthropic API, dashboard, streaming chat with tools/skills/MCP, MTP speculative decoding.
louisevandan/kvasir-netKvasir — control plane for large GGUF models across GPUs/machines; expert-sharded MoE swarm, non-custodial KVR wallet + mobile nodes, OpenAI/Anthropic gateways.
Zanmmi/micro-agent-cli🏆 Ultralight AI Coding Agent 2026 — Blazing CLI, Tiny Codebase, Full Power Free
daimonionnn/amd-rocmfpx-for-winROCmFPX llama.cpp fork for Windows 🏆 — native build, headless OpenAI-compatible server & benchmarks. Tested on AMD Strix Halo (gfx1151), runs on other GPUs too.
epsilonagentx/intel_arc_gpu_llmDocker Compose stack for serving a local, OpenAI-compatible LLM (vLLM on Intel XPU) on an Intel Arc Pro B60 GPU — reproducible config with operator and developer docs.
manishklach/k3-inference-platformProduction-oriented Kimi K3 inference control plane with checkpoint release gates, MoE capacity planning, OpenAI gateway, NVL72 deployment, benchmarks, and observability.
yempik-ai/airgapRun Qwen3.8-27B fully offline on an Apple Silicon Mac (MLX, 4/5/8-bit) and use it as the local backend for Claude Code. Uncensored/abliterated build supported. No API key, no network. Built by yempik.
RELATED Other topics · full topics ranking →
#claude-code
1,555#ai-agents
1,156#llm
1,066#claude
936#python
802#ai
737#developer-tools
723#mcp
719#codex
517Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology