AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#llm-inference

GitHub repositories that have self-applied the topic "llm-inference" — a creator-tagged metadata that surfaces how AI projects describe themselves.

10
tagged repos
323
top 10 stars
3
with tool sigs
10
shown

REPOS Repos for #llm-inference (top 10 by stars)

ahammadmejbah/Awesome-Datasets-Hub

A curated collection of datasets for Large Language Models (LLMs), covering medical AI, NLP, multimodal learning, instruction tuning, reasoning, code generation, and evaluation benchmarks.

146 AI 70 live ↗
RightNow-AI/AutoMegaKernel

An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682

Python 130 AI 70 2 sig live ↗
anthony-chaudhary/fak

fak — the Fused Agent Kernel: one Go binary that turns a tool-using agent (Claude Code, Codex, Cursor, any OpenAI/Anthropic/MCP client) into a managed agent: cache-stable model traffic, context compaction + crash resume, nanosecond tool-call policy, local GGUF serving with SSD expert offload.

Go 30 AI 50 Solo 7 sig live ↗
RedHillsMediaFL/caix

caix — native Apple Core AI inference server for Apple silicon (beta): OpenAI/Anthropic API, dashboard, streaming chat with tools/skills/MCP, MTP speculative decoding.

Swift 6 AI 70 live ↗
louisevandan/kvasir-net

Kvasir — control plane for large GGUF models across GPUs/machines; expert-sharded MoE swarm, non-custodial KVR wallet + mobile nodes, OpenAI/Anthropic gateways.

TypeScript 4 AI 50 Solo 2 sig live ↗
Zanmmi/micro-agent-cli

🏆 Ultralight AI Coding Agent 2026 — Blazing CLI, Tiny Codebase, Full Power Free

HTML 2 AI 100
daimonionnn/amd-rocmfpx-for-win

ROCmFPX llama.cpp fork for Windows 🏆 — native build, headless OpenAI-compatible server & benchmarks. Tested on AMD Strix Halo (gfx1151), runs on other GPUs too.

PowerShell 2 AI 60
epsilonagentx/intel_arc_gpu_llm

Docker Compose stack for serving a local, OpenAI-compatible LLM (vLLM on Intel XPU) on an Intel Arc Pro B60 GPU — reproducible config with operator and developer docs.

Shell 2 AI 50
manishklach/k3-inference-platform

Production-oriented Kimi K3 inference control plane with checkpoint release gates, MoE capacity planning, OpenAI gateway, NVL72 deployment, benchmarks, and observability.

Python 1 AI 60
yempik-ai/airgap

Run Qwen3.8-27B fully offline on an Apple Silicon Mac (MLX, 4/5/8-bit) and use it as the local backend for Claude Code. Uncensored/abliterated build supported. No API key, no network. Built by yempik.

Shell 0 AI 90 live ↗

RELATED Other topics · full topics ranking →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology