AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#inference

GitHub repositories that have self-applied the topic "inference" — a creator-tagged metadata that surfaces how AI projects describe themselves.

10
tagged repos
11
top 10 stars
3
with tool sigs
10
shown

REPOS Repos for #inference (top 10 by stars)

k1n0F/vramsuite

Predictive GPU memory framework for AI inference workflows

Python 4 AI 70
Ar9av/mlx-lm-server

Local MLX tuned models through OpenAI compatible LLM, image generation and audio inference on Apple Silicon — Rust + PyO3 + MLX

Rust 3 AI 50 1 sig
modelmeld/modelmeld

OpenAI- and Anthropic-compatible AI gateway. Capability-based routing. Streaming. BYOK passthrough. No key custody.

Python 2 AI 50 1 sig live ↗
manishklach/k3-inference-platform

Production-oriented Kimi K3 inference control plane with checkpoint release gates, MoE capacity planning, OpenAI gateway, NVL72 deployment, benchmarks, and observability.

Python 1 AI 60
cfregly/gpu-perf-tune

GPU profiling and optimization SKILLS with bundled MCP server

Python 1 AI 45 2 sig
msradam/xk6-llm

Load test LLM inference servers with k6. TTFT, ITL, TPOT, goodput, cost, and energy metrics for any OpenAI-compatible server. Ships to Prometheus and Grafana.

Go 0 AI 90 live ↗
bystray/gonka-mcp-server

MCP server for the GONKA network: run cheap LLM inference through the server (free trial or your own key), get multi-model second opinions, and compare live prices — for any AI agent.

Python 0 AI 70 Solo live ↗
SAGARCHRY0777/inferno

Production-grade distributed ML inference platform — FastAPI gateway, Redis-backed worker pool with dynamic batching, WebSocket result streaming, and a live ops dashboard. Serves YOLO, Whisper, RAG and text models, with an MCP agent server and streaming chat.

TypeScript 0 AI 70
joshuaswarren/sovereign-inference

Provider-neutral access & supply layer for open AI: run open models locally (SIN) and route paid, private, verifiable inference across decentralized providers (SIP-AI). DecentralizeAI hackathon entry.

Python 0 AI 70
iqureshi123/tinyinfer

An LLM inference engine written from scratch in Python and C++/Metal — tokenizer, forward pass, KV cache, and INT4 quantization implemented by hand. No PyTorch, no llama.cpp.

Python 0 AI 70

RELATED Other topics · full topics ranking →

#claude-code

1,564

#ai-agents

1,160

#llm

1,071

#claude

943

#python

806

#ai

741

#developer-tools

728

#mcp

727

#codex

521

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology