#llm-security
GitHub repositories that have self-applied the topic "llm-security" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #llm-security (top 32 by stars)
Security-First Agentic IDE Coding Harness
NovaCode37/claude-security-skillsProduction-ready Claude Code skills for cybersecurity — secret scanning, SAST, prompt-injection testing, HTTP/JWT/dependency auditing. Zero dependencies.
Actenon/actenon-scanFind where agent-controlled intent reaches consequential actions without an authority check. Python + TypeScript + Go. Zero-dependency SAST for AI agents.
keynv-labs/keynvSelf-host secrets manager with an AI-safety layer. Aliases instead of values; AI agents never see real credentials. Cloud option coming.
prat3ik/evalbotEvalBot — local-first chatbot security & quality evaluation (FastAPI + Next.js). Evaluate chatbot answers against your own docs & guidelines with ML/NLP + AI-judge scoring. Apache-2.0.
clay-good/proxilionProxilion is the security layer for the agentic workforce. It turns managed AI agents into governed users by enforcing strict cryptographic boundaries on every API call to SaaS like Google Workspace, Salesforce, or Atlassian.
Krishita17/agent-memory-poisoningAttack taxonomy, toolkit & defenses for persistent-memory poisoning of LLM agents. Author: Krishita Sanjay Choksi.
germankovacevic-lab/agent-audit-gateAn audit gate for AI agent outbound messaging: hold the draft, let a senior agent review it (no private-data leak, prompt-injection resistance), then release. Reference implementation.
KezoSec/rag-poisoning-labA self-contained AI security lab demonstrating document poisoning, indirect prompt injection, and data exfiltration in RAG systems. Explores the "helpfulness paradox" across local and frontier LLMs.
Vadale/project-guardianAI Guardian Firewall — a local, user-space, agent-agnostic firewall that mediates an autonomous AI agent's actions (files, shell, network, services) with a deterministic policy boundary, a tamper-evident audit log, and a human-in-the-loop approval cockpit. No kernel modules. Apache-2.0.
woshilaohei/border-guardBorder Guard — AI-native territory sovereignty & self-evolving security OS. 6-module D-S fusion engine, trajectory detection, anchor detection, fission engine, territory adjudicator. Full design + Python implementation.
mthamil107/ai-bot-shieldDrop-in middleware that detects AI bot traffic. Community-maintained signature database + RFC 9421 Web Bot Auth verification. Node + Python + Go. Apache-2.0.
BrendenKennedy/claude-for-ai-platformsClaude Code scaffold for building AI platforms securely — agent/LLM security, Kubernetes, SRE, observability, identity, and supply chain, grounded in published framework canon (OWASP, NIST, CIS, SLSA). Data-science lanes included.
Niki-1337/proxy-aiOpen-source AI Security Gateway that sanitizes secrets, PII, and internal context before prompts reach external LLMs.
HDHNezherParking-cum-Y638-Intl-Ltd/titanA disciplined 10-stage autonomous agent pipeline (Reflexion, Tree Search, ReAct) for Claude Code and AI coding agents. Lifts local LLMs to senior engineering quality.
guorunjie/agentic-workflow-guardStatic analysis for AI automation workflows. Find prompt-injection paths, overpowered tools, and write-capable agent jobs before they run.
hamodywe/promptfenceStatic analysis for AI agents in GitHub Actions — finds where attacker-controlled text reaches an agent's prompt, and what that agent is allowed to do with it.
MAUROCERON/ai-agent-security-mini-auditFree AI-agent risk self-check, security checklist, and USD 59 launch-readiness mini-audit offer.
SamsonCyber/llm-injection-field-guideDark Promptery: 324 LLM prompt-injection techniques, crosswalked to OWASP/MITRE ATLAS/NIST/CWE.
perpensum/agent-directed-manipulationA reproducible definition separating agent-directed manipulation from legitimate machine-readable self-presentation. Two mechanically decidable axes, 13 conformance cases.
roleplay-sh/ai-agent-social-engineering-researchCurated research on AI agent social engineering, manipulated delegation, prompt injection, tool-use boundaries, and agent security evaluation.
Gowrav-M/agent-skillguardPolicy-as-code admission controller for AI agent skills and MCP tools. SkillBOM, lockfiles, and supply-chain baselines.
ch040602/agent-prompt-injection-zooSource-backed archive of agent prompt-injection incidents, research records, trust-boundary patterns, schemas, and sanitized defensive summaries.
Reeflex-io/reeflexA seatbelt for the AI acting on your systems — deterministic, open-source governance gate for AI-agent actions
Thomas-LEON/agentguard🛡️ Experimental security guardrails for LangChain agent code execution (Alpha)
Builder106/halberdA JSON-RPC firewall for MCP agents — inspects every tools/call between an LLM and its MCP servers, blocking argument injection and capability creep before they reach the host.
MasonNagel5/MCP-Security-ScannerControlled study measuring whether prompt-injection poison hidden in MCP tool descriptions actually changes an LLM's behavior.
0xsl1m/shadowshieldUnified open-source security shield for agentic AI systems — defense-in-depth prompt-injection protection (canary tokens, agent-trace alignment audit, tool-call guarding, PII/secret scanning).
SamsonCyber/agentic-dm-gatewayHermes-inspired security control plane for LLM agents over private DMs: allowlist, PIN, kill switch, rate limits, injection heuristics, secret redaction, audit log.
Mithun-veerabuthiran/SSS-Secure-Shielding-ServicePrivacy-first Chrome extension and Flask backend for protecting AI prompts through real-time PII detection, anonymization, pseudonymization, and redaction before sensitive data reaches AI services.
rbardyla-boop/claude_powerplantTrust-bounded acceptance harness for Claude coding agents — sanitized workspaces, typed tools, isolated oracle evaluation, evidence receipts.
Matik103/sanctum-runtimeOpen-source trust layer for autonomous AI — gate agent, robot, smart home, and industrial actions before they run. Policies, HITL, Ollama/OpenAI, audit. MIT. npm @sanctum-runtime/sdk
RELATED Other topics · full topics ranking →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology