#prompt-injection
GitHub repositories that have self-applied the topic "prompt-injection" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #prompt-injection (top 40 by stars)
Security-First Agentic IDE Coding Harness
NovaCode37/claude-security-skillsProduction-ready Claude Code skills for cybersecurity — secret scanning, SAST, prompt-injection testing, HTTP/JWT/dependency auditing. Zero dependencies.
ralfyishere/agent-zero-trustZero-trust repo intake for AI coding agents — scan the instruction environment before Claude Code, Cursor, Codex, or Gemini touches a repo. Ships its own false-negative ledger.
prat3ik/evalbotEvalBot — local-first chatbot security & quality evaluation (FastAPI + Next.js). Evaluate chatbot answers against your own docs & guidelines with ML/NLP + AI-judge scoring. Apache-2.0.
grnbtqdbyx-create/trace-to-skillCodex Issue Radar and maintainer-readiness tooling for AI coding agents.
Krishita17/agent-memory-poisoningAttack taxonomy, toolkit & defenses for persistent-memory poisoning of LLM agents. Author: Krishita Sanjay Choksi.
KhushiTripathi762/Sentinel-AIAI-powered cybersecurity assistant for Prompt Injection and Phishing URL Detection.
KezoSec/rag-poisoning-labA self-contained AI security lab demonstrating document poisoning, indirect prompt injection, and data exfiltration in RAG systems. Explores the "helpfulness paradox" across local and frontier LLMs.
germankovacevic-lab/agent-audit-gateAn audit gate for AI agent outbound messaging: hold the draft, let a senior agent review it (no private-data leak, prompt-injection resistance), then release. Reference implementation.
calionauta/pi-leakguardseatbelt for my pi.dev agent: blocks secret-file access, redacts creds from output, blocks egress + git leaks.
Vadale/project-guardianAI Guardian Firewall — a local, user-space, agent-agnostic firewall that mediates an autonomous AI agent's actions (files, shell, network, services) with a deterministic policy boundary, a tamper-evident audit log, and a human-in-the-loop approval cockpit. No kernel modules. Apache-2.0.
bharath31/nomineeThe authorization layer for AI agents — allow/deny/ask policy, human approvals, and tamper-evident receipts on every tool call. Zero-dep, framework-neutral, no SaaS.
woshilaohei/border-guardBorder Guard — AI-native territory sovereignty & self-evolving security OS. 6-module D-S fusion engine, trajectory detection, anchor detection, fission engine, territory adjudicator. Full design + Python implementation.
Kenny27lokku/prompt-integrity-validatorLint Your Prompts, Ship Better Agents – Prompt Refiner 2026 Rule Engine
Mughal-Baig/local-ai-agentAgentTrail is a local-first AI agent layer for Ollama/local models that shows its work: semantic search, diff-safe edits, receipts, replay, memory, MCP, reports, and trust controls.
mthamil107/ai-bot-shieldDrop-in middleware that detects AI bot traffic. Community-maintained signature database + RFC 9421 Web Bot Auth verification. Node + Python + Go. Apache-2.0.
BrendenKennedy/claude-for-ai-platformsClaude Code scaffold for building AI platforms securely — agent/LLM security, Kubernetes, SRE, observability, identity, and supply chain, grounded in published framework canon (OWASP, NIST, CIS, SLSA). Data-science lanes included.
Edward0l1/skill-flare-discoverBest AI Agent Skill Finder 2026 – Multi-Registry Install & Security Labels
mohamedzhioua/proofguardKill-tested guard skills for AI coding agents , self-invoking quality gates that catch AI failure modes in security, tests, docs, dependencies, and diffs before the agent says "done." Works with Claude Code, Codex, and Cursor.
TAIPANBOX/tokenfuseTokenFuse — runtime control for AI agents: per-run budgets, loop detection, burn forecast, kill-switch. Observability shows the fire; TokenFuse is the automatic extinguisher.
guorunjie/agentic-workflow-guardStatic analysis for AI automation workflows. Find prompt-injection paths, overpowered tools, and write-capable agent jobs before they run.
hamodywe/promptfenceStatic analysis for AI agents in GitHub Actions — finds where attacker-controlled text reaches an agent's prompt, and what that agent is allowed to do with it.
roleplay-sh/ai-agent-social-engineering-researchCurated research on AI agent social engineering, manipulated delegation, prompt injection, tool-use boundaries, and agent security evaluation.
ch040602/agent-prompt-injection-zooSource-backed archive of agent prompt-injection incidents, research records, trust-boundary patterns, schemas, and sanitized defensive summaries.
perpensum/agent-directed-manipulationA reproducible definition separating agent-directed manipulation from legitimate machine-readable self-presentation. Two mechanically decidable axes, 13 conformance cases.
SamsonCyber/llm-injection-field-guideDark Promptery: 324 LLM prompt-injection techniques, crosswalked to OWASP/MITRE ATLAS/NIST/CWE.
Gowrav-M/agent-skillguardPolicy-as-code admission controller for AI agent skills and MCP tools. SkillBOM, lockfiles, and supply-chain baselines.
MAUROCERON/ai-agent-security-mini-auditFree AI-agent risk self-check, security checklist, and USD 59 launch-readiness mini-audit offer.
Vedansh5545/llm-shieldbenchTrustworthy AI evaluation tool for testing chatbot safety, reliability, hallucination behavior, privacy risk, and instruction-following quality.
marion-official/unicode-cleanupClaude Code slash commands that flag Unicode used outside string literals — a common prompt-injection / homoglyph vector
Bobcatsfan33/loomdbAn agent-native database. Sessions are branches an agent can fork, merge, and rewind; every write records what it was derived from; and taint-and-recall tells you exactly what a poisoned input contaminated. Built on substrate.
SamsonCyber/agentic-dm-gatewayHermes-inspired security control plane for LLM agents over private DMs: allowlist, PIN, kill switch, rate limits, injection heuristics, secret redaction, audit log.
rosscyking1115/redteam-foundryLLM red-team evaluation harness — prompt injection, refusal, leakage and staleness, with cross-judge validation and attack-corpus audits.
mobius-style/mobius-browser-guardAn auditable, bounded mediation layer for Claude in Chrome browser actions — content-blind action gating, with its own failure map published.
Eastern-comptrollership272/crucible-Route coding tasks through an automated 8-agent pipeline for optimized, secure, and production-ready code output in Claude Code.
certior/certior-guardA policy hook for Claude Code
Builder106/halberdA JSON-RPC firewall for MCP agents — inspects every tools/call between an LLM and its MCP servers, blocking argument injection and capability creep before they reach the host.
yeodh10/prompt-guardPrompt Injection Guard - a defense-in-depth (rules + LLM) demo that detects and blocks prompt-injection / jailbreak attempts in LLM user input (Streamlit + Claude).
0xsl1m/shadowshieldUnified open-source security shield for agentic AI systems — defense-in-depth prompt-injection protection (canary tokens, agent-trace alignment audit, tool-call guarding, PII/secret scanning).
guangxiangdebizi/tool-output-spoofing-labBenchmarking schema-valid false tool observations and defense baselines for tool-using LLM agents.
RELATED Other topics · full topics ranking →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology