#ai-safety
GitHub repositories that have self-applied the topic "ai-safety" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #ai-safety (top 40 by stars)
Umbrella for the LLM Dark Patterns Hooks suite — single-purpose Claude Code Stop hooks that suppress sycophancy, paternalism, false-success, permission-loops, training-cutoff confidence at the textual boundary.
Actenon/actenon-scanFind where agent-controlled intent reaches consequential actions without an authority check. Python + TypeScript + Go. Zero-dependency SAST for AI agents.
prat3ik/evalbotEvalBot — local-first chatbot security & quality evaluation (FastAPI + Next.js). Evaluate chatbot answers against your own docs & guidelines with ML/NLP + AI-judge scoring. Apache-2.0.
keynv-labs/keynvSelf-host secrets manager with an AI-safety layer. Aliases instead of values; AI agents never see real credentials. Cloud option coming.
Krishita17/agent-memory-poisoningAttack taxonomy, toolkit & defenses for persistent-memory poisoning of LLM agents. Author: Krishita Sanjay Choksi.
germankovacevic-lab/agent-audit-gateAn audit gate for AI agent outbound messaging: hold the draft, let a senior agent review it (no private-data leak, prompt-injection resistance), then release. Reference implementation.
Bobcatsfan33/PharosThe trust control plane for enterprise AI agents — real-time policy verdicts in under 800ms and litigation-grade evidence of every decision. Pharos decides. Pharos proves.
Vadale/project-guardianAI Guardian Firewall — a local, user-space, agent-agnostic firewall that mediates an autonomous AI agent's actions (files, shell, network, services) with a deterministic policy boundary, a tamper-evident audit log, and a human-in-the-loop approval cockpit. No kernel modules. Apache-2.0.
IgorGanapolsky/mac-yolo-safeguardsOS-level safeguards layer for AI agent loops — runaway-kill, timeouts & resource limits so coding agents (Claude Code, Cursor, Codex, Antigravity) in YOLO mode can't burn tokens or freeze your Mac
attestplane/attestplaneVerifiable audit substrate for AI agents — EU AI Act Article 12 ready. Apache 2.0.
github-yjc/context-chronicleOpenCode memory governance plugin — Knowledge Graph + Hybrid Search + Smart Compaction + Tool Firewall + Stop Gate. 6 MCP tools with slash-command UX. Apache-2.0.
Willbass65/SEAI-Identity-StandardSovereign Embedded Artificial Intelligence — A hardware-rooted identity framework for autonomous AI systems. Open standard for AI birth certificates, hardware attestation, lineage tracking, authority scopes, and revocation.
mvp-imran/AIAgent-SoftwareAgencyAIAgent-SoftwareAgency
ngu-gif/genai-role-playbookGenAI Career Roadmap 2026 🚀 | AI Job Paths & Skills Guide
BryceWDesign/IX-BlackFox-WorldTwinA governed world-model evidence layer for AI agents: simulate bounded scenarios, track assumptions, score prediction-vs-reality error, and produce human-reviewable execution evidence.
gregoryhorn/hermes-loop-engineeringHermes Agent starter kit for safe scheduled, stateful AI-agent loops
mohamedzhioua/proofguardKill-tested guard skills for AI coding agents , self-invoking quality gates that catch AI failure modes in security, tests, docs, dependencies, and diffs before the agent says "done." Works with Claude Code, Codex, and Cursor.
pinalmdave/SwitchyardDetect when Claude silently falls back from Fable 5 to Opus 4.8, log it to a local tamper-evident ledger, and keep your work on the frontier model.
RudrenduPaul/evolveguardRegression-testing CI gate for self-edited Claude Agent Skills (SKILL.md, MEMORY.md) -- golden-transcript record/replay, zero hosted infra.
codewithEshaYoutube/SerenixA Privacy-Preserving Edge AI System for Trustworthy Safety Compliance
Argyronix/lormLORM — Layered Operational Responsibility Model: L0-L5 autonomy levels graded by authorization source, with a trust promotion/demotion lifecycle. Spec + Claude Code plugin (agent skill + policy-driven enforcement hooks).
castroquiles/glapagosGlobal Laboratory for AI Progress and Governance: Open Systems — The open multilateral AI platform for the Americas
africanmarketos591/mvr-coding-agent-twinPublic controlled beta for the MVR Coding Agent Twin, a market-reality co-processor for AI coding agents.
cuiqi5656/agent-quality-benchmarkEvidence-first, reproducible quality benchmarks for AI agents.
Yacineutt/AI-AgenticSafeAI AgenticSafe - stop your AI agent from breaking production at 3 a.m. 8 battle-tested doctrines, 3 real post-mortems, zero dependencies. By Yacine Mahboub, Founder of WEVIA.
api-evangelist/luminosaiLuminos.AI is an AI governance and evaluation platform that tests AI systems — classical machine learning, generative AI, and autonomous agents — for legal, regulatory, and reputational risk.
ss1738/proof-carrying-aiProof-carrying compliance certificates for AI agent actions: machine-checked (Coq, axiom-free) + zero-knowledge proofs that an agent action obeyed a formal policy.
sunnydubey1111/agent-trajectory-sentinelReal-time detection and repair of LLM agent failures — a one-class behavioural monitor at ~200 µs/step, with 2,823 committed traces.
lekshmiparu23-ai/SafeSearchAI-Risk-Aware-ChatbotRisk-Aware AI Chatbot — Real-time query classification using 5-Layer Hybrid LLM + Keyword detection. Built with Jetpack Compose, Llama 3.1 8B, Cerebras API & Firebase.
Vedansh5545/llm-shieldbenchTrustworthy AI evaluation tool for testing chatbot safety, reliability, hallucination behavior, privacy risk, and instruction-following quality.
ntholm86/agent-context-memorySpecification for governance-first context memory in autonomous AI agent systems. Defines the Mandate Gate: a pre-work mandate must exist before an agent session is valid.
rodelgithub/local-llm-bridgeUnlock AI Coding with OpenCode Local Provider 2026 - Auto-Detect Ollama LM Studio
kalyvask/ai-safety-osSame agent action, safe or dangerous by context: email team vs investor; delete scratch vs prod config. A safety layer routes each action: auto, confirm, escalate, block. A reading model auto-executes 0% of unsafe actions vs a baseline's 38% (McNemar p<0.001). Plus a runtime loop: an agent earns or loses autonomy with each counterparty over time.
ramenprotokol/openai-build-week-2026Train and run evidence-first AI coding agents with human approval gates, isolated patches, and independent verification.
justuseapen/polarity-benchDoes following a negated instruction come free with the positive one? A reproducible eval — two pre-registered nulls for frontier Claude.
MasonNagel5/MCP-Security-ScannerControlled study measuring whether prompt-injection poison hidden in MCP tool descriptions actually changes an LLM's behavior.
KristopherKubicki/golem-covenantA covenantal standard for bounded, answerable, revocable AI agents.
Thomas-LEON/agentguard🛡️ Experimental security guardrails for LangChain agent code execution (Alpha)
Jott2121/sabotDo your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.
AviroopFX/AgentGuardA trust and observability layer for AI agents — intercepts risky tool calls, enforces approval workflows, and logs a full audit trail before agents can act.
RELATED Other topics · full topics ranking →
#claude-code
1,555#ai-agents
1,156#llm
1,066#claude
936#python
802#ai
737#developer-tools
723#mcp
719#codex
517Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology