AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#ai-safety

GitHub repositories that have self-applied the topic "ai-safety" — a creator-tagged metadata that surfaces how AI projects describe themselves.

56
tagged repos
51
top 40 stars
6
with tool sigs
40
shown

REPOS Repos for #ai-safety (top 40 by stars)

waitdeadai/llm-dark-patterns

Umbrella for the LLM Dark Patterns Hooks suite — single-purpose Claude Code Stop hooks that suppress sycophancy, paternalism, false-success, permission-loops, training-cutoff confidence at the textual boundary.

Shell 16 AI 100 Solo live ↗
Actenon/actenon-scan

Find where agent-controlled intent reaches consequential actions without an authority check. Python + TypeScript + Go. Zero-dependency SAST for AI agents.

Python 4 AI 70
prat3ik/evalbot

EvalBot — local-first chatbot security & quality evaluation (FastAPI + Next.js). Evaluate chatbot answers against your own docs & guidelines with ML/NLP + AI-judge scoring. Apache-2.0.

Python 3 AI 70
keynv-labs/keynv

Self-host secrets manager with an AI-safety layer. Aliases instead of values; AI agents never see real credentials. Cloud option coming.

TypeScript 3 AI 70 1 sig live ↗
Krishita17/agent-memory-poisoning

Attack taxonomy, toolkit & defenses for persistent-memory poisoning of LLM agents. Author: Krishita Sanjay Choksi.

Python 2 AI 100
germankovacevic-lab/agent-audit-gate

An audit gate for AI agent outbound messaging: hold the draft, let a senior agent review it (no private-data leak, prompt-injection resistance), then release. Reference implementation.

TypeScript 2 AI 100
Bobcatsfan33/Pharos

The trust control plane for enterprise AI agents — real-time policy verdicts in under 800ms and litigation-grade evidence of every decision. Pharos decides. Pharos proves.

TypeScript 2 AI 70
Vadale/project-guardian

AI Guardian Firewall — a local, user-space, agent-agnostic firewall that mediates an autonomous AI agent's actions (files, shell, network, services) with a deterministic policy boundary, a tamper-evident audit log, and a human-in-the-loop approval cockpit. No kernel modules. Apache-2.0.

Rust 2 AI 70
IgorGanapolsky/mac-yolo-safeguards

OS-level safeguards layer for AI agent loops — runaway-kill, timeouts & resource limits so coding agents (Claude Code, Cursor, Codex, Antigravity) in YOLO mode can't burn tokens or freeze your Mac

JavaScript 2 AI 70 Solo 3 sig live ↗
attestplane/attestplane

Verifiable audit substrate for AI agents — EU AI Act Article 12 ready. Apache 2.0.

Python 2 AI 70 1 sig live ↗
github-yjc/context-chronicle

OpenCode memory governance plugin — Knowledge Graph + Hybrid Search + Smart Compaction + Tool Firewall + Stop Gate. 6 MCP tools with slash-command UX. Apache-2.0.

TypeScript 2 AI 45 Solo live ↗
Willbass65/SEAI-Identity-Standard

Sovereign Embedded Artificial Intelligence — A hardware-rooted identity framework for autonomous AI systems. Open standard for AI birth certificates, hardware attestation, lineage tracking, authority scopes, and revocation.

Python 1 AI 100
mvp-imran/AIAgent-SoftwareAgency

AIAgent-SoftwareAgency

Shell 1 AI 100
ngu-gif/genai-role-playbook

GenAI Career Roadmap 2026 🚀 | AI Job Paths & Skills Guide

HTML 1 AI 100
BryceWDesign/IX-BlackFox-WorldTwin

A governed world-model evidence layer for AI agents: simulate bounded scenarios, track assumptions, score prediction-vs-reality error, and produce human-reviewable execution evidence.

Python 1 AI 70 Solo live ↗
gregoryhorn/hermes-loop-engineering

Hermes Agent starter kit for safe scheduled, stateful AI-agent loops

Python 1 AI 70 Solo live ↗
mohamedzhioua/proofguard

Kill-tested guard skills for AI coding agents , self-invoking quality gates that catch AI failure modes in security, tests, docs, dependencies, and diffs before the agent says "done." Works with Claude Code, Codex, and Cursor.

JavaScript 1 AI 70 Solo live ↗
pinalmdave/Switchyard

Detect when Claude silently falls back from Fable 5 to Opus 4.8, log it to a local tamper-evident ledger, and keep your work on the frontier model.

Python 1 AI 70 Solo 1 sig live ↗
RudrenduPaul/evolveguard

Regression-testing CI gate for self-edited Claude Agent Skills (SKILL.md, MEMORY.md) -- golden-transcript record/replay, zero hosted infra.

Python 1 AI 70
codewithEshaYoutube/Serenix

A Privacy-Preserving Edge AI System for Trustworthy Safety Compliance

JavaScript 1 AI 70 Solo live ↗
Argyronix/lorm

LORM — Layered Operational Responsibility Model: L0-L5 autonomy levels graded by authorization source, with a trust promotion/demotion lifecycle. Spec + Claude Code plugin (agent skill + policy-driven enforcement hooks).

Python 1 AI 70 1 sig
castroquiles/glapagos

Global Laboratory for AI Progress and Governance: Open Systems — The open multilateral AI platform for the Americas

HTML 1 AI 70 Solo live ↗
africanmarketos591/mvr-coding-agent-twin

Public controlled beta for the MVR Coding Agent Twin, a market-reality co-processor for AI coding agents.

Python 0 AI 100 Solo 2 sig live ↗
cuiqi5656/agent-quality-benchmark

Evidence-first, reproducible quality benchmarks for AI agents.

Python 0 AI 100
Yacineutt/AI-AgenticSafe

AI AgenticSafe - stop your AI agent from breaking production at 3 a.m. 8 battle-tested doctrines, 3 real post-mortems, zero dependencies. By Yacine Mahboub, Founder of WEVIA.

0 AI 100 Solo live ↗
api-evangelist/luminosai

Luminos.AI is an AI governance and evaluation platform that tests AI systems — classical machine learning, generative AI, and autonomous agents — for legal, regulatory, and reputational risk.

0 AI 100
ss1738/proof-carrying-ai

Proof-carrying compliance certificates for AI agent actions: machine-checked (Coq, axiom-free) + zero-knowledge proofs that an agent action obeyed a formal policy.

Python 0 AI 100
sunnydubey1111/agent-trajectory-sentinel

Real-time detection and repair of LLM agent failures — a one-class behavioural monitor at ~200 µs/step, with 2,823 committed traces.

Python 0 AI 100 Solo live ↗
lekshmiparu23-ai/SafeSearchAI-Risk-Aware-Chatbot

Risk-Aware AI Chatbot — Real-time query classification using 5-Layer Hybrid LLM + Keyword detection. Built with Jetpack Compose, Llama 3.1 8B, Cerebras API & Firebase.

Kotlin 0 AI 100
Vedansh5545/llm-shieldbench

Trustworthy AI evaluation tool for testing chatbot safety, reliability, hallucination behavior, privacy risk, and instruction-following quality.

Python 0 AI 100
ntholm86/agent-context-memory

Specification for governance-first context memory in autonomous AI agent systems. Defines the Mandate Gate: a pre-work mandate must exist before an agent session is valid.

0 AI 100
rodelgithub/local-llm-bridge

Unlock AI Coding with OpenCode Local Provider 2026 - Auto-Detect Ollama LM Studio

HTML 0 AI 100
kalyvask/ai-safety-os

Same agent action, safe or dangerous by context: email team vs investor; delete scratch vs prod config. A safety layer routes each action: auto, confirm, escalate, block. A reading model auto-executes 0% of unsafe actions vs a baseline's 38% (McNemar p<0.001). Plus a runtime loop: an agent earns or loses autonomy with each counterparty over time.

Python 0 AI 100
ramenprotokol/openai-build-week-2026

Train and run evidence-first AI coding agents with human approval gates, isolated patches, and independent verification.

TypeScript 0 AI 90 Solo live ↗
justuseapen/polarity-bench

Does following a negated instruction come free with the positive one? A reproducible eval — two pre-registered nulls for frontier Claude.

Python 0 AI 70
MasonNagel5/MCP-Security-Scanner

Controlled study measuring whether prompt-injection poison hidden in MCP tool descriptions actually changes an LLM's behavior.

Python 0 AI 70
KristopherKubicki/golem-covenant

A covenantal standard for bounded, answerable, revocable AI agents.

HTML 0 AI 70 Solo live ↗
Thomas-LEON/agentguard

🛡️ Experimental security guardrails for LangChain agent code execution (Alpha)

Python 0 AI 70
Jott2121/sabot

Do your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.

Python 0 AI 70 Solo live ↗
AviroopFX/AgentGuard

A trust and observability layer for AI agents — intercepts risky tool calls, enforces approval workflows, and logs a full audit trail before agents can act.

Python 0 AI 70

RELATED Other topics · full topics ranking →

#claude-code

1,555

#ai-agents

1,156

#llm

1,066

#claude

936

#python

802

#ai

737

#developer-tools

723

#mcp

719

#codex

517

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology