AI Dev Impact Lab JA
← Topics ranking · 2026-08
GITHUB TOPIC

#interpretability

GitHub repositories that have self-applied the topic "interpretability" — a creator-tagged metadata that surfaces how AI projects describe themselves.

3
tagged repos
0
top 3 stars
0
with tool sigs
3
shown

REPOS Repos for #interpretability (top 3 by stars)

leo-t-1/persona-agentic-misalignment

Does an LLM's willingness to blackmail depend on its persona — and on whether that persona is prompted or trained into the weights? Two-part study on Gemma-3-12B: 16 prompted personas (n=100) + 10 QLoRA persona adapters, in the Lynch et al. shutdown scenario.

HTML 0 AI 70 Solo live ↗
leo-im/entropy-lens

Token-level entropy trajectories from LLM logprobs. Models can measure their own uncertainty — grounding it in truth requires external verification.

Python 0 AI 70
ursadropsus/znou

Supporting files for Anoikis², Case Studies in GPT-2 Small/EVE Frontier (Gamified Mechanistic Interpretability)

Python 0 AI 45

RELATED Other topics · full topics ranking →

#claude-code

1,576

#ai-agents

1,183

#llm

1,086

#claude

949

#python

829

#ai

746

#mcp

741

#developer-tools

739

#codex

525

Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology