#interpretability
GitHub repositories that have self-applied the topic "interpretability" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #interpretability (top 3 by stars)
Does an LLM's willingness to blackmail depend on its persona — and on whether that persona is prompted or trained into the weights? Two-part study on Gemma-3-12B: 16 prompted personas (n=100) + 10 QLoRA persona adapters, in the Lynch et al. shutdown scenario.
leo-im/entropy-lensToken-level entropy trajectories from LLM logprobs. Models can measure their own uncertainty — grounding it in truth requires external verification.
ursadropsus/znouSupporting files for Anoikis², Case Studies in GPT-2 Small/EVE Frontier (Gamified Mechanistic Interpretability)
RELATED Other topics · full topics ranking →
#claude-code
1,576#ai-agents
1,183#llm
1,086#claude
949#python
829#ai
746#mcp
741#developer-tools
739#codex
525Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology