#interpretability
GitHub Topic「interpretability」がついているAI関連リポジトリの集計。Topicはリポジトリ作者が自己申告するメタタグで、AI関連の文脈で何が「interpretability」と分類されているかを可視化します。
REPOS #interpretability のRepo (TOP 3 / Stars降順)
Does an LLM's willingness to blackmail depend on its persona — and on whether that persona is prompted or trained into the weights? Two-part study on Gemma-3-12B: 16 prompted personas (n=100) + 10 QLoRA persona adapters, in the Lynch et al. shutdown scenario.
leo-im/entropy-lensToken-level entropy trajectories from LLM logprobs. Models can measure their own uncertainty — grounding it in truth requires external verification.
ursadropsus/znouSupporting files for Anoikis², Case Studies in GPT-2 Small/EVE Frontier (Gamified Mechanistic Interpretability)
RELATED 他のTopicも見る · 全Topicランキング →
#claude-code
1,576#ai-agents
1,183#llm
1,086#claude
949#python
829#ai
746#mcp
741#developer-tools
739#codex
525集計対象: 各Repoの最新contentスナップショットの topics_json に小文字一致でマッチしたもの。 算出方法