#reproducible-research
GitHub repositories that have self-applied the topic "reproducible-research" — a creator-tagged metadata that surfaces how AI projects describe themselves.
REPOS Repos for #reproducible-research (top 8 by stars)
Codex-native MATLAB scientific plot generation, GPT-5.6 review, controlled repair, and reproducible evidence
sunnydubey1111/agent-trajectory-sentinelReal-time detection and repair of LLM agent failures — a one-class behavioural monitor at ~200 µs/step, with 2,823 committed traces.
oryxintel/oryxflow-claude-pluginTrustworthy, reproducible AI data analysis in Claude Code. A data-science plugin (skill + slash commands) that stops your coding agent building on stale data and records what produced every result.
rosscyking1115/redteam-foundryLLM red-team evaluation harness — prompt injection, refusal, leakage and staleness, with cross-judge validation and attack-corpus audits.
Jott2121/sabotDo your agent pipeline's own checks catch planted faults? Measured on LangGraph, CrewAI and AutoGen: median 16.7%. One prompt-level change takes it to 55.0%. Pre-registered spec, Apache-2.0 harness, every raw trace published.
noamt/kirKeep It Real — a schema-level deferral channel that makes LLM brevity safe. Code, data, and pre-registered protocols: a 400-conversation replication sweep, a placebo control, and a blinded human study (n=46).
satwiksps/scaffoldscopeControlled coding-agent harness ablations with auditable traces, reproducible evidence bundles, and SWE-bench interoperability.
mool32/metric-autopsyRed-team a computed single-cell metric before you believe it: a metric-agnostic gate system (Claude skill + pip package) catching QC/technical/mathematical artifacts
RELATED Other topics · full topics ranking →
#claude-code
1,564#ai-agents
1,160#llm
1,071#claude
943#python
806#ai
741#developer-tools
728#mcp
727#codex
521Aggregated by case-insensitive match against topics_json of each repo's latest content snapshot. methodology