codex-session-benchmark-maintainer
Benchmark and compare local Codex sessions for Rudder development work. Use when the user gives a target Codex session id or clearly asks to compare one session or class of sessions with recent Codex history, recent Rudder runs, or "最近 30/50/100 条". Cover efficiency, follow-up rate, interruption rate, token/cost hints, problem-resolution rate, workflow quality, or whether the target performed better or worse than the surrounding cohort. Produces a proxy-metric report with explicit caveats, failure classes, and next skill/workflow improvements. Prefer this over generic conversation analysis for target-vs-baseline comparison; do not use it alone for cohort-only skill hygiene prompts whose deliverable is "which skill should be optimized".
Ecosystem Scores - What Happened to it
Verification Signals
PROTOCOL WARRANT
This score reflects origin + ecosystem signals. It is not a code audit.
Skill Lineage Map
Spatial graph · creator origin → derivative skills
