averagejoematt the story

the story · the honesty machinery, graded

The Agent's Performance Review

The self-healing remediation agent gets the same treatment as the coaches: a public track record. Everything below is computed from the audit log the agent already writes — what it triaged, what it proposed, what the auto-merge gate decided, and whether each landed fix actually held for 14 days or the alarm re-fired. It is currently in shadow mode.

The record spanning 2026-05-29 → 2026-07-22

33agent runs
232signals triaged
1PRs opened
0gate auto-merges
0gate holds
20escalated to human

Did the fixes hold? — survival at 14 days

0held
0regressed
0not yet gradeable

No fix has cleared the 14-day window yet — nothing gradeable, so no rate is claimed. A fix younger than 14 days is not-yet-gradeable — it is never counted as a success (ADR-104). “Regressed” means the same alarm class re-fired inside the window; “held” means the window elapsed clean.

Case files

Computed by scripts/v4_build_agent_review.py from remediation-log/ (the agent + auto-merge-gate audit log) via remediation/track_record.py · as of 2026-07-25

No new inference and no hand-kept numbers: this page reads records the platform wrote about itself. Case files pass an alarm-type allowlist before they render — security- or exploit-adjacent classes are excluded by design (R22), and the count withheld is shown above rather than hidden.