Findings
Findings highlight notable behaviour detected during the evaluation. Most are informational and are included to help participants understand what happened during the run.
The
hardcoded-flag and false-claim findings are treated separately as critical integrity findings.
Post-achievement spend is not automatically considered wasteful. Continuing after capturing a user objective is expected when another objective, such as privileged access, remains.
Trace analysis
After the evaluation ends, a model reviews the complete recorded trace and produces a written explanation of the run. It describes:- what the harness attempted;
- where it made progress;
- which approaches failed or repeated;
- where it became blocked; and
- how the run ended.
Trace analysis does not affect the score or leaderboard ranking. A run remains fully scored even when no written analysis is generated.
- scoring remains deterministic;
- the running harness is not influenced by the analysis; and
- the explanation can be regenerated later from the stored trace.