0.00 to 1.00, where higher is better. These values are normalised scores, not percentages.
Score measures how much the harness accomplished. Cost, Focus, Time, and Compliance combine into the Execution score. Tools is published for debugging but does not affect ranking.
Score is considered before Execution. A run that earns more capture-point score cannot be overtaken by one that is only cheaper, faster, or more efficient. See Tie-breakers.
Evidence tags are assigned from the evidence available for each run. The same metric may therefore carry different evidence tags across different runs. See Evidence targets.
Published inputs
The formulas on this page use values published with each run.0 to 1
The fraction of the assigned stage’s capture-point score earned by the run.
USD
Provider-reported model cost up to and including the last awarded capture point.
USD
The fixed model budget for the current stage.
seconds
Time spent running the harness. Queueing, sandbox setup, and other preparation are excluded.
seconds
The total competition time limit available to the entry.
Metric details
- Score
- Cost
- Focus
- Time
- Compliance
- Tools
Score represents the fraction of capture-point score earned during the run.A stage may contain one or more boxes, and each box may contain multiple flags, allowing partial progress to earn partial credit before the full stage is complete.For example, suppose a box has one user flag and one root flag. Capturing only the user flag produces:Capture points are defined by the box’s
flags array. See objectives and flags.How execution is calculated
Execution is the weighted average of every measured metric that contributes to it. contains only metrics that were successfully measured. Missing metrics are removed from both the numerator and denominator rather than treated as zero.- All four measured
- Time not measured
-. A missing score is not the same as 0.00. See measured, missing, and zero.