qwen3-4b-instruct — calibration report

2026-09-05 12:22 · pipeline 2026.09.05 · seed ? · 108 probe generations · 80 repair attempts

Scope: this file. Qwen_Qwen3-4B-Instruct-2507-Q5_K_L.gguf

Everything below was measured on this model, at this size, at this quantization. Quantization changes the logit distribution and every signal here is a property of that distribution — a different quant of the same weights is a different organism and needs its own run. None of it is claimed to transfer to another model.

Entropy signal
usable
p=0.000
Effect seen
d=-0.94
detectable: d≥0.72
Failures
20
of 108 generations
Recoverable
11 / 20
by any strategy tested

The rule

Implementable as written. Every number below was measured on this model, on this run.

  1. Discard and regenerate if plateau_start_quartile > -0.5. This combination catches 73% of failing generations at the cost of flagging 17% of correct ones — balanced accuracy 0.783 against a chance of 0.50. Measured by rebuilding the whole combination inside 5 cross-validation folds and scoring it on records each fold never saw, then checking it beats 40/40 label-shuffled runs (p=0.0244).
  2. On a failed generation, repair in this order: test_retry → temp_retry → skeleton_fill. This ordering is greedy set cover over what actually recovered failures in this run, not a preference — test_retry alone recovered 7 of 20.
  3. When the failure is a logic error, route to test_retry first (3/6, 50%), ahead of temp_retry 33%, skeleton_fill 17%.
  4. When the failure is an assertion error, route to test_retry first (4/14, 29%), ahead of temp_retry 29%, skeleton_fill 21%.
  5. Stop after the ordered chain. 11 of 20 failures were recoverable by any strategy tested; the remaining 9 were not recovered by any of them, so further retries on those are spend without evidence behind it.

Where to cut, and what it costs

One signal, one curve. What changes is where you cut it and what you do with each side. Fitted and scored across cross-validation folds on 108 unique generations (18 failures) — each number is what the rule did on records its own fold never saw. Pick one row; they are three settings of the same dial, not three rules to stack.

Catch wrong — regenerate when plateau_start_quartile > -0.5. Chosen in 4 of 5 folds; only readable once the generation is complete, so the tokens are already spent when it fires.
Failures caught
67%
of all failing generations
Correct work lost
17%
regenerated needlessly
Net tokens
-54,084
per 1000 generations (-27.8%)

Per 1000 generations: 111.2 doomed runs aborted early, 139.2 correct ones thrown away and regenerated. Token counts are this run's own.

Trust right — ship untested when surprisal_max ≤ 1.1439. Chosen in 2 of 5 folds; only readable once the generation is complete, so the tokens are already spent when it fires.
Output covered
46%
shipped without testing
Purity
91.6%
of what ships is correct
Defects shipped
38.6
per 1000 generations

This saves test execution, not generation tokens — the generation is already paid for by the time any signal can be read. Worth it when the test suite is expensive or latency matters more than compute. The defect count is the price.

Both — one cut point, not two, on surprisal_max. Chosen in 2 of 5 folds; only readable once the generation is complete, so the tokens are already spent when it fires.

The fit put both cut points at the same value, so there is no undecided band: every generation lands in trust or in bail. That makes this row the same single threshold as the two above, read from both sides — not a third, richer rule.

ZoneConditionWhat to doShare of output
Trustsurprisal_max ≤ 1.1439ship untested — 91.6% correct45%
Bandbetween the twotest it — the signal cannot decide0%
Bailsurprisal_max > 1.1439regenerate — catches 68% of all failures55%

Net -124,612 tokens per 1000 generations from the bail zone (-64.1% of baseline).

Ship the trust zone untested, regenerate the bail zone, test the band between them. Both cut points come from one fit on one signal, ordered so the zones cannot overlap. The band is what the signal cannot decide — the honest residue, not a gap in the analysis.

No operating point here pays for itself in tokens. The best of them still nets -54,084 tokens per 1000 generations. That is not a flaw in the rule — this run's winning signal is only readable once a generation is complete, so aborting saves nothing that was not already spent, while every wrongly flagged generation pays for a full replacement. Use these rules to spend less on testing, or to decide what to ship, not to spend fewer tokens.

What separates correct from incorrect: plateau_start_quartile

CorrectIncorrect
Every probe generation, placed by its plateau_start_quartile. Vertical marks are group means.
-1 0 1 2 3 plateau_start_quartile Correct n=100 plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = -1.0000 (correct) plateau_start_quartile = 0.0000 (correct) plateau_start_quartile = 0.0000 (correct) plateau_start_quartile = 0.0000 (correct) plateau_start_quartile = 0.0000 (correct) plateau_start_quartile = 0.0000 (correct) plateau_start_quartile = 1.0000 (correct) plateau_start_quartile = 1.0000 (correct) plateau_start_quartile = 1.0000 (correct) plateau_start_quartile = 2.0000 (correct) plateau_start_quartile = 2.0000 (correct) plateau_start_quartile = 2.0000 (correct) plateau_start_quartile = 3.0000 (correct) plateau_start_quartile = 3.0000 (correct) plateau_start_quartile = 3.0000 (correct) plateau_start_quartile = 3.0000 (correct) mean -0.640 Incorrect n=20 plateau_start_quartile = -1.0000 (incorrect) plateau_start_quartile = -1.0000 (incorrect) plateau_start_quartile = -1.0000 (incorrect) plateau_start_quartile = -1.0000 (incorrect) plateau_start_quartile = -1.0000 (incorrect) plateau_start_quartile = -1.0000 (incorrect) plateau_start_quartile = -1.0000 (incorrect) plateau_start_quartile = 0.0000 (incorrect) plateau_start_quartile = 0.0000 (incorrect) plateau_start_quartile = 0.0000 (incorrect) plateau_start_quartile = 0.0000 (incorrect) plateau_start_quartile = 1.0000 (incorrect) plateau_start_quartile = 1.0000 (incorrect) plateau_start_quartile = 1.0000 (incorrect) plateau_start_quartile = 2.0000 (incorrect) plateau_start_quartile = 2.0000 (incorrect) plateau_start_quartile = 2.0000 (incorrect) plateau_start_quartile = 2.0000 (incorrect) plateau_start_quartile = 2.0000 (incorrect) plateau_start_quartile = 2.0000 (incorrect) mean 0.400

Correct generations averaged plateau_start_quartile -0.640; incorrect ones 0.400. The groups are shifted, not merely overlapping — this is the separation the rule above is built on.

A permutation test puts that difference at p=0.0001, against an alpha of 0.05. Reshuffling which generations were labelled correct reproduces a gap this large routinely.

Across 158 candidate statistics the largest effect was plateau_start_quartile at d=-1.015 (high=bad), and it survives correction.

11 of 20 failures were recoverable by at least one repair strategy. Recovery is measured directly — break it, repair it, re-run the tests — so these counts do not depend on any of the statistical argument above.

Power: signal detected. Observed effect d=-0.94; this run could detect d>=0.72.

Every candidate signal, corrected

All 158 numeric statistics were tested against the same pass/fail outcome, so raw p-values are corrected with Benjamini-Hochberg FDR. A statistic is called usable only if its adjusted p clears 0.05 and |d| ≥ 0.2. The shaded band is what this run was too small to detect at all.

too small for this run to detect (|d| < 0.72) -1 -0.5 0 0.5 1 ← higher on incorrect Cohen's d higher on correct → plateau_start_quartile plateau_start_quartile: d=-1.015, p=0.0002, adjusted p=0.0105 — usable -1.01 usable kl_uniform_min kl_uniform_min: d=+0.938, p=0.0001, adjusted p=0.0079 — usable +0.94 usable max_entropy max_entropy: d=-0.938, p=0.0001, adjusted p=0.0079 — usable -0.94 usable full_ent_max full_ent_max: d=-0.908, p=0.0005, adjusted p=0.0198 — usable -0.91 usable q3 q3: d=-0.885, p=0.0016, adjusted p=0.0506 — not usable -0.89 surprisal_early_vs_rest surprisal_early_vs_rest: d=+0.805, p=0.0022, adjusted p=0.0579 — not usable +0.81 kl_cent_std kl_cent_std: d=+0.789, p=0.0034, adjusted p=0.0672 — not usable +0.79 surprisal_max surprisal_max: d=-0.767, p=0.0030, adjusted p=0.0672 — not usable -0.77 mass5_min mass5_min: d=+0.745, p=0.0069, adjusted p=0.0869 — not usable +0.74 rank_std rank_std: d=-0.736, p=0.0042, adjusted p=0.0737 — not usable -0.74 rank_max rank_max: d=-0.731, p=0.0075, adjusted p=0.0869 — not usable -0.73 mass1_early_vs_rest mass1_early_vs_rest: d=-0.708, p=0.0060, adjusted p=0.0869 — not usable -0.71 margin_early_vs_rest margin_early_vs_rest: d=-0.688, p=0.0077, adjusted p=0.0869 — not usable -0.69 surprisal_std surprisal_std: d=-0.680, p=0.0094, adjusted p=0.0928 — not usable -0.68 kl_uniform_std kl_uniform_std: d=-0.678, p=0.0087, adjusted p=0.0916 — not usable -0.68 full_ent_std full_ent_std: d=-0.666, p=0.0104, adjusted p=0.0967 — not usable -0.67 rank_early_vs_rest rank_early_vs_rest: d=+0.656, p=0.0116, adjusted p=0.1018 — not usable +0.66 early_vs_rest early_vs_rest: d=+0.631, p=0.0142, adjusted p=0.1122 — not usable +0.63 kl_uniform_early_vs_rest kl_uniform_early_vs_rest: d=-0.631, p=0.0142, adjusted p=0.1122 — not usable -0.63 full_ent_early_vs_rest full_ent_early_vs_rest: d=+0.625, p=0.0150, adjusted p=0.1129 — not usable +0.62 mass5_late_minus_early mass5_late_minus_early: d=+0.616, p=0.0077, adjusted p=0.0869 — not usable +0.62 q_range q_range: d=-0.604, p=0.0253, adjusted p=0.1438 — not usable -0.60 mass1_std mass1_std: d=-0.578, p=0.0255, adjusted p=0.1438 — not usable -0.58 full_ent_late_minus_early full_ent_late_minus_early: d=-0.577, p=0.0293, adjusted p=0.1438 — not usable -0.58 late_vs_early_half late_vs_early_half: d=-0.576, p=0.0295, adjusted p=0.1438 — not usable -0.58 kl_uniform_late_minus_early kl_uniform_late_minus_early: d=+0.576, p=0.0296, adjusted p=0.1438 — not usable +0.58 kl_step_late_minus_early kl_step_late_minus_early: d=+0.572, p=0.0291, adjusted p=0.1438 — not usable +0.57 mass5_std mass5_std: d=-0.570, p=0.0310, adjusted p=0.1438 — not usable -0.57 total_variation total_variation: d=-0.569, p=0.0271, adjusted p=0.1438 — not usable -0.57 kl_cent_mean kl_cent_mean: d=+0.552, p=0.0331, adjusted p=0.1438 — not usable +0.55 first_token_entropy first_token_entropy: d=-0.550, p=0.0335, adjusted p=0.1438 — not usable -0.55 full_ent_first full_ent_first: d=-0.550, p=0.0335, adjusted p=0.1438 — not usable -0.55 kl_uniform_first kl_uniform_first: d=+0.550, p=0.0335, adjusted p=0.1438 — not usable +0.55 mass1_min mass1_min: d=+0.546, p=0.0350, adjusted p=0.1438 — not usable +0.55 mass1_first mass1_first: d=+0.544, p=0.0350, adjusted p=0.1438 — not usable +0.54 surprisal_first surprisal_first: d=-0.544, p=0.0350, adjusted p=0.1438 — not usable -0.54 phi_first phi_first: d=+0.544, p=0.0351, adjusted p=0.1438 — not usable +0.54 margin_first margin_first: d=+0.543, p=0.0352, adjusted p=0.1438 — not usable +0.54 margin_std margin_std: d=-0.540, p=0.0373, adjusted p=0.1473 — not usable -0.54 kl_step_early_vs_rest kl_step_early_vs_rest: d=-0.538, p=0.0355, adjusted p=0.1438 — not usable -0.54 n_spikes n_spikes: d=+0.537, p=0.0439, adjusted p=0.1576 — not usable +0.54 n_semantic n_semantic: d=-0.523, p=0.0393, adjusted p=0.1478 — not usable -0.52 n_tokens n_tokens: d=-0.523, p=0.0393, adjusted p=0.1478 — not usable -0.52 osc_sign_changes osc_sign_changes: d=-0.520, p=0.0434, adjusted p=0.1576 — not usable -0.52 rank_mean rank_mean: d=-0.507, p=0.0464, adjusted p=0.1629 — not usable -0.51 mass1_late_minus_early mass1_late_minus_early: d=+0.500, p=0.0557, adjusted p=0.1806 — not usable +0.50 surprisal_mean surprisal_mean: d=-0.487, p=0.0560, adjusted p=0.1806 — not usable -0.49 kl_uniform_mean kl_uniform_mean: d=+0.486, p=0.0552, adjusted p=0.1806 — not usable +0.49 mean_entropy mean_entropy: d=-0.486, p=0.0552, adjusted p=0.1806 — not usable -0.49 full_ent_mean full_ent_mean: d=-0.480, p=0.0584, adjusted p=0.1845 — not usable -0.48 margin_mean margin_mean: d=+0.474, p=0.0654, adjusted p=0.2006 — not usable +0.47 phi_mean phi_mean: d=+0.469, p=0.0667, adjusted p=0.2006 — not usable +0.47 mass1_mean mass1_mean: d=+0.468, p=0.0673, adjusted p=0.2006 — not usable +0.47 mono_down_frac mono_down_frac: d=-0.463, p=0.0832, adjusted p=0.2347 — not usable -0.46 margin_late_minus_early margin_late_minus_early: d=+0.454, p=0.0802, adjusted p=0.2336 — not usable +0.45 surprisal_late_minus_early surprisal_late_minus_early: d=-0.449, p=0.0874, adjusted p=0.2423 — not usable -0.45 surprisal_f10_max surprisal_f10_max: d=+0.443, p=0.0813, adjusted p=0.2336 — not usable +0.44 kl_step_mean kl_step_mean: d=+0.433, p=0.0893, adjusted p=0.2433 — not usable +0.43 q_curvature q_curvature: d=+0.423, p=0.0957, adjusted p=0.2563 — not usable +0.42 q_monotone_down q_monotone_down: d=-0.421, p=0.1446, adjusted p=0.3218 — not usable -0.42 first_to_max_ratio first_to_max_ratio: d=+0.415, p=0.1073, adjusted p=0.2826 — not usable +0.41 surprisal_f10_mean surprisal_f10_mean: d=+0.408, p=0.1129, adjusted p=0.2924 — not usable +0.41 kl_step_max kl_step_max: d=+0.404, p=0.1152, adjusted p=0.2936 — not usable +0.40 kl_step_std kl_step_std: d=-0.390, p=0.1314, adjusted p=0.3169 — not usable -0.39 surprisal_min surprisal_min: d=-0.387, p=0.1343, adjusted p=0.3169 — not usable -0.39 mass5_f10_max mass5_f10_max: d=+0.385, p=0.1404, adjusted p=0.3169 — not usable +0.39 first10_max first10_max: d=+0.381, p=0.1392, adjusted p=0.3169 — not usable +0.38 full_ent_f10_max full_ent_f10_max: d=+0.381, p=0.1392, adjusted p=0.3169 — not usable +0.38 mass1_max mass1_max: d=+0.379, p=0.1389, adjusted p=0.3169 — not usable +0.38 kl_uniform_max kl_uniform_max: d=+0.378, p=0.1392, adjusted p=0.3169 — not usable +0.38 full_ent_min full_ent_min: d=-0.377, p=0.1395, adjusted p=0.3169 — not usable -0.38 first10_std first10_std: d=+0.364, p=0.1585, adjusted p=0.3478 — not usable +0.36 mass20_late_minus_early mass20_late_minus_early: d=+0.361, p=0.1816, adjusted p=0.3726 — not usable +0.36 tail_mass_late_minus_early tail_mass_late_minus_early: d=-0.361, p=0.1816, adjusted p=0.3726 — not usable -0.36 margin_min margin_min: d=+0.359, p=0.1697, adjusted p=0.3673 — not usable +0.36 mass1_f10_mean mass1_f10_mean: d=-0.345, p=0.1816, adjusted p=0.3726 — not usable -0.34 phi_first10_mean phi_first10_mean: d=-0.345, p=0.1816, adjusted p=0.3726 — not usable -0.34 rank_f10_max rank_f10_max: d=+0.339, p=0.3446, adjusted p=0.5760 — not usable +0.34 rank_f10_mean rank_f10_mean: d=+0.339, p=0.3446, adjusted p=0.5760 — not usable +0.34 margin_f10_mean margin_f10_mean: d=-0.337, p=0.1926, adjusted p=0.3852 — not usable -0.34 mass20_f10_max mass20_f10_max: d=+0.322, p=0.2292, adjusted p=0.4527 — not usable +0.32 rank_f5_mean rank_f5_mean: d=+0.316, p=0.3536, adjusted p=0.5760 — not usable +0.32 margin_max margin_max: d=+0.306, p=0.2452, adjusted p=0.4783 — not usable +0.31 trend_rho trend_rho: d=+0.299, p=0.2489, adjusted p=0.4786 — not usable +0.30 tail_entropy tail_entropy: d=-0.295, p=0.2605, adjusted p=0.4786 — not usable -0.29 head_entropy head_entropy: d=-0.293, p=0.2585, adjusted p=0.4786 — not usable -0.29 q1 q1: d=-0.293, p=0.2589, adjusted p=0.4786 — not usable -0.29 mass5_early_vs_rest mass5_early_vs_rest: d=-0.287, p=0.1859, adjusted p=0.3766 — not usable -0.29 kl_step_f10_mean kl_step_f10_mean: d=-0.277, p=0.2718, adjusted p=0.4936 — not usable -0.28 surprisal_f5_mean surprisal_f5_mean: d=+0.272, p=0.3080, adjusted p=0.5468 — not usable +0.27 q4_over_mean q4_over_mean: d=+0.270, p=0.2972, adjusted p=0.5336 — not usable +0.27 kl_cent_slope kl_cent_slope: d=+0.261, p=0.3241, adjusted p=0.5636 — not usable +0.26 q4 q4: d=-0.261, p=0.3246, adjusted p=0.5636 — not usable -0.26 mass5_mean mass5_mean: d=+0.259, p=0.2597, adjusted p=0.4786 — not usable +0.26 kl_cent_f10_mean kl_cent_f10_mean: d=+0.249, p=0.3475, adjusted p=0.5760 — not usable +0.25 rank_late_minus_early rank_late_minus_early: d=-0.242, p=0.3537, adjusted p=0.5760 — not usable -0.24 first10_mean first10_mean: d=+0.240, p=0.3608, adjusted p=0.5760 — not usable +0.24 kl_uniform_f10_mean kl_uniform_f10_mean: d=-0.240, p=0.3608, adjusted p=0.5760 — not usable -0.24 full_ent_f10_mean full_ent_f10_mean: d=+0.240, p=0.3609, adjusted p=0.5760 — not usable +0.24 kl_step_slope kl_step_slope: d=-0.232, p=0.3944, adjusted p=0.6232 — not usable -0.23 q_argmax q_argmax: d=+0.209, p=0.4867, adjusted p=0.7539 — not usable +0.21 mass5_f10_mean mass5_f10_mean: d=-0.206, p=0.4443, adjusted p=0.6950 — not usable -0.21 tail_mass_f10_max tail_mass_f10_max: d=+0.206, p=0.5879, adjusted p=0.7939 — not usable +0.21 mass20_f10_mean mass20_f10_mean: d=-0.204, p=0.5840, adjusted p=0.7939 — not usable -0.20 tail_mass_f10_mean tail_mass_f10_mean: d=+0.204, p=0.5840, adjusted p=0.7939 — not usable +0.20 mass20_slope mass20_slope: d=-0.201, p=0.6879, adjusted p=0.8558 — not usable -0.20 tail_mass_slope tail_mass_slope: d=+0.201, p=0.6879, adjusted p=0.8558 — not usable +0.20 early_slope early_slope: d=+0.188, p=0.5269, adjusted p=0.7707 — not usable +0.19 full_ent_slope full_ent_slope: d=+0.188, p=0.5269, adjusted p=0.7707 — not usable +0.19 kl_uniform_slope kl_uniform_slope: d=-0.188, p=0.5269, adjusted p=0.7707 — not usable -0.19 kl_step_min kl_step_min: d=-0.182, p=0.5317, adjusted p=0.7707 — not usable -0.18 spike_quartile spike_quartile: d=+0.181, p=0.5581, adjusted p=0.7739 — not usable +0.18 mass5_f5_mean mass5_f5_mean: d=-0.181, p=0.5584, adjusted p=0.7739 — not usable -0.18 mass1_f5_mean mass1_f5_mean: d=-0.180, p=0.4927, adjusted p=0.7558 — not usable -0.18 mass1_slope mass1_slope: d=-0.176, p=0.5525, adjusted p=0.7739 — not usable -0.18 margin_f5_mean margin_f5_mean: d=-0.174, p=0.5083, adjusted p=0.7707 — not usable -0.17 margin_slope margin_slope: d=-0.168, p=0.5565, adjusted p=0.7739 — not usable -0.17 spike_relative_pos spike_relative_pos: d=+0.167, p=0.5210, adjusted p=0.7707 — not usable +0.17 rank_slope rank_slope: d=-0.165, p=0.6249, adjusted p=0.8027 — not usable -0.17 mass5_slope mass5_slope: d=-0.156, p=0.8115, adjusted p=0.9158 — not usable -0.16 q1_over_mean q1_over_mean: d=+0.152, p=0.5551, adjusted p=0.7739 — not usable +0.15 kl_cent_f10_max kl_cent_f10_max: d=-0.152, p=0.6035, adjusted p=0.8023 — not usable -0.15 mass1_f10_max mass1_f10_max: d=-0.150, p=0.6143, adjusted p=0.8023 — not usable -0.15 margin_f10_max margin_f10_max: d=-0.149, p=0.6133, adjusted p=0.8023 — not usable -0.15 kl_uniform_f10_max kl_uniform_f10_max: d=-0.147, p=0.6195, adjusted p=0.8023 — not usable -0.15 kl_step_f10_max kl_step_f10_max: d=-0.146, p=0.6194, adjusted p=0.8023 — not usable -0.15 surprisal_slope surprisal_slope: d=+0.128, p=0.6534, adjusted p=0.8259 — not usable +0.13 mass5_first mass5_first: d=-0.125, p=0.6364, adjusted p=0.8109 — not usable -0.12 mass5_max mass5_max: d=+0.115, p=1.0000, adjusted p=1.0000 — not usable +0.12 mass20_std mass20_std: d=-0.107, p=0.7826, adjusted p=0.8960 — not usable -0.11 tail_mass_std tail_mass_std: d=-0.107, p=0.7826, adjusted p=0.8960 — not usable -0.11 first3_mean first3_mean: d=+0.098, p=0.7064, adjusted p=0.8567 — not usable +0.10 tail_mass_first tail_mass_first: d=-0.097, p=0.7095, adjusted p=0.8567 — not usable -0.10 mass20_first mass20_first: d=+0.097, p=0.7103, adjusted p=0.8567 — not usable +0.10 q_early_drop q_early_drop: d=-0.094, p=0.6939, adjusted p=0.8565 — not usable -0.09 kl_cent_f5_mean kl_cent_f5_mean: d=-0.080, p=0.7612, adjusted p=0.8940 — not usable -0.08 kl_step_f5_mean kl_step_f5_mean: d=-0.077, p=0.7694, adjusted p=0.8940 — not usable -0.08 first5_mean first5_mean: d=+0.077, p=0.7695, adjusted p=0.8940 — not usable +0.08 full_ent_f5_mean full_ent_f5_mean: d=+0.077, p=0.7695, adjusted p=0.8940 — not usable +0.08 kl_uniform_f5_mean kl_uniform_f5_mean: d=-0.077, p=0.7695, adjusted p=0.8940 — not usable -0.08 mass20_min mass20_min: d=+0.073, p=0.8650, adjusted p=0.9330 — not usable +0.07 tail_mass_max tail_mass_max: d=-0.073, p=0.8650, adjusted p=0.9330 — not usable -0.07 spike_token_idx spike_token_idx: d=-0.064, p=0.8053, adjusted p=0.9154 — not usable -0.06 kl_cent_max kl_cent_max: d=-0.064, p=0.8565, adjusted p=0.9330 — not usable -0.06 q2 q2: d=-0.057, p=0.8312, adjusted p=0.9314 — not usable -0.06 mass20_early_vs_rest mass20_early_vs_rest: d=-0.056, p=0.8848, adjusted p=0.9330 — not usable -0.06 tail_mass_early_vs_rest tail_mass_early_vs_rest: d=+0.056, p=0.8848, adjusted p=0.9330 — not usable +0.06 tail_mass_min tail_mass_min: d=+0.052, p=1.0000, adjusted p=1.0000 — not usable +0.05 mass20_mean mass20_mean: d=+0.049, p=0.8931, adjusted p=0.9330 — not usable +0.05 tail_mass_mean tail_mass_mean: d=-0.049, p=0.8931, adjusted p=0.9330 — not usable -0.05 mass20_f5_mean mass20_f5_mean: d=-0.038, p=0.8970, adjusted p=0.9330 — not usable -0.04 tail_mass_f5_mean tail_mass_f5_mean: d=+0.038, p=0.8970, adjusted p=0.9330 — not usable +0.04 kl_cent_late_minus_early kl_cent_late_minus_early: d=-0.037, p=0.8918, adjusted p=0.9330 — not usable -0.04 spike_height spike_height: d=-0.035, p=0.8976, adjusted p=0.9330 — not usable -0.04 q_slope q_slope: d=-0.029, p=0.9092, adjusted p=0.9389 — not usable -0.03 path_ratio path_ratio: d=+0.023, p=0.9335, adjusted p=0.9577 — not usable +0.02 kl_cent_early_vs_rest kl_cent_early_vs_rest: d=-0.020, p=0.9403, adjusted p=0.9585 — not usable -0.02 osc_rate osc_rate: d=-0.015, p=0.9537, adjusted p=0.9659 — not usable -0.01
Table view — all signals
SignalDirectiondpp adjustedUsable
kl_uniform_minhigh=good+0.9380.00010.0079yes
max_entropyhigh=bad-0.9380.00010.0079yes
plateau_start_quartilehigh=bad-1.0150.00020.0105yes
full_ent_maxhigh=bad-0.9080.00050.0198yes
q3high=bad-0.8850.00160.0506no
surprisal_early_vs_resthigh=good+0.8050.00220.0579no
surprisal_maxhigh=bad-0.7670.00300.0672no
kl_cent_stdhigh=good+0.7890.00340.0672no
rank_stdhigh=bad-0.7360.00420.0737no
mass1_early_vs_resthigh=bad-0.7080.00600.0869no
mass5_minhigh=good+0.7450.00690.0869no
rank_maxhigh=bad-0.7310.00750.0869no
margin_early_vs_resthigh=bad-0.6880.00770.0869no
mass5_late_minus_earlyhigh=good+0.6160.00770.0869no
kl_uniform_stdhigh=bad-0.6780.00870.0916no
surprisal_stdhigh=bad-0.6800.00940.0928no
full_ent_stdhigh=bad-0.6660.01040.0967no
rank_early_vs_resthigh=good+0.6560.01160.1018no
early_vs_resthigh=good+0.6310.01420.1122no
kl_uniform_early_vs_resthigh=bad-0.6310.01420.1122no
full_ent_early_vs_resthigh=good+0.6250.01500.1129no
q_rangehigh=bad-0.6040.02530.1438no
mass1_stdhigh=bad-0.5780.02550.1438no
total_variationhigh=bad-0.5690.02710.1438no
kl_step_late_minus_earlyhigh=good+0.5720.02910.1438no
full_ent_late_minus_earlyhigh=bad-0.5770.02930.1438no
late_vs_early_halfhigh=bad-0.5760.02950.1438no
kl_uniform_late_minus_earlyhigh=good+0.5760.02960.1438no
mass5_stdhigh=bad-0.5700.03100.1438no
kl_cent_meanhigh=good+0.5520.03310.1438no
first_token_entropyhigh=bad-0.5500.03350.1438no
full_ent_firsthigh=bad-0.5500.03350.1438no
kl_uniform_firsthigh=good+0.5500.03350.1438no
mass1_firsthigh=good+0.5440.03500.1438no
mass1_minhigh=good+0.5460.03500.1438no
surprisal_firsthigh=bad-0.5440.03500.1438no
phi_firsthigh=good+0.5440.03510.1438no
margin_firsthigh=good+0.5430.03520.1438no
kl_step_early_vs_resthigh=bad-0.5380.03550.1438no
margin_stdhigh=bad-0.5400.03730.1473no
n_semantichigh=bad-0.5230.03930.1478no
n_tokenshigh=bad-0.5230.03930.1478no
osc_sign_changeshigh=bad-0.5200.04340.1576no
n_spikeshigh=good+0.5370.04390.1576no
rank_meanhigh=bad-0.5070.04640.1629no
kl_uniform_meanhigh=good+0.4860.05520.1806no
mean_entropyhigh=bad-0.4860.05520.1806no
mass1_late_minus_earlyhigh=good+0.5000.05570.1806no
surprisal_meanhigh=bad-0.4870.05600.1806no
full_ent_meanhigh=bad-0.4800.05840.1845no
margin_meanhigh=good+0.4740.06540.2006no
phi_meanhigh=good+0.4690.06670.2006no
mass1_meanhigh=good+0.4680.06730.2006no
margin_late_minus_earlyhigh=good+0.4540.08020.2336no
surprisal_f10_maxhigh=good+0.4430.08130.2336no
mono_down_frachigh=bad-0.4630.08320.2347no
surprisal_late_minus_earlyhigh=bad-0.4490.08740.2423no
kl_step_meanhigh=good+0.4330.08930.2433no
q_curvaturehigh=good+0.4230.09570.2563no
first_to_max_ratiohigh=good+0.4150.10730.2826no
surprisal_f10_meanhigh=good+0.4080.11290.2924no
kl_step_maxhigh=good+0.4040.11520.2936no
kl_step_stdhigh=bad-0.3900.13140.3169no
surprisal_minhigh=bad-0.3870.13430.3169no
mass1_maxhigh=good+0.3790.13890.3169no
first10_maxhigh=good+0.3810.13920.3169no
full_ent_f10_maxhigh=good+0.3810.13920.3169no
kl_uniform_maxhigh=good+0.3780.13920.3169no
full_ent_minhigh=bad-0.3770.13950.3169no
mass5_f10_maxhigh=good+0.3850.14040.3169no
q_monotone_downhigh=bad-0.4210.14460.3218no
first10_stdhigh=good+0.3640.15850.3478no
margin_minhigh=good+0.3590.16970.3673no
mass1_f10_meanhigh=bad-0.3450.18160.3726no
mass20_late_minus_earlyhigh=good+0.3610.18160.3726no
phi_first10_meanhigh=bad-0.3450.18160.3726no
tail_mass_late_minus_earlyhigh=bad-0.3610.18160.3726no
mass5_early_vs_resthigh=bad-0.2870.18590.3766no
margin_f10_meanhigh=bad-0.3370.19260.3852no
mass20_f10_maxhigh=good+0.3220.22920.4527no
margin_maxhigh=good+0.3060.24520.4783no
trend_rhohigh=good+0.2990.24890.4786no
head_entropyhigh=bad-0.2930.25850.4786no
q1high=bad-0.2930.25890.4786no
mass5_meanhigh=good+0.2590.25970.4786no
tail_entropyhigh=bad-0.2950.26050.4786no
kl_step_f10_meanhigh=bad-0.2770.27180.4936no
q4_over_meanhigh=good+0.2700.29720.5336no
surprisal_f5_meanhigh=good+0.2720.30800.5468no
kl_cent_slopehigh=good+0.2610.32410.5636no
q4high=bad-0.2610.32460.5636no
rank_f10_maxhigh=good+0.3390.34460.5760no
rank_f10_meanhigh=good+0.3390.34460.5760no
kl_cent_f10_meanhigh=good+0.2490.34750.5760no
rank_f5_meanhigh=good+0.3160.35360.5760no
rank_late_minus_earlyhigh=bad-0.2420.35370.5760no
first10_meanhigh=good+0.2400.36080.5760no
kl_uniform_f10_meanhigh=bad-0.2400.36080.5760no
full_ent_f10_meanhigh=good+0.2400.36090.5760no
kl_step_slopehigh=bad-0.2320.39440.6232no
mass5_f10_meanhigh=bad-0.2060.44430.6950no
q_argmaxhigh=good+0.2090.48670.7539no
mass1_f5_meanhigh=bad-0.1800.49270.7558no
margin_f5_meanhigh=bad-0.1740.50830.7707no
spike_relative_poshigh=good+0.1670.52100.7707no
early_slopehigh=good+0.1880.52690.7707no
full_ent_slopehigh=good+0.1880.52690.7707no
kl_uniform_slopehigh=bad-0.1880.52690.7707no
kl_step_minhigh=bad-0.1820.53170.7707no
mass1_slopehigh=bad-0.1760.55250.7739no
q1_over_meanhigh=good+0.1520.55510.7739no
margin_slopehigh=bad-0.1680.55650.7739no
spike_quartilehigh=good+0.1810.55810.7739no
mass5_f5_meanhigh=bad-0.1810.55840.7739no
mass20_f10_meanhigh=bad-0.2040.58400.7939no
tail_mass_f10_meanhigh=good+0.2040.58400.7939no
tail_mass_f10_maxhigh=good+0.2060.58790.7939no
kl_cent_f10_maxhigh=bad-0.1520.60350.8023no
margin_f10_maxhigh=bad-0.1490.61330.8023no
mass1_f10_maxhigh=bad-0.1500.61430.8023no
kl_step_f10_maxhigh=bad-0.1460.61940.8023no
kl_uniform_f10_maxhigh=bad-0.1470.61950.8023no
rank_slopehigh=bad-0.1650.62490.8027no
mass5_firsthigh=bad-0.1250.63640.8109no
surprisal_slopehigh=good+0.1280.65340.8259no
mass20_slopehigh=bad-0.2010.68790.8558no
tail_mass_slopehigh=good+0.2010.68790.8558no
q_early_drophigh=bad-0.0940.69390.8565no
first3_meanhigh=good+0.0980.70640.8567no
tail_mass_firsthigh=bad-0.0970.70950.8567no
mass20_firsthigh=good+0.0970.71030.8567no
kl_cent_f5_meanhigh=bad-0.0800.76120.8940no
kl_step_f5_meanhigh=bad-0.0770.76940.8940no
first5_meanhigh=good+0.0770.76950.8940no
full_ent_f5_meanhigh=good+0.0770.76950.8940no
kl_uniform_f5_meanhigh=bad-0.0770.76950.8940no
mass20_stdhigh=bad-0.1070.78260.8960no
tail_mass_stdhigh=bad-0.1070.78260.8960no
spike_token_idxhigh=bad-0.0640.80530.9154no
mass5_slopehigh=bad-0.1560.81150.9158no
q2high=bad-0.0570.83120.9314no
kl_cent_maxhigh=bad-0.0640.85650.9330no
mass20_minhigh=good+0.0730.86500.9330no
tail_mass_maxhigh=bad-0.0730.86500.9330no
mass20_early_vs_resthigh=bad-0.0560.88480.9330no
tail_mass_early_vs_resthigh=good+0.0560.88480.9330no
kl_cent_late_minus_earlyhigh=bad-0.0370.89180.9330no
mass20_meanhigh=good+0.0490.89310.9330no
tail_mass_meanhigh=bad-0.0490.89310.9330no
mass20_f5_meanhigh=bad-0.0380.89700.9330no
tail_mass_f5_meanhigh=good+0.0380.89700.9330no
spike_heighthigh=bad-0.0350.89760.9330no
q_slopehigh=bad-0.0290.90920.9389no
path_ratiohigh=good+0.0230.93350.9577no
kl_cent_early_vs_resthigh=bad-0.0200.94030.9585no
osc_ratehigh=bad-0.0150.95370.9659no
mass5_maxhigh=good+0.1151.00001.0000no
tail_mass_minhigh=good+0.0521.00001.0000no

Which set of signals covers the most failures

The best single statistic is rarely the whole story — different failures can carry different signatures. This is a greedy set cover over detection rules: each step adds the signal catching the most failures not already caught, and is refused if it only buys recall by flagging correct work.

Failures caught
73%
cross-validated
Correct flagged
17%
the cost of the gate
Balanced accuracy
0.783
chance = 0.50
Signals used
1.2
typical, across folds
Label-shuffle null: beats 40 of 40 shuffles (p=0.0244). The entire selection procedure was re-run on 40 sets of shuffled pass/fail labels, which land at 0.491 — chance. Cross-validation shows the procedure generalises; only this shows there was a pattern to find.

p cannot go below 0.024 with 40 shuffles.

Which signals held up

A signal picked in one or two folds is a fitting artefact, not part of this model's pattern.

0 1 2 3 4 5 folds selected in (of 5) plateau_start_quartile plateau_start_quartile: selected in 5 of 5 folds 5/5 · stable kl_cent_std kl_cent_std: selected in 1 of 5 folds 1/5 · artefact
The combination, as shipped
SignalFlag whenThreshold
plateau_start_quartileabove-0.5

Fires if any row matches. Fitted on all records, so its raw score is optimistic — the cross-validated numbers above are the ones to trust.

What repairs failures

Direct measurement, not correlation: each failed generation was re-attempted with every strategy and re-tested. Nothing in this section depends on the entropy question above.

0 1 2 3 4 5 6 7 failures recovered (of 20) test_retry test_retry: recovered 7 of 20 failures, 4 of them recovered by nothing else 7 (35%) · 4 only here temp_retry temp_retry: recovered 6 of 20 failures, 1 of them recovered by nothing else 6 (30%) · 1 only here skeleton_fill skeleton_fill: recovered 4 of 20 failures, 1 of them recovered by nothing else 4 (20%) · 1 only here decompose 0 (0%)

Routing by error type

logic errors
25%50%75%100%test_retrytest_retry on logic errors: 3/6 recovered (50%)3/6temp_retrytemp_retry on logic errors: 2/6 recovered (33%)2/6skeleton_fillskeleton_fill on logic errors: 1/6 recovered (17%)1/6decompose0/6
assertion errors
25%50%75%100%temp_retrytemp_retry on assertion errors: 4/14 recovered (29%)4/14test_retrytest_retry on assertion errors: 4/14 recovered (29%)4/14skeleton_fillskeleton_fill on assertion errors: 3/14 recovered (21%)3/14decompose0/14

What shows promise

Leads, not findings. Each one names the run that would settle it.

q3

d=-0.885, raw p=0.0016, adjusted p=0.0506. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
surprisal_early_vs_rest

d=+0.805, raw p=0.0022, adjusted p=0.0579. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
surprisal_max

d=-0.767, raw p=0.0030, adjusted p=0.0672. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_cent_std

d=+0.789, raw p=0.0034, adjusted p=0.0672. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
rank_std

d=-0.736, raw p=0.0042, adjusted p=0.0737. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass1_early_vs_rest

d=-0.708, raw p=0.0060, adjusted p=0.0869. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass5_min

d=+0.745, raw p=0.0069, adjusted p=0.0869. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
rank_max

d=-0.731, raw p=0.0075, adjusted p=0.0869. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
margin_early_vs_rest

d=-0.688, raw p=0.0077, adjusted p=0.0869. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass5_late_minus_early

d=+0.616, raw p=0.0077, adjusted p=0.0869. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_uniform_std

d=-0.678, raw p=0.0087, adjusted p=0.0916. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
surprisal_std

d=-0.680, raw p=0.0094, adjusted p=0.0928. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
full_ent_std

d=-0.666, raw p=0.0104, adjusted p=0.0967. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
rank_early_vs_rest

d=+0.656, raw p=0.0116, adjusted p=0.1018. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
early_vs_rest

d=+0.631, raw p=0.0142, adjusted p=0.1122. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_uniform_early_vs_rest

d=-0.631, raw p=0.0142, adjusted p=0.1122. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
full_ent_early_vs_rest

d=+0.625, raw p=0.0150, adjusted p=0.1129. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
q_range

d=-0.604, raw p=0.0253, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass1_std

d=-0.578, raw p=0.0255, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
total_variation

d=-0.569, raw p=0.0271, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_step_late_minus_early

d=+0.572, raw p=0.0291, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
full_ent_late_minus_early

d=-0.577, raw p=0.0293, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
late_vs_early_half

d=-0.576, raw p=0.0295, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_uniform_late_minus_early

d=+0.576, raw p=0.0296, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass5_std

d=-0.570, raw p=0.0310, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_cent_mean

d=+0.552, raw p=0.0331, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
first_token_entropy

d=-0.550, raw p=0.0335, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
full_ent_first

d=-0.550, raw p=0.0335, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_uniform_first

d=+0.550, raw p=0.0335, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass1_first

d=+0.544, raw p=0.0350, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass1_min

d=+0.546, raw p=0.0350, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
surprisal_first

d=-0.544, raw p=0.0350, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
phi_first

d=+0.544, raw p=0.0351, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
margin_first

d=+0.543, raw p=0.0352, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_step_early_vs_rest

d=-0.538, raw p=0.0355, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
margin_std

d=-0.540, raw p=0.0373, adjusted p=0.1473. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
n_semantic

d=-0.523, raw p=0.0393, adjusted p=0.1478. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
n_tokens

d=-0.523, raw p=0.0393, adjusted p=0.1478. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
osc_sign_changes

d=-0.520, raw p=0.0434, adjusted p=0.1576. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
n_spikes

d=+0.537, raw p=0.0439, adjusted p=0.1576. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
rank_mean

d=-0.507, raw p=0.0464, adjusted p=0.1629. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.

To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.

Settings this run used

ParameterValueSource
min_p0.0registry (card-verified)
presence_penalty0.0registry (card-verified)
repeat_penalty1.0registry (card-verified)
temp0.7registry (card-verified)
top_k20registry (card-verified)
top_p0.8registry (card-verified)
thinkingFalseregistry (card-verified)

Generated by the Local Model Calibration Kit. Raw records for every generation and repair attempt are in the run’s _raw.jsonl; calibrate.py reanalyze rebuilds this report from them without re-running the model.