Scope: this file.Qwen_Qwen3-4B-Instruct-2507-Q5_K_L.gguf
Everything below was measured on this model, at this size, at this quantization. Quantization changes the logit distribution and every signal here is a property of that distribution — a different quant of the same weights is a different organism and needs its own run. None of it is claimed to transfer to another model.
Entropy signal
usable
p=0.000
Effect seen
d=-0.94
detectable: d≥0.72
Failures
20
of 108 generations
Recoverable
11 / 20
by any strategy tested
The rule
Implementable as written. Every number below was measured on this model, on this run.
Discard and regenerate if plateau_start_quartile > -0.5. This combination catches 73% of failing generations at the cost of flagging 17% of correct ones — balanced accuracy 0.783 against a chance of 0.50. Measured by rebuilding the whole combination inside 5 cross-validation folds and scoring it on records each fold never saw, then checking it beats 40/40 label-shuffled runs (p=0.0244).
On a failed generation, repair in this order: test_retry → temp_retry → skeleton_fill. This ordering is greedy set cover over what actually recovered failures in this run, not a preference — test_retry alone recovered 7 of 20.
When the failure is a logic error, route to test_retry first (3/6, 50%), ahead of temp_retry 33%, skeleton_fill 17%.
When the failure is an assertion error, route to test_retry first (4/14, 29%), ahead of temp_retry 29%, skeleton_fill 21%.
Stop after the ordered chain. 11 of 20 failures were recoverable by any strategy tested; the remaining 9 were not recovered by any of them, so further retries on those are spend without evidence behind it.
Where to cut, and what it costs
One signal, one curve. What changes is where you cut it and what you do with each side. Fitted and scored across cross-validation folds on 108 unique generations (18 failures) — each number is what the rule did on records its own fold never saw. Pick one row; they are three settings of the same dial, not three rules to stack.
Catch wrong — regenerate when plateau_start_quartile > -0.5. Chosen in 4 of 5 folds; only readable once the generation is complete, so the tokens are already spent when it fires.
Failures caught
67%
of all failing generations
Correct work lost
17%
regenerated needlessly
Net tokens
-54,084
per 1000 generations (-27.8%)
Per 1000 generations: 111.2 doomed runs aborted early, 139.2 correct ones thrown away and regenerated. Token counts are this run's own.
Trust right — ship untested when surprisal_max ≤ 1.1439. Chosen in 2 of 5 folds; only readable once the generation is complete, so the tokens are already spent when it fires.
Output covered
46%
shipped without testing
Purity
91.6%
of what ships is correct
Defects shipped
38.6
per 1000 generations
This saves test execution, not generation tokens — the generation is already paid for by the time any signal can be read. Worth it when the test suite is expensive or latency matters more than compute. The defect count is the price.
Both — one cut point, not two, on surprisal_max. Chosen in 2 of 5 folds; only readable once the generation is complete, so the tokens are already spent when it fires.
The fit put both cut points at the same value, so there is no undecided band: every generation lands in trust or in bail. That makes this row the same single threshold as the two above, read from both sides — not a third, richer rule.
Zone
Condition
What to do
Share of output
Trust
surprisal_max ≤ 1.1439
ship untested — 91.6% correct
45%
Band
between the two
test it — the signal cannot decide
0%
Bail
surprisal_max > 1.1439
regenerate — catches 68% of all failures
55%
Net -124,612 tokens per 1000 generations from the bail zone (-64.1% of baseline).
Ship the trust zone untested, regenerate the bail zone, test the band between them. Both cut points come from one fit on one signal, ordered so the zones cannot overlap. The band is what the signal cannot decide — the honest residue, not a gap in the analysis.
No operating point here pays for itself in tokens. The best of them still nets -54,084 tokens per 1000 generations. That is not a flaw in the rule — this run's winning signal is only readable once a generation is complete, so aborting saves nothing that was not already spent, while every wrongly flagged generation pays for a full replacement. Use these rules to spend less on testing, or to decide what to ship, not to spend fewer tokens.
What separates correct from incorrect: plateau_start_quartile
CorrectIncorrect
Every probe generation, placed by its plateau_start_quartile. Vertical marks are group means.
Correct generations averaged plateau_start_quartile -0.640; incorrect ones 0.400. The groups are shifted, not merely overlapping — this is the separation the rule above is built on.
A permutation test puts that difference at p=0.0001, against an alpha of 0.05. Reshuffling which generations were labelled correct reproduces a gap this large routinely.
Across 158 candidate statistics the largest effect was plateau_start_quartile at d=-1.015 (high=bad), and it survives correction.
11 of 20 failures were recoverable by at least one repair strategy. Recovery is measured directly — break it, repair it, re-run the tests — so these counts do not depend on any of the statistical argument above.
Power: signal detected. Observed effect d=-0.94; this run could detect d>=0.72.
Every candidate signal, corrected
All 158 numeric statistics were tested against the same pass/fail outcome, so raw p-values are corrected with Benjamini-Hochberg FDR. A statistic is called usable only if its adjusted p clears 0.05 and |d| ≥ 0.2. The shaded band is what this run was too small to detect at all.
Table view — all signals
Signal
Direction
d
p
p adjusted
Usable
kl_uniform_min
high=good
+0.938
0.0001
0.0079
yes
max_entropy
high=bad
-0.938
0.0001
0.0079
yes
plateau_start_quartile
high=bad
-1.015
0.0002
0.0105
yes
full_ent_max
high=bad
-0.908
0.0005
0.0198
yes
q3
high=bad
-0.885
0.0016
0.0506
no
surprisal_early_vs_rest
high=good
+0.805
0.0022
0.0579
no
surprisal_max
high=bad
-0.767
0.0030
0.0672
no
kl_cent_std
high=good
+0.789
0.0034
0.0672
no
rank_std
high=bad
-0.736
0.0042
0.0737
no
mass1_early_vs_rest
high=bad
-0.708
0.0060
0.0869
no
mass5_min
high=good
+0.745
0.0069
0.0869
no
rank_max
high=bad
-0.731
0.0075
0.0869
no
margin_early_vs_rest
high=bad
-0.688
0.0077
0.0869
no
mass5_late_minus_early
high=good
+0.616
0.0077
0.0869
no
kl_uniform_std
high=bad
-0.678
0.0087
0.0916
no
surprisal_std
high=bad
-0.680
0.0094
0.0928
no
full_ent_std
high=bad
-0.666
0.0104
0.0967
no
rank_early_vs_rest
high=good
+0.656
0.0116
0.1018
no
early_vs_rest
high=good
+0.631
0.0142
0.1122
no
kl_uniform_early_vs_rest
high=bad
-0.631
0.0142
0.1122
no
full_ent_early_vs_rest
high=good
+0.625
0.0150
0.1129
no
q_range
high=bad
-0.604
0.0253
0.1438
no
mass1_std
high=bad
-0.578
0.0255
0.1438
no
total_variation
high=bad
-0.569
0.0271
0.1438
no
kl_step_late_minus_early
high=good
+0.572
0.0291
0.1438
no
full_ent_late_minus_early
high=bad
-0.577
0.0293
0.1438
no
late_vs_early_half
high=bad
-0.576
0.0295
0.1438
no
kl_uniform_late_minus_early
high=good
+0.576
0.0296
0.1438
no
mass5_std
high=bad
-0.570
0.0310
0.1438
no
kl_cent_mean
high=good
+0.552
0.0331
0.1438
no
first_token_entropy
high=bad
-0.550
0.0335
0.1438
no
full_ent_first
high=bad
-0.550
0.0335
0.1438
no
kl_uniform_first
high=good
+0.550
0.0335
0.1438
no
mass1_first
high=good
+0.544
0.0350
0.1438
no
mass1_min
high=good
+0.546
0.0350
0.1438
no
surprisal_first
high=bad
-0.544
0.0350
0.1438
no
phi_first
high=good
+0.544
0.0351
0.1438
no
margin_first
high=good
+0.543
0.0352
0.1438
no
kl_step_early_vs_rest
high=bad
-0.538
0.0355
0.1438
no
margin_std
high=bad
-0.540
0.0373
0.1473
no
n_semantic
high=bad
-0.523
0.0393
0.1478
no
n_tokens
high=bad
-0.523
0.0393
0.1478
no
osc_sign_changes
high=bad
-0.520
0.0434
0.1576
no
n_spikes
high=good
+0.537
0.0439
0.1576
no
rank_mean
high=bad
-0.507
0.0464
0.1629
no
kl_uniform_mean
high=good
+0.486
0.0552
0.1806
no
mean_entropy
high=bad
-0.486
0.0552
0.1806
no
mass1_late_minus_early
high=good
+0.500
0.0557
0.1806
no
surprisal_mean
high=bad
-0.487
0.0560
0.1806
no
full_ent_mean
high=bad
-0.480
0.0584
0.1845
no
margin_mean
high=good
+0.474
0.0654
0.2006
no
phi_mean
high=good
+0.469
0.0667
0.2006
no
mass1_mean
high=good
+0.468
0.0673
0.2006
no
margin_late_minus_early
high=good
+0.454
0.0802
0.2336
no
surprisal_f10_max
high=good
+0.443
0.0813
0.2336
no
mono_down_frac
high=bad
-0.463
0.0832
0.2347
no
surprisal_late_minus_early
high=bad
-0.449
0.0874
0.2423
no
kl_step_mean
high=good
+0.433
0.0893
0.2433
no
q_curvature
high=good
+0.423
0.0957
0.2563
no
first_to_max_ratio
high=good
+0.415
0.1073
0.2826
no
surprisal_f10_mean
high=good
+0.408
0.1129
0.2924
no
kl_step_max
high=good
+0.404
0.1152
0.2936
no
kl_step_std
high=bad
-0.390
0.1314
0.3169
no
surprisal_min
high=bad
-0.387
0.1343
0.3169
no
mass1_max
high=good
+0.379
0.1389
0.3169
no
first10_max
high=good
+0.381
0.1392
0.3169
no
full_ent_f10_max
high=good
+0.381
0.1392
0.3169
no
kl_uniform_max
high=good
+0.378
0.1392
0.3169
no
full_ent_min
high=bad
-0.377
0.1395
0.3169
no
mass5_f10_max
high=good
+0.385
0.1404
0.3169
no
q_monotone_down
high=bad
-0.421
0.1446
0.3218
no
first10_std
high=good
+0.364
0.1585
0.3478
no
margin_min
high=good
+0.359
0.1697
0.3673
no
mass1_f10_mean
high=bad
-0.345
0.1816
0.3726
no
mass20_late_minus_early
high=good
+0.361
0.1816
0.3726
no
phi_first10_mean
high=bad
-0.345
0.1816
0.3726
no
tail_mass_late_minus_early
high=bad
-0.361
0.1816
0.3726
no
mass5_early_vs_rest
high=bad
-0.287
0.1859
0.3766
no
margin_f10_mean
high=bad
-0.337
0.1926
0.3852
no
mass20_f10_max
high=good
+0.322
0.2292
0.4527
no
margin_max
high=good
+0.306
0.2452
0.4783
no
trend_rho
high=good
+0.299
0.2489
0.4786
no
head_entropy
high=bad
-0.293
0.2585
0.4786
no
q1
high=bad
-0.293
0.2589
0.4786
no
mass5_mean
high=good
+0.259
0.2597
0.4786
no
tail_entropy
high=bad
-0.295
0.2605
0.4786
no
kl_step_f10_mean
high=bad
-0.277
0.2718
0.4936
no
q4_over_mean
high=good
+0.270
0.2972
0.5336
no
surprisal_f5_mean
high=good
+0.272
0.3080
0.5468
no
kl_cent_slope
high=good
+0.261
0.3241
0.5636
no
q4
high=bad
-0.261
0.3246
0.5636
no
rank_f10_max
high=good
+0.339
0.3446
0.5760
no
rank_f10_mean
high=good
+0.339
0.3446
0.5760
no
kl_cent_f10_mean
high=good
+0.249
0.3475
0.5760
no
rank_f5_mean
high=good
+0.316
0.3536
0.5760
no
rank_late_minus_early
high=bad
-0.242
0.3537
0.5760
no
first10_mean
high=good
+0.240
0.3608
0.5760
no
kl_uniform_f10_mean
high=bad
-0.240
0.3608
0.5760
no
full_ent_f10_mean
high=good
+0.240
0.3609
0.5760
no
kl_step_slope
high=bad
-0.232
0.3944
0.6232
no
mass5_f10_mean
high=bad
-0.206
0.4443
0.6950
no
q_argmax
high=good
+0.209
0.4867
0.7539
no
mass1_f5_mean
high=bad
-0.180
0.4927
0.7558
no
margin_f5_mean
high=bad
-0.174
0.5083
0.7707
no
spike_relative_pos
high=good
+0.167
0.5210
0.7707
no
early_slope
high=good
+0.188
0.5269
0.7707
no
full_ent_slope
high=good
+0.188
0.5269
0.7707
no
kl_uniform_slope
high=bad
-0.188
0.5269
0.7707
no
kl_step_min
high=bad
-0.182
0.5317
0.7707
no
mass1_slope
high=bad
-0.176
0.5525
0.7739
no
q1_over_mean
high=good
+0.152
0.5551
0.7739
no
margin_slope
high=bad
-0.168
0.5565
0.7739
no
spike_quartile
high=good
+0.181
0.5581
0.7739
no
mass5_f5_mean
high=bad
-0.181
0.5584
0.7739
no
mass20_f10_mean
high=bad
-0.204
0.5840
0.7939
no
tail_mass_f10_mean
high=good
+0.204
0.5840
0.7939
no
tail_mass_f10_max
high=good
+0.206
0.5879
0.7939
no
kl_cent_f10_max
high=bad
-0.152
0.6035
0.8023
no
margin_f10_max
high=bad
-0.149
0.6133
0.8023
no
mass1_f10_max
high=bad
-0.150
0.6143
0.8023
no
kl_step_f10_max
high=bad
-0.146
0.6194
0.8023
no
kl_uniform_f10_max
high=bad
-0.147
0.6195
0.8023
no
rank_slope
high=bad
-0.165
0.6249
0.8027
no
mass5_first
high=bad
-0.125
0.6364
0.8109
no
surprisal_slope
high=good
+0.128
0.6534
0.8259
no
mass20_slope
high=bad
-0.201
0.6879
0.8558
no
tail_mass_slope
high=good
+0.201
0.6879
0.8558
no
q_early_drop
high=bad
-0.094
0.6939
0.8565
no
first3_mean
high=good
+0.098
0.7064
0.8567
no
tail_mass_first
high=bad
-0.097
0.7095
0.8567
no
mass20_first
high=good
+0.097
0.7103
0.8567
no
kl_cent_f5_mean
high=bad
-0.080
0.7612
0.8940
no
kl_step_f5_mean
high=bad
-0.077
0.7694
0.8940
no
first5_mean
high=good
+0.077
0.7695
0.8940
no
full_ent_f5_mean
high=good
+0.077
0.7695
0.8940
no
kl_uniform_f5_mean
high=bad
-0.077
0.7695
0.8940
no
mass20_std
high=bad
-0.107
0.7826
0.8960
no
tail_mass_std
high=bad
-0.107
0.7826
0.8960
no
spike_token_idx
high=bad
-0.064
0.8053
0.9154
no
mass5_slope
high=bad
-0.156
0.8115
0.9158
no
q2
high=bad
-0.057
0.8312
0.9314
no
kl_cent_max
high=bad
-0.064
0.8565
0.9330
no
mass20_min
high=good
+0.073
0.8650
0.9330
no
tail_mass_max
high=bad
-0.073
0.8650
0.9330
no
mass20_early_vs_rest
high=bad
-0.056
0.8848
0.9330
no
tail_mass_early_vs_rest
high=good
+0.056
0.8848
0.9330
no
kl_cent_late_minus_early
high=bad
-0.037
0.8918
0.9330
no
mass20_mean
high=good
+0.049
0.8931
0.9330
no
tail_mass_mean
high=bad
-0.049
0.8931
0.9330
no
mass20_f5_mean
high=bad
-0.038
0.8970
0.9330
no
tail_mass_f5_mean
high=good
+0.038
0.8970
0.9330
no
spike_height
high=bad
-0.035
0.8976
0.9330
no
q_slope
high=bad
-0.029
0.9092
0.9389
no
path_ratio
high=good
+0.023
0.9335
0.9577
no
kl_cent_early_vs_rest
high=bad
-0.020
0.9403
0.9585
no
osc_rate
high=bad
-0.015
0.9537
0.9659
no
mass5_max
high=good
+0.115
1.0000
1.0000
no
tail_mass_min
high=good
+0.052
1.0000
1.0000
no
Which set of signals covers the most failures
The best single statistic is rarely the whole story — different failures can carry different signatures. This is a greedy set cover over detection rules: each step adds the signal catching the most failures not already caught, and is refused if it only buys recall by flagging correct work.
Failures caught
73%
cross-validated
Correct flagged
17%
the cost of the gate
Balanced accuracy
0.783
chance = 0.50
Signals used
1.2
typical, across folds
Label-shuffle null: beats 40 of 40 shuffles (p=0.0244). The entire selection procedure was re-run on 40 sets of shuffled pass/fail labels, which land at 0.491 — chance. Cross-validation shows the procedure generalises; only this shows there was a pattern to find.
p cannot go below 0.024 with 40 shuffles.
Which signals held up
A signal picked in one or two folds is a fitting artefact, not part of this model's pattern.
The combination, as shipped
Signal
Flag when
Threshold
plateau_start_quartile
above
-0.5
Fires if any row matches. Fitted on all records, so its raw score is optimistic — the cross-validated numbers above are the ones to trust.
What repairs failures
Direct measurement, not correlation: each failed generation was re-attempted with every strategy and re-tested. Nothing in this section depends on the entropy question above.
Routing by error type
logic errorsassertion errors
What shows promise
Leads, not findings. Each one names the run that would settle it.
q3
d=-0.885, raw p=0.0016, adjusted p=0.0506. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
surprisal_early_vs_rest
d=+0.805, raw p=0.0022, adjusted p=0.0579. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
surprisal_max
d=-0.767, raw p=0.0030, adjusted p=0.0672. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_cent_std
d=+0.789, raw p=0.0034, adjusted p=0.0672. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
rank_std
d=-0.736, raw p=0.0042, adjusted p=0.0737. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass1_early_vs_rest
d=-0.708, raw p=0.0060, adjusted p=0.0869. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass5_min
d=+0.745, raw p=0.0069, adjusted p=0.0869. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
rank_max
d=-0.731, raw p=0.0075, adjusted p=0.0869. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
margin_early_vs_rest
d=-0.688, raw p=0.0077, adjusted p=0.0869. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass5_late_minus_early
d=+0.616, raw p=0.0077, adjusted p=0.0869. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_uniform_std
d=-0.678, raw p=0.0087, adjusted p=0.0916. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
surprisal_std
d=-0.680, raw p=0.0094, adjusted p=0.0928. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
full_ent_std
d=-0.666, raw p=0.0104, adjusted p=0.0967. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
rank_early_vs_rest
d=+0.656, raw p=0.0116, adjusted p=0.1018. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
early_vs_rest
d=+0.631, raw p=0.0142, adjusted p=0.1122. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_uniform_early_vs_rest
d=-0.631, raw p=0.0142, adjusted p=0.1122. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
full_ent_early_vs_rest
d=+0.625, raw p=0.0150, adjusted p=0.1129. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
q_range
d=-0.604, raw p=0.0253, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass1_std
d=-0.578, raw p=0.0255, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
total_variation
d=-0.569, raw p=0.0271, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_step_late_minus_early
d=+0.572, raw p=0.0291, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
full_ent_late_minus_early
d=-0.577, raw p=0.0293, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
late_vs_early_half
d=-0.576, raw p=0.0295, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_uniform_late_minus_early
d=+0.576, raw p=0.0296, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass5_std
d=-0.570, raw p=0.0310, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_cent_mean
d=+0.552, raw p=0.0331, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
first_token_entropy
d=-0.550, raw p=0.0335, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
full_ent_first
d=-0.550, raw p=0.0335, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_uniform_first
d=+0.550, raw p=0.0335, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass1_first
d=+0.544, raw p=0.0350, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
mass1_min
d=+0.546, raw p=0.0350, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
surprisal_first
d=-0.544, raw p=0.0350, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
phi_first
d=+0.544, raw p=0.0351, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
margin_first
d=+0.543, raw p=0.0352, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
kl_step_early_vs_rest
d=-0.538, raw p=0.0355, adjusted p=0.1438. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
margin_std
d=-0.540, raw p=0.0373, adjusted p=0.1473. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
n_semantic
d=-0.523, raw p=0.0393, adjusted p=0.1478. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
n_tokens
d=-0.523, raw p=0.0393, adjusted p=0.1478. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
osc_sign_changes
d=-0.520, raw p=0.0434, adjusted p=0.1576. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
n_spikes
d=+0.537, raw p=0.0439, adjusted p=0.1576. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
rank_mean
d=-0.507, raw p=0.0464, adjusted p=0.1629. Clears p<0.05 on its own but not after correcting for the 158 statistics tested against the same outcome. With that many tests, one or two hits at p<0.05 is what noise looks like.
To settle it: Pre-register it: decide in advance to test this one statistic on a fresh run and report that result whichever way it goes. Re-scanning and picking the new winner is how the original mistake happened.
Settings this run used
Parameter
Value
Source
min_p
0.0
registry (card-verified)
presence_penalty
0.0
registry (card-verified)
repeat_penalty
1.0
registry (card-verified)
temp
0.7
registry (card-verified)
top_k
20
registry (card-verified)
top_p
0.8
registry (card-verified)
thinking
False
registry (card-verified)
Generated by the Local Model Calibration Kit. Raw records for every generation and repair attempt are in the run’s _raw.jsonl; calibrate.py reanalyze rebuilds this report from them without re-running the model.