Keep the step inside the stable range, and every evaluation is feasible.
An explicit diffusion step on a graph is stable only while h < 2/lambda_max: that is connection B. We turned the limit into a search box. Instead of letting NSGA-II try any step size, we restricted h to [h_min, 0.98 h*(s)], with h*(s) = 2/lambda_max(s) computed on the training data. On all six graph problems we tried, every evaluation in the restricted box was feasible and the search pulled ahead early.
- Feasible evaluations, restricted
- 6 of 6
- problems with a mean feasible fraction of exactly 1.000, against 0.852 to 0.940 on the full range.
- Evaluations to target, new problems
- 2.3x to 4.5x fewer
- per-arm medians on the three new graphs (e 51.5 to 22.5, f 47 to 10.5, g 93 to 38).
- Anytime HV at 50 evaluations
- higher on 6 of 6
- problems where the mean anytime hypervolume after 50 evaluations is higher with the restriction.
LF: linear map, full h rangeLR: linear map, h in [h_min, 0.98 h*(s)]
Mean feasible fraction of the 300 evaluations
e3-ary treeT7, seeds 1-30
0.870 to 1.000
fspiderT7, seeds 1-30
0.852 to 1.000
gsmall worldT7, seeds 1-30
0.900 to 1.000
aX4 spec B, 2-D gridT6, seeds 21-70
0.899 to 1.000
bpath graphT6, seeds 21-70
0.940 to 1.000
dcycle graphT6, seeds 21-70
0.936 to 1.000
Median evaluations to reach the target box
e3-ary treeT7, reached 29 and 30 of 30
51.5 to 22.5
fspiderT7, reached 30 and 30 of 30
47 to 10.5
gsmall worldT7, reached 28 and 30 of 30
93 to 38
aX4 spec B, 2-D gridT6, reached 50 and 50 of 50
62 to 46.5
bpath graphT6, reached 41 and 45 of 50
148 to 124.5
dcycle graphT6, reached 0 and 0 of 50
no run reached the box
NSGA-II, population 20 x 15 generations, exactly 300 evaluations per run, paired optimizer seeds. Problems e, f and g are new graph families (T7, seeds 1-30); the restriction was chosen after a, b and d were known, so only e, f and g test it. Rows a, b and d use T6's new optimizer seeds 21-70 on the same fixed datasets as the earlier runs. Evaluations to target are censored at 301; on d no run of either arm reached T5's target box, which was kept, not changed. These process numbers are descriptive. Sources: docs/research/chart-selection-2026-10/T7/results.json (table) and docs/research/chart-selection-2026-10/T6/results.json (secondary_new_seeds_21_70.per_arm).
The numbers behind the chart
| Problem | Feasible fraction, LF to LR | Evals to target, LF to LR (reached) | Anytime HV at 50, LF to LR | Source |
|---|---|---|---|---|
| e 3-ary tree | 0.870 to 1.000 | 51.5 to 22.5 (29 and 30 of 30) | 0.08068 to 0.08372 | T7, seeds 1-30 |
| f spider | 0.852 to 1.000 | 47 to 10.5 (30 and 30 of 30) | 0.03179 to 0.03541 | T7, seeds 1-30 |
| g small world | 0.900 to 1.000 | 93 to 38 (28 and 30 of 30) | 0.03077 to 0.03380 | T7, seeds 1-30 |
| a X4 spec B, 2-D grid | 0.899 to 1.000 | 62 to 46.5 (50 and 50 of 50) | 0.03659 to 0.03923 | T6, seeds 21-70 |
| b path graph | 0.940 to 1.000 | 148 to 124.5 (41 and 45 of 50) | 0.01398 to 0.01509 | T6, seeds 21-70 |
| d cycle graph | 0.936 to 1.000 | 301 to 301 (0 and 0 of 50) | 0.15896 to 0.16211 | T6, seeds 21-70 |
Evaluations to target fall with the restriction on 5 of the 5 problems where a target was reached (d: no run reached the box). On the new problems the paired median difference in anytime HV at 50, LR minus LF (uncorrected, descriptive), is e +0.00278 (p 0.000608), f +0.00247 (p 1.64e-7), g +0.00225 (p 0.00000922).
What the restriction is
The search coordinate for the step is u in [0, 1]. The full-range arm, LF, maps it linearly onto the whole box h in [1e-3, 1]. The restricted arm, LR, maps it linearly onto [h_min, cap(s)], with cap(s) = min(h_max, 0.98 h*(s)). Everything else is unchanged: the other four inputs, the objective, the optimizer and the 300-evaluation budget.
LF reproduces the raw arm: on a and b (T3), its held-out HV differs from raw only at float rounding, with the same evaluations to target on every seed.
Why it is practical
The cap needs lambda_max(s), not the objective. Computing it costs no objective evaluations, so the box is known before the search starts, and no evaluation lands on a step beyond the stability limit. On T6's new seeds the restricted search also had the smaller spread across seeds on held-out HV, on all three problems.
Below the lead: what the restriction does not show
The final score: a real gain on one problem only
The preregistered endpoint was held-out hypervolume after all 300 evaluations, read by one rule: a real effect if Holm p < 0.05; no effect if not significant and the 95% CI lies inside +-1% of the base median; inconclusive otherwise. By that rule the restriction is a real effect only on problem a, X4 spec B: T3 +1.14% (Holm p 0.0101) and T6 +0.52% (Holm p 0.0447). On new optimizer seeds the gain on a is less than half the first estimate and its CI includes zero, and those seeds share the same fixed datasets. Everywhere else it is inconclusive or no effect.
real effectno effectinconclusive+-1% of the base median
aX4 spec B, 2-D gridT3, seeds 1-20, LR - LF
+1.14% [+0.48%, +1.61%]
Holm p 0.0101, real effect
aX4 spec B, 2-D gridT6, seeds 21-70, LR - LF
+0.52% [-0.13%, +0.88%]
Holm p 0.0447, real effect
bpath graphT3, seeds 1-20, LR - LF
+0.74% [-0.20%, +4.10%]
Holm p 0.0962, inconclusive
bpath graphT6, seeds 21-70, LR - LF
+0.19% [-0.74%, +1.32%]
Holm p 0.618, inconclusive
dcycle graphT5, seeds 1-20, LR - raw
-0.18% [-0.96%, +0.28%]
p 0.430, no effect
dcycle graphT6, seeds 21-70, LR - LF
+0.17% [-0.43%, +1.25%]
Holm p 0.618, inconclusive
e3-ary treeT7, seeds 1-30, LR - LF
+0.04% [-0.09%, +0.62%]
Holm p 0.698, no effect
fspiderT7, seeds 1-30, LR - LF
-0.64% [-2.58%, +3.69%]
Holm p 0.698, inconclusive
gsmall worldT7, seeds 1-30, LR - LF
+1.00% [+0.14%, +2.37%]
Holm p 0.0702, inconclusive
Each row is one preregistered comparison, scaled by its own base median so the +-1% band lines up. The base is LF, except T5's d row, which compares LR with raw (T5 did not run LF; T3 found raw equal to LF up to the rounding of h); T5's single primary comparison has no Holm family. T6's seeds 21-70 are independent of seeds 1-20 in optimizer randomness only. Sources: T3/results.json primary, T5/results.json d.primary, T6/results.json primary_new_seeds_21_70, T7/results.json primary.
a: real, smaller than first measured
T3, seeds 1-20: +0.000471 (+1.14% of the base), CI [+0.000196, +0.000663], 17 wins and 3 losses, Holm p 0.0101: real effect.
T6, seeds 21-70: +0.000214 (+0.52% of the base), CI [-0.000053, +0.000363], 31 wins and 19 losses, Holm p 0.0447: real effect.
T6 calls its a result real by the rule but only barely: the CI of the median difference includes 0.
b and d: not shown either way
On b and on d the comparisons are inconclusive, except T5's 20-seed d test, which reads no effect: no difference larger than 1% of raw was detected, smaller effects are not excluded, and the CI also allows a loss of up to about 1%. The seed-to-seed spread on b and d was larger than T6's power note assumed, so its design was less sensitive there than planned. Non-detection is not absence.
e, f, g: no real effect
No effect on e; inconclusive on f and g. g leans positive (+1.00% of LF, uncorrected p 0.0234) but does not survive Holm (0.0702). On f the restricted fronts were less often feasible on the held-out data (held-out feasible fraction -0.144, uncorrected p 0.000952). One plausible mechanism, not tested: the cap is computed on the training data, and the held-out realizations have their own lambda_max.
Does the unstable share predict the gain? Not here.
The hypothesis was that the restriction helps most where the full range wastes more of its h interval on unstable steps. Across the six problems it gets no visible support: Spearman rho = 0.086, n = 6, descriptive only. The highest-share problem, f, has a negative point estimate, and the shares span only 0.47 to 0.77.
| Problem | Unstable share of the raw h range | Relative held-out effect (median diff / base median) |
|---|---|---|
| d cycle graph | 0.470 | -0.18% |
| b path graph | 0.485 | +0.74% |
| e 3-ary tree | 0.656 | +0.04% |
| g small world | 0.686 | +1.00% |
| a X4 spec B, 2-D grid | 0.720 | +1.14% |
| f spider | 0.766 | -0.64% |
Source: T7/results.json descriptive_share_vs_effect. a and b from T3 (LR vs LF, seeds 1-20), d from T5 (LR vs raw, seeds 1-20), e, f and g from T7 (seeds 1-30).
Changing the coordinate, not the range: no detectable benefit
T3 and T4 kept the full h range and bent the coordinate instead: nonlinear maps of u onto [h_min, h_max] whose knot tracks the stability limit h*(s). Five map shapes, two problems, 20 seeds each. No aligned coordinate gave a detectable held-out gain on either problem. That is inconclusive, not excluded: every coordinate CI reaches past +-1% of LF. On a, the part due to alignment beyond plain nonlinearity (knot at h*(s) vs a fixed knot) reads no effect for 4 of the 5 shapes.
| Map shape (test) | a: coordinate | a: alignment beyond nonlinearity | b: coordinate | b: alignment beyond nonlinearity |
|---|---|---|---|---|
| Hermite in log h, knot u = 0.5 (T3) | inconclusive (Holm 1) | no effect (Holm 0.783) | inconclusive (Holm 0.0962) | inconclusive (Holm 1) |
| F1 power (T4) | inconclusive (Holm 1) | no effect (Holm 1) | inconclusive (Holm 0.851) | inconclusive (Holm 1) |
| F2 piecewise linear (T4) | inconclusive (Holm 1) | inconclusive (Holm 1) | inconclusive (Holm 1) | inconclusive (Holm 1) |
| F3 Hermite, knot u = 0.8 (T4) | inconclusive (Holm 1) | no effect (Holm 1) | inconclusive (Holm 1) | inconclusive (Holm 1) |
| F4 logistic in log h (T4) | inconclusive (Holm 1) | no effect (Holm 1) | inconclusive (Holm 1) | inconclusive (Holm 1) |
On b every aligned coordinate has a negative median, and T3's CI lies entirely below zero (Holm 0.0962). Only these five shapes and 20 seeds were tested. Before all of this, T1 found one small chart effect: the log chart on a, +0.000345 held-out HV (0.84% of raw), 15 wins and 5 losses, Holm p 0.0365. Its negative control is too imprecise to rule out a generic chart effect of that size, and "chart choice matters more than optimizer choice" is not established. Sources: T3/results.json, T4/results.json and T1/results.json primary.
Picking the chart automatically
With a pilot run (T2)
Reproducible: the same seeds give byte-identical scores, where CASCADE Phase A changes k on every run. But it mostly picks the domain restriction, and it costs more than the run it chooses for. Its choices are not shown to be safe.
- a: restricted theta 20; 607 evaluations against a 300-evaluation run
- b: restricted theta 16, log chart 2, CASCADE k = 0.0177 1, clipped-envelope theta 1; 607 evaluations against a 300-evaluation run
- c: raw 13, CASCADE k = 0.0177 7; 427 evaluations against a 300-evaluation run
Without one (T5)
Free: no objective evaluations for the decision, 912 lambda_max probes per pick where there is a limit. It almost always picks the restricted range where there is a limit, and raw where there is none. It cannot predict whether the restriction will help the final score: on the new problem d, its pick read no effect.
- a: LR 20 of 20 seeds; 0 evaluations to decide (3 for an optional spec check)
- b: LR 20 of 20 seeds; 0 evaluations to decide (3 for an optional spec check)
- c: raw 20 of 20 seeds; 0 evaluations to decide (3 for an optional spec check)
- d: LR 19, log chart 1 of 20 seeds; 0 evaluations to decide (3 for an optional spec check)
Sources: T2/results.json per_problem, T5/results.json choices and cost. Problem c is a negative control (DTLZ2, no stability limit).
The seven tests, one line each
Not established
Chart selection T1: is the chart effect real?
A chart effect on held-out HV is supported only for the log chart on problem a, and it is small (+0.000345, 0.84% of raw, Holm p 0.0365). The negative control is too imprecise to rule out a generic chart effect of that size, and "chart choice matters more than optimizer choice" is not established.
docs/research/chart-selection-2026-10/T1/FINDINGS.md
Not affordable
Chart selection T2: a pilot-based chart picker
The measured picker is reproducible, where CASCADE Phase A is not, but what it mostly picks is a domain restriction, not a chart. It is not affordable at this budget (427-607 evaluations against a 300-evaluation main run), and harmlessness is not established.
docs/research/chart-selection-2026-10/T2/FINDINGS.md
Inconclusive
Chart selection T3: aligned coordinate vs restricted domain
The only effect supported by the preregistered rule is the domain restriction on problem a (+0.000471, 1.14% of LF, Holm p 0.010). The stability-aligned coordinate is inconclusive on a and leans worse on b; a coordinate contribution is neither shown nor excluded.
docs/research/chart-selection-2026-10/T3/FINDINGS.md
Inconclusive
Chart selection T4: four more coordinate shapes
Across a power map, a piecewise-linear map, a Hermite map with the knot at u = 0.8 and a logistic map in log h, no stability-aligned coordinate on the full domain gave a detectable held-out HV gain at 20 seeds: 0 of 8 comparisons is a real effect, and none is excluded at the 1% margin. The coordinate effects are inconclusive, not excluded.
docs/research/chart-selection-2026-10/T4/FINDINGS.md
No effect
Chart selection T5: a zero-evaluation picker, tested prospectively
The picker costs 0 objective evaluations for its decision and almost always picks LR, the stability-restricted domain, wherever the problem has a limit. On the new problem d, LR - raw on held-out HV was "no effect" by the preregistered rule (median -0.00029, CI inside +-1% of raw, p 0.43). The rule tells which arm makes every step stable, not whether the restriction will help held-out HV.
docs/research/chart-selection-2026-10/T5/FINDINGS.md
Marginal
Chart selection T6: more seeds for the restriction
On 50 new optimizer seeds (21-70) with the same fixed datasets, the held-out gain on a is a real effect by the preregistered rule, but only barely: Holm p 0.0447, +0.000214 (0.52% of LF), and the bootstrap CI of the median difference includes 0. b and d are inconclusive.
docs/research/chart-selection-2026-10/T6/FINDINGS.md
No real effect
Chart selection T7: three new graph problems
On three new graph families, restricting h to [h_min, 0.98 h*(s)] showed no real effect on held-out HV at 300 evaluations: no effect on e, inconclusive on f and g. Across the six problems in the series, the held-out benefit is real only on a, and the unstable share does not order the effects (Spearman rho 0.086, n = 6, descriptive).
docs/research/chart-selection-2026-10/T7/FINDINGS.md
Method and limits
Each test fixed its protocol (PREREGISTRATION.md, or RULE.md for the pickers) before its scored runs. From T3 on, every run executed in a fresh process, the whole execution was repeated in reverse order, and the merge refused to compute statistics unless the two executions matched bit for bit. Every check had a known-bad control that had to fail. Deviations are listed in each FINDINGS.md.
All runs are offline pymoo NSGA-II; no GlobalMOO calls were made. Problems: a X4 spec B, 2-D grid, b path graph, d cycle graph, e 3-ary tree, f spider, g small world, each connection B with the five inputs (s, h, n, p, c), plus c, a DTLZ2 negative control. The primary endpoint throughout is held-out HV at 300 evaluations: hypervolume scored on noise realizations held out from training.
Sources
Write-ups and data: docs/research/chart-selection-2026-10/, folders T1 to T7, each with FINDINGS.md, PREREGISTRATION.md (or RULE.md) and results.json. Every number on this page is read at build time from those results.json files or computed from them (the x-fold ratios and the percentages of a base median); the copy the page imports, lib/chart-selection/data.json, is checked value by value against them in lib/chart-selection.test.ts.
This series and the other nulls on the site are listed together on Evidence.