GlobalMOO vs pymoo NSGA-II, October 2026
We had a headline. Then we tested it fairly.
What GlobalMOO is
GlobalMOO is a hosted inverse-design service. You give it the input box, it picks a set of initial cases, you report their outputs and set one objective per output, and it then proposes one input at a time until the target is satisfied. We wanted it for expensive evaluations: when one run of a model costs seconds to hours, a service that learns from every result could, in principle, reach a good answer in fewer runs. That was the hope. This page is how we checked it.
The June claim, and why it was not a fair test
On 2026-06-12 the site reported GlobalMOO at chi-squared 0.00035 against pymoo's 0.061: "99.4% lower chi-squared". Tracing the code that produced it showed three ways the two engines were treated differently:
- Budgets not verified equal. pymoo got 180 measured evaluations. GlobalMOO was configured for up to 200 loop iterations, plus initial cases that were neither counted nor scored. Its actual total was not recorded.
- Different objectives. Fragility was on for GlobalMOO and off for pymoo, so the two engines optimized different objective vectors.
- Different scoring. GlobalMOO's best value came from its loop history only; its initial cases were excluded.
Separately, the benchmark itself had a weakness. Its optimum sat at the exact centre of the box (k = 0, chi2 = 0), and GlobalMOO's initial design contained the exact box centre in every run we have: 9 of 9 X4 replicates, 2 of 2 accidental runs and 10 of 10 runs in the rerun. That is a confound of the June benchmark. It is not a measured cause of the historical figure: the historical score used the loop history only and excluded the initial cases.
How we tested it fairly
- Preregistered before any scored run. The design, the statistics and the reading rule were fixed and hashed first (revision 2.4).
- Reviewed by two independent auditors. Sol and Astra reviewed the preregistration over several rounds and approved revision 2.4 with no findings.
- Equal measured budgets. 200 objective evaluations per run for each engine, with every GlobalMOO initial case counted.
- Optimum moved off the centre. Each pair shifted the chi2 optimum in k by its own random amount (up to 0.024 either way); both engines in a pair saw the same shift.
- Fragility off for both. Both engines optimized the same objectives.
- A reproducibility bug fixed. The June code cached operators under Python object ids that get reused, so its spectral outputs depended on the process's allocation history. The cache is now cleared before every evaluation, in both engines.
- Ten paired runs, pair seeds 11 to 20, scored by the best chi2 over every evaluation of each run.
The result
Preregistered reading: inconclusive.The preregistered test, a two-sided Wilcoxon signed-rank test over the 10 pairs, gave p = 0.105. The rule needed p < 0.05 to call an effect real, so this study makes no "real effect" claim in either direction.
Descriptive direction, not a test result
pymoo's median best chi2 was 0.00320 against GlobalMOO's 0.0428. GlobalMOO was lower in 2 of 10 pairs. The bootstrap 95% CI of the median difference (GlobalMOO minus pymoo, +0.0375) is [+0.0087, +0.0745], entirely above zero.
Method: the June problem with its optimum shifted off the centre; the best chi2 over all evaluations of each run; fragility off; 10 pairs x 200 evaluations. It does not cover today's chi2 problem, other budgets or problems, or any general ranking of the two optimizers. At 10 pairs the test's simulated power is 0.60 at effect size 0.8, so a moderate real difference could read as inconclusive.
pymoo NSGA-IIGlobalMOO
Best chi2 over 200 evaluations, log scale
11GlobalMOO best at evaluation 179 (guided iteration)
0.01961 vs 0.003643
12GlobalMOO best at evaluation 169 (guided iteration)
0.0002889 vs 0.03368
13GlobalMOO best at evaluation 103 (initial design)
0.2285 vs 0.03469
14GlobalMOO best at evaluation 46 (initial design)
0.01480 vs 0.2860
15GlobalMOO best at evaluation 15 (initial design)
0.0005570 vs 0.09492
16GlobalMOO best at evaluation 40 (initial design)
0.003562 vs 0.01753
17GlobalMOO best at evaluation 55 (initial design)
0.03189 vs 0.07359
18GlobalMOO best at evaluation 156 (guided iteration)
0.002838 vs 0.02499
19GlobalMOO best at evaluation 80 (initial design)
0.0006259 vs 0.07508
20GlobalMOO best at evaluation 199 (guided iteration)
0.00008595 vs 0.05082
Each row is one pair: both engines on the same shifted problem with the same noise seed. Values are the best chi2 of each run, shown to four significant figures. Source: docs/research/globalmoo-equal-budget-2026-10/runs/summary.json and the per-run records.
What we learned about how GlobalMOO behaves
For 4 inputs the service opens with a 141-case initial design, box centre included, and the API has no setting for its size. At a 200-evaluation budget that leaves 59 guided iterations. In 6 of 10 runs (seeds 13, 14, 15, 16, 17, 19) GlobalMOO's best point was one of its initial cases, not a guided suggestion. The centre was in the initial design in 10 of 10 runs.
An earlier toy experiment (X4, connection B) points the same way. Reparameterizing the step size to stay inside its stable range removed the blow-up penalties that had kept the service's first run stuck in an unstable region. But in all 6 replicates on the reformulated problem, the training hypervolume after its 151 initial cases equals the hypervolume at evaluation 250, to every printed digit: the 99 native guided iterations before the service's stop added nothing measurable. Source: docs/research/nnc-openai-crossanalysis-2026-10/X4/GLOBALMOO-B2.md.
An incident, recorded
During setup, a test that deliberately removed safety gates made real GlobalMOO calls by accident: 400 evaluations measured, 600 booked. The run was noticed because it was slow, and stopped. Pairs 1 and 2 were seen before the real run, so seeds 1 to 3 were excluded and the scored run used seeds 11 to 20 under a fresh cap. The mutation tool now disables the paid client before it changes anything, and every control, mutation and dry-run process runs with the key file path pointed at a file that does not exist and with a flag that makes the client refuse to start. The full note is INCIDENT-2026-10-09.md in the research folder.
We tested the hybrid idea
The way we actually want to use GlobalMOO is a hybrid, for expensive evaluations on rugged problems with many local optima: let GlobalMOO map the search space first, then run pymoo where it looks promising. We tested that under the same safeguards as the rerun, preregistered and approved by Sol and Astra before any scored run. Every scored approach spent exactly 300 measured evaluations:
- pymoo alone: NSGA-II for all 300.
- The hybrid: one GlobalMOO run of 200 evaluations (its 141 fixed initial cases plus 59 guided iterations), then 100 NSGA-II evaluations seeded from the best 20 distinct stage-1 points.
- The cheap-sweep control: 200 Latin hypercube points, then the same pymoo stage under the same rule.
- GlobalMOO alone (200 evaluations), recorded as a descriptive reference only.
The problems were two rugged benchmarks at 4 inputs, ZDT4 and DTLZ1, each with its optimum shifted off the box centre per seed. On both, an offline pilot showed that pymoo struggles at this budget. There were 10 paired seeds (101-110), scored by IGD+ against the known front, lower is better, with Holm applied across the two problems.
Preregistered reading: all 4 primary comparisons inconclusive.Every 95% CI of the median difference includes zero, and no Holm-adjusted p is below 0.05. The study supports no claim that the hybrid beats pymoo alone, or that GlobalMOO's sweep beats a cheap Latin hypercube sweep, on either problem.
| Comparison | Median IGD+, base vs hybrid | 95% CI of median d | Hybrid lower | Holm p | Reading |
|---|---|---|---|---|---|
| Hybrid vs pymoo alone, ZDT4 | 3.317 vs 4.158 | [-0.676, +2.633] | 4 of 10 | 0.387 | inconclusive |
| Hybrid vs pymoo alone, DTLZ1 | 11.77 vs 6.488 | [-8.516, +3.152] | 6 of 10 | 0.387 | inconclusive |
| Hybrid vs cheap sweep, ZDT4 | 4.468 vs 4.158 | [-1.914, +1.906] | 6 of 10 | 0.922 | inconclusive |
| Hybrid vs cheap sweep, DTLZ1 | 11.13 vs 6.488 | [-6.540, +0.888] | 7 of 10 | 0.387 | inconclusive |
Descriptive direction, not a test result
On DTLZ1 the hybrid had the lowest median of the three scored approaches (6.488, against 11.77 for pymoo alone and 11.13 for the cheap sweep), and it won 6 of 10 pairs against pymoo alone and 7 of 10 against the cheap sweep. On ZDT4 pymoo alone had the lower median (3.317 against the hybrid's 4.158), and the hybrid won 4 of 10 pairs.
The narrowing step barely narrowed. The preregistered rule took the bounding box of the 20 selected stage-1 points and widened it by 10% of the original range. On DTLZ1 that box kept about 99% of the original volume (median volume fraction 0.987 for the hybrid, 0.975 for the cheap sweep), and in 7 of 20 DTLZ1 runs it was the whole original box (hybrid seeds 103, 104, 106, 107; cheap-sweep seeds 103, 109, 110). So stage 2 was in practice NSGA-II seeded with stage 1's best points, not a search inside a smaller region, and the DTLZ1 direction cannot be credited to narrowing.
Limits: two shifted standard benchmarks, not a real problem; one budget regime (300 evaluations per run, a 200-evaluation GlobalMOO stage, a 100-evaluation pymoo stage); GlobalMOO's 141 fixed initial cases are 70% of its stage, so its "sweep" is mostly that fixed design; power is modest (0.78 at effect size 1.0, or 0.61 at the first Holm step), so moderate effects can read as inconclusive; not a general ranking of the optimizers. GlobalMOO evaluations charged: 4000. Source: docs/research/globalmoo-hybrid-2026-10/FINDINGS.md and docs/research/globalmoo-hybrid-2026-10/runs/.
In plain terms
- The June headline compared the two engines with budgets not verified equal, and with different objectives and scoring. It does not tell you which one is better.
- Given the same 200 evaluations on the same problem with its optimum moved, the preregistered test could not tell them apart (p = 0.105).
- Descriptively, the direction favours pymoo: GlobalMOO was lower in only 2 of 10 pairs. That is a description, not a test result.
- At this budget, 141 of GlobalMOO's 200 evaluations went to the opening design the service chooses.
- The hybrid idea, a GlobalMOO sweep and then pymoo, was tested on two rugged benchmarks. All 4 preregistered comparisons were inconclusive, and the narrowing step barely narrowed.
Where this is useful
Each line sits in one tier: measured, true under named assumptions, or an untested proposal.
Demonstrated
With equal measured budgets, every evaluation counted and the optimum moved off the centre, this study found no GlobalMOO advantage on the June problem (
docs/research/globalmoo-equal-budget-2026-10/FINDINGS.md).Demonstrated
The service ran reliably here: 10 of 10 runs completed 200 evaluations with 0 failed requests and 0 retries, in 41 to 51 s each (
docs/research/globalmoo-equal-budget-2026-10/FINDINGS.md).Applies under assumptions
If your budget is close to the size of the opening design, most of it goes to that design. This assumes the service sends the same 141 cases for 4 inputs that it sent in every run we recorded; the API exposes no size setting.
Applies under assumptions
If one evaluation takes milliseconds, the service round trip dominates. In X4, with a 4 to 5 ms objective, the round trip was about 100x the evaluation cost; this assumes similar network latency to ours.
Demonstrated
On two shifted rugged benchmarks at 300 evaluations, the hybrid showed no preregistered advantage over pymoo alone or over a cheap Latin hypercube sweep: all 4 comparisons were inconclusive (
docs/research/globalmoo-hybrid-2026-10/FINDINGS.md).Worth trying
The hybrid on a real expensive problem, or with a stage-2 rule that actually shrinks the search box. Not tested here.
Worth trying
Budgets well above the opening design, where guided iterations are most of the run. Not tested here.
Every number on this page is rendered from the run records in docs/research/globalmoo-equal-budget-2026-10/runs/, from public/globalmoo_vs_fallback_baseline.json (the June figures), or quoted from the FINDINGS files named above. The June figures stay on the synergy page and the simulator as history; the Evidence ledger lists this result with the other null and negative ones.