Modular Norm RandOpt

Population-Efficient Ensembling
through Architecture-Aware Perturbations

Kirato YoshiharaHiroaki Hamade

The University of Osaka

Overview

A model’s architecture tells us how to explore its weights.
Use that structure to build stronger ensembles from fewer candidates.

RandOpt adds random Gaussian noise to pretrained weights, selects the best K candidates on a task, and combines their answers by voting.

Modular Norm RandOpt shapes these perturbations using each module’s natural norm, its place in the architecture, and calibrated sensitivity. The selection and voting steps stay the same.

Modular Norm RandOpt

Repeat independently for i = 1, …, 9

Repeat independently for i = 1 through 9: draw fresh Gaussian noise Zᵢ, normalize its direction with its module's natural norm and apply architecture- and sensitivity-based scaling to obtain Δθᵢ, then add it to the same pretrained weights θ to create one candidate θ′ᵢ. Each matrix illustrates a parameter tensor. Keep each completed candidate while drawing the next independent noise. Once all nine candidates are ready, retain candidates 3, 5, and 7 according to their selection scores, then combine their answers A, A, and B into A by plurality voting. Strengths, matrix entries, candidate identities, and answers are schematic. Fresh noise Zi Normalize + scale Shaped perturbation Δθi + Pretrained weights θ θ′i = θ + Δθi N candidates θ′1θ′2θ′3θ′4θ′5θ′6θ′7θ′8θ′9 Keep K models New input A A B A Sample noiseGaussian directions Modular perturbationsModule-wise norms + calibrated scales Select top KScore N candidates; keep K models VoteCombine their answers Repeat independently for i = 1 through 9: draw fresh Gaussian noise Zᵢ, normalize its direction with its module's natural norm and apply architecture- and sensitivity-based scaling to obtain Δθᵢ, then add it to the same pretrained weights θ to create one candidate θ′ᵢ. Each matrix illustrates a parameter tensor. Keep each completed candidate while drawing the next independent noise. Once all nine candidates are ready, retain candidates 3, 5, and 7 according to their selection scores, then combine their answers A, A, and B into A by plurality voting. Strengths, matrix entries, candidate identities, and answers are schematic. Fresh noise Zi Sample noiseDraw Gaussian directions Normalize + scale Shaped perturbation Δθi + Pretrained weights θ Modular perturbationsModule-wise norms + calibrated scales θ′i = θ + Δθi N candidates θ′1θ′2θ′3θ′4θ′5θ′6θ′7θ′8θ′9 Keep K models Select top KScore the population; keep an ensemble New input A A B A VoteCombine the selected models’ answers
One independent noise draw per candidate. Schematic example: N = 9, K = 3.

Scale each tensor in its own geometry.

RandOpt uses one noise scale across tensors. Modular Norm RandOpt gives each tensor its own size, measured in its own norm.

01Allocate

Where the tensor scales come from.

Computed once, then shared by every candidate.

  1. 01Start with the model

    Mass is an architectural weight, independent of parameter count. Only active tensors contribute to the total.

  2. 02Allocate to module groups

    Use the fixed group weights in Table A2. Longer colored marks mean more mass.

    Embedding1
    Attention½
    MLP½
    Head1
    Norm1/10
    Other1/10
  3. 03Split across repeated layers

    Attention and MLP each share their mass equally across 28 layers in Qwen2.5-1.5B. Expand one layer.

    Attention · 28 equal shares
    Layer 12 highlighted · mass 1/56
    MLP · 28 equal shares
    Layer 12 highlighted · mass 1/56
  4. 04Divide among tensors

    QKV counts as three shares; gate–up as two. Within each logical module, split its share equally over its physical tensors.

    Attention · QKV : O = 3 : 1
    QKVO
    MLP · gate/up : down = 2 : 1
    gateupdown
  5. 05Apply sensitivity correction

    ρp\rho_p is set once by a Jacobian-based calibration on 64 Countdown prompts and bounded to [12,2][\tfrac12,\,2]. See Section 3.2 and Appendix A.

    Worked example · Q, layer 12

    Tensor mass
    mp=MgLg×wuvwv×1num_p=\dfrac{M_g}{L_g}\times\dfrac{w_u}{\sum_v w_v}\times\dfrac1{n_u}
    mQ=1/228×14×1=1224m_{\mathrm Q}=\dfrac{1/2}{28}\times\dfrac14\times1=\dfrac1{224}
    sQ=mMmQρQs_{\mathrm Q}=\dfrac{m_{\mathcal M}}{m_{\mathrm Q}}\,\rho_{\mathrm Q}=3.2×224×ρQ=3.2\times224\times\rho_{\mathrm Q}=716.8ρQ=716.8\,\rho_{\mathrm Q}

    All six groups: mM=1+12+12+1+110+110=3.2m_{\mathcal M}=1+\tfrac12+\tfrac12+1+\tfrac1{10}+\tfrac1{10}=3.2

02Sample

Independent noise, every time.

Draw fresh Gaussian noise for each candidate and each tensor. Every candidate starts from the same pretrained weights.

Z1,p
Z2,p
Z3,p
A fresh draw for each candidate and each tensor.

03Measure

Measure what the tensor does.

Match the norm to the tensor’s role: the strongest linear direction, the largest embedding row, or the largest coordinate.

Linear mapsSpectral normThe strongest amplification direction.
Z2=σmax(Z)\lVert Z\rVert_2=\sigma_{\max}(Z)
=maxv2=1Zv2=\max_{\lVert v\rVert_2=1}\lVert Zv\rVert_2
EmbeddingsMaximum-row ℓ₂The largest Euclidean norm of any row.
maxrZr,:2\max_r\lVert Z_{r,:}\rVert_2
=maxrcZrc2=\max_r\sqrt{\textstyle\sum_c Z_{rc}^{\,2}}
1-D parametersℓ∞ normThe largest absolute coordinate.
z=maxjzj\lVert z\rVert_\infty=\max_j|z_j|

04Scale

Normalize the direction. Set its size.

Normalize to unit natural norm, then set the size to R/sₚ. The direction is preserved; the scale follows the architecture.

NoiseRaw noise
NormalizeUnit norm
ScaleSize R/sₚ
The direction stays the same; its natural-norm size changes.
RandOptΔθi,p=σZi,p\Delta\theta_{i,p}=\sigma Z_{i,p}One coefficient across tensors.

05Unchanged

Keep the same selection and voting.

Evaluate N candidates, keep the top K, and combine their answers by plurality voting. Only the perturbation geometry changes.

N candidates θ′1θ′2θ′3θ′4θ′5θ′6θ′7θ′8θ′9 Keep K models New input A A B A
The same selection and voting procedure as RandOpt.

Stronger ensembles. From fewer candidates.

Qwen2.5-1.5B-Instruct

Population scaling on Countdown and GSM8K Qwen2.5-1.5B-Instruct. Four panels: Countdown above GSM8K; K=10 on the left and K=25 on the right. Population sizes 25, 50, 100, 200, 300 are equally spaced categories. Curves show means; shaded bands show one sample standard deviation across seeds 42, 43, 44. Modular Norm RandOpt at N=100 exceeds RandOpt at N=300 on Countdown by 1.53 and 0.93 percentage points; at N=25 it exceeds RandOpt at N=300 on GSM8K by 2.60 and 3.11 points. These are comparisons of means, not statistical significance claims. Populations are nested prefixes of the same 300 candidates per run. RandOptModular Norm RandOpt Ensemble accuracy (%) Countdown · K=10. 3× fewer candidates; mean accuracy difference +1.53 percentage points. Countdown · K=10253035402550100200300 RandOpt: 30.00 ± 3.08% at N=25RandOpt: 31.27 ± 2.39% at N=50RandOpt: 32.80 ± 1.04% at N=100RandOpt: 33.47 ± 1.22% at N=200RandOpt: 33.33 ± 1.21% at N=300Modular Norm RandOpt: 32.93 ± 1.10% at N=25Modular Norm RandOpt: 32.67 ± 2.02% at N=50Modular Norm RandOpt: 34.87 ± 2.27% at N=100Modular Norm RandOpt: 36.93 ± 1.80% at N=200Modular Norm RandOpt: 35.13 ± 1.85% at N=300 3× fewer candidates Modular Norm RandOpt N=100 vs RandOpt N=300 · +1.53 pt Countdown · K=25. 3× fewer candidates; mean accuracy difference +0.93 percentage points. Countdown · K=252550100200300 RandOpt: 32.80 ± 2.69% at N=25RandOpt: 34.33 ± 1.55% at N=50RandOpt: 36.67 ± 1.55% at N=100RandOpt: 36.67 ± 1.50% at N=200RandOpt: 36.87 ± 1.27% at N=300Modular Norm RandOpt: 29.00 ± 2.25% at N=25Modular Norm RandOpt: 35.07 ± 1.86% at N=50Modular Norm RandOpt: 37.80 ± 3.22% at N=100Modular Norm RandOpt: 39.73 ± 1.10% at N=200Modular Norm RandOpt: 38.73 ± 1.86% at N=300 3× fewer candidates Modular Norm RandOpt N=100 vs RandOpt N=300 · +0.93 pt GSM8K · K=10. ≥12× fewer candidates; mean accuracy difference +2.60 percentage points. GSM8K · K=106466687072742550100200300 RandOpt: 66.64 ± 0.53% at N=25RandOpt: 66.79 ± 1.14% at N=50RandOpt: 67.12 ± 1.01% at N=100RandOpt: 68.16 ± 1.36% at N=200RandOpt: 68.36 ± 1.39% at N=300Modular Norm RandOpt: 70.96 ± 0.84% at N=25Modular Norm RandOpt: 71.42 ± 0.26% at N=50Modular Norm RandOpt: 71.17 ± 0.79% at N=100Modular Norm RandOpt: 72.10 ± 0.30% at N=200Modular Norm RandOpt: 72.91 ± 0.22% at N=300 ≥12× fewer candidates Modular Norm RandOpt N=25 vs RandOpt N=300 · +2.60 pt GSM8K · K=25. ≥12× fewer candidates; mean accuracy difference +3.11 percentage points. GSM8K · K=252550100200300 RandOpt: 66.44 ± 0.69% at N=25RandOpt: 67.73 ± 0.32% at N=50RandOpt: 68.11 ± 0.61% at N=100RandOpt: 68.39 ± 0.87% at N=200RandOpt: 68.54 ± 0.77% at N=300Modular Norm RandOpt: 71.65 ± 0.15% at N=25Modular Norm RandOpt: 72.55 ± 0.60% at N=50Modular Norm RandOpt: 73.16 ± 0.66% at N=100Modular Norm RandOpt: 73.39 ± 0.50% at N=200Modular Norm RandOpt: 73.34 ± 0.54% at N=300 ≥12× fewer candidates Modular Norm RandOpt N=25 vs RandOpt N=300 · +3.11 pt Candidate population size N Population scaling on Countdown and GSM8K Qwen2.5-1.5B-Instruct. Four panels in one column: Countdown K=10, Countdown K=25, GSM8K K=10, GSM8K K=25. Population sizes 25, 50, 100, 200, 300 are equally spaced categories. Curves show means; shaded bands show one sample standard deviation across seeds 42, 43, 44. Modular Norm RandOpt at N=100 exceeds RandOpt at N=300 on Countdown by 1.53 and 0.93 percentage points; at N=25 it exceeds RandOpt at N=300 on GSM8K by 2.60 and 3.11 points. These are comparisons of means, not statistical significance claims. Populations are nested prefixes of the same 300 candidates per run. RandOptModular Norm RandOpt Countdown · K=10. 3× fewer candidates; mean accuracy difference +1.53 percentage points. Countdown · K=10253035402550100200300Accuracy (%)Candidate population N RandOpt: 30.00 ± 3.08% at N=25RandOpt: 31.27 ± 2.39% at N=50RandOpt: 32.80 ± 1.04% at N=100RandOpt: 33.47 ± 1.22% at N=200RandOpt: 33.33 ± 1.21% at N=300Modular Norm RandOpt: 32.93 ± 1.10% at N=25Modular Norm RandOpt: 32.67 ± 2.02% at N=50Modular Norm RandOpt: 34.87 ± 2.27% at N=100Modular Norm RandOpt: 36.93 ± 1.80% at N=200Modular Norm RandOpt: 35.13 ± 1.85% at N=300 3× fewer candidates MN N=100 vs RandOpt N=300 · +1.53 pt Countdown · K=25. 3× fewer candidates; mean accuracy difference +0.93 percentage points. Countdown · K=25253035402550100200300Accuracy (%)Candidate population N RandOpt: 32.80 ± 2.69% at N=25RandOpt: 34.33 ± 1.55% at N=50RandOpt: 36.67 ± 1.55% at N=100RandOpt: 36.67 ± 1.50% at N=200RandOpt: 36.87 ± 1.27% at N=300Modular Norm RandOpt: 29.00 ± 2.25% at N=25Modular Norm RandOpt: 35.07 ± 1.86% at N=50Modular Norm RandOpt: 37.80 ± 3.22% at N=100Modular Norm RandOpt: 39.73 ± 1.10% at N=200Modular Norm RandOpt: 38.73 ± 1.86% at N=300 3× fewer candidates MN N=100 vs RandOpt N=300 · +0.93 pt GSM8K · K=10. ≥12× fewer candidates; mean accuracy difference +2.60 percentage points. GSM8K · K=106466687072742550100200300Accuracy (%)Candidate population N RandOpt: 66.64 ± 0.53% at N=25RandOpt: 66.79 ± 1.14% at N=50RandOpt: 67.12 ± 1.01% at N=100RandOpt: 68.16 ± 1.36% at N=200RandOpt: 68.36 ± 1.39% at N=300Modular Norm RandOpt: 70.96 ± 0.84% at N=25Modular Norm RandOpt: 71.42 ± 0.26% at N=50Modular Norm RandOpt: 71.17 ± 0.79% at N=100Modular Norm RandOpt: 72.10 ± 0.30% at N=200Modular Norm RandOpt: 72.91 ± 0.22% at N=300 ≥12× fewer candidates MN N=25 vs RandOpt N=300 · +2.60 pt GSM8K · K=25. ≥12× fewer candidates; mean accuracy difference +3.11 percentage points. GSM8K · K=256466687072742550100200300Accuracy (%)Candidate population N RandOpt: 66.44 ± 0.69% at N=25RandOpt: 67.73 ± 0.32% at N=50RandOpt: 68.11 ± 0.61% at N=100RandOpt: 68.39 ± 0.87% at N=200RandOpt: 68.54 ± 0.77% at N=300Modular Norm RandOpt: 71.65 ± 0.15% at N=25Modular Norm RandOpt: 72.55 ± 0.60% at N=50Modular Norm RandOpt: 73.16 ± 0.66% at N=100Modular Norm RandOpt: 73.39 ± 0.50% at N=200Modular Norm RandOpt: 73.34 ± 0.54% at N=300 ≥12× fewer candidates MN N=25 vs RandOpt N=300 · +3.11 pt
Mean ± sample SD over 3 seeds. Each N uses a nested prefix of the same 300 candidates per run.
Candidate counts are equally spaced on the x-axis. Highlighted comparisons are differences in mean accuracy.

Performance across seven tasks.

N = 100 · K = 25 · Mean over 3 seeds · Scores (%)

Qwen2.5-1.5B-Instruct scores across seven tasks. Three colored series show mean scores on a shared 0–100% axis, with guidelines every 20 points. Interactive values are available when JavaScript is enabled.

Comparison with iterative baselines

Qwen2.5-1.5B-Instruct

Countdown

1,500 test examples

1500 test examples. Points show mean accuracy over three seeds. The horizontal axis is linear in total model–prompt evaluations: one model generating and scoring one prompt. Main-run totals include search, checkpoint selection, and final evaluation, following Appendix D.2. MN-RandOpt is shown at N=100, K=25; N=3000, K=25; and N=3000, K=1. Iterative methods and the MN K=1 control return single models. Missing scores are omitted, not plotted as zero. Accuracy (%)10%20%30%40%0200k400k600k38.40% at 57,500 evals10.46× fewer evals vs ES-at-Scale · +2.73 ppN=100

GSM8K

1,319 test examples

1319 test examples. Points show mean accuracy over three seeds. The horizontal axis is linear in total model–prompt evaluations: one model generating and scoring one prompt. Main-run totals include search, checkpoint selection, and final evaluation, following Appendix D.2. MN-RandOpt is shown at N=100, K=25; N=3000, K=25; and N=3000, K=1. Iterative methods and the MN K=1 control return single models. Missing scores are omitted, not plotted as zero. Accuracy (%)60%65%70%75%0200k400k600k73.16% at 52,975 evals11.35× fewer evals vs ES-at-Scale · +0.05 ppN=100

Total model–prompt evaluations

One evaluation is one model generating and scoring one prompt. Totals include search, checkpoint selection, and final evaluation (Appendix D.2).
Points show mean accuracy over 3 seeds. Iterative baselines use K=1.