LLM Compression by Block Removal with Constrained Binary Optimization

TL;DR

Hugging Face announced a pruning technique that casts transformer block removal as an Ising‑glass constrained binary optimization problem, enabling fast, high‑quality depth compression and delivering up to a 23‑point MMLU gain on Llama‑3.3‑70B‑Instruct at 50 % compression.

Why block selection is a many‑body problem

Existing depth‑pruning methods rank individual blocks using heuristics such as magnitude or sensitivity, treating each block as independent (a mean‑field approximation). In reality, the impact of removing a block depends on which other blocks are also removed, creating pairwise couplings analogous to spin interactions in a magnet. Ignoring these couplings discards valuable information, especially when large numbers of blocks are pruned.

Turning block selection into an energy‑minimization problem

  • Attach a binary variable (x_i) to each transformer block (0 = keep, 1 = remove).

  • Perform a second‑order Taylor expansion of the model loss with respect to (x), yielding an approximate Hessian (H).

  • Diagonal entries of (H) capture individual block importance; off‑diagonal entries capture pairwise couplings.

  • The pruning objective becomes:

    [ \min_{x}; x^{\top} H^{0} x \quad \text{s.t.}; \sum_i x_i = M ]

    This is a constrained binary optimization (CBO) problem, mathematically identical to finding low‑energy states of an Ising glass with a fixed magnetization (the number of removed blocks). Crucially, low‑energy configurations strongly correlate with high downstream benchmark scores, making energy a cheap proxy for model quality.

Efficient evaluation of candidate configurations

The full Hessian is computed once from forward and backward passes on a small calibration dataset. Afterward, evaluating any block‑removal configuration requires only a matrix‑vector multiplication, eliminating the need for full model inference or benchmark evaluation. The same Hessian can be reused for different compression targets (M).

Solving the optimization problem

  • Exact brute‑force: For modest configuration spaces (e.g., removing 8 of 80 blocks in Llama‑3.3‑70B), enumerating billions of configurations on a single GPU is feasible; the hardest case took ~2 days.
  • Quantum‑inspired solvers: Larger spaces are tackled by converting the CBO to a QUBO (adding a penalty for the cardinality constraint) and feeding it to classical, quantum, or quantum‑inspired optimizers such as tabu search, quantum annealing, and QAOA. An open‑source tabu solver consistently reaches the lowest‑energy states within seconds on the hardest verified instances.
  • Practical goal: Rather than insisting on the global ground state, the pipeline extracts a spectrum of low‑energy states, providing multiple high‑quality pruning candidates with minimal additional cost.

Importance of the low‑energy spectrum

Because energy is an imperfect proxy, the absolute ground state is not always the best pruned model. By sampling several low‑lying excited states, practitioners obtain a diverse set of configurations. In experiments on Llama‑3.1‑8B‑Instruct, the 17th excited state—removing an early block—outperformed the ground state after light retraining, disproving the assumption that optimal pruning always removes a contiguous late‑layer chunk.

Empirical results

Model Blocks removed MMLU
Llama‑3.3‑70B‑Instruct (original) 0 82.2
CBO (ours) 32 / 80 76.6
Block influence (baseline) 32 / 80 59.3
CBO (ours) 40 / 80 76.9
Block influence (baseline) 40 / 80 54.0
  • At 50 % depth compression (40/80 blocks removed), CBO retains an MMLU of ~77, while the strongest baseline falls to the mid‑50s.
  • Similar gains appear on Qwen3‑14B (≈10 MMLU points at 12/40 blocks removed) and on lighter compression settings where all methods converge.

Generalization to heterogeneous architectures

The Ising formulation does not depend on block homogeneity; it only requires a coupling matrix. Applying CBO to NVIDIA‑Nemotron‑3‑Nano‑30B‑A3B‑FP8—a hybrid model interleaving Mamba2, attention, and MoE layers—produced pruned configurations that outperformed block‑influence baselines on AIME25 and GPQA without retraining. The results confirm that redundancy varies across expert layers, and that exploring the coupled configuration space uncovers the most disposable components.

Fit with Multiverse Computing’s stack

Recasting pruning as an Ising Hamiltonian leverages Multiverse Computing’s existing classical and quantum‑inspired optimization infrastructure. Block removal composes naturally with other compression techniques (quantization, low‑rank SVD, width pruning, knowledge‑distillation healing), forming a unified pipeline rather than a competing approach.

Accessing the research

Sources