Block removal for LLM pruning framed as Ising optimization
Researchers reformulate transformer block removal as constrained binary optimization mapped to an Ising spin-glass model, letting them rank pruning candidates without benchmarking each one. At 50% compression of Llama-3.3-70B-Instruct, it gains nearly 23 percentage points on MMLU over the best competing method.