Length-Bucketed Batching: How to Stop Wasting GPU Cycles on Padding for Small Language Models
The padding problem: why naive batching wastes compute
Processing one ticket per forward pass is described in the article as “the single largest source of waste in the whole pipeline.” The reason is memory bandwidth. At batch size 1, a small model is memory-bandwidth bound rather than compute-bound: the hardware streams every weight out of memory to serve a single sequence, then does it all again for the next one. The arithmetic units are left mostly twiddling their thumbs. The article notes this applies on both GPU and CPU, and CPU is where a 0.5B model most often actually runs.
Batching fixes the bandwidth problem by amortising the weight read across many sequences at once. But naive batching introduces a second problem. Sequences in a batch must be padded to a common length, and real-world text has a long tail. If the longest item in your queue is a few hundred tokens and the median is well under a hundred, padding every batch to the global maximum means most of the computation is spent on filler.
Think about a support inbox. A handful of tickets carry long pasted logs and multi-paragraph histories. The bulk are two or three sentences. Pad them all to match the worst case and you have built a pipeline that spends the majority of its arithmetic on padding tokens that contribute nothing to the answer.

Length-bucketed batching explained
The fix is almost embarrassingly simple. Sort sequences by token length, group them into batches, and pad each batch only to its local maximum rather than the global one.
That is the whole technique. It is a scheduling change, not a modelling change. The weights are untouched, the tokeniser is untouched, and the predictions should be identical. What changes is how much padding each batch carries. A batch of short tickets pads to the longest short ticket; a batch of long ones pads to the longest long one. The waste is contained within each bucket instead of being smeared across the entire run.
Benchmark setup
The benchmarks use Qwen2.5-0.5B-Instruct in float16 via Hugging Face Transformers, running on an M2 MacBook Air with 24GB of RAM and a 16-core Neural Engine, with execution on the CPU. That is a deliberately realistic edge-deployment scenario: a half-billion-parameter model on laptop-class hardware, which is precisely the configuration where a 0.5B model tends to end up in production.
The task is support ticket processing, carried over from the first article in the series, and the constrained scoring technique from that entry is carried forward too. Each item therefore costs exactly one forward pass, which keeps the comparison clean.

What the results show
The batched version runs the same data twice. The first pass uses arbitrary order, isolating the effect of batching alone. The second sorts by length, showing what the sort adds on top. Both runs track how much of the processed token budget went to padding.
Because the two batched runs perform identical arithmetic per real token and differ only in the padding they carry, the gap between them is a direct measurement of the sort’s benefit rather than a confounded comparison. The author reports a large throughput increase on identical hardware and an identical model, purely from scheduling, and processing the same 600 tickets in a fraction of the wall-clock time with identical predictions. Exact throughput and speedup figures are not published alongside the piece, so treat the magnitude as indicative rather than a benchmark you can quote. For your own workload, measure it, and the methodology above tells you exactly how.
The KV cache wrinkle
Combining this with the prefix caching from article two is where things get fiddly:
- The KV cache built in that earlier piece has a batch dimension of 1.
- Reusing it across a batch means expanding every key and value tensor along that dimension.
- You then have to crop it back correctly afterwards.
That is doable, but it is not free and it is not automatic. The author’s advice is to do it deliberately when the prefix is long, and to verify predictions against the unbatched path rather than assuming the optimisations compose cleanly. Two techniques that each preserve output in isolation can still interact badly once you start reshaping tensors by hand.

The verification principle
Which brings us to the line that ought to be pinned above every optimisation desk:
“An optimization that changes your predictions is not an optimization, it is a regression with a stopwatch attached.”
None of the three techniques in this series makes the model smarter. Each was accepted only after its output was checked against the slower path it replaced. That discipline is what separates a genuine speedup from a quietly degraded system that happens to finish faster. If you cannot demonstrate output equivalence, you have not optimised anything. You have changed the product and measured the wrong thing.
Practical takeaways for UK engineers
Reach for length-bucketed batching whenever you are looping over items and the lengths vary. That covers most document processing, ticket triage, classification and extraction workloads. If your sequences are already near-uniform in length, sorting buys you little and you can skip it.
It sits naturally alongside constrained output and prefix caching. Constrained output keeps each item to one forward pass; prefix caching avoids recomputing a shared system prompt; length bucketing ensures the batch you build is not mostly padding. Together they attack three different sources of waste without touching the model.
The UK angle here is not cosmetic. Plenty of British teams run small models on CPU-only instances, on-premise boxes or modest cloud tenancies, often because data residency or cost rules out shipping sensitive material to a large managed endpoint. NHS trusts, local councils and smaller public sector suppliers are the obvious examples, where citizen or patient data is easier to keep inside a controlled environment than to send out. On that kind of hardware, a scheduling win is the cheapest capacity you will ever buy, because you are not paying for a new accelerator or a bigger instance. You are paying for a sort.
Before trusting any throughput claim, including your own, measure three things: wall-clock time per item, the proportion of the token budget spent on padding, and whether outputs match the unbatched baseline exactly. The first tells you if it is faster. The second tells you why. The third tells you whether the result is still correct.
For teams running small models on CPUs, edge boxes or modest cloud instances, the appeal here is that the gain comes from scheduling rather than silicon. You do not need a new GPU to benefit. You need to sort a list.
One Comment
Comments are closed.