At Akavion we make training runs fit. For teams with a job that needs more memory than their GPU has, we use a quantum method that runs on standard GPUs to store the job's training state more compactly, so that the same GPU holds a bigger training run.
Why memory
During training, a GPU has to keep several things in memory at once: the model's weights, the gradients and optimizer state for the parameters being trained, and the activations saved for the backward pass. When they add up to more than the GPU's memory, the run won't start. Teams usually get around it by picking settings that fit (a smaller batch or a shorter sequence), and they give up some training quality to do it. We wrote about that in The wrong shortage.
We think training state has a pattern, and that a pattern takes fewer bytes to store than the values it describes. Physicists worked out how to do this decades ago, for quantum systems too large to write down. These methods are called tensor networks. They find the pattern in a large object and store the pattern, and they run on ordinary computers. We wrote about where that comes from in It starts with the math.
We have two methods so far. One works on the model's weights and the other on the optimizer's state, and which one works best for a training run depends on the run itself.
Compressing the weights
a weight matrix16-bit4-bit + correction 4-bit weightslost precisioncorrection Fig. 1 · Four bits for each weight, plus a correction. Sketch, not to scale.We store most of each large weight matrix in 4 bits. The format is NF4, from the public QLoRA work, and we use it as published. A weight stored this way takes 4 bits where most training runs use 16. The model as a whole shrinks by less than that, because some layers stay at full precision and our correction takes space of its own.
Four bits store a number less precisely than sixteen. Across billions of weights the small errors add up, and the model gets worse. Our method measures what was lost, looks for the pattern in it, and trains a correction that stores that pattern. During training the 4-bit weights stay fixed and the correction learns.
If you know QLoRA, this will sound familiar. It also freezes a 4-bit base and trains a small correction. Ours is a different correction, and it comes from the tensor-network math. Akavion's work is in applying it to training, which means deciding what to compress and how far while keeping quality inside a limit the team sets. To your training code, the model is represented differently, and the data pipeline, the loss and the objective stay as they are.
Compressing the optimizer
directionkept exact sizetensor train Fig. 2 · Two values for each weight. One is kept whole. Sketch, not to scale.When every weight in a model is being trained, the optimizer becomes one of the largest things in memory. Adam, the most common optimizer family, keeps two running values for each weight. One sets the direction of the next update and the other sets its size, and in a standard mixed-precision setup they come to about half of the training state before activations.
Our optimizer method compresses them. The direction values are kept exact, since training is most sensitive to them. The size values are stored as a tensor train, a chain of small tensors that reconstructs the original state.
Which one fits a run
The two suit different jobs. If the base model is frozen in 4 bits and only a small correction is trained, the optimizer has little to keep track of, so the saving is in the weights. If every weight is being trained, the optimizer's state is large, and that is where the optimizer method saves memory.
What it costs
Compression comes at a cost, and we'll keep saying so.
compress lesscompress more memory savedspeedquality illustrative Fig. 3 · Compress more and you save more memory and give up some speed and quality.On the weights, most of the memory saving comes from the 4-bit format. Our correction is there to reduce the quality trade-off that 4-bit storage brings, and it takes a little memory and compute of its own.
QualityA compressed model is a slightly different model, and we measure how different on text it hasn't seen. The gap gets smaller as models get larger.
SpeedWith compressed weights, the GPU has to unpack them and apply the correction on every step, so a compressed run is slower than an uncompressed one on the same card, and small models pay the most. The optimizer method adds a little work to each step, and in our runs so far the cost has been negligible. The measured figure will be in our research notes.
You choose the trade-off, and we make your run fit.The more you compress, the more memory you save and the more speed and quality you give up. How far to go depends on your model and how much quality you can spare, and we work that out with you at the start of a pilot. You choose the trade-off, and we make your run fit.
How we measure
We decide what counts as success before a test runs, and we keep the failures in the record. Every result is compared against two other runs of the same model, on the same data and the same hardware. One is the uncompressed model in 16-bit, which is how most teams train today. The other is the simplest alternative anyone can use for free: plain 4-bit for the weights, ordinary compression at the same size for the optimizer. We always include that second comparison, because it shows what our method adds.
Memory is reported as three separate numbers: the size of the compressed model, the peak while it's being built and the peak once it's running. Quality is held-out NLL, a standard score for how well a model predicts text it hasn't seen, run across several random seeds and checked against a limit set before the run. For speed we count tokens per second once the run has warmed up.
What's next
Next is activations. They grow with batch size and sequence length, and for long-sequence and reinforcement-learning runs they can be where the memory runs out first. Our research on compressing activations is underway, and we'll update this page when the results are in.
Run a pilot with us
A pilot starts with a training job you care about and a quality limit you set. We fit the method to your model, run it against the two baselines, and send you the same tables we look at. Results from your run are shared with you and aren't published.
If you've been sizing runs to what fits, tell us what you train. Perhaps the run fits after all.
Last updated October 2026