How it works.
We store a training job's state more compactly, so the same GPU holds a bigger training run.
The weights
Four bits for each weight, plus a correction.
Most of each large weight matrix is stored in 4 bits, in the public NF4 format. Our correction stores the pattern in what 4 bits lost.
The optimizer
Two values for each weight. One is kept whole.
Adam keeps a direction and a size for every weight. Direction stays exact. Size is stored as a tensor train, a chain of small tensors.
How we measure
Three numbers, two baselines.
Memory, speed and quality are measured on the same run and compared with the uncompressed 16-bit model and the simplest free alternative. The target is set before the run starts.
What we need from you
You set the quality limit. We make the run fit.
Tell us the training job and the quality you need to keep. We choose the method and how far to compress. Your data pipeline, loss and objective stay as they are.