How it works.

We store a training job's state more compactly, so the same GPU holds a bigger training run.

The weights

a weight matrix16-bit4-bit + correction 4-bit weightslost precisioncorrection

Four bits for each weight, plus a correction.

Most of each large weight matrix is stored in 4 bits, in the public NF4 format. Our correction stores the pattern in what 4 bits lost.

The optimizer

directionkept exact sizetensor train

Two values for each weight. One is kept whole.

Adam keeps a direction and a size for every weight. Direction stays exact. Size is stored as a tensor train, a chain of small tensors.

How we measure

16-bitpeak memory16-bitthroughput16-bitquality

Three numbers, two baselines.

Memory, speed and quality are measured on the same run and compared with the uncompressed 16-bit model and the simplest free alternative. The target is set before the run starts.

What we need from you

compress lesscompress more memory savedspeedquality illustrative

You set the quality limit. We make the run fit.

Tell us the training job and the quality you need to keep. We choose the method and how far to compress. Your data pipeline, loss and objective stay as they are.