How we measure.

densethroughputdensepeak memorydensequalitythroughput · peak memory · held-out quality, against dense precision

Three axes, one control.

Every method is measured on throughput, peak memory and held-out quality, against the same control: the same model trained at dense precision on the same hardware. Every axis is oriented so that higher is better, so a cost reads as a negative number and nobody has to guess.

What counts as a result.

A method has a result when all three numbers are measured on the same run. A memory number on its own is a planning value, not a result. Proxies don't substitute: reconstruction error is not held-out loss, and timing one component is not training throughput.

Across sizes.

We run the same suite across a range of model sizes so the effect reads as a trend rather than a point, and repeat seeds where the result is close. Every run is recorded, including the ones that failed.

Where results are today.

We share results with the teams we're working with, and publish here as the measurements come in.