Every year, more quantization papers promise nearly lossless 4-bit models. But a small bit width does not always mean a simple deployment: rotations, low-rank branches, and hardware-specific formats can turn a compact model into a complicated inference pipeline.
From the paper: a survey of ≤4-bit activation quantization across eight conferences (2024–2026). Red: extra operators.Green: no extra operator.Full figure
Can intelligence live in a body made of plain 4-bit operations?
Naive low-bit matrix multiplication. No rotation, no SVD branch, no smoothing. Move the complexity into training, and keep the deployed operator simple.
W4A4 / W1.58A4 / PLAIN LOW-BIT GEMM
03 / THE EVIDENCE
Start with the results.
None = no additional inference operator /Named operators = additional transformations. Highlighted rows are ours.
IMAGE GENERATION / 0.6B
DiT-XL/2
ImageNet 256 W4A4 and W1.58A4
Plain W4A4 reaches 6.48 FID-10K, versus 6.78 for full precision. With W1.58A4, FID-50K improves from 5.57 with direct QAT to 3.30.
W1.58A4 visual comparisons
ImageNet samples / panel 01. Rows: FP, direct QAT, RobuQ, JustQuant.ImageNet samples / panel 02. Original comparisons from the paper appendix.
W / A denotes weight / activation bits. All rows from the corresponding main-paper table are retained, including higher-performing baselines on individual metrics. MP = mixed precision; Norm = normalization.
TEXT-TO-IMAGE / 12B
FLUX.1
Naive W4A4 / group size 64 18 hours on 4 H200 GPUs
On schnell, plain W4A4 reaches 0.7149 GenEval, above the tested rotation and SVD baselines. On dev, ImageReward recovers to 0.956, close to the BF16 reference of 0.958.
FLUX / 18 high-resolution comparisons
Sunlit leaves01 / 181024 × 1024
Full precisionBF16ConvRotW4A4 + RotationSVDQuantW4A4 + Smooth + SVDJustQuantNaive W4A4
FLUX.1-schnell: 4 sampling steps. MJHQ metrics and GenEval scores.
Method
MJHQ FID ↓
CLIP-IQA ↑
ImageReward ↑
Single ↑
Count ↑
Color ↑
Overall ↑
Additional operators
FP
19.15
0.936
0.959
0.9844
0.5938
0.7872
0.6677
--
ConvRot
17.86
0.936
0.947
0.9812
0.6406
0.7819
0.6817
Rotation
SVDQuant
18.51
0.938
0.968
0.9906
0.6219
0.8138
0.6775
SVD
Naive INT4
18.37
0.929
0.902
0.9844
0.5813
0.7846
0.6753
None
QAT
17.98
0.925
0.912
0.9844
0.6625
0.7872
0.6825
None
JustQuant
17.62
0.943
0.993
0.9938
0.6844
0.8351
0.7149
None
FLUX.1-dev / first-page comparison
Full precisionBF16PTQNaive W4A4QATNaive W4A4JustQuantNaive W4A4
FLUX.1-dev: 50 sampling steps. MJHQ metrics and GenEval scores.
Method
MJHQ FID ↓
CLIP-IQA ↑
ImageReward ↑
Single ↑
Count ↑
Color ↑
Overall ↑
Additional operators
FP
20.26
0.952
0.958
0.9875
0.7281
0.7926
0.6707
--
ConvRot
19.78
0.953
0.948
0.9844
0.7094
0.7979
0.6662
Rotation
SVDQuant
19.71
0.948
0.961
0.9969
0.6938
0.8032
0.6687
SVD
Naive INT4
24.29
0.923
0.754
0.9844
0.7094
0.7952
0.6465
None
QAT
22.54
0.917
0.520
0.9500
0.5312
0.6888
0.5302
None
QAD
20.44
0.955
0.898
0.9906
0.6906
0.7819
0.6585
None
JustQuant
20.21
0.951
0.956
0.9906
0.7188
0.8032
0.6643
None
The FP / ConvRot / SVDQuant / JustQuant comparisons use original 1024 × 1024 PNGs from our image archive. The dev comparison is from the paper. W4A4 applies to transformer-block linear layers; text encoders, VAE, and boundary modules remain at higher precision. Plain operators still require suitable quantization, packing, and GEMM kernels.
DIFFUSION LANGUAGE GENERATION / 105M
ELF-B
OpenWebText Four evaluation seeds
The same idea extends beyond images. Generation perplexity stays close to full precision under both judge models, for both W4A4 and W1.58A4, without Hadamard or SVD operators.
ELF-B (105M): OpenWebText, W4A4. Generation perplexity over four seeds.
Method
Qwen3-8B-Base PPL ↓
Delta vs. FP
GPT-2-Large PPL ↓
Delta vs. FP
Additional operators
FP
14.729 ± 0.152
--
19.029 ± 0.305
--
--
JustQuant
14.854 ± 0.344
+0.125
19.180 ± 0.417
+0.152
None
RobuQ
16.751 ± 0.384
+2.022
22.089 ± 0.529
+3.061
Hadamard, SVD
Direct QAT
19.625 ± 0.327
+4.896
25.011 ± 0.465
+5.983
None
ELF-B (105M): OpenWebText, W1.58A4. Generation perplexity over four seeds.
Method
Qwen3-8B-Base PPL ↓
Delta vs. FP
GPT-2-Large PPL ↓
Delta vs. FP
Additional operators
FP
14.823 ± 0.208
--
19.254 ± 0.253
--
--
JustQuant
14.626 ± 0.482
-0.197
19.038 ± 0.634
-0.216
None
RobuQ
18.537 ± 0.265
+3.714
25.707 ± 0.414
+6.453
Hadamard, SVD
Direct QAT
16.776 ± 0.255
+1.953
22.051 ± 0.391
+2.797
None
Mean and variability over four seeds, reported as in the paper. The small W1.58A4 advantage over FP is within the reported variance.
DIFFUSION LANGUAGE REASONING / 8B
LLaDA
NVFP4 W4A4 Quantization-format transfer
Progressive distillation also transfers to NVFP4. GSM8K greedy accuracy reaches 0.6543, while held-out task scores remain broadly comparable to ordinary QAD.
LLaDA-8B: GSM8K, NVFP4 W4A4. All scores are higher-is-better.
Method
Greedy ↑
pass@1 ↑
pass@4 ↑
pass@16 ↑
FP
0.6748
0.6307
0.8539
0.9325
PTQ
0.6035
0.5704
0.8122
0.9113
QAT
0.6138
0.5836
0.8203
0.9163
QAD
0.6475
0.6015
0.8287
0.9196
JustQuant
0.6543
0.6137
0.8455
0.9249
LLaDA-8B: held-out transfer after GSM8K-only distillation, NVFP4 W4A4.
Method
PIQA ↑
ARC-C ↑
C-EVAL ↑
HumanEval ↑
FP
0.8300
0.8662
0.6440
0.4512
PTQ
0.8280
0.8495
0.6420
0.4268
QAT
0.8150
0.8200
0.6050
0.3512
QAD
0.8320
0.8562
0.6440
0.4512
JustQuant
0.8420
0.8562
0.6420
0.4451
Distilled only on GSM8K. This is an NVFP4 format-transfer experiment, not evidence of hardware-independent INT4 deployment. The paper does not report an additional-operator column for this experiment.
04 / THE SHIP OF THESEUS
Change the material. Preserve the intelligence.
If every plank of a ship is replaced, what makes it the same ship? For a model, we ask a practical version: can its behavior survive a new, low-bit body?
01Original materialThe original ship
02Planks being replacedDifferent planks, same structure
03New materialThe ship carries on
Original planks Replacement planksNew material. The same ship. A metaphor for preserving behavior in a low-bit body.
Theseus QAD is progressive distillation. We begin by matching short low-bit segments to a full-precision teacher, then gradually merge them into longer paths, ending with full-model supervision. The student is already low-bit; what grows is the scope of distillation.
One block. Longer paths. The whole model. Additional guidance during training, not additional operators at inference.
The idea is simple. The details matter.
The schedule, objectives, ablations, and limitations are in the paper. The companion article is now available, and the project repository is open.