Skip to project
Kaicheng Yang

MODEL QUANTIZATION / RESEARCH PROJECT

JustQuant

Less machinery. More intelligence.

01 / THE DEPLOYMENT GAP

4-bit on paper.
Complicated in practice.

Every year, more quantization papers promise nearly lossless 4-bit models. But a small bit width does not always mean a simple deployment: rotations, low-rank branches, and hardware-specific formats can turn a compact model into a complicated inference pipeline.

Paper survey of eight conferences from ICLR 2024 to ICML 2026. Stacked bars count activation-quantization papers at 4 bits or below: green means no extra operator and red means extra operators. The shares using extra operators are 50%, 50%, 75%, 50%, 100%, 100%, 67%, and 80%, respectively.
From the paper: a survey of ≤4-bit activation quantization across eight conferences (2024–2026). Red: extra operators. Green: no extra operator.Full figure

WATCH / JUSTQUANT

The idea in motion.

Open video page

02 / A SIMPLE QUESTION

Can intelligence live in a body
made of plain 4-bit operations?

Naive low-bit matrix multiplication. No rotation, no SVD branch, no smoothing. Move the complexity into training, and keep the deployed operator simple.

W4A4 / W1.58A4 / PLAIN LOW-BIT GEMM

03 / THE EVIDENCE

Start with the results.

None = no additional inference operator / Named operators = additional transformations. Highlighted rows are ours.

IMAGE GENERATION / 0.6B

DiT-XL/2

ImageNet 256
W4A4 and W1.58A4

Plain W4A4 reaches 6.48 FID-10K, versus 6.78 for full precision. With W1.58A4, FID-50K improves from 5.57 with direct QAT to 3.30.

DiT-XL/2: ImageNet 256, 10K samples, 50 sampling steps, CFG 1.5.
Method W / A Training / steps FID ↓ sFID ↓ IS ↑ Additional operators
FP 32 / 32 - / 7000k 6.78 20.56 243.70 --
SVDQuant 4 / 4 PTQ / 0k 81.28 67.27 29.17 Smooth, SVD
ConvRot 4 / 4 PTQ / 0k 18.69 32.06 139.54 Rotation
QuaRot 4 / 4 PTQ / 0k 53.31 56.74 53.12 Rotation
VETA-DiT 4 / 4 PTQ / 14.4k 9.87 25.52 202.55 Smooth, Rotation
TreeQ 4 / 4 PTQ / 36.9k 6.92 20.86 219.66 MP, SVD, Rotation
RobuQ 4 / 4 QAT / 20k 8.20 22.94 222.02 Norm, SVD, Rotation
QAD 4 / 4 QAD / 20k 188.26 190.47 3.89 None
QAT 4 / 4 QAT / 20k 10.72 26.55 196.53 None
JustQuant 4 / 4 QAT / 20k 6.48 19.31 238.31 None
TerDiT 1.58 / 32 QAT / 1750k 8.23 20.57 172.23 Norm
JustQuant 1.58 / 32 QAT / 110k 6.56 19.84 249.65 None
RobuQ 1.58 / 4 QAT / 100k 7.76 20.81 195.98 Norm, SVD, Rotation
QAT 1.58 / 4 QAT / 100k 9.93 20.33 155.85 None
JustQuant 1.58 / 4 QAT / 110k 7.23 19.88 223.99 None
DiT-XL/2: ImageNet 256, 50K samples, 250 sampling steps, CFG 1.5.
Method W / A Training / steps FID ↓ sFID ↓ IS ↑ Additional operators
FP 32 / 32 - / 7000k 2.27 4.55 277.83 --
SVDQuant 4 / 4 PTQ / 0k 60.01 37.85 39.26 Smooth, SVD
ConvRot 4 / 4 PTQ / 0k 7.98 13.36 187.96 Rotation
VETA-DiT 4 / 4 PTQ / 14.4k 2.99 6.27 247.05 Smooth, Rotation
RobuQ 4 / 4 QAT / 20k 2.71 5.20 265.02 Norm, SVD, Rotation
QAT 4 / 4 QAT / 20k 5.15 8.34 220.53 None
JustQuant 4 / 4 QAT / 20k 3.32 5.79 256.56 None
TerDiT 1.58 / 32 QAT / 1750k 4.34 4.99 183.49 Norm
JustQuant 1.58 / 32 QAT / 110k 3.05 4.59 266.83 None
RobuQ 1.58 / 4 QAT / 100k 3.21 4.86 217.04 Norm, SVD, Rotation
QAT 1.58 / 4 QAT / 100k 5.57 5.57 166.06 None
JustQuant 1.58 / 4 QAT / 110k 3.30 4.51 241.95 None

W / A denotes weight / activation bits. All rows from the corresponding main-paper table are retained, including higher-performing baselines on individual metrics. MP = mixed precision; Norm = normalization.

TEXT-TO-IMAGE / 12B

FLUX.1

Naive W4A4 / group size 64
18 hours on 4 H200 GPUs

On schnell, plain W4A4 reaches 0.7149 GenEval, above the tested rotation and SVD baselines. On dev, ImageReward recovers to 0.956, close to the BF16 reference of 0.958.

FLUX.1-schnell: 4 sampling steps. MJHQ metrics and GenEval scores.
Method MJHQ FID ↓ CLIP-IQA ↑ ImageReward ↑ Single ↑ Count ↑ Color ↑ Overall ↑ Additional operators
FP 19.15 0.936 0.959 0.9844 0.5938 0.7872 0.6677 --
ConvRot 17.86 0.936 0.947 0.9812 0.6406 0.7819 0.6817 Rotation
SVDQuant 18.51 0.938 0.968 0.9906 0.6219 0.8138 0.6775 SVD
Naive INT4 18.37 0.929 0.902 0.9844 0.5813 0.7846 0.6753 None
QAT 17.98 0.925 0.912 0.9844 0.6625 0.7872 0.6825 None
JustQuant 17.62 0.943 0.993 0.9938 0.6844 0.8351 0.7149 None
FLUX.1-dev: 50 sampling steps. MJHQ metrics and GenEval scores.
Method MJHQ FID ↓ CLIP-IQA ↑ ImageReward ↑ Single ↑ Count ↑ Color ↑ Overall ↑ Additional operators
FP 20.26 0.952 0.958 0.9875 0.7281 0.7926 0.6707 --
ConvRot 19.78 0.953 0.948 0.9844 0.7094 0.7979 0.6662 Rotation
SVDQuant 19.71 0.948 0.961 0.9969 0.6938 0.8032 0.6687 SVD
Naive INT4 24.29 0.923 0.754 0.9844 0.7094 0.7952 0.6465 None
QAT 22.54 0.917 0.520 0.9500 0.5312 0.6888 0.5302 None
QAD 20.44 0.955 0.898 0.9906 0.6906 0.7819 0.6585 None
JustQuant 20.21 0.951 0.956 0.9906 0.7188 0.8032 0.6643 None

The FP / ConvRot / SVDQuant / JustQuant comparisons use original 1024 × 1024 PNGs from our image archive. The dev comparison is from the paper. W4A4 applies to transformer-block linear layers; text encoders, VAE, and boundary modules remain at higher precision. Plain operators still require suitable quantization, packing, and GEMM kernels.

DIFFUSION LANGUAGE GENERATION / 105M

ELF-B

OpenWebText
Four evaluation seeds

The same idea extends beyond images. Generation perplexity stays close to full precision under both judge models, for both W4A4 and W1.58A4, without Hadamard or SVD operators.

ELF-B (105M): OpenWebText, W4A4. Generation perplexity over four seeds.
Method Qwen3-8B-Base PPL ↓ Delta vs. FP GPT-2-Large PPL ↓ Delta vs. FP Additional operators
FP 14.729 ± 0.152 -- 19.029 ± 0.305 -- --
JustQuant 14.854 ± 0.344 +0.125 19.180 ± 0.417 +0.152 None
RobuQ 16.751 ± 0.384 +2.022 22.089 ± 0.529 +3.061 Hadamard, SVD
Direct QAT 19.625 ± 0.327 +4.896 25.011 ± 0.465 +5.983 None
ELF-B (105M): OpenWebText, W1.58A4. Generation perplexity over four seeds.
Method Qwen3-8B-Base PPL ↓ Delta vs. FP GPT-2-Large PPL ↓ Delta vs. FP Additional operators
FP 14.823 ± 0.208 -- 19.254 ± 0.253 -- --
JustQuant 14.626 ± 0.482 -0.197 19.038 ± 0.634 -0.216 None
RobuQ 18.537 ± 0.265 +3.714 25.707 ± 0.414 +6.453 Hadamard, SVD
Direct QAT 16.776 ± 0.255 +1.953 22.051 ± 0.391 +2.797 None

Mean and variability over four seeds, reported as in the paper. The small W1.58A4 advantage over FP is within the reported variance.

DIFFUSION LANGUAGE REASONING / 8B

LLaDA

NVFP4 W4A4
Quantization-format transfer

Progressive distillation also transfers to NVFP4. GSM8K greedy accuracy reaches 0.6543, while held-out task scores remain broadly comparable to ordinary QAD.

LLaDA-8B: GSM8K, NVFP4 W4A4. All scores are higher-is-better.
Method Greedy ↑ pass@1 ↑ pass@4 ↑ pass@16 ↑
FP 0.6748 0.6307 0.8539 0.9325
PTQ 0.6035 0.5704 0.8122 0.9113
QAT 0.6138 0.5836 0.8203 0.9163
QAD 0.6475 0.6015 0.8287 0.9196
JustQuant 0.6543 0.6137 0.8455 0.9249
LLaDA-8B: held-out transfer after GSM8K-only distillation, NVFP4 W4A4.
Method PIQA ↑ ARC-C ↑ C-EVAL ↑ HumanEval ↑
FP 0.8300 0.8662 0.6440 0.4512
PTQ 0.8280 0.8495 0.6420 0.4268
QAT 0.8150 0.8200 0.6050 0.3512
QAD 0.8320 0.8562 0.6440 0.4512
JustQuant 0.8420 0.8562 0.6420 0.4451

Distilled only on GSM8K. This is an NVFP4 format-transfer experiment, not evidence of hardware-independent INT4 deployment. The paper does not report an additional-operator column for this experiment.

04 / THE SHIP OF THESEUS

Change the material.
Preserve the intelligence.

If every plank of a ship is replaced, what makes it the same ship? For a model, we ask a practical version: can its behavior survive a new, low-bit body?

Original material: 0 of six schematic hull sections replaced, with the same outline and sails.
01Original materialThe original ship
Planks being replaced: 3 of six schematic hull sections replaced, with the same outline and sails.
02Planks being replacedDifferent planks, same structure
New material: 6 of six schematic hull sections replaced, with the same outline and sails.
03New materialThe ship carries on
Original planks Replacement planksNew material. The same ship. A metaphor for preserving behavior in a low-bit body.

Theseus QAD is progressive distillation. We begin by matching short low-bit segments to a full-precision teacher, then gradually merge them into longer paths, ending with full-model supervision. The student is already low-bit; what grows is the scope of distillation.

Theseus QAD: teacher-guided quantized segments grow from one block to two, four, and finally the whole model. Local dense supervision becomes global supervision.
One block. Longer paths. The whole model. Additional guidance during training, not additional operators at inference.

The idea is simple.
The details matter.

The schedule, objectives, ablations, and limitations are in the paper. The companion article is now available, and the project repository is open.