UE5M3 FP4 Block Scaling for Stable Language Model Pretraining
Stable 4-bit floating-point (FP4) pretraining is difficult because the E2M1 payload represents only a narrow range of magnitudes. NVIDIA's Transformer Engine \nv{} recipe addresses this with current-tensor scaling, a randomized Hadamard transform (RHT), and bfloat16 (BF16) final layers, adding work outside the FP4 matrix multiplications. We instead pair E2M1 payloads with unsigned E5M3 (\ue{}) block scales. Their wider range permits periodic tensor scaling, while our recipe applies selective stochastic rounding to backward gradients, omits RHT, and uses FP4 in all eligible internal linears.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%stabilityai/stable-diffusion-3.5-large →
“Fuzzy title match (0.73): “UE5M3 FP4 Block Scaling for Stable Language Model Pretrainin” ≈ “stabilityai/stable-diffusion-3.5-large””
- FuzzySimilar title/name (fuzzy) · 59%stabilityai/stable-diffusion-xl-base-1.0 →
“Fuzzy title match (0.73): “UE5M3 FP4 Block Scaling for Stable Language Model Pretrainin” ≈ “stabilityai/stable-diffusion-xl-base-1.0””
- FuzzySimilar title/name (fuzzy) · 59%CompVis/stable-diffusion-v1-4 →
“Fuzzy title match (0.73): “UE5M3 FP4 Block Scaling for Stable Language Model Pretrainin” ≈ “CompVis/stable-diffusion-v1-4””
- PossiblePossibly related (embedding) · 50%A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation →
- FuzzySimilar title/name (fuzzy) · 59%DLR-RM/stable-baselines3 →
“Fuzzy title match (0.73): “UE5M3 FP4 Block Scaling for Stable Language Model Pretrainin” ≈ “DLR-RM/stable-baselines3””
- LinkedLinked via arxiv author · 85%Robert Hu →
“UE5M3 FP4 Block Scaling for Stable Language Model Pretraining”
- LinkedLinked via arxiv author · 85%Carlo Luschi →
“UE5M3 FP4 Block Scaling for Stable Language Model Pretraining”
- LinkedLinked via arxiv author · 85%Paul Balanca →
“UE5M3 FP4 Block Scaling for Stable Language Model Pretraining”
