Publications
- Near-Lossless MXFP4 Compression for Accelerated LLM Serving: Jointly Tuning Distribution Transforms, poster, PyTorch Conference Europe 2026
- dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats, G. Franco, I. Colbert, P. Monteagudo-Lago, F. Marty, N. Fraser, arXiv 2026
- Evaluation on the Hub: Better Best Practices for Data and Model Measurements, Hugging Face team, ACL Anthology, EMNLP 2022
Talks
- Quantization on AMD Instinct, 2025, Paris first inference & vLLM meet-up
- XXX TODO, 2025, AMD hackathon in Paris
ROCm technical blog posts
- Production-Ready MXFP4 Online Rotation with Fused Kernels on AMD Instinct™ MI355X, August 2026
- Serving NVFP4 Models on AMD Instinct™ MI355 Accelerators, July 2026
- Programming Tensor Descriptors in Composable Kernel (CK), March 2026
- Advanced MXFP4 Quantization: Combining Fine-Tuned Rotations with SmoothQuant for Near-Lossless Compression, February 2026
- High-Accuracy MXFP4, MXFP6, and Mixed-Precision Models on AMD GPUs, October 2025