I am a Senior Software Engineer at AMD, specialized in model optimization and deep learning inference, working on the AMD Quark open-source model compression toolkit.
I also contribute to vLLM and SGLang with a focus on quantization and serving on AMD Instinct GPUs, as well as lm-evaluation-harness for evaluation.
Previously working at Hugging Face, I used to contribute to Optimum, Transformers and text-generation-inference libraries.
What’s new?
- October 2026: I was named a community reviewer for the vLLM project!