Written by Rahul Nair, usually unrahul online and rahulunair on GitHub.
Speculative decoding at 512K on Intel Arc Pro: why MTP wins early and DFlash2 wins late
intel
arc
arc-pro
xe2
xpu
sglang
llm-inference
speculative-decoding
agentic-coding
Qwen3.8-27B on Intel Arc Pro: 2.82 to 59.5 tok/s
intel
arc
arc-pro
xe2
xpu
sglang
llm-inference
quantization
agentic-coding
No matching items