Written by Rahul Nair, usually unrahul online and rahulunair on GitHub.
Planning capacity for an LLM server on Intel Arc Pro
intel
arc
arc-pro
xe2
xpu
sglang
llm-inference
speculative-decoding
agentic-coding
Speculative decoding at long context on Intel Arc Pro
intel
arc
arc-pro
xe2
xpu
sglang
llm-inference
speculative-decoding
agentic-coding
Routing Qwen3.8-27B onto the right Intel XPU paths
intel
arc
arc-pro
xe2
xpu
sglang
llm-inference
quantization
agentic-coding
No matching items