Samama Usman
I'm a backend-focused full-stack engineer who spent the last couple of years shipping products, with a couple of hackathon wins along the way. Now I'm all in on ML systems and LLM inference, making models run faster and cheaper at the serving layer: request scheduling and continuous batching, prefill/decode disaggregation, distributed KV cache and prefix reuse across replicas, all to keep a GPU fleet saturated.
Latest writing
All posts →- Moving to a faster card made decoding slower
A GPU with 2.11x the memory bandwidth decoded Qwen2.5-1.5B at 0.44x the throughput. Measuring why, on one card with the hardware held fixed, found the GPU idle 41% of every token.