A GPU with 2.11x the memory bandwidth decoded Qwen2.5-1.5B at 0.44x the throughput. Measuring why, on one card with the hardware held fixed, found the GPU idle 41% of every token.