== models--google--gemma-4-E2B-it bytes read per decode token: 4.597 GB bound @ 277.0 GB/s measured stream: 60.3 tok/s == models--google--gemma-4-E2B-it-qat-w4a16-ct bytes read per decode token: 1.862 GB bound @ 277.0 GB/s measured stream: 148.8 tok/s