non-default args: {'host': '127.0.0.1', 'model': 'google/gemma-4-E2B-it-qat-w4a16-ct', 'dtype': 'float16', 'max_model_len': 16384, 'gpu_memory_utilization': 0.9, 'max_num_seqs': 8} Casting torch.bfloat16 to torch.float16. Using MarlinLinearKernel for CompressedTensorsWNA16 Model loading took 8.02 GiB memory and 88.981559 seconds Available KV cache memory: 4.66 GiB GPU KV cache size: 519,681 tokens, Maximum concurrency for 16,384 tokens per request: 31.72x init engine (profile, create kv cache, warmup model) took 61.77 s (compilation: 2.59 s)