Model loading took 9.75 GiB memory and 82.751806 seconds GPU KV cache size: 723,484 tokens, Maximum concurrency for 8,192 tokens per request: 88.32x