== MCP server ======================================================== rig gpu-2B-cloudrun-devops-agent initialize ok in 3219 ms == Server info ======================================================= name Self-Hosted vLLM DevOps Agent protocol 2025-11-25 == Tools ============================================================= tools/list 27 tools in 2 ms ... == tools/call cloudrun_query_gemma4_with_stats ======================= arguments {"prompt":"In one sentence, what is a TPU?"} latency 1333 ms isError false result: ### 📊 Performance Stats - **Model:** `/mnt/models/gemma-4-E2B-it` - **Time to First Token (TTFT):** `0.086s` - **Total Generation Time:** `0.688s` - **Tokens per Second:** `53.11 tokens/s` - **Total Tokens (approx.):** `32` ### 💬 Model Response A TPU (Tensor Processing Unit) is a specialized type of integrated circuit designed to accelerate machine learning workloads, particularly those involving tensor operations common in deep learning. == Server log (stderr) =============================================== ... INFO - 📡 Automatically discovered vLLM at: https://.a.run.app 2026-09-11 11:41:36,754 - httpx - INFO - HTTP Request: GET https://.a.run.app/health "HTTP/1.1 200 OK" 2026-09-11 11:41:37,399 - httpx2 - INFO - HTTP Request: GET https://.a.run.app/v1/models "HTTP/1.1 200 OK" 2026-09-11 11:41:37,479 - httpx2 - INFO - HTTP Request: POST https://.a.run.app/v1/chat/completions "HTTP/1.1 200 OK"