### sagemaker-gemma Serving `google/gemma-4-E2B-it` with vLLM on a **SageMaker real-time endpoint**, managed entirely through the aws CLI. | Setting | Value | | --- | --- | | Region | `us-east-2` | | Endpoint | `gemma-4-e2b` | | Instance types (priority order) | `ml.g6.xlarge`, `ml.g6.2xlarge`, `ml.g6.4xlarge` | | Max model length | `8192` | | Image | `newest SageMaker vLLM image (find_vllm_image)` | | Execution role | `sagemaker-gemma-execution-role` | **Order of work:** check_quotas → deploy_endpoint → get_endpoint_status until `InService` (about 10 minutes once an instance is placed) → verify_model_health → query_model → delete_endpoint.