📡 **Model** — `local-llamacpp-1650ti-2b-q4_0` - **Name:** `google/gemma-4-E2B-it-qat-q4_0-gguf` - **Path:** `/home/xbill/models/gemma-4-E2B-it-qat-q4_0/gemma-4-E2B_q4_0-it.gguf` - **On disk:** 3.35 GB - **Quantization slot:** `q4_0` — but the dominant tensor type is **Q6_K**. Both embedding tensors are Q6_K (2.257 GB of 3.334 GB); only the ~1.08 GB transformer body is actually Q4_0. - **Resident on GPU:** ~1.31 GiB. `per_layer_token_embd` (1.93 GB, 58% of the file) is `TENSOR_READ_LAZY` and is served by GET_ROWS out of the mmap. Run `inspect_gguf.py` to re-derive the split from the artifact rather than trusting these numbers.