[](https://hamel.dev/notes/llm/inference/inference.html#cb12-1)bash start_server.sh --model-id TheBloke/Llama-2-7B-GPTQ --quantize gptq