tokenize prompt run decoder get logits from LM Head apply temperature filter with top-k or top-p sample or choose token append token repeat