manage context control temperature choose decoding strategy stop generation correctly handle repetition optimize inference cost