observe environment update internal state evaluate possible actions choose the best action execute action repeat