add position run multi-head attention preserve input through residuals normalize representations apply feed-forward updates repeat safely across many layers