These two modes of generating answers open a huge opportunity for RLMs: 1. They can programmatically explore using regexes, find operations using regular Python 2. They can create small variables to save work (they are inside a REPL, so old work is never lost) 3. They can recursively call agents to summarize. 4. Subagents can be parallel or sequential. The LLM intelligently decides this. A reason the RLM may want to call subagents sequentially is if it needs to do a running summary of a long context text that needs prior information. 5. They can also use external tools, but you have to expose them through your sandbox layer (Deno, for example) To understand how RLMs work in more visual detail, how they can be implemented from scratch, and see some real trajectories where it attacks real world problems, check out this video tutorial: Please accept cookies to access this content Check out my open-source implementation of RLMs; it comes with a TUI log viewer for recursive traces. Here is the full system prompt that I used for my RLM implementation. This will reveal a lot! **Click here to reveal the full System Prompt (it's hidden because it's long)**. You can find the author-recommended prompt in the RLM paper (linked below). The prompt here was repurposed from the paper's prompt, with a few additional few-shot examples and instructions that reduced failure states on open-source models (tested on Minimax-M2.7, GLM-5.1) markdown