In practical RLM implementations, we will be able to parallelize these calls. Multiple subagents working in parallel on orthogonal tasks is not just super cool, but it actually gets a ton of stuff done really fast. **Notice what just happened.** * The LLM assigned 3 subagents the task of managing fruits, countries, and animals * The subagents (as we saw previously) will return the answers calling `FINAL` in their own local REPL * That outputs lands directly inside the `FRUIT_DICT`, `ANIMAL_DICT` and `COUNTRY_DICT` dictionaries of the main agent's REPL * **The subagent outputs are entered into the REPL, they are not loaded directly into the context of the LLM (like how CodeAct or ReAct subagents worked).** To view the subagent outputs, the main agent needs to inspect it deliberately with `print` statements. The main agent did not even need to: * Load the entire subagent output into context * Read any of the fruit names * Generate the final output token by token from memory * It **composed** an answer by forming the key symbols through recursive calls and delivering the final output as a composition. ![](https://assets.insightmediagroup.io/media/wp-content/uploads/2026/05/image-188-1024x576.png) The Basic RLM architecture with Deno and Pyodide ### 3.4 The RLM's Output Space * RLMs can choose two ways to return their FINAL output. * One, it can compose answers into Python variables and return them (like the example above) * Or it can generate a response on its own autoregressively, just like a normal LLM In the case below, the output was autoregressively generated. markdown