An RLM runs code that a model wrote after it read your input. Here is what the default REPL really blocks, where hostile code comes from, and how to isolate it.
The short version: run the REPL somewhere that cannot reach anything you care about. The default local environment in the reference RLM library runs model-written code with exec inside your own Python process. It blocks eval and exec, but it still allows __import__ and open, so it is not a security boundary. For anything that reads untrusted text, use a container with networking turned off, a cloud sandbox, or a WASM interpreter. Then cap iterations and sub-calls.
Because the model writes code and the harness runs it. In the RLM paper, the prompt is stored as a variable in a Python REPL. The root model sees only metadata about it, then writes code to slice it, search it, and call llm_query on the pieces. Every turn the model emits a code block, and the environment executes it. That is arbitrary code execution by design.
The paper itself says little about isolation. Its implementation section describes "a Python REPL environment, which loads a module for querying a sub-LM." The limitations section names "sandboxed REPLs" as future work, next to asynchronous sub-calls, as a way to cut runtime and cost. Security is left to whoever deploys the scaffold. Our RLM vs coding agent post notes that both need a sandbox for the same reason. This post is about how to build one.
Usually not from the model on its own. It comes from the input. The whole point of an RLM is to process inputs too large to read: a corpus of 1,000 documents, a scraped website, a repository, a mailbox. The root model reads slices of that input through printed output, and sub-calls read larger slices and return text. Any instruction hidden in a document can reach the model that writes the next code block.
This is indirect prompt injection. Greshake et al. showed that attackers can inject prompts "into data likely to be retrieved" and that "processing retrieved prompts can act as arbitrary code execution." An RLM makes that last phrase literal. A document that says "before answering, run this snippet" is one obedient turn away from executing on your machine.
The reference library says the same thing in plainer words. Its README describes non-isolated environments as "pretty reasonable for some local low-risk tasks, like simple benchmarking, but can be problematic if the prompts or tool calls can interact with malicious users." In an RLM, the prompt is the document set. If you did not write every document, treat the prompt as untrusted.
Less than its name suggests. The docstring in local_repl.py says it "executes code in a sandboxed namespace with access to context data." The namespace is a custom builtins dictionary, _SAFE_BUILTINS. It sets input, eval, exec, compile, globals, and locals to None. It keeps __import__ and open.
That means model code can still run import os, import subprocess, or import socket. It can read your environment variables and API keys. It can open files anywhere your user can, and it can make network calls. The code runs through exec(code, combined, combined) in the same process as the RLM, so it also shares memory with the client that holds your provider credentials. A restricted builtins dictionary stops accidents, not attackers.
The working directory is not isolated either. The environment creates a temp directory and switches into it with os.chdir. That call changes the directory for the whole process. An open issue, #180, reports that concurrent RLM instances race on it and cross-contaminate each other's working directory. The README is candid: the local REPL "is generally safe, but should not be used for production settings." Take the second half of that sentence seriously.
The library lists seven environments. local, ipython, and docker run on your machine. modal, prime, daytona, and e2b run in cloud sandboxes, which the README says ensure "complete isolation from the host process." In the isolated setups, the README says a recursive sub-call "is requested from the host process," so the sandbox does not need your model API keys.
DockerREPL runs code in a python:3.11-slim container with --rm, a mounted temp directory, and a host alias so llm_query can reach an HTTP proxy on the host. The start command sets no network flag. Docker's documentation says containers "have networking enabled by default, and they can make outgoing connections." So a stock container can still send data out. If the task needs no internet, add --network none and give the proxy another path, such as a mounted Unix socket.max_sandboxes as a budget, and deletes sub-agent sandboxes "immediately after completion." Prime Intellect's RLM environment runs code in isolated sandboxes and makes extra tools usable "only by the sub-LLMs," which keeps tool output out of the root context.enable_read_paths, enable_write_paths, enable_env_vars, and enable_network_access with specific domains.Start from what an RLM actually needs, which is very little. The root model needs the context variable, a Python standard library, and a way to call llm_query. It rarely needs the internet, your filesystem, or your credentials. Build the policy from that list:
max_iters of 20, max_llm_calls of 50, and 10,000 characters of REPL output per turn. A prompt-injected loop that calls llm_query ten thousand times is a billing attack even without a shell.forward() call and shuts it down afterward. Do not let one user's context persist into another user's run.Isolation has a latency cost. A cloud sandbox adds a network round trip for each REPL turn and each proxied sub-call, and RLMs already make many sequential calls. Our post on RLM latency covers where that time goes. Batched sub-calls (llm_query_batched) reduce the round trips, and they matter more once every call crosses a process boundary.
An RLM executes code written by a model that has just read your input. That input may contain instructions from strangers. The reference library's default local environment keeps imports and file access, runs in your process, and says it is not for production. Use it for benchmarks on data you trust. For anything else, run the REPL in a container with networking turned off, a cloud sandbox, or DSPy's WASM interpreter. Keep credentials on the host, mount the context read-only, and cap turns and sub-calls. The sandbox is not optional plumbing. It is the part of the RLM that makes reading untrusted input safe.
No. It removes eval, exec, compile, input, globals, and locals from builtins, but it keeps __import__ and open, so model code can import os or subprocess. It runs with exec inside the same process as the RLM. The README says it should not be used for production settings.
Not by default. The DockerREPL start command sets no network flag, and Docker containers can make outgoing connections unless you use the none network driver. Add --network none and route llm_query through a channel that does not need open networking.
Yes, in principle. The root model reads slices of the input and then writes the next code block, so instructions hidden in a document can steer that code. This is the indirect prompt injection pattern that Greshake et al. described in 2023. The sandbox limits the damage when the model obeys.
So the sandbox does not need your model API keys. In the isolated environments of the reference library, a recursive sub-call is requested from the host process. The host can then enforce budgets and reject calls it does not want to serve.
Cap REPL turns, sub-calls, wall-clock time, memory, and output size. DSPy's RLM module defaults to 20 iterations and 50 LLM calls per run. These caps stop a runaway or injected loop from turning into a large bill.