Both run the same REPL loop from the Recursive Language Models paper. They differ in how deep sub-calls go, where code runs, and what you can tune.
The short version: dspy.RLM is a DSPy module that runs the RLM loop inside a typed signature, with one level of sub-calls through llm_query, a WASM sandbox by default, and prompts an optimizer can tune. The reference library, rlms from the paper authors, is a standalone inference engine: it adds true recursive child RLMs through rlm_query, many sandbox backends, cost and time budgets, and a training harness. Pick DSPy when the RLM is one step in a larger DSPy program. Pick the reference library when you want the full paradigm as the paper describes it.
The idea comes from Recursive Language Models by Alex L. Zhang, Tim Kraska and Omar Khattab. The paper treats a long prompt "as part of an external environment" and lets the model "programmatically examine, decompose, and recursively call itself over snippets of the prompt." The authors keep the reference code at alexzhang13/rlm, installed with pip install rlms, and describe it as "an extensible inference engine and training environment."
DSPy ships its own version in dspy/predict/rlm.py, next to ReAct and ProgramOfThought. The reference README lists "DSPy.RLM" first in its section on RLMs used in the wild, so the two projects know about each other. They share the core loop: context lives in a REPL variable, the model writes Python, and code can call a sub-model. The differences are in what wraps that loop.
The reference library replaces a chat call. You build RLM(backend="openai", backend_kwargs={"model_name": ...}) and call rlm.completion(prompt). The README frames it as a drop-in for llm.completion(prompt, model). Input is a prompt, and output is text.
DSPy starts from a signature instead. The DSPy RLM guide says dspy.RLM "takes the same signature you'd hand Predict or ChainOfThought." Input fields become REPL variables. Output fields define what the model must return, with types. The model ends a run by calling SUBMIT(...), and DSPy parses each value to its declared type. On a type error it feeds the message back for another try.
The reference library ends a run in a different way. In its current system prompt, the REPL holds an answer dict, and the run stops when the model sets answer["ready"] = True. The result is free text in answer["content"], with no schema check.
Only one level, and this is the largest difference. DSPy injects two functions into the sandbox: llm_query(prompt) and llm_query_batched(prompts). Each is a plain model call that returns text. A sub-call cannot open its own REPL. The dspy.RLM source has no rlm_query and no depth setting.
The reference library has both kinds of call. Its prompt describes llm_query as "a single LLM completion call (no REPL, no iteration)" and rlm_query as a function that "spawns a recursive RLM sub-call," where "the child gets its own REPL environment." rlm_query_batched runs several children in parallel, capped by max_concurrent_subcalls (default 4). Depth is set with max_depth. The docstring says that "when depth >= max_depth," a call "falls back to plain LM completion." The default max_depth is 1, so out of the box the reference library also behaves like a one-level system. You raise the limit to get deeper trees. Our post on how deep an RLM should recurse covers when that is worth it.
DSPy has three limits, all set in the constructor: max_iters (default 20), max_llm_calls (default 50) and max_output_chars (default 10_000). Each prompt in a batch counts as one call against max_llm_calls. When the cap is hit, the error goes back to the model so it can finish in plain Python. If the loop runs out of iterations without a SUBMIT, a separate extract predictor reads the history and fills the outputs.
The reference library has more stop conditions: max_iterations (default 30), max_budget in USD, max_timeout in seconds, max_tokens and max_errors. The docstring notes that max_budget "requires cost-tracking backend (e.g., OpenRouter)." It has no direct cap on the number of sub-calls.
For a cheaper sub-model, DSPy uses one argument, sub_lm, which falls back to dspy.settings.lm. The reference library uses other_backends and other_backend_kwargs, and its sub-call functions take an optional model argument. It also has separate sampling_args and sub_sampling_args, which matters if you want different temperatures for root and sub-calls.
DSPy defaults to PythonInterpreter, which the guide says runs code "in a Deno and Pyodide WASM sandbox with no filesystem or network access." You can pass interpreter_factory= to swap it. LocalInterpreter runs normal CPython, and the guide warns it is "for trusted code." Each call builds a fresh interpreter and shuts it down at the end. Large inputs such as DataFrames can load once through SandboxSerializable, so the model works with the real object.
The reference library defaults the other way. Its local environment runs code "in the same process as the RLM itself," and the README says it "should not be used for production settings." For isolation it supports ipython, docker, modal, prime, daytona and e2b. It also offers persistent=True for multi-turn sessions and compaction=True to summarize long root history. DSPy has neither option. So DSPy is safer by default, and the reference library has more ways to run remote. Our guide to sandboxing the RLM REPL compares the risks.
This is where each project plays to its origin. In DSPy, the action and extract steps "are ordinary dspy.Predict instances," so optimizers can tune the prompts. The guide names GEPA and MIPROv2 and says tuning "improves the loop's behavior, not just the task instructions." You get prompt optimization with no weight updates.
The reference library goes after the weights. Its training/ folder exposes rlm.RLM as a verifiers environment that plugs into Prime Intellect's prime-rl, with an OOLONG example. The paper reports that the authors "post-train the first model around the RLM," called RLM-Qwen3-8B. DSPy has no training harness for the loop.
DSPy also marks the class with @experimental. The guide says to "pin a version if you depend on it." The reference repo changes fast too, so pin both.
Use dspy.RLM when you already build in DSPy, want typed outputs, want a safe sandbox by default, and plan to tune prompts with an optimizer. Accept that sub-calls are single model calls, one level deep. Use the reference rlms library when you need child RLMs through rlm_query, cloud sandboxes, dollar or time budgets, or RL training. Treat it as the version closest to the paper. Both are young and both change often, so read the source for the release you install.
The paper is by Alex L. Zhang, Tim Kraska and Omar Khattab. The reference code lives in the alexzhang13/rlm repository, while dspy.RLM lives in the stanfordnlp/dspy repository. The reference README lists DSPy.RLM among projects that use RLMs.
The README says to run pip install rlms, and it requires Python 3.11 or later. You then import it with from rlm import RLM. DSPy users get dspy.RLM with the normal dspy install.
Yes. Pass a dspy.LM as sub_lm and llm_query and llm_query_batched will use it. If you leave it unset, sub-calls use dspy.settings.lm, the same model that drives the loop.
It does not return nothing. A second predictor called extract reads the variable metadata and the full REPL history and fills the output fields directly. You get the best answer the trajectory supports.
No. The default local environment runs code in the same process as the RLM through Python exec, and the README says not to use it in production. Pick docker, modal, prime, daytona or e2b for isolation.