Back to Blog
October 7, 2026

The short version: in the reference RLM library you do not write the asyncio code yourself. The REPL already has llm_query_batched(prompts, model=None), and the library runs that batch with asyncio.gather behind an asyncio.Semaphore of 16. Recursive children go through rlm_query_batched, which uses a thread pool of 4 by default. You write your own asyncio fan-out only when you build your own harness, and then the pattern is a semaphore, asyncio.to_thread and gather.

The paper and the code disagree on this point, so it helps to read both. The sections below separate what the library documents from what you build yourself.

Does the RLM paper run sub-calls in parallel?

No. The RLM paper by Zhang, Kraska and Khattab says that in its experiments "all LM calls are blocking / sequential." Its limitations appendix has a heading that reads "RLMs without asynchronous LM calls are slow." The authors write that they "implemented all sub-LM queries naively as blocking / sequential calls" and that they are "confident that this can be resolved with a robust implementation."

The paper also names the fix as future work. It says "asynchronous sub-calls and sandboxed REPLs can potentially significantly reduce the runtime and inference cost of RLMs, but further contribute to this complexity." The paper does not report a measured speedup from parallel sub-calls, so this article does not give one.

The cost of sequential calls is visible in the reproduction study. It measured DeepSeek v3.2 at 3.6 seconds for a base S-NIAH query, 89.3 seconds at depth 1 and 344.5 seconds at depth 2. Our post on whether RLMs are slower than a single long-context call covers those numbers.

What does the reference library do today?

The current code is ahead of the paper. The README of the authors' repository lists "single LM calls (llm_query / llm_query_batched)" and "parallel batched sub-calls bounded by max_concurrent_subcalls." So the helper exists, and its name is llm_query_batched.

The signature in the local REPL source is _llm_query_batched(self, prompts: list[str], model: str | None = None) -> list[str]. The REPL exposes it to the model as llm_query_batched. The docstring says it returns a "list of responses in the same order as input prompts."

The system prompt teaches the root model to use it. The example in the prompt builds a list and then runs answers = llm_query_batched(prompts). The prompt describes the function as "much faster than sequential llm_query calls for independent queries."

Where is the asyncio code inside the library?

It is in the LM handler, in a method named _handle_batched. The REPL sends the whole prompt list to the handler in one request. The handler then does three things, all quoted from the source:

  • It builds a limit: sem = asyncio.Semaphore(handler.batch_max_concurrent). The default for batch_max_concurrent is 16.
  • It wraps each prompt: async with sem: return await client.acompletion(prompt).
  • It collects the results: return await asyncio.gather(*tasks, return_exceptions=True), started with asyncio.run(run_all()).

The comment in the code explains the last flag: "one failed call doesn't abort the whole batch; failures are surfaced per-prompt as error completions." That matches the Python documentation for gather. With return_exceptions=True, "exceptions are treated the same as successful results, and aggregated in the result list." The same page says the order of results "corresponds to the order of awaitables," which is why chunk 3 always lands in slot 3.

One detail limits your control. In rlm/core/rlm.py, the handler is built as LMHandler(client, other_backend_client=other_backend_client). That call does not pass batch_max_concurrent, so 16 applies unless you construct the handler yourself. The constructor argument that RLM does expose is max_concurrent_subcalls, and it controls a different path.

Do recursive child RLMs use asyncio too?

No, they use threads. _rlm_query_batched in the same REPL file runs each child RLM in a ThreadPoolExecutor. The pool size is min(self.max_concurrent_subcalls, len(prompts)), and max_concurrent_subcalls defaults to 4. The docstring gives the reason: the sub-calls "are independent and I/O-bound." If recursion is not configured, the function falls back to llm_query_batched.

DSPy takes the thread route for plain sub-calls as well. In dspy/predict/rlm.py, llm_query_batched(prompts) submits each prompt to a ThreadPoolExecutor with max_workers=8. A failed model call returns a string that starts with [ERROR] in that slot. See our comparison of dspy.RLM and the reference library for the other differences.

So there are two limits in the reference library, and they stack. A depth-2 run can start 4 children at a time, and each child can send a batch of up to 16 concurrent plain calls. Check that product against your provider rate limit before you raise either number.

How do you write the pattern yourself?

You need this only when you write your own harness, or when your sub-call function is synchronous and has no batched form. The following is an illustrative sketch. It uses only standard asyncio APIs, and call_model stands for your own blocking function. It is not code from the RLM library.

  1. Set a limit: sem = asyncio.Semaphore(8).
  2. Wrap one call: async def one(p): async with sem: return await asyncio.to_thread(call_model, p).
  3. Fan out: answers = await asyncio.gather(*(one(p) for p in prompts), return_exceptions=True).
  4. Check each slot with isinstance(a, BaseException) before you aggregate.

Each part has a documented job. The asyncio docs say a semaphore "manages an internal counter," and that acquire() blocks at zero until another task calls release(). The docs call async with "the preferred way to use a Semaphore." They describe asyncio.to_thread as a way to run "IO-bound functions/methods that would otherwise block the event loop." If your client has a native async method, await it directly and drop to_thread, as the library does with acompletion.

asyncio.TaskGroup is the newer option, added in Python 3.11. Its failure rule is different. The docs say that the first time a task in the group fails, "the remaining tasks in the group are cancelled." That is right for work where one failure makes the rest useless. It is wrong for a chunk fan-out where 19 good answers out of 20 still help. The library chose gather with return_exceptions=True, and for a map over chunks you should too.

What goes wrong when you add concurrency?

  • A second event loop. The docs for asyncio.run say it "cannot be called when another asyncio event loop is running in the same thread." If your application already runs a loop, model-written REPL code that calls asyncio.run in that thread will raise an error. The batched helper avoids this because its loop runs on the handler side.
  • Batches that are too wide. The system prompt tells the model that batching "only parallelizes" and does not relax the budget. It gives "a useful rough ceiling" of about 20 prompts per batch and about 100K characters per prompt. It calls "tiny-prompt mega-batches" the anti-pattern.
  • Dependent calls. Parallel calls cannot see each other. A running-summary loop, where call 2 needs the output of call 1, must stay sequential. Only independent chunks belong in a batch.
  • Per-call timing. The handler sets each call's execution_time to the batch total divided by the number of prompts. The code labels this "approximate per-prompt time." Do not read it as a real latency for one call.
  • The serial floor. Root turns stay sequential, because each turn reads the output of the turn before it. Concurrency shortens the leaves of the tree only. Extra depth adds more serial turns, which is one reason to keep recursion shallow.

The bottom line

The paper measured a sequential implementation and named async sub-calls as future work. The library has since added them. In the reference library, call llm_query_batched for independent plain calls and rlm_query_batched for independent children. The first uses asyncio with a cap of 16. The second uses threads with a cap of 4. In DSPy, llm_query_batched uses 8 threads. Write your own semaphore, to_thread and gather code only for a custom harness, and keep failures per slot. Both repositories change often, so read the source for the release you install.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, HTML version. States that all LM calls were blocking and sequential, and names asynchronous sub-calls as future work. arxiv.org/html/2512.24601
  2. Wang, D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, HTML version. Execution times of 3.6, 89.3 and 344.5 seconds for DeepSeek v3.2 on S-NIAH. arxiv.org/html/2603.02615
  3. Zhang, A. L., et al. "rlm" GitHub repository, README. Lists llm_query_batched, rlm_query_batched and max_concurrent_subcalls. raw.githubusercontent.com/alexzhang13/rlm/main/README.md
  4. Zhang, A. L., et al. "rlm/environments/local_repl.py" source. Signatures of the batched helpers and the thread pool for recursive children. raw.githubusercontent.com/alexzhang13/rlm/main/rlm/environments/local_repl.py
  5. Zhang, A. L., et al. "rlm/core/lm_handler.py" source. The asyncio.gather and asyncio.Semaphore code in _handle_batched, and the batch_max_concurrent default of 16. raw.githubusercontent.com/alexzhang13/rlm/main/rlm/core/lm_handler.py
  6. Zhang, A. L., et al. "rlm/core/rlm.py" source. The max_concurrent_subcalls default of 4 and the LMHandler construction. raw.githubusercontent.com/alexzhang13/rlm/main/rlm/core/rlm.py
  7. Zhang, A. L., et al. "rlm/utils/prompts.py" source. System prompt text on llm_query_batched and the rough ceiling of 20 prompts per batch. raw.githubusercontent.com/alexzhang13/rlm/main/rlm/utils/prompts.py
  8. Stanford NLP. "dspy/predict/rlm.py" source. llm_query_batched with a ThreadPoolExecutor and max_workers of 8. raw.githubusercontent.com/stanfordnlp/dspy/main/dspy/predict/rlm.py
  9. Python Software Foundation. "Coroutines and Tasks." Python 3 documentation. Behaviour of asyncio.gather, asyncio.TaskGroup and asyncio.to_thread. docs.python.org/3/library/asyncio-task.html
  10. Python Software Foundation. "Synchronization Primitives." Python 3 documentation. Behaviour of asyncio.Semaphore. docs.python.org/3/library/asyncio-sync.html
  11. Python Software Foundation. "Runners." Python 3 documentation. The rule that asyncio.run cannot be called when another event loop runs in the same thread. docs.python.org/3/library/asyncio-runner.html
FAQ

Frequently asked questions

Is there a batched sub-call function in the reference RLM library?

Yes. The REPL exposes llm_query_batched(prompts, model=None) for plain model calls and rlm_query_batched(prompts, model=None) for recursive children. Both return a list of strings in the same order as the input prompts.

How many sub-calls run at once by default?

For llm_query_batched, the LM handler uses an asyncio.Semaphore with batch_max_concurrent, which defaults to 16. For rlm_query_batched, a thread pool uses max_concurrent_subcalls, which defaults to 4. In DSPy, llm_query_batched uses a thread pool with max_workers set to 8.

Should I use asyncio.gather or asyncio.TaskGroup for a chunk fan-out?

Use gather with return_exceptions=True when partial results are useful. The Python docs say a TaskGroup cancels the remaining tasks the first time one task fails. The reference library uses gather so that one failed call does not abort the batch.

Did the RLM paper measure the speedup from parallel sub-calls?

No. The paper says all LM calls in its experiments were blocking and sequential. It says asynchronous sub-calls can potentially reduce runtime, but it reports no measured speedup.

Can every RLM sub-call run in parallel?

No. Only independent calls can, such as the same question over separate chunks. A loop where each call needs the previous answer must stay sequential, and so must the turns of the root model.