The RLM paper ran every sub-call one after another and called async calls future work. The current reference library ships a batched helper that uses asyncio for you.
The short version: in the reference RLM library you do not write the asyncio code yourself. The REPL already has llm_query_batched(prompts, model=None), and the library runs that batch with asyncio.gather behind an asyncio.Semaphore of 16. Recursive children go through rlm_query_batched, which uses a thread pool of 4 by default. You write your own asyncio fan-out only when you build your own harness, and then the pattern is a semaphore, asyncio.to_thread and gather.
The paper and the code disagree on this point, so it helps to read both. The sections below separate what the library documents from what you build yourself.
No. The RLM paper by Zhang, Kraska and Khattab says that in its experiments "all LM calls are blocking / sequential." Its limitations appendix has a heading that reads "RLMs without asynchronous LM calls are slow." The authors write that they "implemented all sub-LM queries naively as blocking / sequential calls" and that they are "confident that this can be resolved with a robust implementation."
The paper also names the fix as future work. It says "asynchronous sub-calls and sandboxed REPLs can potentially significantly reduce the runtime and inference cost of RLMs, but further contribute to this complexity." The paper does not report a measured speedup from parallel sub-calls, so this article does not give one.
The cost of sequential calls is visible in the reproduction study. It measured DeepSeek v3.2 at 3.6 seconds for a base S-NIAH query, 89.3 seconds at depth 1 and 344.5 seconds at depth 2. Our post on whether RLMs are slower than a single long-context call covers those numbers.
The current code is ahead of the paper. The README of the authors' repository lists "single LM calls (llm_query / llm_query_batched)" and "parallel batched sub-calls bounded by max_concurrent_subcalls." So the helper exists, and its name is llm_query_batched.
The signature in the local REPL source is _llm_query_batched(self, prompts: list[str], model: str | None = None) -> list[str]. The REPL exposes it to the model as llm_query_batched. The docstring says it returns a "list of responses in the same order as input prompts."
The system prompt teaches the root model to use it. The example in the prompt builds a list and then runs answers = llm_query_batched(prompts). The prompt describes the function as "much faster than sequential llm_query calls for independent queries."
It is in the LM handler, in a method named _handle_batched. The REPL sends the whole prompt list to the handler in one request. The handler then does three things, all quoted from the source:
sem = asyncio.Semaphore(handler.batch_max_concurrent). The default for batch_max_concurrent is 16.async with sem: return await client.acompletion(prompt).return await asyncio.gather(*tasks, return_exceptions=True), started with asyncio.run(run_all()).The comment in the code explains the last flag: "one failed call doesn't abort the whole batch; failures are surfaced per-prompt as error completions." That matches the Python documentation for gather. With return_exceptions=True, "exceptions are treated the same as successful results, and aggregated in the result list." The same page says the order of results "corresponds to the order of awaitables," which is why chunk 3 always lands in slot 3.
One detail limits your control. In rlm/core/rlm.py, the handler is built as LMHandler(client, other_backend_client=other_backend_client). That call does not pass batch_max_concurrent, so 16 applies unless you construct the handler yourself. The constructor argument that RLM does expose is max_concurrent_subcalls, and it controls a different path.
No, they use threads. _rlm_query_batched in the same REPL file runs each child RLM in a ThreadPoolExecutor. The pool size is min(self.max_concurrent_subcalls, len(prompts)), and max_concurrent_subcalls defaults to 4. The docstring gives the reason: the sub-calls "are independent and I/O-bound." If recursion is not configured, the function falls back to llm_query_batched.
DSPy takes the thread route for plain sub-calls as well. In dspy/predict/rlm.py, llm_query_batched(prompts) submits each prompt to a ThreadPoolExecutor with max_workers=8. A failed model call returns a string that starts with [ERROR] in that slot. See our comparison of dspy.RLM and the reference library for the other differences.
So there are two limits in the reference library, and they stack. A depth-2 run can start 4 children at a time, and each child can send a batch of up to 16 concurrent plain calls. Check that product against your provider rate limit before you raise either number.
You need this only when you write your own harness, or when your sub-call function is synchronous and has no batched form. The following is an illustrative sketch. It uses only standard asyncio APIs, and call_model stands for your own blocking function. It is not code from the RLM library.
sem = asyncio.Semaphore(8).async def one(p): async with sem: return await asyncio.to_thread(call_model, p).answers = await asyncio.gather(*(one(p) for p in prompts), return_exceptions=True).isinstance(a, BaseException) before you aggregate.Each part has a documented job. The asyncio docs say a semaphore "manages an internal counter," and that acquire() blocks at zero until another task calls release(). The docs call async with "the preferred way to use a Semaphore." They describe asyncio.to_thread as a way to run "IO-bound functions/methods that would otherwise block the event loop." If your client has a native async method, await it directly and drop to_thread, as the library does with acompletion.
asyncio.TaskGroup is the newer option, added in Python 3.11. Its failure rule is different. The docs say that the first time a task in the group fails, "the remaining tasks in the group are cancelled." That is right for work where one failure makes the rest useless. It is wrong for a chunk fan-out where 19 good answers out of 20 still help. The library chose gather with return_exceptions=True, and for a map over chunks you should too.
asyncio.run in that thread will raise an error. The batched helper avoids this because its loop runs on the handler side.execution_time to the batch total divided by the number of prompts. The code labels this "approximate per-prompt time." Do not read it as a real latency for one call.The paper measured a sequential implementation and named async sub-calls as future work. The library has since added them. In the reference library, call llm_query_batched for independent plain calls and rlm_query_batched for independent children. The first uses asyncio with a cap of 16. The second uses threads with a cap of 4. In DSPy, llm_query_batched uses 8 threads. Write your own semaphore, to_thread and gather code only for a custom harness, and keep failures per slot. Both repositories change often, so read the source for the release you install.
Yes. The REPL exposes llm_query_batched(prompts, model=None) for plain model calls and rlm_query_batched(prompts, model=None) for recursive children. Both return a list of strings in the same order as the input prompts.
For llm_query_batched, the LM handler uses an asyncio.Semaphore with batch_max_concurrent, which defaults to 16. For rlm_query_batched, a thread pool uses max_concurrent_subcalls, which defaults to 4. In DSPy, llm_query_batched uses a thread pool with max_workers set to 8.
Use gather with return_exceptions=True when partial results are useful. The Python docs say a TaskGroup cancels the remaining tasks the first time one task fails. The reference library uses gather so that one failed call does not abort the batch.
No. The paper says all LM calls in its experiments were blocking and sequential. It says asynchronous sub-calls can potentially reduce runtime, but it reports no measured speedup.
No. Only independent calls can, such as the same question over separate chunks. A loop where each call needs the previous answer must stay sequential, and so must the turns of the root model.