The RLM paper and its reference library pass text to sub-calls. We read the code, the follow-up papers, and the provider vision docs to see what it takes to send images instead.
The short version: Not out of the box. The RLM paper and the reference library treat the context as text in a Python REPL, and llm_query is documented and typed as taking a string. You can still route a sub-call to a vision language model, but you have to build the image path yourself: keep image bytes or file references in the REPL, and pass provider-format image blocks to a VLM from code. One follow-up paper, RVLM, has already done this and is the best worked example.
Nothing. The RLM paper by Zhang, Kraska, and Khattab frames the whole method as a way to let language models "process arbitrarily long prompts," and its abstract makes no mention of images or other modalities. The prompt is loaded into the REPL as a variable, the root model writes code to slice it, and sub-calls read the slices. Every piece of that loop is text.
That is a scope choice, not a hard limit of the idea. The core move (keep the input outside the model, let code pick what each call sees) does not care what the input is made of. But no experiment in the paper runs a sub-call on an image, so there is no published RLM result that tells you how well a vision sub-call works inside the loop.
Not as a feature. We read the source of the official rlm repository (package version 0.1.3) and searched it for image, vision, and multimodal handling. The only "image" hits are Docker, Modal, and Daytona sandbox images. There is no image loader, no image primitive in the REPL, and no mention of vision in the README or docs.
The REPL functions are text-first by design. In rlm/environments/local_repl.py, llm_query(prompt, model=None) is annotated as prompt: str, and the system prompt tells the root model to use it "for simple extraction, summarization, or Q&A over a chunk of text." Context loading is text or JSON as well: a string is written to a .txt file, and a dict or list is written with json.dump and read back into context.
There is one unadvertised opening. The client classes accept str | list[dict]. The OpenAI client passes a list of messages straight to chat.completions.create, and the Anthropic client passes the non-system messages straight to messages.create. Neither inspects the content of each message. So, from reading the code, a list of messages that holds image content blocks should reach a vision model through llm_query on those two backends. The Gemini client is different: it wraps every message content in types.Part(text=content), so an image block does not survive on that backend. We did not run this path, the type hints say str, and nothing in the docs promises it. Treat it as a hack you must test, not a supported API.
Yes. RVLM (Mayumu et al., March 2026) says directly that it builds on the RLM paradigm and extends the REPL with images. It adds a context_images list to the REPL namespace, where each image is a dictionary with data (base64 or URL), media_type, and detail. It adds describe_image, which "sends a single image to a sub-VLM and returns a textual description," and llm_query_with_images, which "constructs multimodal messages and dispatches to a sub-VLM." It also exposes crop, enhance, and difference-map functions so the root can zoom into regions before it asks.
The design choice that matters most: in RVLM, only the first user message is multimodal. Every later root turn is text, and the images are reached only through the REPL functions. The authors say this cuts per-iteration API cost by about 70 to 80 percent. We could not find a measurement behind that figure in the paper, so read it as an estimate. RVLM was evaluated on brain MRI (BraTS 2023 Meningioma) and chest X-ray (MIMIC-CXR) with Gemini 2.5 Flash as both root and sub-VLM, and the abstract reports consistency and qualitative findings, not a benchmark score against a baseline.
A second example goes further from text. TimeRLM (Zumarraga et al., August 2026) applies RLM-style recursion to long time series for anomaly localization. The series lives in the REPL as a JSON variable, and one variant gives the sandbox vision libraries so the model can "render time-series as a plot and perceive it visually." That is a useful pattern: the data is not an image, but the model turns a slice of it into one when a look beats a table of numbers.
The work is small if you keep the RVLM shape. The pieces are:
vlm_query(prompt, images) that builds a provider message and calls a vision model. For OpenAI, the images and vision guide shows an image_url content part that takes a URL or a data:image/jpeg;base64,... data URL, with a detail setting of low, high, original, or auto. For Anthropic, the vision guide shows an image content block with a base64, URL, or Files API file_id source.Register the function in the same way the library registers llm_query, or run the RLM with a custom setup code block that defines it. If your VLM calls go to a separate client, the usage counters in the library will not see them, so log tokens yourself. The same sandbox concerns from our post on sandboxing the RLM REPL apply, with one more: a REPL that can read image paths can read other files too.
Images are priced as input tokens, and the counts add up fast across many sub-calls. Anthropic states that an image costs ceil(width / 28) x ceil(height / 28) visual tokens, and its table lists 1,296 tokens for a 1000x1000 image. On models in its high-resolution tier, the long edge is capped at 2576 px and 4,784 visual tokens. OpenAI states that vision models "convert image inputs into billable input tokens" and that image tokens also count toward tokens-per-minute limits.
Request limits also shape how you batch. Anthropic allows 600 images per API request on most models (100 on models with a 200k-token context window), but a stricter per-image dimension limit applies once a request has more than 20 images, and the 32 MB request size limit can hit first. OpenAI lists up to 1,500 images and 512 MB per request. Anthropic also warns that base64 images in a multi-turn call are resent in the payload on every turn. That is the strongest reason to follow the RVLM rule and never put images in the root history: in an RLM the root runs many turns, and each resend is paid in full.
Parallel fan-out makes this worse. If the root calls a vision sub-call on every page of a 400-page scanned document in a batch, the rate limit will bite before the context window does. Our post on prompt caching for sub-calls covers the text side of this cost, and the same logic applies to a shared image prefix. The sub-model also now has to be a VLM, which narrows the choice of cheap leaf models compared with picking a base model for a text RLM.
RLM sub-calls can use a vision language model, but the paper does not test it and the reference library does not ship it. The clean way is to keep images as REPL data, add a function that sends selected images or crops to a VLM, and return text to the root. RVLM is a working reference for that design, and TimeRLM shows the same idea for data you render into plots. Until someone publishes a controlled comparison, any claim that vision sub-calls beat a single long multimodal call is untested.
The reference library loads context as a string or as JSON, so a PDF has to be converted first. You can extract the text for the normal RLM loop and render figure pages to images that a separate vision function reads on demand. The library does not do either step for you.
Not officially. It is typed and documented as taking a string. From reading the code, a list of messages with image blocks should pass through the OpenAI and Anthropic clients unchanged, but the Gemini client wraps every message as text, and we did not test the path.
It does not need to be if sub-calls return text. RVLM used Gemini 2.5 Flash for both roles, but its design keeps images out of the root history after the first message. A text-only root that calls a vision leaf is a valid setup as long as the root gets useful metadata about each image.
Not a shared one that we could find. RVLM reports qualitative and consistency results on brain MRI and chest X-ray data, and TimeRLM built its own synthetic time-series benchmark called AnomalyXL. Neither compares a vision RLM against a single long multimodal call on a common task.