Back to Blog
September 22, 2026

The short version: Not out of the box. The RLM paper and the reference library treat the context as text in a Python REPL, and llm_query is documented and typed as taking a string. You can still route a sub-call to a vision language model, but you have to build the image path yourself: keep image bytes or file references in the REPL, and pass provider-format image blocks to a VLM from code. One follow-up paper, RVLM, has already done this and is the best worked example.

What does the RLM paper say about images?

Nothing. The RLM paper by Zhang, Kraska, and Khattab frames the whole method as a way to let language models "process arbitrarily long prompts," and its abstract makes no mention of images or other modalities. The prompt is loaded into the REPL as a variable, the root model writes code to slice it, and sub-calls read the slices. Every piece of that loop is text.

That is a scope choice, not a hard limit of the idea. The core move (keep the input outside the model, let code pick what each call sees) does not care what the input is made of. But no experiment in the paper runs a sub-call on an image, so there is no published RLM result that tells you how well a vision sub-call works inside the loop.

Does the reference library support multimodal sub-calls?

Not as a feature. We read the source of the official rlm repository (package version 0.1.3) and searched it for image, vision, and multimodal handling. The only "image" hits are Docker, Modal, and Daytona sandbox images. There is no image loader, no image primitive in the REPL, and no mention of vision in the README or docs.

The REPL functions are text-first by design. In rlm/environments/local_repl.py, llm_query(prompt, model=None) is annotated as prompt: str, and the system prompt tells the root model to use it "for simple extraction, summarization, or Q&A over a chunk of text." Context loading is text or JSON as well: a string is written to a .txt file, and a dict or list is written with json.dump and read back into context.

There is one unadvertised opening. The client classes accept str | list[dict]. The OpenAI client passes a list of messages straight to chat.completions.create, and the Anthropic client passes the non-system messages straight to messages.create. Neither inspects the content of each message. So, from reading the code, a list of messages that holds image content blocks should reach a vision model through llm_query on those two backends. The Gemini client is different: it wraps every message content in types.Part(text=content), so an image block does not survive on that backend. We did not run this path, the type hints say str, and nothing in the docs promises it. Treat it as a hack you must test, not a supported API.

Has anyone built an RLM with vision sub-calls?

Yes. RVLM (Mayumu et al., March 2026) says directly that it builds on the RLM paradigm and extends the REPL with images. It adds a context_images list to the REPL namespace, where each image is a dictionary with data (base64 or URL), media_type, and detail. It adds describe_image, which "sends a single image to a sub-VLM and returns a textual description," and llm_query_with_images, which "constructs multimodal messages and dispatches to a sub-VLM." It also exposes crop, enhance, and difference-map functions so the root can zoom into regions before it asks.

The design choice that matters most: in RVLM, only the first user message is multimodal. Every later root turn is text, and the images are reached only through the REPL functions. The authors say this cuts per-iteration API cost by about 70 to 80 percent. We could not find a measurement behind that figure in the paper, so read it as an estimate. RVLM was evaluated on brain MRI (BraTS 2023 Meningioma) and chest X-ray (MIMIC-CXR) with Gemini 2.5 Flash as both root and sub-VLM, and the abstract reports consistency and qualitative findings, not a benchmark score against a baseline.

A second example goes further from text. TimeRLM (Zumarraga et al., August 2026) applies RLM-style recursion to long time series for anomaly localization. The series lives in the REPL as a JSON variable, and one variant gives the sandbox vision libraries so the model can "render time-series as a plot and perceive it visually." That is a useful pattern: the data is not an image, but the model turns a slice of it into one when a look beats a table of numbers.

What would you build to add vision sub-calls yourself?

The work is small if you keep the RVLM shape. The pieces are:

  1. Keep images out of the root prompt. Load file paths or base64 strings into a REPL variable, and give the root only metadata (count, file names, sizes, page numbers). This is the RLM rule applied to pixels: the root sees a handle, not the payload.
  2. Add a VLM function to the REPL. Write something like vlm_query(prompt, images) that builds a provider message and calls a vision model. For OpenAI, the images and vision guide shows an image_url content part that takes a URL or a data:image/jpeg;base64,... data URL, with a detail setting of low, high, original, or auto. For Anthropic, the vision guide shows an image content block with a base64, URL, or Files API file_id source.
  3. Return text only. The VLM answers in text, and that text goes back into a REPL variable like any other sub-call output. The root model keeps reasoning over strings.
  4. Add image tools the root can call in code. Crop, resize, and page-split functions let the root send a small region instead of a full scan. This is the image version of chunking a long document.

Register the function in the same way the library registers llm_query, or run the RLM with a custom setup code block that defines it. If your VLM calls go to a separate client, the usage counters in the library will not see them, so log tokens yourself. The same sandbox concerns from our post on sandboxing the RLM REPL apply, with one more: a REPL that can read image paths can read other files too.

What limits and costs change when sub-calls see images?

Images are priced as input tokens, and the counts add up fast across many sub-calls. Anthropic states that an image costs ceil(width / 28) x ceil(height / 28) visual tokens, and its table lists 1,296 tokens for a 1000x1000 image. On models in its high-resolution tier, the long edge is capped at 2576 px and 4,784 visual tokens. OpenAI states that vision models "convert image inputs into billable input tokens" and that image tokens also count toward tokens-per-minute limits.

Request limits also shape how you batch. Anthropic allows 600 images per API request on most models (100 on models with a 200k-token context window), but a stricter per-image dimension limit applies once a request has more than 20 images, and the 32 MB request size limit can hit first. OpenAI lists up to 1,500 images and 512 MB per request. Anthropic also warns that base64 images in a multi-turn call are resent in the payload on every turn. That is the strongest reason to follow the RVLM rule and never put images in the root history: in an RLM the root runs many turns, and each resend is paid in full.

Parallel fan-out makes this worse. If the root calls a vision sub-call on every page of a 400-page scanned document in a batch, the rate limit will bite before the context window does. Our post on prompt caching for sub-calls covers the text side of this cost, and the same logic applies to a shared image prefix. The sub-model also now has to be a VLM, which narrows the choice of cheap leaf models compared with picking a base model for a text RLM.

The bottom line

RLM sub-calls can use a vision language model, but the paper does not test it and the reference library does not ship it. The clean way is to keep images as REPL data, add a function that sends selected images or crops to a VLM, and return text to the root. RVLM is a working reference for that design, and TimeRLM shows the same idea for data you render into plots. Until someone publishes a controlled comparison, any claim that vision sub-calls beat a single long multimodal call is untested.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, v3 May 2026. Defines the RLM as a text prompt held in a REPL; no multimodal inputs. arxiv.org/abs/2512.24601
  2. Zhang, A. L., et al. "rlm: General plug-and-play inference library for Recursive Language Models." GitHub, v0.1.3. Source for llm_query typing, context loading, and client message handling. github.com/alexzhang13/rlm
  3. Mayumu, N., Khan, Z., Stephens, M., Mukala, P., Oroumchian, F. "RVLM: Recursive Vision-Language Models with Adaptive Depth." arXiv:2603.24224, March 2026. REPL image primitives, sub-VLM calls, and the text-only root history design. arxiv.org/abs/2603.24224
  4. Zumarraga, N., et al. "TimeRLM: Recursive Language Models Enable Precise Anomaly Localization in Long-Context Time-Series." arXiv:2608.03391, August 2026. RLM recursion over time series with a plot-rendering vision variant. arxiv.org/abs/2608.03391
  5. OpenAI. "Images and vision." OpenAI API docs, accessed September 2026. image_url content parts, detail levels, per-request limits, and image token billing. developers.openai.com/api/docs/guides/images-vision
  6. Anthropic. "Vision." Claude API docs, accessed September 2026. Image content blocks, per-request image limits, and the visual token formula. platform.claude.com/docs/en/build-with-claude/vision
FAQ

Frequently asked questions

Can I load a PDF with figures into an RLM as context?

The reference library loads context as a string or as JSON, so a PDF has to be converted first. You can extract the text for the normal RLM loop and render figure pages to images that a separate vision function reads on demand. The library does not do either step for you.

Does llm_query accept an image in the reference library?

Not officially. It is typed and documented as taking a string. From reading the code, a list of messages with image blocks should pass through the OpenAI and Anthropic clients unchanged, but the Gemini client wraps every message as text, and we did not test the path.

Should the root model also be a vision model?

It does not need to be if sub-calls return text. RVLM used Gemini 2.5 Flash for both roles, but its design keeps images out of the root history after the first message. A text-only root that calls a vision leaf is a valid setup as long as the root gets useful metadata about each image.

Is there a benchmark for multimodal RLMs?

Not a shared one that we could find. RVLM reports qualitative and consistency results on brain MRI and chest X-ray data, and TimeRLM built its own synthetic time-series benchmark called AnomalyXL. Neither compares a vision RLM against a single long multimodal call on a common task.