Back to Blog
October 3, 2026

The short version: No official one. As of October 2026 the RLM authors have published the code, two sets of model weights and a benchmark, but not the trajectories they trained on. Hugging Face holds a few small community uploads, and none of them is a documented, licensed copy of the paper's training set. The practical route is to generate your own trajectories with the reference library's logger, and the paper gives the recipe.

What did the RLM authors actually release?

Four things are public, and none of them is a trajectory corpus.

  • The inference and training code. The alexzhang13/rlm repository describes itself as "both an extensible inference engine and training environment". It includes a logger, a trajectory visualizer and a training/ folder.
  • RLM-Qwen3-8B weights. The mit-oasys/rlm-qwen3-8b-v0.1 model card carries an MIT license tag. Its full description of the training data is one sentence: "The model was trained on trajectories produced using a fixed system prompt." It links no dataset.
  • A LoRA adapter for a larger model. mit-oasys/rlm-qwen3-30b-a3b-v0.1 is an Apache-2.0 adapter of about 51 MB for Qwen3-30B-A3B-Instruct-2507, "trained as a Recursive Language Model policy on a mixed long-context environment suite using RL".
  • A benchmark. The only dataset under the mit-oasys organization is oolong-pairs. It holds 20 question templates with ground-truth answers at 11 context lengths. It is evaluation data, and its card warns that it does not even contain the context text.

The alexzhang13 account lists no datasets on Hugging Face. So the weights that came out of the trajectories are public, and the trajectories are not.

What does the paper say about the trajectories it trained on?

Enough to rebuild them. Appendix A of the RLM paper lays out the pipeline behind RLM-Qwen3-8B:

  • The teacher was Qwen3-Coder-480B-A35B running as an RLM, with Qwen3-8B answering the sub-calls.
  • The tasks were 750 English LongBenchPro tasks, which produced 2,250 candidate trajectories.
  • Trajectories that scored exactly 0.0, or that never went past one turn, were removed. That left 1,072.
  • Each root turn became its own supervised sample: the input is the full history, the output is what the root model wrote at that step.
  • Turns longer than about 100k characters were dropped to fit the Qwen3-8B context limit.
  • Training ran in the prime-rl library at batch size 64 for 300 steps, which took 48 H100 hours.

The main text rounds the corpus to "1,000 filtered trajectories". The abstract reports that the resulting model beats base Qwen3-8B by 28.3% on average. The appendix gives no download link for the data. It does state the idea that makes the corpus small: leaf sub-calls are ordinary LLM requests, so only the root model's behavior needs teaching. The small-model efficiency post covers what that training bought.

Are there community RLM trajectory datasets on Hugging Face?

A few, and each comes with a limit you should know before you train on it.

HotCopyAI/rlm-trajectories-seed is the best documented. It is Apache-2.0 and its card calls it "a 12-row seed corpus of synthetic Recursive Language Model trajectories". The rows are hand-crafted, the generated code is JavaScript for the vendor's own CLI, and the card says the token and cost figures are illustrative. Twelve synthetic rows in a different harness will not train a root model for the Python reference library.

Two uploads from one Hugging Face user, dated February 2026, look like an attempt to repeat the paper's recipe. rlm-longbenchpro-raw has 512 rows with question, response, execution time and token usage for Qwen/Qwen3-Coder-Next. rlm-longbenchpro-sft-filtered has 2,760 prompt and completion pairs. Neither has a written dataset card, and neither declares a license. Nothing on the pages says how the runs were filtered or which system prompt produced them. Treat them as something to inspect, not something to ship a model on.

The reproduction study says "code and data are available", and that is true for evaluation. Its repository stores per-condition result files in JSON and CSV for DeepSeek v3.2 and Kimi K2 on S-NIAH and OOLONG. Those files record how each run scored. They are not a training corpus.

How do you generate your own RLM trajectories?

The reference library already records everything a training set needs. Follow the paper's pipeline in six steps.

  1. Pick tasks with a scorer. You need a number for each run so you can reject bad ones. The paper used LongBenchPro. Any long-context task set with ground truth works.
  2. Pick a teacher. Use a model stronger than the student. The base model requirements apply to the teacher first.
  3. Turn on the logger. Pass logger=RLMLogger(log_dir="./logs") to RLM. Each completion then writes a .jsonl file. The README says the completion's metadata field "holds the full trajectory (run config + all iterations and sub-calls)".
  4. Reject failures. Drop runs that score zero and runs that end in one turn.
  5. Split by root turn. Write one sample for each iteration, with the history as input and the root output as target. Drop samples that exceed the student's context limit.
  6. Patch the template mistakes. This step is easy to skip and the paper says it mattered. In the teacher's output, 16% of turns misused FINAL and 13% wrongly returned a REPL variable through FINAL_VAR. A programmatic fix for those patterns led to "much better performance" in the student.

Hold out an evaluation set before you start, and score the student on tasks the teacher never saw. The evaluation guide lists the ablations to run. The paper's own test was strict on this point: it trained on LongBenchPro and measured on four unrelated benchmarks.

Do you need a trajectory dataset at all?

Not if you train with reinforcement learning. The repository's training/ folder exposes rlm.RLM as a verifiers environment that plugs into prime-rl. The trainer produces its own rollouts and scores them against a reward, so no teacher corpus exists at any point. The repository ships an OOLONG example environment and a config named rlm-qwen3-30b-example.toml. The published 30B adapter came from this path.

The paper used the same approach for a second experiment. It RL-trained a Qwen3-4B model on the MRCRv2 split with 32k to 64k tokens and 2 needles, for 150 steps at batch size 128 with 4 rollouts per example. The trained RLM then generalized to the split with 512K to 1M tokens and 8 needles.

The authors point the same way for future work. They expect that larger scale and "ideally on-policy and online" rollouts will be necessary to get the most from RLM training. A static trajectory file is the cheap first step. RL is where the authors expect the gains.

Why can't you reuse trajectories from another harness?

Because a trajectory encodes its scaffold. The RLM-Qwen3-8B card says the model "assumes the environment/scaffold from our RLM repo". The 30B adapter card says the same thing more bluntly: it is "not a drop-in chat model" and "expects the RLM system prompt and REPL scaffolding".

Three things are baked into every sample: the system prompt, the names of the sub-call functions, and the final-answer convention. Change any of them and the student learns a protocol your runtime does not speak. The paper met this problem inside one harness. The prompt written for GPT-5 led to "different, undesirable behavior" in Qwen3-Coder until the authors added a sentence. The fine-tuned 8B model also needed a slightly different prompt because its window is 32k, not 272k.

So collect trajectories in the exact harness you will deploy, with the exact prompt. Data from a JavaScript CLI or from another library's RLM module does not transfer.

The bottom line

There is no official public dataset of RLM trajectories. The authors released code, an 8B model, a 30B LoRA adapter and a benchmark. The community uploads are either tiny and synthetic or undocumented and unlicensed. The gap is smaller than it looks: about 1,000 filtered trajectories were enough for the paper's result, the logger that captures them is in the repository, and the filtering recipe is in Appendix A. Generate the data in your own harness, patch the template mistakes, and test on tasks the teacher never saw.

References & Further Reading

  1. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, v3, May 2026. Abstract, fine-tuning setup, Appendix A training details (750 LongBenchPro tasks, 2,250 candidates, 1,072 after filtering, FINAL and FINAL_VAR error rates, 48 H100 hours), MRCRv2 RL training, Appendix B prompt notes. arxiv.org/abs/2512.24601
  2. Zhang, A. L. et al. "rlm: Recursive Language Models." GitHub repository README and training README, accessed October 2026. RLMLogger and trajectory metadata, visualizer, verifiers and prime-rl training harness, OOLONG example environment. github.com/alexzhang13/rlm
  3. MIT OASYS. "RLM-Qwen3-8B-v0.1." Hugging Face model card, accessed October 2026. MIT license tag, statement that the model was trained on trajectories produced using a fixed system prompt and assumes the RLM repo scaffold. huggingface.co/mit-oasys/rlm-qwen3-8b-v0.1
  4. MIT OASYS. "RLM Qwen3-30B-A3B v0.1." Hugging Face model card, accessed October 2026. Apache-2.0 LoRA adapter trained with RL on a mixed long-context environment suite; not a drop-in chat model. huggingface.co/mit-oasys/rlm-qwen3-30b-a3b-v0.1
  5. MIT OASYS. "Oolong-Pairs." Hugging Face dataset card, accessed October 2026. Benchmark of 20 question templates at 11 context lengths; the only dataset under the organization. huggingface.co/datasets/mit-oasys/oolong-pairs
  6. HotCopy. "HotCopy RLM Trajectories (Seed)." Hugging Face dataset card, 2026. Apache-2.0, 12-row synthetic seed corpus, schema and stated limits. huggingface.co/datasets/HotCopyAI/rlm-trajectories-seed
  7. Shamsi, S. "rlm-longbenchpro-sft-filtered" and "rlm-longbenchpro-raw." Hugging Face datasets, February 2026. Community uploads with 2,760 and 512 rows; no written card and no declared license. huggingface.co/datasets/ShayanShamsi/rlm-longbenchpro-sft-filtered
  8. Wang, D. "Think, But Don't Overthink: Reproducing Recursive Language Models." arXiv:2603.02615, 2026. Reproduction with DeepSeek v3.2 and Kimi K2; released code and evaluation result files. arxiv.org/abs/2603.02615
FAQ

Frequently asked questions

Did the RLM authors release the trajectories used to train RLM-Qwen3-8B?

No. The model weights are on Hugging Face at mit-oasys/rlm-qwen3-8b-v0.1, and the model card says only that the model was trained on trajectories produced using a fixed system prompt. The card links no dataset, and the only dataset under the mit-oasys organization is the OOLONG-Pairs benchmark. The paper's Appendix A describes how the trajectories were collected and filtered.

How many trajectories does it take to fine-tune a small model as an RLM?

The paper's result used about 1,000. The authors collected 2,250 candidate trajectories from 750 English LongBenchPro tasks and kept 1,072 after removing zero-score and single-turn runs. Each root turn then became a separate training sample, so the number of samples is larger than the number of trajectories. The authors say much larger scale will be needed to get the most from RLM training.

What license do the released RLM models use?

The RLM-Qwen3-8B model card on Hugging Face carries an MIT license tag. The RLM Qwen3-30B-A3B adapter card carries an Apache-2.0 tag and is a LoRA adapter, so it also needs the base Qwen3-30B-A3B-Instruct-2507 model. Check each card and the base model's license before commercial use.

Can I use the reference RLM library's logs as training data?

Yes, after processing. Passing RLMLogger with a log_dir makes each completion write a JSONL file, and the completion metadata holds the full trajectory with all iterations and sub-calls. You still need to score each run, reject failures, split by root turn, and patch template mistakes before the logs are usable for supervised fine-tuning.

Is the 12-row HotCopy dataset enough to fine-tune an RLM?

No. Its own card describes it as a synthetic seed corpus of hand-crafted trajectories and says it is not suitable for benchmarking or for cost estimates. The generated code is JavaScript for a different harness, so it does not match the Python REPL protocol of the reference library. It is useful as a reading example of the trajectory shape.