The paper trained an 8B model on about 1,000 trajectories, and people want that file. It is not public. The weights, the logger and the recipe are.
The short version: No official one. As of October 2026 the RLM authors have published the code, two sets of model weights and a benchmark, but not the trajectories they trained on. Hugging Face holds a few small community uploads, and none of them is a documented, licensed copy of the paper's training set. The practical route is to generate your own trajectories with the reference library's logger, and the paper gives the recipe.
Four things are public, and none of them is a trajectory corpus.
training/ folder.The alexzhang13 account lists no datasets on Hugging Face. So the weights that came out of the trajectories are public, and the trajectories are not.
Enough to rebuild them. Appendix A of the RLM paper lays out the pipeline behind RLM-Qwen3-8B:
The main text rounds the corpus to "1,000 filtered trajectories". The abstract reports that the resulting model beats base Qwen3-8B by 28.3% on average. The appendix gives no download link for the data. It does state the idea that makes the corpus small: leaf sub-calls are ordinary LLM requests, so only the root model's behavior needs teaching. The small-model efficiency post covers what that training bought.
A few, and each comes with a limit you should know before you train on it.
HotCopyAI/rlm-trajectories-seed is the best documented. It is Apache-2.0 and its card calls it "a 12-row seed corpus of synthetic Recursive Language Model trajectories". The rows are hand-crafted, the generated code is JavaScript for the vendor's own CLI, and the card says the token and cost figures are illustrative. Twelve synthetic rows in a different harness will not train a root model for the Python reference library.
Two uploads from one Hugging Face user, dated February 2026, look like an attempt to repeat the paper's recipe. rlm-longbenchpro-raw has 512 rows with question, response, execution time and token usage for Qwen/Qwen3-Coder-Next. rlm-longbenchpro-sft-filtered has 2,760 prompt and completion pairs. Neither has a written dataset card, and neither declares a license. Nothing on the pages says how the runs were filtered or which system prompt produced them. Treat them as something to inspect, not something to ship a model on.
The reproduction study says "code and data are available", and that is true for evaluation. Its repository stores per-condition result files in JSON and CSV for DeepSeek v3.2 and Kimi K2 on S-NIAH and OOLONG. Those files record how each run scored. They are not a training corpus.
The reference library already records everything a training set needs. Follow the paper's pipeline in six steps.
logger=RLMLogger(log_dir="./logs") to RLM. Each completion then writes a .jsonl file. The README says the completion's metadata field "holds the full trajectory (run config + all iterations and sub-calls)".FINAL and 13% wrongly returned a REPL variable through FINAL_VAR. A programmatic fix for those patterns led to "much better performance" in the student.Hold out an evaluation set before you start, and score the student on tasks the teacher never saw. The evaluation guide lists the ablations to run. The paper's own test was strict on this point: it trained on LongBenchPro and measured on four unrelated benchmarks.
Not if you train with reinforcement learning. The repository's training/ folder exposes rlm.RLM as a verifiers environment that plugs into prime-rl. The trainer produces its own rollouts and scores them against a reward, so no teacher corpus exists at any point. The repository ships an OOLONG example environment and a config named rlm-qwen3-30b-example.toml. The published 30B adapter came from this path.
The paper used the same approach for a second experiment. It RL-trained a Qwen3-4B model on the MRCRv2 split with 32k to 64k tokens and 2 needles, for 150 steps at batch size 128 with 4 rollouts per example. The trained RLM then generalized to the split with 512K to 1M tokens and 8 needles.
The authors point the same way for future work. They expect that larger scale and "ideally on-policy and online" rollouts will be necessary to get the most from RLM training. A static trajectory file is the cheap first step. RL is where the authors expect the gains.
Because a trajectory encodes its scaffold. The RLM-Qwen3-8B card says the model "assumes the environment/scaffold from our RLM repo". The 30B adapter card says the same thing more bluntly: it is "not a drop-in chat model" and "expects the RLM system prompt and REPL scaffolding".
Three things are baked into every sample: the system prompt, the names of the sub-call functions, and the final-answer convention. Change any of them and the student learns a protocol your runtime does not speak. The paper met this problem inside one harness. The prompt written for GPT-5 led to "different, undesirable behavior" in Qwen3-Coder until the authors added a sentence. The fine-tuned 8B model also needed a slightly different prompt because its window is 32k, not 272k.
So collect trajectories in the exact harness you will deploy, with the exact prompt. Data from a JavaScript CLI or from another library's RLM module does not transfer.
There is no official public dataset of RLM trajectories. The authors released code, an 8B model, a 30B LoRA adapter and a benchmark. The community uploads are either tiny and synthetic or undocumented and unlicensed. The gap is smaller than it looks: about 1,000 filtered trajectories were enough for the paper's result, the logger that captures them is in the repository, and the filtering recipe is in Appendix A. Generate the data in your own harness, patch the template mistakes, and test on tasks the teacher never saw.
No. The model weights are on Hugging Face at mit-oasys/rlm-qwen3-8b-v0.1, and the model card says only that the model was trained on trajectories produced using a fixed system prompt. The card links no dataset, and the only dataset under the mit-oasys organization is the OOLONG-Pairs benchmark. The paper's Appendix A describes how the trajectories were collected and filtered.
The paper's result used about 1,000. The authors collected 2,250 candidate trajectories from 750 English LongBenchPro tasks and kept 1,072 after removing zero-score and single-turn runs. Each root turn then became a separate training sample, so the number of samples is larger than the number of trajectories. The authors say much larger scale will be needed to get the most from RLM training.
The RLM-Qwen3-8B model card on Hugging Face carries an MIT license tag. The RLM Qwen3-30B-A3B adapter card carries an Apache-2.0 tag and is a LoRA adapter, so it also needs the base Qwen3-30B-A3B-Instruct-2507 model. Check each card and the base model's license before commercial use.
Yes, after processing. Passing RLMLogger with a log_dir makes each completion write a JSONL file, and the completion metadata holds the full trajectory with all iterations and sub-calls. You still need to score each run, reject failures, split by root turn, and patch template mistakes before the logs are usable for supervised fine-tuning.
No. Its own card describes it as a synthetic seed corpus of hand-crafted trajectories and says it is not suitable for benchmarking or for cost estimates. The generated code is JavaScript for a different harness, so it does not match the Python REPL protocol of the reference library. It is useful as a reading example of the trajectory shape.