Community research initiative · 2026

Flex-RSI

Benchmarking Recursive Self‑Improvement of Frontier Agents

Study how agents improve across rounds with a flexible, budget-aware benchmark designed to be practical to run and easy to extend.

At a glance

Robots, panoramas, gigapixel photos and 3D scenes: the same improvement protocol across very different tasks.

Action · RoboDojoArrange largest number
ActionActive search
Reconstruction3D tracking
Action · RoboDojoInsert tubes
PerceptionVisual search
Action · RoboCasaOpen cabinet
Reconstruction3D tracking
Action · RoboDojoThrow bottles into dustbin

Overview

How far can a frontier agent improve itself?

Flex-RSI measures how well today's frontier agent systems improve on their own: given a downstream task, the agent iteratively refines its own harness, from its solution to its tools and prompts, to raise its score within a fixed budget. Tasks span visual perception and reasoning, robot control, 3D reconstruction and more.

We are actively expanding Flex-RSI with more tasks, tracks and models.

Shared starting point

Begin from a documented task setup and starter solution, then let the agent improve its own approach.

Declared budgets

Set limits for time, tokens, API spend or GPU hours, and report the budget with each run.

Held-out evaluation

Use validation to guide iteration and reserve test data for the final score.

Traceable progress

Record edits, experiments, model calls and outcomes so each run can be understood and repeated.

Help shape the benchmark

Bring a task, a model baseline, an evaluation adapter, or feedback on the protocol. Small, reproducible contributions can make RSI experiments easier for more people to run and compare. We will open-source the benchmark for broader evaluation once its infrastructure is complete.

Demos

From starting points to improved behavior

Compare initial and self-improved solutions, then follow selected robot-control episodes across rounds. Each example is paired with its task metric and evaluation split.

Leaderboard

Results

Standings are kept per task, so models are compared only where they were measured. The overview shows which models have been run on what; each category ranks its tasks on the hidden test split and follows every run's validation curve.

Tracks

One protocol, many kinds of tasks

The current examples span visual search, active search, 3D tracking, and robot control. Flex-RSI is designed to grow as contributors bring new tasks and domains.

Perception

Seeing and understanding: recognising objects, reading fine detail and making sense of whole scenes in images and video.

Reasoning

Drawing conclusions over several steps: spatial, temporal and logical reasoning about what is observed and what is known.

Action

Acting in physical and simulated worlds: planning, controlling robots and moving to gather the information a task needs.

Reconstruction

Recovering the structure of the world: 3D and 4D geometry, correspondence and motion from images and video.

Creation

Making and editing content: images, scenes and other media produced to precise instructions.

Bio

Scientific problems in the life sciences, from sequences and molecules to cells and experiments.

Safety

Staying safe and within bounds: improving without breaking the rules, budgets and boundaries an agent is given.

…

More domains as the benchmark grows.

Contribute

How to contribute

  1. Start

    Choose a task and model, then declare the run budget.

  2. Iterate

    The agent revises its solution, prompts, or tools over repeated rounds.

  3. Accept

    Use validation to check whether a change improves on the starting point.

  4. Stop

    End at the declared budget or round limit, and record resources used.

  5. Test

    Score the selected final solution once on held-out test data.

Submission channel opening soon.

Citation

BibTeX

@misc{flexrsi2026,
  title  = {Flex-RSI: Benchmarking Recursive Self-Improvement
            of Frontier Agents},
  author = {To be announced},
  year   = {2026},
  note   = {Project page: https://flex-rsi.github.io}
}