Perception
Seeing and understanding: recognising objects, reading fine detail and making sense of whole scenes in images and video.
Community research initiative · 2026
Benchmarking Recursive Self‑Improvement of Frontier Agents
Study how agents improve across rounds with a flexible, budget-aware benchmark designed to be practical to run and easy to extend.
At a glance
Robots, panoramas, gigapixel photos and 3D scenes: the same improvement protocol across very different tasks.
Overview
Flex-RSI measures how well today's frontier agent systems improve on their own: given a downstream task, the agent iteratively refines its own harness, from its solution to its tools and prompts, to raise its score within a fixed budget. Tasks span visual perception and reasoning, robot control, 3D reconstruction and more.
We are actively expanding Flex-RSI with more tasks, tracks and models.
Begin from a documented task setup and starter solution, then let the agent improve its own approach.
Set limits for time, tokens, API spend or GPU hours, and report the budget with each run.
Use validation to guide iteration and reserve test data for the final score.
Record edits, experiments, model calls and outcomes so each run can be understood and repeated.
Bring a task, a model baseline, an evaluation adapter, or feedback on the protocol. Small, reproducible contributions can make RSI experiments easier for more people to run and compare. We will open-source the benchmark for broader evaluation once its infrastructure is complete.
Demos
Compare initial and self-improved solutions, then follow selected robot-control episodes across rounds. Each example is paired with its task metric and evaluation split.
Leaderboard
Standings are kept per task, so models are compared only where they were measured. The overview shows which models have been run on what; each category ranks its tasks on the hidden test split and follows every run's validation curve.
Tracks
The current examples span visual search, active search, 3D tracking, and robot control. Flex-RSI is designed to grow as contributors bring new tasks and domains.
Seeing and understanding: recognising objects, reading fine detail and making sense of whole scenes in images and video.
Drawing conclusions over several steps: spatial, temporal and logical reasoning about what is observed and what is known.
Acting in physical and simulated worlds: planning, controlling robots and moving to gather the information a task needs.
Recovering the structure of the world: 3D and 4D geometry, correspondence and motion from images and video.
Making and editing content: images, scenes and other media produced to precise instructions.
Scientific problems in the life sciences, from sequences and molecules to cells and experiments.
Staying safe and within bounds: improving without breaking the rules, budgets and boundaries an agent is given.
More domains as the benchmark grows.
Contribute
Choose a task and model, then declare the run budget.
The agent revises its solution, prompts, or tools over repeated rounds.
Use validation to check whether a change improves on the starting point.
End at the declared budget or round limit, and record resources used.
Score the selected final solution once on held-out test data.
Submission channel opening soon.
Citation
@misc{flexrsi2026,
title = {Flex-RSI: Benchmarking Recursive Self-Improvement
of Frontier Agents},
author = {To be announced},
year = {2026},
note = {Project page: https://flex-rsi.github.io}
}