Where AI Self-Improvement (RSI) Stands: 8 Takeaways / AI 自我改进(RSI)走到哪一步了:8 条 take away
On October 3, NICE hosted an online workshop, “Mechanisms, Evidence and Limits of AI Self-Improvement,” with four researchers who work on this topic: Chengsong Huang (PhD student, Washington University in St. Louis), Shilong Liu (postdoc at Princeton University, joining Columbia University as an assistant professor in fall 2027), Da Yin (NeoCognition) and Cheng Qian (PhD student, UIUC). They discussed the main approaches to RSI, its current bottlenecks, how to evaluate it, and what it means for academia and industry. I summarized the discussion in the 8 takeaways below.
1. RSI improves the whole system, not just the model
When people say “self-improvement” today, what gets improved is a whole agent system. At least three parts of it can be changed.
- Weights. Training changes the model itself.
- Harness. The tools, memory and execution flow wrapped around the model. Changing it does not require retraining the model.
- Environment and data. Huang’s view is that whatever algorithm you use to optimize an agent, you need training data that is harder and closer to the downstream task, and for agents that data is the environment. A stronger model can synthesize better environments, and better environments in turn train a stronger model.
The three advance in turns. Qian calls it an “upward spiral”: components of the harness get trained into the model’s own abilities, and the stronger model then needs a new harness. Liu’s research trajectory followed the same order: first let the model build its own tools, then let it explore in environments, and finally train the model on the weakly supervised data that the exploration produced.
If all three can be changed, they should be changed together. Huang argues that co-optimizing them can never do worse than optimizing each one separately; the only cost is that the optimization itself becomes harder.
Yin offered an analogy. In snooker it is hard to say which cue is the best. Results come from a player’s long practice with their own cue, and even a top player may not reach the same level with a different one. Models and harnesses work the same way: every company has its own infrastructure and data distribution, so company A’s training recipe may not work well on company B’s harness.
There is also the question of who does the improving. Qian divides self-improvement into three levels: revising the current answer, accumulating experience across tasks, and getting better at generating and selecting improvements. Their AI for AI work sits mainly at the third level. In the past, people designed a harness or a set of skills to make the executor more capable. Now a builder AI builds the harness, writes the environments and does the debugging for a target AI, and the relationship between the two resembles that of an advisor and a PhD student.
2. What RSI can do today, and what it cannot do yet
RSI works this time because models can now run the whole improvement loop on their own, and after each change they can find out quickly and cheaply whether the change was right. Tasks where that is not possible mark its current limits: writing, which people have to judge, and scientific questions that need experiments in the physical world.
Why now? Qian gave three reasons:
- Foundation models can handle every step of the loop on their own: reading code, proposing changes, running experiments, analyzing feedback and recording what they learned. No person is needed to connect the steps.
- Harnesses made improvement cheap. Without retraining the model, you can change how a tool is wrapped, the memory policy or the execution flow, and test the change directly on the existing model.
- There are more executable environments. Reinforcement learning has its gyms, and many new benchmarks are no longer “one input, one output” but sandboxes in which an agent can act freely. In environments like these a system gets more direct feedback.
Liu’s view is that a person plus a computer has always been an RSI system. The person’s bandwidth is just too narrow, and a person cannot run 24 hours a day, seven days a week. The loop only became fast once LLMs took over some of the person’s steps.
Huang pointed to the most important condition: reward has to be cheap. When his team used AI to optimize test-time scaling algorithms, the obstacle was the long wait for each reward. They then precomputed the results and turned the setup into a static simulation. Once reward came back quickly, optimizing the algorithm “became a very simple thing.” If the problem is then recast as writing code or text (code as policy), a language model can take it over.
So what can it not do yet? Liu’s assessment is that when the environment is verifiable, a model will do well given some time. The hard cases are those where results are difficult or expensive to verify, such as science and engineering problems in the physical world. Huang’s example is writing: whether a revision is better still has to be decided by people in an A/B test, each round of feedback takes a long time, and so the loop is slow. In his view, RSI has little trouble with any problem that can be written as code and gives reward easily. The bottleneck is that many real-world settings cannot be modeled that way.
For Liu, the central problem of RSI is how to get good reward from the real world. A connecting layer is still missing between today’s AI systems and real problems in work and daily life. He considers this the most important piece and also the hardest, but “once the connection is made, it can develop well on its own.”
He added that real-world problems never come with a perfect reward anyway. People also rely on many proxy tasks to filter out wrong answers layer by layer. So there is no need to reach 100% in one step; getting a little better each time is enough.
3. The environment is the overlooked knob
An agent learns from the trajectories it produces by interacting with an environment. Yet people have only optimized the agent side, while the environment side is expensive and never changes. This is the starting point of EnvHarness, a paper Huang wrote during an internship at Google.
First, the cost. The 89 environments in Terminal-Bench 2.0 each took an expert 15.9 hours on average, about 1,400 hours in total. The 33 environments in APEX-Agents are heavier: about 485 hours each, about 16,000 hours in total, or $1.6 million at $100 an hour. Training needs thousands or tens of thousands of environments, and people cannot build that many by hand.
Second, the environments do not change. Once built, an environment has one fixed difficulty: a weak model never solves it, a strong model finds it too easy, and neither learns anything. Even if the difficulty is right at the start, the environment stops being useful after a few rounds of training make the model stronger.
EnvHarness does not build new environments. It wraps a harness around an existing one, the same idea as an agent harness wrapped around a frozen model. The task itself and the verifier that decides whether the task is done stay untouched. Only four things are adjusted: the initial state, the actions, the observations and the state transitions. It has three components:
- Stage changes the initial state. Take the task “put a clean mug on the table”: hide the mug in a drawer first and the task gets harder; put the mug in the agent’s hand in advance and it gets easier.
- Contract changes the rules. Disable the shortcut action that goes straight to a location, and the agent has to learn to navigate.
- Chain links several short tasks into one long task. Finishing one does not end the episode; the agent goes on to the next.
The harness is written by another agent, called EnvRigger. It reads the trajectories the policy produced, finds systematic weaknesses, writes a piece of harness code, and then has the policy run a new batch of trajectories in the modified environment to validate the change. If the task has become unsolvable, or no longer poses any challenge, the change is rejected.
The talk included an example from SWE-bench. The policy kept submitting patches without running the failing test, so EnvRigger wrote a fake pre-commit hook that blocks such submissions. From this the policy learned a skill: run the tests both before and after a change. In this experiment EnvRigger and the policy being trained used the same model, so the gain was not distilled from a stronger model.
4. The biggest technical bottleneck is attribution
The middle step of self-improvement is diagnosis: see a failure, find its cause, and propose a change that can be tested. Qian considers this the main bottleneck right now.
The difficulty is that the same surface symptom can have entirely different causes. A failure may come from reasoning, from a tool interface, from missing information or interference from the memory module, or even from a misreading of the training objective, and each calls for a different intervention. Until now this step has relied on human experience.
The weakness of today’s models is that they work bottom-up. They look at what went wrong case by case and then patch each case. A person works top-down: first judge where the overall approach is flawed, then verify. Attribution by patching one case at a time overfits easily.
Qian suggested two directions.
- Diagnosis needs evidence. Look at intermediate states, at the gap between the expected result and the actual feedback, and at actual resource use, and infer the cause from these. The system cannot be treated as a black box.
- Attribution needs controlled experiments. As in an ablation study in a paper, replace or remove one of the model, the tools, the memory or the execution flow, or replay the run, and see how the result changes. Today this work is done mostly by people.
His conclusion: what the system needs to be taught is not a trick for some benchmark but a more abstract way of thinking, namely how to test a hypothesis.
EnvRigger, described above, can be seen as a small-scale implementation of this idea. It looks for systematic weaknesses across many trajectories instead of targeting a single case.
5. Gains from self-iteration saturate quickly, and a higher score does not mean more capability
Qian described a pattern from their experiments. They had a model improve its own harness repeatedly and measured the effect after each round. On the held-out set, the score rose for the first few rounds and stopped rising by the fourth or fifth. So more iteration does not mean better results. But the model cannot tell when to stop. In Qian’s words, cost control “may simply not be on its mind”: when the results stop improving, it keeps adding rounds. Optimization that keeps polishing small details overfits more and more, and in the end the score can even drop.
People work differently. They try an idea on a small portion of the data first and commit large-scale resources only after it proves effective. That small portion should also be as diverse as possible, so that it exposes as many problems as possible. Qian thinks this staged allocation of resources is exactly what systems need to learn.
Huang added that RSI easily overfits to the task at hand. What it finds may be a shortcut, or even a way to hack the task outright. So one has to tell apart an agent that has really become stronger from one that has only become better at this kind of problem.
An audience member asked whether self-improvement that cannot loop forever still counts as RSI. Huang answered that a model has a finite number of parameters, so its ability must have a ceiling, and “hoping that RSI can improve without limit is unrealistic.” What matters is how fast a system approaches the ceiling, and whether it can sense how far it is from saturation so that it can stop early.
Qian added that the number of iterations should not be a threshold in the definition. A loop stops when the budget runs out or when marginal gains shrink. What matters is whether the overall direction is upward.
6. A train/test split is not enough to evaluate RSI
For a system that finds shortcuts on its own, setting aside a test set does not prevent data leakage or reward hacking. How should it be evaluated, then? Each of the four speakers offered an approach.
Use one-time evaluations tied to a point in time. This is Huang’s proposal. One kind is prediction: put the old and new versions of a model into a real market at the same time, or have them predict future events. Nobody knows the answers in advance, so nothing can leak. The other kind is human A/B testing. The price is high cost and slow results, which means an RSI algorithm cannot be measured online.
Or step back and use fully controlled synthetic environments. The compromise Liu described is to build your own environments and fill them with synthetic data the model has never seen. But he thinks evaluation ultimately has to land in real workflows such as finance, law and medicine, where people have already defined many metrics.
Environments should be realistic, controllable and dynamic. These are Qian’s three criteria. Controllable means, first of all, reproducible: the larger the environment, the less stable the results when different people test the same model. Because RSI runs for many rounds, controllable also means that the gain from each round can be measured separately. Dynamic means that the environment reveals constraints or user intent step by step, which tests how a model adapts when results differ from what it expected. Realistic and controllable pull against each other, however: real-world signals are messy and hard to turn into a structured environment, while an environment built entirely by hand or by an LLM looks fake.
Decide first which learning challenge you want to test. Yin argues that an RSI benchmark cannot be just a subset of an older benchmark. It needs a learning challenge of its own, and efficiency and cost should be part of the metrics. When his team built ApprenticeBench, preventing reward hacking still depended on a lot of people reading agent trajectories. He thinks that detecting this behavior at scale could be a benchmark in itself.
The host, Dawei Li, added an observation: these problems closely resemble the ones met earlier when evaluating the reasoning ability of large models, namely memorization, data leakage and overfitting to one domain. Lessons from that work, such as continuously updated evaluations and synthetic data, can be borrowed here.
7. Efficiency is an underrated metric
Yin stressed repeatedly that most RSI research today focuses on frontier capability, but in deployment it turns out that agents do not get more practiced the more they work.
He looks at RSI in the setting of enterprise deployment. A new hire reads the documentation, learns the software, practices hands-on and gets feedback from a manager, and eventually becomes proficient. For AI to enter a company’s workflow, it has to go through the same process. ApprenticeBench, recently released by his company NeoCognition, tests exactly this. The job it chose is construction finance accounting, an ordinary white-collar accounting role that does not need a particularly smart model.
People are also unpracticed at first. But after doing the same kind of task dozens or hundreds of times, they always get faster, moving from System 2, which requires thinking, to System 1, which does not. In their evaluation, models showed no such shift at all. As he put it: “It makes no sense that after an agent has done 100 tasks, the time and cost it spends are about the same as before, or even higher than at the start.”
He proposed three directions:
- Index and organize learned experience better, so that it can be retrieved quickly when a similar situation comes up.
- Turn the repeated procedures in the work into reusable skills or tools.
- Learn the job with a frontier model first, then compress the workflow knowledge it gained into a smaller, more specialized model.
This is not only an enterprise problem. Yin said that when training one generation of models after another, people likewise want to distill the reusable knowledge from earlier iterations so that later ones need fewer experiments. Questions of efficiency and cost like these happen to be a direction that academia can afford to work on and that is still at the frontier.
8. The human role moves up a level
Once AI takes over solving problems that are already well defined, the value of people lies in finding the problem and defining it clearly.
Liu’s view is that people have always been moving up one layer of abstraction at a time: from assembly to C++ and Python, and then to neural network frameworks. In the early days everyone still wrote gradient descent by hand. The lower layers still matter, but most people can put their effort where it is closer to applications and creates more value. This is not necessarily a bad change.
Huang has reviewed papers written by a fully automated research system. His impression is that the defining feature of a purely AI-generated paper is that it “uses experiments to prove something that is not necessarily an important question.” AI is very good at incremental improvement on well-defined problems, such as pushing a benchmark score a little higher.
Real research is more about discovering a problem and then defining it clearly: what the input is, what the output is, and what the metric is. The rest can be handed to AI. “We have gone from being a PhD student to being a PI, from someone who writes code to a project manager, from someone who solves problems to someone who finds them.”
Qian said that human insight should go into the meta level. The question used to be how to be a good PhD student. Now the question is how to be a good advisor: how to give downstream agents a more stable harness, environment and feedback mechanism, so that their self-iteration can keep running reliably.