AI Models Are Improving Themselves—and the Pace Is Accelerating
by Brian Pollack
There is a new model every other day, and there is a reason the pace keeps picking up. The models are now improving themselves. The industry term is RSI, or recursive self-improvement. Google has been the loudest about it, but every major lab is running some version of the same loop.
How Recursive Self-Improvement Works

Here is how it works at a high level.
Imagine a thousand-question university entrance exam. You take it and get immediate feedback. The feedback tells you that you got 80% of the questions wrong, and it tells you exactly which ones. Now you go study those. Before, you had to study for the entire test. Now you know precisely where you are weak. You take the same test again and you are down to 30% wrong. Repeat, and this time you only have 30% of the questions left to study.
AI companies are doing the same thing at scale. Using their own models and agents, they run a test of sorts with 100,000 questions. The model is then changed, adjusted, and fine-tuned. In fact, the agent might create 20 different versions of those changes. Each version takes the test. The best one is kept, and the rest are deleted. Then the agents repeat the loop. That slightly improved model generates 20 more changes, finds more training material in its weak areas, and gets a little better with every pass.
From User Feedback to AI Judges

At first, this was based on you pressing thumbs up and thumbs down to vote on which sessions were good and which were bad. Now those sessions are fed through judging agents that rate every one of them, like a teacher checking all of your homework every single time.
And once you have judges, you can push a model in any direction you want. The judges can vote for models that are more polite, less polite, more upbeat, more verbose, less sycophantic, and so on.
Why Businesses Should Plan for a Moving Target

For anyone buying or building on this, that is the part worth hearing. The model you evaluated last quarter is not the model you are running today, and the one you test next quarter will have been through thousands of rounds of this. Plan for a moving target rather than a fixed one. Now is not the time to lock yourself into a vendor, model, or system.
By the way, there is a lot of controversy around this topic because some labs, especially those in China, are “borrowing” other models as judges and teachers. This could be a hidden time bomb, and I’m writing about that next week.
Does all of this “survival of the best model” sound familiar? If not, I would suggest reading a little-known author named Charles Darwin.
Have questions about software development, AI, or turning an idea into a reliable product? Schedule a call with me. I’m always happy to share what I’ve learned and discuss what might work for your situation.
Originally published on Protovate.AI
Protovate builds practical AI-powered software for complex, real-world environments. Led by Brian Pollack and a global team with more than 30 years of experience, Protovate helps organizations innovate responsibly, improve efficiency, and turn emerging technology into solutions that deliver measurable impact.
Over the decades, the Protovate team has worked with organizations including NASA, Johnson & Johnson, Microsoft, Walmart, Covidien, Singtel, LG, Yahoo, and Lowe’s.
About the Author