Last month, OpenAI published an unusual statistic about how its researchers work. By mid-August, its research organization was using the equivalent of 3.1 agent-workdays for every human workday. Researchers were working with more agents concurrently, while also contributing more code and running more experiments. OpenAI also said it had reached what it calls the “automated research intern” milestone: a system capable of completing well-defined research tasks that might otherwise take a skilled researcher a few days. The name makes the system sound more autonomous than it is. Researchers still decide what problems are worth pursuing, which experiments matter, and what should actually make its way into a model. Even among successful four-to-eight-hour agent tasks, more than half still required at least one human intervention. But that qualification is what makes the development more interesting, not less. The researcher does not need to disappear for the research process to change. The agent only needs to take over enough of the work between deciding what to investigate and understanding what happened. OpenAI’s own data suggests that transition is already underway.

There is a natural tendency to frame this as another labor story: AI is becoming capable enough to automate increasingly sophisticated knowledge work, and research is simply the next profession to encounter it. I think that misses the more useful part. AI research has always been heavily automated. One of my first jobs out of college was as a machine learning engineer working on hyperparameter optimization, where the entire objective was to automate a part of the experimental process that would be slow and tedious to perform manually. Training pipelines automate data processing. Schedulers allocate compute. Evaluation suites run the same battery of tests against every checkpoint. Distributed systems turn one experiment into hundreds. The difference now is not that automation has entered machine learning research. It is the level of abstraction at which it can operate.
From automation to delegation
In its analysis, OpenAI uses a taxonomy developed by Epoch AI that breaks frontier AI R&D into a fairly intuitive sequence: decide what to work on, design an experiment, build the code and data required to run it, execute the experiment, analyze the result, and communicate what happened. Historically, software was excellent at automating the middle of that loop once a researcher had specified exactly what should happen. A training job could run without supervision. A sweep could test hundreds of configurations. A benchmark could score every resulting checkpoint. The researcher still had to translate the actual research question into those procedures, investigate failures, interpret the results, and decide what to do next.
Modern agents can increasingly operate on the less structured parts of that work. Instead of being told to execute a particular function with a particular set of arguments, the model can be asked to implement an evaluation, investigate why an experiment is failing, compare multiple trajectories, fix infrastructure blocking a run, monitor an experiment, or trace a surprising result back through the code that produced it. None of these capabilities independently amount to an automated scientist. They do, however, change how much work a single researcher can keep moving at once. A person who previously had to choose which experiment was worth spending an afternoon implementing can increasingly hand off several of them, check in later, and spend more time deciding which results are worth following.
That changes where the bottleneck sits. More experiments are not automatically better experiments, and OpenAI is careful about this in its own write-up. Faster code generation does not mean scientific progress increases at the same rate, and eventually some other constraint—compute, experimental design, interpretation, or simply the ability to identify good questions—becomes limiting. But a research organization does not need to automate those final constraints for the system to become materially more productive. It only needs to make the parts around them cheap enough that researchers can spend a greater percentage of their time on the decisions that still require judgment.
There are increasingly formal versions of the same pattern. Systems such as GEPA treat execution traces as input to an optimization loop. GEPA samples trajectories—including reasoning, tool calls, and tool outputs—uses natural-language reflection to diagnose what went wrong, and proposes and tests prompt updates based on those failures. The important idea is broader than any particular optimization algorithm. An experiment no longer needs to produce only a score. It can produce enough context for another model to help determine what should change before the next experiment runs.
That matters because agents leave unusually rich evidence behind them. A normal benchmark might tell you that a model got 63 percent of a task set correct. An agent trajectory can tell you what the model believed, which tool it selected, what the tool returned, where context disappeared, whether it noticed that an assumption had failed, how it recovered, and whether the final failure came from the model, the harness, the environment, or the instructions surrounding it. Once models are capable enough to reason over that data, trace analysis itself becomes another task that can be delegated.
Training at agent scale
Xiaomi’s MiMo-V2.6 training run shows the same idea much closer to the training loop itself. In September, the team publicly streamed six days of reinforcement-learning training before releasing and open-sourcing the MiMo-V2.6 series on September 22. By the end of the run, MiMo-V2.6-Pro and Flash had each completed 30 RL steps, with the number of tokens per training step reaching 3.5–3.7 billion. Xiaomi used a fully asynchronous architecture and mixed coding, general-agent, visual, and cybersecurity tasks across multiple harnesses, while also scaling the compute spent grading and assigning credit to the resulting trajectories.

The raw scale is impressive, but I find the machinery around it more interesting. An agent enters an environment and attempts a task. It reasons, calls tools, receives feedback, changes its approach, and eventually produces some outcome. The system evaluates that trajectory, determines how successful it was, and uses that result as training signal before repeating the process again across thousands of other attempts. Once the environments reset reliably, the tasks produce meaningful signal, and the graders can tell genuine success from superficial completion, the thing being scaled is not simply compute. It is experience.
That creates a very different engineering problem from training a static model against a fixed corpus. The environment now matters. The harness matters. The tools exposed to the model matter. The amount and structure of context matter. The grader matters. Observability matters because a failed rollout that only produces a zero is much less useful than one where the underlying trajectory makes the failure legible. The research system begins to look less like a single training job and more like a large distributed software system built around repeatedly creating situations in which the model can act, measuring what happens, and feeding those results back into the next version.
This is the part of the MiMo work I find particularly valuable to see publicly. Most of the time, the output of frontier AI research arrives at the end. A model gets released, a paper explains what changed, and a benchmark table shows how much better the result was. The training infrastructure that made the improvement possible stays mostly invisible. Watching the run exposes how much of modern post-training is really about building the machinery around the model: environments, graders, evaluation systems, asynchronous execution, and enough observability to understand whether a capability is improving for the reason you think it is.
The loop is conceptually simple. Run the model. Preserve what happened. Evaluate the outcome. Understand why it succeeded or failed. Change something. Run it again. What is changing is how much of that loop can now execute without a person manually carrying information from one stage into the next.
What this looks like in our own R&D
We see a much smaller version of the same pattern in our work at Pensar. We build autonomous offensive-security agents, so a large part of our R&D process already revolves around evaluating long-running model behavior rather than individual responses. Some of the traces come from production engagements that our team reviews for quality or from customer issues we want to understand more deeply. Others come from our internal evaluations, which range from conventional web application and cloud exploitation environments to more complicated operational-technology and cyber-physical systems. The environments are different, but the research problem is usually the same: an agent ran for some period of time, produced an enormous amount of state and behavior, and we need to understand what it tells us about the model and the system around it.
A few years ago, reviewing that data manually was reasonable. At my previous company, we would run red-team engagements with modified open-source models against large enterprise environments, and when a model did something strange I could usually open the session data and read through the trace myself. The interactions were short enough that the researcher could remain the analysis layer. That stops scaling once the model is capable of working autonomously for hours. Some of our pentests now continue testing for hours or days depending on the size of the client’s infrastructure, and our longest-horizon evaluations produce enough behavior that reading every action manually would eliminate much of the leverage we gained by running the agent autonomously in the first place.
We increasingly use research agents to help with that analysis. They inspect the session data, identify points in the trajectory that deserve attention, and help narrow what might be hours of autonomous behavior down to a much smaller number of interesting failures. Sometimes the issue is straightforward tool selection. We have experimented with smaller and more specialized models that understand the general objective but choose the wrong tool or fail to use information that was discovered earlier in the engagement. Other failures only become obvious over a longer horizon. An agent might successfully reach part of an OT environment but fail to recognize that it should pivot deeper into the operational network. Another model might appear technically capable of completing the same task but refuse because of its guardrails, even when the rules of engagement and authorization are explicitly provided.

Those differences are why a benchmark score is not enough for the kind of work we are doing. We test a mixture of open- and closed-source models, and the same harness can expose very different strengths and weaknesses between them. One model maintains context well but struggles with tool selection. Another is more willing to act but fails to explore deeply enough. A model that performs extremely well in a standard software environment can behave very differently once the task involves a cyber-physical system, multiple trust boundaries, or a long period of autonomous operation. The useful question is rarely just which model completed more challenges. It is why one did and another did not, and whether the failure belongs to the model, the context strategy, the prompt, the available tools, or the broader agent architecture.
Research agents help compress that search space, but they do not remove the need for judgment. Once a trace has been analyzed, someone still has to decide what the finding means for the product. We might discover a recurring failure mode, but fixing it is only one possible use of engineering time. We may instead care more about reducing inference cost, improving latency, extracting more capability from a model that is already working well, or building something needed by customers we are supporting. Those decisions contain product and business context that does not exist in the trace. Leadership decides what the company is trying to accomplish; the research system helps us understand what is preventing the agent from getting there.
The same boundary appears in OpenAI’s description of its own research agents. Humans still decide which questions matter. The automated part keeps expanding underneath that decision layer. I think that is the more realistic near-term model for research automation: not a system that independently determines where an organization should go, but one that dramatically increases how much evidence, experimentation, and implementation can happen between the moments where a person needs to choose the direction.
Evals are core infrastructure
Most evaluation systems are still described as measurement. You change a model or a harness, run a set of tasks, compare the new result against the previous one, and decide whether you made progress. That remains important, but it undersells what an eval becomes once you preserve the trajectory and have systems capable of reasoning about it. A failed evaluation can become a regression test. A strong trajectory can become an example of desired behavior. A repeated failure can tell you that the issue is not the model weights at all but the context, tool interface, system prompt, or workflow surrounding them. Eventually the same data can be curated into SFT examples and other post-training datasets, or used to design tasks and environments for reinforcement learning.
That is the direction we are building toward internally as well. Today, most of our eval infrastructure exists to compare models and measure changes to our harness. Over time, we want more of those failures and successes to feed directly into the next stage of improvement. If a model repeatedly fails in a particular class of environment, the useful output is not simply a lower score on a dashboard. It is a set of trajectories that explain the weakness well enough to change the system, construct better training examples, or design a more targeted evaluation that tells us whether we actually fixed it.
Thinking about evals this way also changes how I think about safety and capabilities work. Both benefit from the same underlying infrastructure. A capabilities researcher can use automated systems to run more experiments, inspect more trajectories, and find more ways to improve performance. A safety researcher can use the same systems to test more behaviors, look for failure modes, and validate whether a mitigation actually holds over long-horizon interaction rather than a handful of carefully selected prompts. The amount of useful work a research group can perform becomes less tightly connected to the number of people available to manually execute every experiment.
This does not mean that every experiment should run autonomously or that more experimentation is automatically good. It means a small group of experts can direct substantially more testing when the expensive parts of execution and first-pass analysis are handled by software. The expert is still there. What changes is the amount of evidence they can generate and inspect before making the next decision.
The same idea applies well outside frontier model labs. Software engineers are already using agents to write features, fix bugs, perform migrations, review code, and operate infrastructure. Most companies are adopting the execution layer faster than the evaluation layer. They are getting very good at producing more agent-generated work without building much infrastructure for understanding where those agents repeatedly fail, whether a task was actually completed correctly, or how those failures should change the way the agent operates next time.
That is probably the most useful lesson to take from what frontier AI labs are doing. The model is only one component of the system. The environment around it can be engineered too. You can measure the workflow, preserve traces, evaluate outcomes, find recurring failure modes, improve the tools and context available to the agent, and turn those failures into tests the next version has to pass. If agents are becoming part of how your engineers build software, then the agent workflow itself becomes something worth testing with the same seriousness as the software it produces.
There is a more dramatic way to tell this story, where automated research immediately becomes recursive self-improvement and the models begin autonomously building smarter versions of themselves. I do not think that framing is necessary yet. OpenAI’s researchers still choose the research agenda. Xiaomi’s researchers built the environments and graders around its reinforcement-learning run. We still decide what matters when one of our research agents finds an interesting failure. There are plenty of constraints left, and improving one part of the research loop simply makes the next bottleneck easier to see.
What is already happening is more concrete. The systems being researched are beginning to participate in the process used to improve them. They can implement experiments, operate inside training environments, generate trajectories, inspect failures, and help researchers decide where to look next. Every one of those steps used to consume human attention. Each one that becomes cheap increases the number of iterations a research organization can perform before a person needs to intervene.
That is enough to matter. The next generation of models will not be built only by finding better architectures or spending more compute. They will also be built by organizations that get better at constructing the feedback systems around them: environments that expose meaningful behavior, evals that tell them when something has changed, agents that can analyze the resulting evidence, and researchers who know which results are worth turning into the next experiment.
The models are getting better. Increasingly, so is the machinery we use to make them better.

