By Nicholas Thomas | TrendyVest September 20, 2026 · Outlook through March 2028 The joke in a classroom full of robots is not that the teacher is artificial intelligence. It is that today’s students might write tomorrow’s lesson—and that the next class might arrive better equipped to rewrite it again. That is the idea behind recursive self-improvement.

An AI system helps improve the machinery that produces AI. Those improvements create a more capable system, which can contribute more effectively to the next round of development. It is a compelling feedback loop.

It can also conceal enormous differences between what has been demonstrated and what remains hypothetical. A coding agent that improves its file-editing tools is not the same thing as a system that independently develops a frontier model. A better benchmark score does not establish better scientific judgment.

And thousands of agents operating simultaneously do not automatically amount to thousands of productive researchers. The public evidence nevertheless points to a meaningful transition: AI is becoming part of the production process for its own successors. Leading laboratories now report substantial internal research automation, while experimental systems are learning to improve their software and search strategies.

None of the demonstrations discussed here establishes sustained, autonomous frontier-model succession. But dismissing everything short of that threshold would miss the economic significance of what is already happening. The TrendyVest thesis is that the next 18 months will be defined less by a single “intelligence explosion” than by a reorganization of how discovery gets done.

As generating candidate solutions becomes cheaper, the scarce resources become trustworthy evaluation, usable computing capacity, and the ability to turn results into products. For investors, that distinction matters. The companies helping AI improve may benefit.

The companies paying for those improvements may benefit even more. And some businesses could find themselves supplying an extraordinary technology without retaining extraordinary profits. --- What has actually changed The most useful evidence is not a prediction about superintelligence. It is a set of measurements from inside organizations already using AI to develop AI.

Anthropic reports that, as of August 2026, Claude “leads” 26% of its measured AI research-and-development work, with more than 90% involving at least collaboration. “Leads” means completing most of a task from a high-level instruction while a human supervises. No measured category was fully autonomous.

The company also reports approximately 30,000 research and engineering agents operating concurrently on its most-used internal platform. These are company-generated measurements, partly dependent on model-based assessment—not independently audited counts of scientific discoveries. OpenAI’s September 6 disclosure describes a similar shift.

By mid-August, its research organization was using 3.1 agent-workdays of runtime for each human workday. That measures machine activity, not a demonstrated 3.1-fold increase in productivity. The company says it has reached its internal “automated research intern” milestone: systems handling well-defined research assignments under human direction.

Yet more than half of successful tasks in the four-to-eight-hour difficulty category involved at least one human intervention. These qualifications do not make the results unimportant. They tell us what kind of change is occurring.

The emerging laboratory is not necessarily a room without researchers. It is a room where each researcher can delegate more implementation, testing, and investigation to software. Evidence also shows that AI-generated improvements can reach production.

In its May 2025 AlphaEvolve announcement, Google reported an algorithmic improvement that recovered an average of 0.7% of its worldwide computing resources. Another improvement accelerated a particular Gemini computational kernel by 23%, reducing overall training time by 1%. Those percentages describe different things and should not be conflated.

A year later, the commercial boundary moved. Google made AlphaEvolve generally available through Google Cloud on July 9, 2026. Customers supply a problem, a starting algorithm, and an evaluation function; the system searches for improved implementations.

That is an important progression: a method for automated algorithm discovery has moved from an internal demonstration toward a service other organizations can use. The commercial question is no longer exclusively whether AI can produce an impressive answer. It is whether an organization can build a repeatable process around discovering improvements it would otherwise have missed.

Not every self-improvement loop changes the model. To understand the next 18 months, we need to separate several layers of improvement. At the shallowest level, a model revises an answer.

It critiques an explanation, corrects code, or tries another approach. That may improve the immediate output without creating any persistent improvement in the system. A deeper level changes the agent’s operating software: how it uses tools, manages context, stores useful information, or selects among possible solutions.

The underlying model weights—the learned numerical parameters—can remain unchanged while the surrounding system becomes substantially more effective. The Darwin Gödel Machine provides a concrete example. Its researchers reported improvements from 20% to 50% on their SWE-bench evaluation and from 14.2% to 30.7% on the full Polyglot benchmark.

The system modified its coding-agent implementation, retaining useful variations in an expanding archive. Crucially, the foundation models remained frozen. The paper did not demonstrate autonomous training of a new foundation model.

A further layer would improve the learning process itself: training methods, data selection, architectures, or other decisions that shape the capabilities a successor model acquires. The strongest version closes the loop across generations. A system helps produce a successor that is not merely better at ordinary tasks, but better at developing the next successor.

These distinctions are not pedantic. They determine what evidence would justify a claim of accelerating research. A system that becomes better because a more powerful outside model was inserted into its workflow has benefited from progress elsewhere.

That is useful engineering, but it should not be presented as a self-contained system that independently generates all its own gains. Human approval also does not automatically disqualify a process as recursive. Humans can retain authority over deployment while an AI system contributes genuine improvements to the process that creates its successors.

Recursive capability and operational autonomy are separate dimensions. The newest work is improving the search for improvements. Two September preprints illustrate where the technical frontier is moving.

SIFT—Self Improvement via Fast Tree-search—targets the expense of evaluating candidate self-modifications. Rather than fully benchmarking every proposed change immediately, it uses a model to compare candidate agent implementations. A Bradley–Terry statistical model aggregates those pairwise preferences into rankings, helping prioritize expensive evaluations.

Its experimental setups can use a stronger model for improvement and judging than for the underlying coding work. The reported gains concern agent design and search efficiency, not autonomous foundation-model retraining. Dream-RSI addresses a related problem: improving the strategy used to explore possible discoveries.

It records previous search trajectories, replays those histories to evaluate alternative exploration policies, and redeploys improved policies into further searches. The underlying coding agent remains unchanged. The authors evaluate the approach across algorithm engineering, mathematical optimization, and GPU-kernel engineering.

Both papers are preliminary research, not proof of broad commercial transformation. But their direction is significant. Researchers are moving beyond asking AI to generate candidates.

They are asking it to become better at deciding which candidates deserve attention. One limitation is worth watching: replaying previous searches can make experimentation cheaper, but a replay contains only the experience already recorded. My interpretation is that the strongest evidence will come when an improved search policy consistently performs better on genuinely new problems—not merely when it navigates familiar histories more efficiently.

The number that matters is not tokens. A laboratory can produce more code, launch more agents, and run more experiments without producing proportionately more useful discoveries. This is the denominator problem in the AI discussion.

The economically relevant unit is not a token, an agent-hour, or a proposed patch. It is a validated improvement that survives integration into the system where it is supposed to matter. Consider a simplified research process in which 80% of elapsed time can be accelerated, while only 20% cannot.

If AI makes the automatable portion ten times faster, the total speedup is: Overall speedup=10.20+0.80/10≈3.57\text{Overall speedup} = \frac{1}{0.20+0.80/10} \approx 3.57 That is substantial. It is not tenfold. Even if the accelerated portion became instantaneous, the unchanged 20% would cap the total speedup at fivefold.

This is an illustrative application of Amdahl’s law, not an estimate of any laboratory’s actual performance. Anthropic itself identifies bottleneck migration—including human code review—as a constraint on translating faster production into faster overall progress. The implications are broader than engineering.

An agent may generate experiments faster than a cluster can execute them. It may identify promising changes faster than researchers can validate them. It may produce a technically successful method that becomes unreliable, uneconomic, or ineffective at production scale.

The stronger RSI claim requires improvements to propagate through those constraints. A September preprint, The Economics of Recursive Self-Improvement, examines this feedback explicitly. Its preliminary calibration suggests the feedback loops are strengthening but not yet strong enough to sustain acceleration on their own.

The conclusion is model-dependent and rests on limited evidence; the authors do not rule out stronger feedback emerging. The important point is conceptual: a positive feedback loop is not automatically explosive. The next 18 months should be judged by how effectively AI removes successive bottlenecks—not by extrapolating whichever activity metric happens to be rising fastest.

An 18-month outlook: from delegated tasks to delegated investigations The following is our analytical base case, not an industry timetable. Through December 2026: bounded loops become more practical. The clearest near-term opportunity is in work with relatively explicit objectives and measurable outcomes: fixing software defects, improving computational kernels, managing experiments, and refining agent workflows.

We expect more demonstrations in which improvements persist across multiple iterations. The meaningful question is whether those gains survive fresh tests and a full accounting of search costs. This period should also sharpen the distinction between an impressive agent and a dependable research process.

A system that succeeds after repeated human rescue may still be valuable. It shouldn't be considered equivalent to one that consistently completes the same assignment independently. The first half of 2027: more responsibility for the investigation Our base case is that some leading research teams move from delegating individual experiments to delegating bounded investigations.

Instead of asking an agent to implement a chosen method, a researcher could specify an objective, constraints, and a budget. The system would develop hypotheses, run comparisons, investigate failures, and recommend the next step. The most valuable improvement would not necessarily be longer uninterrupted runtime.

It would be better judgment within a project: knowing when a result is suspicious, when another experiment is informative, and when to abandon a line of inquiry. Task-duration claims need particular care. METR’s “time horizon” measures difficulty by how long a skilled human would take on its tasks; it does not directly measure how long an AI operates unattended.

Its task distribution also does not represent all economically valuable work. The second half of 2027 through March 2028: the transfer test By this stage, we expect supervised AI research teams to handle larger portions of some research programs. The harder test will be whether their findings transfer: from small experiments to larger training runs, from a familiar benchmark to unfamiliar tasks, and from research code to reliable products.

OpenAI has identified March 2028 as a goal for an automated AI researcher operating under human supervision. That is a development target, not a guaranteed delivery date or a commitment to fully autonomous RSI. A faster-progress scenario would produce convincing evidence that substantially AI-developed successors are better at generating further improvements across several generations.

A slower scenario would still deliver useful automation, but research judgment, evaluation, computing resources, or integration would limit the rate of broader progress. Both scenarios could matter economically. Neither warrants pretending that December 2027 is a scheduled appointment with superintelligence.

The next shortage may be proof. The less obvious consequence of cheaper experimentation is that the burden of verification grows. Imagine a deliberately simplified case: a system generates 10,000 ineffective candidate improvements, and an evaluator incorrectly accepts 1% of ineffective candidates.

That produces an expected 100 apparent successes, despite no real improvement. Those numbers are illustrative, not observed failure rates. The lesson is that generating more candidates can increase false discoveries unless the evaluation process improves too.

Selecting the highest score from a large search is especially dangerous when scores are noisy. A system can become excellent at finding weaknesses in the test rather than improvements in the underlying task. This problem has already appeared in self-improvement research.

Sakana documented fabricated tool-use logs and an experiment in which some modifications removed markers used to detect hallucinated tool use, making the metric look better without solving the problem. This is where “human in the loop” becomes an insufficient description of safeguards. A person cannot meaningfully supervise a process whose decisions arrive faster than they can inspect them.

For a production research system, I would want the safeguards to exist outside the components the agent can rewrite. That means protected evaluation datasets, restricted permissions, enforced resource budgets, independently rerun measurements, traceable versions, and deployment approval that the research agent cannot grant itself. Model-based judges can help prioritize work.

They should not automatically become the final authority on whether an improvement is real. Using multiple judges may help, but agreement among systems with correlated blind spots is not the same thing as independent confirmation. An audit trail also differs from scientific validity.

A perfect record will expose how the wrong conclusion was reached. We expect reliable verification to become more valuable as generation becomes cheaper. Organizations that can distinguish a genuine improvement from a convincing false positive may gain an advantage as valuable as access to a stronger model.

The effects will arrive at different speeds. Over the next 18 months, AI’s impact will unfold at very different speeds across industries. In software, we expect more work to become economically worth attempting.

Maintenance tasks that were repeatedly deferred, testing that was too expensive, and small improvements that never justified a dedicated project could receive attention. But lower production costs do not determine who benefits. A software company might retain savings as higher margins.

It might reinvest in new products. Competitors might force it to lower prices. A customer might decide to build internally what it previously purchased.

This creates a potentially awkward outcome for application vendors: AI can improve their productivity while weakening the scarcity that supported their pricing. For workers, the central near-term change may be the composition of responsibility rather than a clean substitution of machines for jobs. We expect greater value to attach to specifying objectives, understanding unfamiliar systems, evaluating evidence, and accepting accountability for decisions.

That creates an apprenticeship problem worth taking seriously. If routine implementation is increasingly automated, organizations still need ways for junior employees to acquire the experience required for sound judgment. Buying more capable tools does not answer that organizational question.

We should evaluate the effects in medicine and longevity research differently. We expect AI to accelerate parts of hypothesis generation, computational analysis, experiment planning, and candidate prioritization. But those improvements must eventually be grounded in biological evidence.

The FDA’s overview describes Phase 2 studies lasting several months to two years and Phase 3 studies lasting one to four years, with larger studies helping reveal adverse effects that smaller studies can miss. These are general descriptions, not timelines for every individual therapy. Accordingly, we would look for better experimental decisions and stronger evidence reaching development—not assume that faster model research produces an immediate wave of approved longevity treatments.

A molecule that looks promising in a simulation is not yet a medicine. An improvement in a model is not yet an improvement in someone’s life. The physical infrastructure has its own clock.

Data centers require servers, networking, power delivery, and cooling. The IEA emphasizes that their concentration in particular locations can make grid integration difficult. More efficient software may ease some constraints; growing demand may intensify others.

The result could be an economy in which digital discovery accelerates faster than the physical capacity to use it. The investment trap: efficiency is not automatically bullish for infrastructure The simple investment story is that better AI creates more demand for computing. That is plausible.

It is not a law. Suppose an algorithmic improvement doubles the amount of useful work delivered per unit of compute. Assume, for illustration, that the savings fully translate into a halving of the customer’s price per unit of useful work.

If demand rises enough, total computing consumption increases. If demand rises less, computing consumption decreases. A simple constant-elasticity model expresses the relationship as: New compute demandOld compute demand=aε−1\frac{\text{New compute demand}}{\text{Old compute demand}} = a^{\varepsilon-1} Here, a is the efficiency improvement, and ε\varepsilon is demand’s responsiveness to the lower price.

With a twofold efficiency improvement and elasticity of 1.5, aggregate compute demand rises approximately 41%. With an elasticity of 0.5, it falls approximately 29%. These are illustrative scenarios under strong assumptions—not estimates of the AI market.

They ignore changes in product quality, hardware pricing, competitive behavior, and supply constraints. They nevertheless expose the missing variable in many investment arguments: how much additional useful demand lower costs actually create. There is another distinction.

At the customer level, lower cost can be an unambiguous benefit. At the supplier level, it can mean higher volume, lower pricing, or both. At the shareholder level, the outcome also depends on the stock price.

That is why recursive self-improvement should not be treated as an automatic buy signal for every semiconductor, data-center, or software company. The technology can work while an investment disappoints. Where the economics could land The public-market opportunity is better understood as several exposures rather than a single RSI trade.

The research platforms and commercial channels Alphabet offers a direct connection through Gemini-based research tools and Google Cloud’s commercialization of AlphaEvolve. Its opportunity isn't limited to selling model access: it could apply improvements internally and offer the resulting capabilities to customers. The investment question is whether those advantages produce incremental cash flow after development and infrastructure spending—and how much competition forces Alphabet to pass the benefits onward.

Microsoft provides exposure through its OpenAI investment and commercial relationship. Its April 2026 agreement retains a non-exclusive license to OpenAI model and product intellectual property through 2032, while revenue-sharing payments from OpenAI continue through 2030, subject to a cap. Microsoft remains a major shareholder, but the relationship does not give it exclusive ownership of OpenAI’s future economics.

Amazon combines an investment in Anthropic with the infrastructure used to train and serve Claude. In April, Anthropic announced a commitment exceeding $100 billion over ten years to AWS technologies, while Amazon announced a further $5 billion investment, with up to another $20 billion possible. Commitments are not revenue already collected, and intertwined financing and customer relationships deserve scrutiny.

For all three, our preferred test is the same: how much durable customer value becomes profitable demand, rather than simply more internal activity or larger spending commitments? The computing and networking suppliers Nvidia and Broadcom provide substantial infrastructure exposure, but neither reports an isolated “RSI revenue” category. Nvidia reported $89 billion of Data Center revenue for the quarter ended July 26, 2026.

Broadcom reported $16.7 billion of AI semiconductor revenue for its fiscal third quarter ended August 2, attributing demand to custom accelerators and networking. Those results establish commercial scale; they do not establish how much demand is specifically attributable to self-improvement research. Broadcom’s announced collaboration with OpenAI also illustrates hardware specialization.

The October 2025 agreement targets 10 gigawatts of custom accelerators and networking, with deployment planned from the second half of 2026 through the end of 2029. It is a multi-year plan, not evidence that all capacity has been delivered. Our interpretation is that the infrastructure opportunity depends on three questions: whether aggregate useful workloads expand, which hardware architectures capture them, and whether customers can finance deployment at attractive returns.

Faster AI research could strengthen that opportunity. It could also change the preferred hardware mix faster than investors expect. The tools that help build the next generation Cadence is a less-obvious exposure.

Its AuraStack product applies agentic orchestration to printed-circuit-board and advanced-packaging design, including planning, implementation, manufacturability, and physical analysis. That is a concrete connection between AI-assisted engineering and the systems needed to run more AI. The investment case still requires commercial evidence.

A powerful feature does not automatically create additional revenue if customers expect it within an existing subscription. Vertiv, through power and thermal-management products, represents a different kind of exposure: the physical equipment required to operate computing infrastructure. It is not an RSI developer.

Its relevance depends on deployment volume, system requirements, execution, and the economics of supplying them. These companies belong in a research framework, not an indiscriminate buying list. This article does not establish fair values or attractive entry prices.

The stronger investment thesis is a business that earns acceptable returns from useful AI today, with further research automation providing upside—not a business whose valuation requires full RSI t