Are AI-Evolved systems ready for primetime — Not yet!

Exposing the Achilles’ heel of AI-evolved systems

Diagram. An initial program P and a workload W feed AI evolution, which produces an evolved program P′. Both P and P′ go into AIChilles, which searches new inputs W′ and reports weakness types: system crashes, slowdowns, memory blowups, or P′ performing worse than the original P.

Fig. 1. AI evolution optimizes a program for a higher KPI on a fixed workload W. AIChilles [1] takes the initial program P and the evolved program P′, searches new inputs W′, and exposes where P′ crashes, slows down, blows up memory, or performs worse than P.

AI-driven system evolution promises to revolutionize how we innovate computer systems. Traditionally, optimizing system heuristics takes specialized expertise and huge engineering effort. AI evolution provides a new possibility: why not give AI agents a system program, an evaluator, and enough iterations, and let them automatically discover better implementations?

Recent work (e.g., AdaEvolve [2], Engram [3], OpenEvolve [4], CoCoEvolve (Snowflake) [5], AlphaEvolve (Google DeepMind) [6], SFR-AutoR&D (Salesforce) [7], ArchAgent (Google) [8]) claims impressive improvements across a wide range of systems tasks. We got excited and wanted to see: are these improvements robust and generalizable to the broader system? However, we soon found that the promise was not as good as it seemed.

AIChilles can automatically expose weaknesses in AI-evolved systems

When we started testing these systems, we noticed some weaknesses. The evolved programs often found a solution that worked well on the benchmark, but failed badly on new input workloads. We also observed that the program could crash at runtime, become much slower as the workload scaled, consume substantially more memory, or return a worse solution.

Motivated by this observation, we designed AIChilles (paper, code) [1], an agentic system for finding the “Achilles’ heel” of AI-evolved system programs. AIChilles takes as input an original program and its AI-evolved version. It infers the valid workload space, then searches for concrete workloads where the evolved program behaves much worse than the original program.

Fig. 2. AIChilles searches beyond the workloads used by the original evaluator. It finds valid inputs where the AI-evolved program behaves much worse relative to the baseline. Here, evolved programs have scalability issues on larger input workloads.

We reported our findings to teams working on AI-evolution frameworks (AdaEvolve [2] and Engram [3]), and they acknowledged the issues found by AIChilles. They agree that when an evaluator rewards only benchmark score, AI evolution may exploit that objective aggressively and reveal gaps in the evaluator (see blog).

Do specifications and multiagent workflows fix this?

More recently, systems such as SkySynth [9] try to address this concern with explicit specifications and structured multi-agent workflows. Instead of asking one agent to optimize the program, SkySynth first asks an AI agent to propose the specification. Different agents then generate implementations, tests, and checks against these requirements before accepting a solution. The authors claim that this structured workflow could avoid the reward-hacking problem in evolved programs.

This made us wonder: perhaps explicit specifications and multiple testing agents could eliminate the failures we had seen before? So we tested SkySynth with AIChilles on its LLM-router application (AIChilles currently supports Python-based programs). Unfortunately, we still found evidence that the evolved program violated its own specification!

We ran AIChilles on the SkySynth-generated program, and found that it violated its own generated specification under new workloads. The specification explicitly required every request to end in one of two states: answered or refused. But under a severe capacity shortage, the evolved router repeatedly retried some requests without ever answering or refusing them.

We also found the full multi-agent evolution loop was expensive. A single run on one input configuration could take more than eight hours (and often exceed the Claude Pro token session limit). This makes it hard to run such a multi-agent workflow as “just-in-time.”

The new “bitter” lesson for AI-evolved systems

When an AI tells us it has improved a system, we should not only ask, “How much better is the score?” We should also ask, “Is this score actually a faithful metric, and does it represent what we want to improve in the system?”

Before we blame AI agents for reward hacking, we should first check whether we gave them an incomplete signal to optimize in the first place.

If you want to run AIChilles on your own AI-evolved system, please let us know!

Or feel free to open an issue on the code repo and we will test them!

References

  1. AIChilles: Automatically Uncovering Hidden Weaknesses in AI-Evolved Systems. arXiv:2606.15834, 2026. Yajie Zhou, Ao Li, Ashwin Silla, Zaoxing Liu, Vyas Sekar. [arXiv] [code]
  2. AdaEvolve: Adaptive LLM Driven Zeroth-Order Optimization. arXiv:2602.20133, 2026. Mert Cemri, Shubham Agrawal, Akshat Gupta, et al. [arXiv]
  3. Improving Coherence and Persistence in Agentic AI for System Optimization (Engram). Proceedings of the ACM Conference on AI and Agentic Systems, 2026. Pantea Karimi, Kimia Noorbakhsh, Mohammad Alizadeh, Hari Balakrishnan. [doi]
  4. OpenEvolve. Open-source project, GitHub. [code]
  5. CoCoEvolve: Evolutionary Optimization for AI Systems. Snowflake Engineering Blog. [blog]
  6. AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms. Google DeepMind blog. [blog]
  7. SFR-AutoR&D. Salesforce AI Research. [project]
  8. ArchAgent v2: A Case Study with the Data Prefetching Championship. Google Research. [paper]
  9. Building Specialized Systems We Can Trust with Agents (SkySynth). SkyDiscover blog. [blog]