跳到内容
← All posts

The Challenge: Vibe-Coding a Self-Evolving Agent

Vibe-coding a ReAct agent is easy; getting it to rewrite its own prompts, tools, and memory from failure feedback is not. This post documents one full attempt — and the methodology that survived.

Neonity17 min read
agentsvibe-codingllmengineering

Summary: Vibe coding (a term Karpathy coined in 2025) lives in the comfort zone of “describe the requirement → let the model write code → run it → describe what’s broken.” That works great for CRUD and utility functions. But when the goal is “an agent that improves itself,” the comfort zone starts to break down. This post documents one full challenge: vibe-coding an agent that rewrites its own prompt, extends its own tools, and evolves its own memory from failure feedback — what felt great, where it derailed, and the methodology that finally converged.

What vibe coding is, and where its sweet spot lies

Vibe coding means: you describe the requirement, skim the diff, run it — the AI writes almost all of the code. Its precondition is fast verifiability: the build passes, tests pass, you can see the UI change with your own eyes. As long as verification is fast, the risk stays manageable even if you don’t fully understand every line.

From the previous post, a ReAct agent (Thought → Action → Observation) is essentially a “fixed skeleton + prompt + tool list.” The skeleton is written once; everything after that is tuning prompts and adding tools — exactly vibe coding’s sweet spot: one-way, stateless, verifiable at every step. So “vibe-coding an agent” isn’t really a challenge; you can have a working one in half an hour.

The real challenge is the next word: self-evolving.

Self-evolution: from “doing things” to “rewriting itself”

An ordinary agent’s errors stop inside the ReAct loop — one step failed, try a different Action. A self-evolving agent’s errors propagate to “modifying its future self”: a failure is not just a signal for the current task, but training signal for rewriting its own prompt, tools, memory — and ultimately its code.

The difference looks like “just one more loop,” but it changes the nature of the system:

  • An ordinary agent’s errors are local: one Observation was wrong, correct it.
  • A self-evolving agent’s errors are global: one prompt change affects every future task. Change it right and it’s evolution; change it wrong and it’s systematic drift.

In other words, self-evolution upgrades “verification” from single-task to cross-task differential comparison. That is the core difficulty of the whole challenge.

Four levels of evolution

I split self-evolution into four levels. The deeper you go, the higher the payoff and the higher the danger:

Level What changes Cost Risk
L1 Prompt self-optimization Feed failure samples back, rewrite the system prompt Very low Low
L2 Tool self-extension When a tool is missing, write and register it Medium Medium
L3 Memory evolution Cluster memory by failure mode, evict stale entries Medium Medium
L4 Code self-modification Edit its own source, rebuild, regression-test, merge High High

L1 is the most common and most effective starting point: feed “failure samples + expected output” back to the model, let it summarize the cause and rewrite the prompt. L4 is the most dangerous — it moves the entire practice of software engineering inside the loop: edit code, run the build, run regressions, decide merge or rollback. Any failure at any step can take the agent down outright.

The challenge, as it actually happened

To dodge the classic “you can’t evaluate it” trap, I picked a task that could be scored automatically: a “data-analysis agent” — input a CSV file plus a question, output a conclusion and plotting code. A fixed eval set of 20 questions, each scored automatically against expected output quality.

Four steps:

  1. Skeleton phase (the peak of vibe coding): I described the requirement in natural language and let the model generate the agent framework plus the eval script. Two hours from zero to working. I didn’t read most of the code — it ran, that was enough.
  2. First eval: 11 of 20 questions correct.
  3. First evolution round: fed the 9 failures back to the model, asked it to revise the system prompt. 11 → 15 in one round. Dramatic.
  4. Second round derails: the model started “answering the eval set” — the prompt began containing memorized rules like “if the question mentions sales, answer X.” Training-set score kept climbing, but it was obviously overfitting. I had to add a held-out test set and a prompt-length budget to pull the behavior back on track.

Step 4 is the most instructive part of the whole challenge: the most dangerous moment of self-evolution is when it looks like it’s working — the score is rising, but for the wrong reason.

A minimal self-evolution loop

Here is the minimal viable loop (pseudocode) that I converged on:

def evolve_once(agent, failures, eval_set):
    # 1. Summarize, don't paste raw errors.
    #    Feeding raw errors makes the model "memorize answers";
    #    summarizing failure modes is what changes behavior.
    feedback = summarize(failures)

    # 2. Generate several candidate improvements, prompt layer only.
    candidates = agent.suggest_changes(feedback, budget=3)

    # 3. Differential eval: every candidate vs. the current version
    #    on the SAME eval_set.
    for c in candidates:
        score_new = eval(agent.with_prompt(c), eval_set)
        if score_new > agent.score:      # accept only improvements
            agent.apply(c)
            agent.score = score_new

    return agent

Three constraints that matter:

  • Eval first: without evaluation, self-evolution is random drift — you can’t even tell whether it’s evolving or degrading.
  • Differential acceptance: every candidate must be compared against the current version on the same eval set, and only strictly better ones are accepted.
  • Summarize first: have the model condense failures into “patterns” before touching the prompt. Skip that step and the model will “memorize answers” instead.

The traps I fell into

Trap Symptom Fix
Eval leakage 100% on training set, 40% on new questions Fixed training set + held-out test set
Feedback noise Wrongly attributing unrelated failures to the prompt Summarize failure modes before editing
Runaway loop Self-evolution never stops, token budget burns Round limit + per-round token quota
Overfitting the eval Prompt grows “if X, answer Y” rules Require reasoning traces in output; grade the process
Uncontrolled code edits One bad change kills the whole agent Sandbox + git snapshot + auto-rollback on failure

What the methodology distilled into

Four takeaways from this challenge:

  1. Vibe coding suits the skeleton and the tools — not the feedback loop itself. You have to draw the loop yourself; the model is a worker inside it, not the loop’s designer.
  2. The first step of self-evolution is not writing code — it’s writing the eval. Get the eval wrong and everything after is self-deception.
  3. Every level of evolution needs differential comparison and rollback. Without differentials, evolution is drift; without rollback, one bad edit is the end.
  4. Budget is the guardrail. Tokens, rounds, prompt length, tool count — cap them all. The biggest cost risk in self-evolution is not a single call, it’s an uncontrolled loop.

Wrap-up

Vibe coding and self-evolution are a natural contradiction: vibe coding’s comfort zone is “don’t read the code, trust that output is verifiable,” while self-evolution demands “precise eval, strict differentials, controlled boundaries.” Reconciling the two — building the skeleton at vibe-coding pace while keeping the evolution loop under engineering discipline — means you’ve thought hard enough about verification.

That is worth more than the agent itself: agents go stale; verification methodology doesn’t.

Next time: since evaluation is where self-evolution most often dies, let’s talk about how to design a reliable eval set for agents — and why most teams die of eval leakage long before they die of an insufficiently strong model.