Skip to main content
AI Agent Reflection Loops: Fix Output Before Returning

AI Agent Reflection Loops: Fix Output Before Returning

AI Integration
••6 min read•By Daily Miranda Pardo

Three out of ten outputs from your agent are wrong. Not obviously wrong — a response in the wrong format, an empty field that should have a value, a reference that doesn't exist in the source document.

The user sees the error. You see it too, in the logs. But the agent already returned the result.

The problem isn't the model. It's the architecture. Most agents work like this: request an output, return it. No review. No validation. No second pair of eyes.

A reflection loop changes that.

Why the First Attempt Isn't Enough

Imagine a data extraction pipeline over PDF contracts. The agent extracts dates, parties, amounts, and key clauses. In development, on five test contracts, accuracy is 97%.

In production, with 150 contracts in different formats, it starts failing. Not always. Not predictably. But it fails.

The most common failure patterns in extraction and generation agents:

  • Wrong format: JSON with unescaped quotes or trailing commas that break the parse
  • Incomplete fields: extracts nine of ten fields because the tenth was on a different page
  • Soft hallucinations: fills an empty field with a plausible but false value
  • Truncated reasoning: reaches a correct conclusion through incorrect steps

These errors share something: the model knows how to evaluate them if you ask it, but it doesn't do so on its own because nobody asked.

The Reflection Pattern: Actor + Critic

The architecture is straightforward. Two roles in a loop:

Actor: generates the output from the original input.

Critic: evaluates that output against an explicit set of criteria.

If the critic approves, the output goes to the user. If not, the critic's feedback goes back to the actor for revision.

Input → Actor → Output v1 → Critic
                                ↓
                          Meets criteria?
                         /               \
                        Yes               No
                        ↓                 ↓
                    Return         Feedback → Actor → Output v2 → ...

The critic isn't a human. It's the same model (or a smaller, cheaper one) evaluating the output with a specific prompt: "Review this output against these criteria. If something fails, tell me exactly what and how to fix it."

TypeScript Implementation

import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic();

interface ReflectionResult {
  output: string;
  iterations: number;
  approved: boolean;
}

async function withReflection(
  actorPrompt: string,
  input: string,
  criteria: string[],
  maxIterations = 3
): Promise<ReflectionResult> {
  let currentOutput = "";
  let iterations = 0;
  let approved = false;

  while (iterations < maxIterations && !approved) {
    const actorContext =
      iterations === 0
        ? `${actorPrompt}\n\nInput:\n${input}`
        : `${actorPrompt}\n\nInput:\n${input}\n\nPrevious attempt with feedback:\n${currentOutput}`;

    const actorResponse = await client.messages.create({
      model: "claude-sonnet-4-6",
      max_tokens: 2048,
      messages: [{ role: "user", content: actorContext }],
    });

    currentOutput =
      actorResponse.content[0].type === "text"
        ? actorResponse.content[0].text
        : "";

    const criticPrompt = `Evaluate the following output against these criteria:

${criteria.map((c, i) => `${i + 1}. ${c}`).join("\n")}

Output to evaluate:
${currentOutput}

Respond with JSON: { "approved": boolean, "feedback": string, "issues": string[] }
If "approved" is true, "feedback" and "issues" can be empty.`;

    const criticResponse = await client.messages.create({
      model: "claude-haiku-4-5-20251001",
      max_tokens: 512,
      messages: [{ role: "user", content: criticPrompt }],
    });

    const criticText =
      criticResponse.content[0].type === "text"
        ? criticResponse.content[0].text
        : '{"approved":false,"feedback":"Evaluation error","issues":[]}';

    try {
      const evaluation = JSON.parse(
        criticText.match(/\{[\s\S]*\}/)?.[0] ?? "{}"
      );
      approved = evaluation.approved === true;

      if (!approved) {
        currentOutput = `REVIEWER FEEDBACK: ${evaluation.feedback}\nIssues: ${evaluation.issues?.join(", ")}\n\nPrevious output:\n${currentOutput}`;
      }
    } catch {
      approved = true;
    }

    iterations++;
  }

  return { output: currentOutput, iterations, approved };
}

The key design decision: the critic uses Claude Haiku (cheaper and faster) while the actor uses Claude Sonnet. The critic doesn't need to generate creative content — it just needs to evaluate against criteria. Haiku does that perfectly at one-tenth the cost.

Writing Effective Critic Criteria

The quality of your reflection loop depends on the quality of the criteria. Vague criteria produce vague evaluations.

Weak criteria (don't work):

  • "The output should be correct"
  • "The response should be complete"

Strong criteria (work in production):

  • "The JSON must parse without errors. Verify no unescaped quotes, trailing commas, or invalid characters."
  • "All these fields must be present and non-null: start_date, end_date, total_amount, parties"
  • "Every date referenced must appear in the source text. If you cite a date not in the input, it's a hallucination."
  • "The total_amount field must be a number, not a string with a currency symbol."

In DAILYMP's AI integration service, we define these criteria with the client before writing any code. They're the most important artifact of the project.

When to Use Reflection (and When Not To)

Use it when:

  • Errors have real costs: a misextracted field creates manual work, a hallucination reaches a client
  • The domain has verifiable correctness criteria (structured data, specific formats)
  • Additional latency (1-3 seconds per iteration) is acceptable

Skip it when:

  • Latency is critical (chatbots requiring responses under 500ms)
  • Output is inherently subjective (creative summaries, marketing copy)
  • Iteration cost exceeds error cost

In practice, a single revision iteration resolves 80% of common failures. You rarely need more than two.

Production Pattern: Observability

Always add per-iteration logging to detect failure patterns:

console.log({
  event: "reflection_iteration",
  iteration: iterations,
  approved,
  model_actor: "claude-sonnet-4-6",
  model_critic: "claude-haiku-4-5-20251001",
  input_length: input.length,
  output_length: currentOutput.length,
});

If 40% of your contracts need two or more iterations, you have a problem in the actor's prompt, not in the loop. The loop is telling you something you couldn't see before.

We integrate this pattern in our AI agent automation service for document extraction pipelines, contract analysis, and report generation. The same agents that hit 97% accuracy in dev maintain that in production.

The Mistake 90% of Teams Make

Build the agent. Test in dev. Deploy. When it fails in prod, adjust the prompt and test again.

The problem: that cycle is reactive. Every failure reaches the user before you see it.

The reflection loop inverts this: the agent evaluates itself before returning the result. Failures stay inside the system.

You don't need a permanent QA team reviewing outputs. You need well-defined criteria and a critic that applies them.

Building an agent that makes errors the model itself could correct? In 30 minutes we review the architecture and see if a reflection loop solves the problem.

Talk about my agent →

Share article

Repetitive processes in your business?

Download the free AI Automation Map — the 5 most time-consuming processes and how to fix them.

No spam. Just the PDF. Unsubscribe anytime.

Written by Daily Miranda Pardo

I help businesses automate processes, build AI agents and connect intelligent systems.