1. Introduction
In Part 1, I covered ReAct — a pattern where an agent alternates between reasoning and action. In Part 2, I covered Plan-and-Execute, where a separate Planner builds an N-step plan. Both patterns share one flaw: the agent doesn’t learn from its mistakes. Every run starts from scratch. The Reflexion pattern fixes this: the agent makes an attempt, evaluates the result, reflects — and uses the accumulated experience on the next try.
2. What is Reflexion
History
In December 2022, Anthropic published Constitutional AI — an approach where a language model critiques its own responses and rewrites them to reduce harmfulness. This was the first large-scale example of the self-critique pattern: generate → self-critique → revise. But Constitutional AI used critique for training (finetuning via RLAIF), not for inference.
In March 2023, Madaan et al. published Self-Refine — iterative improvement through self-feedback. The same LLM plays three roles: generator, critic, and refiner. Result: ~20% average improvement across 7 tasks (Madaan et al., 2023). But there’s a catch — on reasoning tasks (Math Reasoning), the improvement is 0%: the model cannot reliably determine whether its reasoning is correct or not.
And here’s where it gets interesting. In the same month of March 2023, Shinn et al. published Reflexion — solving the core problem of Self-Refine by adding episodic memory and external evaluation. Instead of a single critique-refine cycle — a multi-trial process where the agent accumulates verbal “lessons” and uses them on subsequent attempts. Result: 91% pass@1 on HumanEval (vs. 80% for GPT-4) and +22% absolute on AlfWorld (Shinn et al., 2023).
Analogy with a developer
Imagine: a junior writes code, runs tests — they fail. What do they do? They don’t rewrite from scratch — they read the error, understand the cause, and note “next time I’ll check the nil edge case.” That’s reflection: not just fixing, but extracting a lesson for the future. ReAct is a junior without memory: making the same mistakes every time. Plan-and-Execute is a junior with a plan but without retrospection. Reflexion is a junior who keeps an error diary.
Formal definition
Reflexion operates in three phases, repeated cyclically:
- Act: The Actor (LLM) generates actions and receives observations from the environment.
- Evaluate: The Evaluator assesses the result — with a scalar score or free-form text.
- Reflect: The Self-Reflection model generates verbal feedback — what went wrong and how to fix it. The reflection is stored in episodic memory.
On the next attempt, the Actor receives the contents of episodic memory in context — and can avoid repeating past mistakes.
3. What problem it solves
ReAct’s problem: no learning from mistakes
ReAct, as I showed in Part 1, alternates Thought-Action-Observation in a single loop. If the task isn’t solved — the agent simply starts over. Past experience? Lost. Every run is the first and last.
Plan-and-Execute’s problem: no retrospection
Plan-and-Execute builds a plan and executes it. The Reviser adjusts the plan when needed — but within a single run. Between runs — clean slate. No learning from past mistakes.
Self-Refine’s problem: no memory between attempts
Self-Refine does critique-refine in a single LLM call. There’s improvement on generation tasks (style, format) — but on reasoning tasks 0%, because the model cannot reliably assess the correctness of its own reasoning without an external arbiter (Madaan et al., 2023, Table 1: Math Reasoning).
What Reflexion provides
| Problem | Reflexion’s solution |
|---|---|
| No learning from mistakes | Episodic memory stores verbal lessons between attempts |
| Unreliable self-assessment | Evaluator — external arbiter (tests, compiler, environment) |
| No retrospection | Each attempt is enriched with reflections from past ones |
| Myopic fixes | Reflection focuses on the cause of the error, not the symptom |
4. Architecture
Actor → Evaluator → Reflector → Memory → retry
sequenceDiagram
participant U as User
participant A as Actor
participant E as Evaluator
participant R as Reflector
participant M as Episodic Memory
U->>A: Task + reflections from memory
A->>A: Generate actions (ReAct loop)
A->>E: Trajectory (actions + observations)
alt Result is correct
E-->>U: ✅ Success
else Result has errors
E->>R: Trajectory + evaluation
R->>R: Generate verbal reflection
"I failed because..."
R->>M: Store reflection
M-->>A: Enriched context for next attempt
Note over A,M: Retry with accumulated experience
end
Components
Actor — an LLM generating actions. Can be a ReAct agent or a simple ChatModel. The key difference from regular ReAct: the Actor receives episodic memory contents in the system prompt, allowing it to account for past mistakes.
Evaluator — assesses the Actor’s result. Can be:
- Deterministic: unit tests, compiler, game environment — gives an objective score
- LLM-based: another language model assesses quality — less reliable but applicable for open-ended tasks
Why does this matter? It’s the external verifier that solves Self-Refine’s problem, where the model cannot assess its own correctness.
Self-Reflection model — an LLM generating verbal reflection. Receives: the Actor’s trajectory (actions + observations), the Evaluator’s assessment, past reflections from memory. Generates text like: “I made a mistake handling the empty list edge case. Next time I need to check len() > 0 before accessing an element.”
Episodic Memory — a store of reflections. Simple structure: a list of text strings injected into the Actor’s context on the next attempt. The more attempts — the richer the memory.
5. Evolution of self-critique
Three self-critique patterns emerged within 3 months — each solving the previous one’s problem:
flowchart TD
CAI["🛡️ Constitutional AI
Dec 2022 • Bai/Anthropic
Self-critique → revise
for training (finetuning)
Goal: harmlessness"]
SRF["🔄 Self-Refine
Mar 2023 • Madaan/CMU
Same LLM: generate → critique → refine
for inference (1 call)
~20% improvement on generation"]
REF["🧠 Reflexion
Mar 2023 • Shinn/Princeton+NEU
Actor → Evaluator → Reflector
for inference (N attempts)
+ Episodic Memory"]
HUANG["⚠️ Huang et al.
Oct 2023 • ICLR 2024
LLMs Cannot Self-Correct
Reasoning Yet
Without external feedback — doesn't work"]
RRR["🚀 Reflect, Retry, Reward
May 2025 • Bensal/Writer
RL-trained reflections
1.5B-7B beats 10x models"]
CAI -->|"added
inference-time
critique"| SRF
SRF -->|"added
episodic memory
+ external eval"| REF
REF -->|"showed
limitations
without verifier"| HUANG
HUANG -->|"RL-training
better reflections"| RRR
style CAI fill:#2e7d32,color:#fff
style SRF fill:#1565c0,color:#fff
style REF fill:#e65100,color:#fff
style HUANG fill:#c62828,color:#fff
style RRR fill:#6a1b9a,color:#fff
Comparison of three patterns
| Aspect | Constitutional AI | Self-Refine | Reflexion |
|---|---|---|---|
| Date | Dec 2022 | Mar 2023 | Mar 2023 |
| Goal | Harmlessness (safety) | Output quality | Output quality + learning |
| Critique | Self-critique | Self-feedback | External evaluator + self-reflection |
| Memory | No | No | Episodic memory |
| Attempts | 1 | 1 (iterations inside) | N (multi-trial) |
| Application | Finetuning (offline) | Inference (online) | Inference (online) |
| Result | Improved harmlessness | +20% on generation, 0% on reasoning | +11% HumanEval (91% vs 80% GPT-4), +22% AlfWorld |
| Evaluation type | RLAIF (RL from AI judge) | Self-judge | External verifier |
The pattern is clear: each subsequent pattern adds what the previous one lacked. Constitutional AI had no memory and no multi-trial capability. Self-Refine added inference-time critique, but without memory and without an external verifier. Reflexion closed the loop: the external verifier solves the unreliable self-assessment problem, and episodic memory enables learning between attempts.
6. When it works / when it doesn’t
Why a dedicated section on limitations? Because in October 2023, Huang et al. published “Large Language Models Cannot Self-Correct Reasoning Yet” — and proved that self-correction without external feedback degrades results.
flowchart TD
START[Agent task] --> Q{External
verifier available?}
Q -->|Yes| Q2{Low initial
accuracy?}
Q -->|No| FAIL["❌ Reflexion won't help
Self-correction degrades results
(Huang et al., 2023)"]
Q2 -->|Yes| WORKS["✅ Reflexion works
+11-22% improvement
(HumanEval, AlfWorld)"]
Q2 -->|No| WARN["⚠️ May not be worth it
Risk of degradation on simple tasks
Diminishing returns"]
WORKS --> REC1["Recommendation:
max 3-5 attempts,
evaluate cost/benefit"]
WARN --> REC2["Recommendation:
no more than 2 attempts,
monitor quality"]
style FAIL fill:#c62828,color:#fff
style WORKS fill:#2e7d32,color:#fff
style WARN fill:#e65100,color:#fff
Numbers from Huang et al.
Without an external verifier (intrinsic self-correction), quality drops across all models:
| Model | GSM8K (before → after) | CommonSenseQA (before → after) |
|---|---|---|
| GPT-3.5 | 75.9 → 74.7 | 75.8 → 41.8 |
| GPT-4 | 95.5 → 89.0 | 82.0 → 80.0 |
| Llama-2-70b | 62.0 → 36.5 | 64.0 → 36.5 |
The reason: LLMs are more likely to change a correct answer to an incorrect one than vice versa. The fundamental problem is that the model cannot reliably assess the correctness of its own reasoning (Huang et al., 2023, Figure 1).
When Reflexion works
| Condition | Why |
|---|---|
| External verifier (tests, compiler, env) | Objective assessment → accurate reflection |
| Low initial accuracy | Room for improvement |
| Multi-trial scenario | Memory accumulates lessons |
| Tasks with objective success criterion | Clear error signal |
When Reflexion does NOT work
| Condition | Why |
|---|---|
| No external verifier | Model can’t assess its own correctness |
| High initial accuracy | Risk of degradation (correct → incorrect) |
| Open-ended tasks without criteria | Nothing to evaluate → inaccurate reflection |
| Simple tasks | Cost of reflection isn’t justified |
7. Implementation variants in Eino
Eino doesn’t provide a ready-made reflection.NewAgent() — unlike react.NewAgent() from Part 1 or planexecute.NewAgent() from Part 2. But that’s actually a good thing: Reflexion is not a separate agent type, but a composition pattern that can be implemented in several ways.
Variant A: compose.Graph with a cycle
Eino Graph supports cycles via AddEdge from a node to itself + WithMaxRunSteps to limit iterations. This is the most natural implementation of Reflexion:
flowchart LR
START(("START")) --> Actor["🎬 Actor
(react.Agent)"]
Actor --> Evaluator["🔍 Evaluator
(Lambda: run tests)"]
Evaluator --> Branch{"Pass?"}
Branch -->|Yes| END(("END"))
Branch -->|No| Reflector["🪞 Reflector
(ChatModel)"]
Reflector --> Memory["💾 Memory
(State: []string)"]
Memory --> Actor
style START fill:#2e7d32,color:#fff
style END fill:#2e7d32,color:#fff
style Branch fill:#e65100,color:#fff
style Actor fill:#1565c0,color:#fff
style Evaluator fill:#1565c0,color:#fff
style Reflector fill:#6a1b9a,color:#fff
style Memory fill:#ad1457,color:#fff
Pros: explicit retry cycle, branch support, checkpoint via WithCheckPointStore, can embed a ReAct agent via ExportGraph().
Cons: Graph API uses implicit data passing (entire output → entire input), need to manage state carefully.
Variant B: compose.Workflow (linear, no cycle)
Workflow is a declarative graph with explicit field mapping. Problem: Workflow does not support cycles — always AllPredecessor. For Reflexion, this is fatal: no retry-loop.
But if the cycle is implemented externally (in Go code), and Workflow is used for a single Actor → Evaluator → Reflector iteration — it works:
// External retry-loop
for attempt := 0; attempt < maxAttempts; attempt++ {
result, err := workflow.Invoke(ctx, input)
if result.Passed { break }
memory = append(memory, result.Reflection)
// Inject memory into the next call
input.Reflections = memory
}
Pros: explicit field mapping, type safety, easier to test a single iteration.
Cons: no built-in cycle — have to implement manually, no checkpoint between iterations.
Variant C: deer-go pattern (State Graph)
deer-go is a Go implementation of ByteDance’s DEER-flow on Eino Graph. It uses Goto-based routing: each node writes the next target node to state.Goto, and an agentHandOff function directs execution.
For Reflexion, you could add a Critic node and a Critic → Actor edge (retry). This extends the existing architecture, but deer-go doesn’t contain ready-made reflection patterns — only a re-planning loop.
Pros: ready state graph architecture, checkpoint, human-in-the-loop via InterruptAndRerun.
Cons: more complex, requires HTTP server and MCP tools, overkill for simple scenarios.
Variant comparison
| Aspect | Graph + cycle | Workflow + external loop | deer-go |
|---|---|---|---|
| Retry cycle | Built-in | External (Go code) | Via Goto |
| Checkpoint | ✅ WithCheckPointStore | ❌ Manual | ✅ Built-in |
| Complexity | Medium | Low | High |
| Field mapping | Implicit | Explicit | Via State |
| Production readiness | High | Medium | High |
My choice for the example: compose.Graph — native cycle support, checkpoint, and direct embedding of react.Agent via ExportGraph().
8. Code example
Implementing Reflexion on compose.Graph: Actor (ReAct agent) generates Go code, Evaluator runs tests, Reflector analyzes errors, Memory accumulates reflections.
flowchart TD
START(("START")) --> Actor["🎬 Actor
react.Agent
+ write_code tool"]
Actor --> Evaluator["🔍 Evaluator
Lambda: go test"]
Evaluator --> Branch{"Tests pass?"}
Branch -->|Yes| END(("END ✅"))
Branch -->|No| Reflector["🪞 Reflector
ChatModel: analyze
test failures"]
Reflector --> Memory["💾 Append reflection
to episodic memory"]
Memory --> Actor
style START fill:#2e7d32,color:#fff
style END fill:#2e7d32,color:#fff
style Branch fill:#e65100,color:#fff
package main
import (
"context"
"fmt"
"github.com/cloudwego/eino/compose"
"github.com/cloudwego/eino/components/tool"
"github.com/cloudwego/eino/components/tool/utils"
"github.com/cloudwego/eino-ext/components/model/openai"
"github.com/cloudwego/eino/flow/agent/react"
"github.com/cloudwego/eino/schema"
)
// reflexionState — shared state for a single Reflexion execution.
// Created fresh with each Invoke.
type reflexionState struct {
// Task — the original task (e.g., "write a function that sorts a list")
Task string
// Reflections — accumulated verbal reflections from past attempts
Reflections []string
// Attempt — current attempt number (1-based)
Attempt int
// MaxAttempts — maximum number of attempts
MaxAttempts int
// Code — generated code (Actor output)
Code string
// TestResult — test run result (Evaluator output)
TestResult string
// Passed — flag: did tests pass?
Passed bool
}
// writeCodeTool — Actor's tool: "writes" code to a file.
// In a real application, this would write to disk.
type writeCodeTool struct{}
func (t *writeCodeTool) Info(ctx context.Context) (*schema.ToolInfo, error) {
return &schema.ToolInfo{
Name: "write_code",
Desc: "Write Go code to solve the task. The code will be tested automatically.",
}, nil
}
func (t *writeCodeTool) InvokableRun(ctx context.Context, args string, opts ...tool.Option) (string, error) {
// In a real application: write args to a .go file
return fmt.Sprintf("Code written (%d bytes)", len(args)), nil
}
func main() {
ctx := context.Background()
// 1. Create model for Actor and Reflector
chatModel, err := openai.NewChatModel(ctx, &openai.ChatModelConfig{
Model: "gpt-4o",
})
if err != nil {
panic(err)
}
// 2. Create ReAct agent as Actor
// ExportGraph() allows embedding it into compose.Graph
codeTool := utils.NewTool(&writeCodeTool{}, nil)
actor, err := react.NewAgent(ctx, &react.AgentConfig{
ToolCallingModel: chatModel,
ToolsConfig: compose.ToolsNodeConfig{
Tools: []tool.BaseTool{codeTool},
},
MaxStep: 5,
})
if err != nil {
panic(err)
}
// 3. Build the Reflexion Graph
g := compose.NewGraph[string, string](
compose.WithGenLocalState(func(ctx context.Context) *reflexionState {
return &reflexionState{
MaxAttempts: 3,
}
}),
compose.WithMaxRunSteps(20), // cycle limit
)
// Actor: embed ReAct agent via ExportGraph
actorGraph, actorOpts := actor.ExportGraph()
g.AddGraphNode("actor", actorGraph, actorOpts...)
g.AddEdge(compose.START, "actor")
// Evaluator: Lambda node that "runs tests"
g.AddLambdaNode("evaluator",
compose.InvokableLambda(func(ctx context.Context, code string) (string, error) {
// In a real application: exec.Command("go", "test", "./...")
// Here: simulation
if len(code) > 10 {
return "PASS: all tests passed", nil
}
return "FAIL: TestSortEmpty - expected [], got nil", nil
}),
)
g.AddEdge("actor", "evaluator")
// Reflector: ChatModel analyzes errors
g.AddChatModelNode("reflector", chatModel)
g.AddEdge("evaluator", "reflector")
// Conditional Branch: pass → END, fail → back to Actor
g.AddBranch("evaluator", compose.NewGraphBranch(
func(ctx context.Context, testResult string) (string, error) {
// Read state to check attempt count
if err := compose.ProcessState[*reflexionState](ctx,
func(ctx context.Context, s *reflexionState) error {
s.Attempt++
s.TestResult = testResult
// Simple heuristic criterion
s.Passed = len(testResult) > 4 && testResult[:4] == "PASS"
},
); err != nil {
return "", err
}
var target string
_ = compose.ProcessState[*reflexionState](ctx,
func(ctx context.Context, s *reflexionState) error {
if s.Passed || s.Attempt >= s.MaxAttempts {
target = compose.END
} else {
target = "reflector" // reflect first, then retry
}
return nil
},
)
return target, nil
},
map[string]bool{compose.END: true, "reflector": true},
))
// Reflector → Actor: retry with reflection in context
g.AddEdge("reflector", "actor")
// END
g.AddEdge(compose.END, compose.END)
// 4. Compile and run
runnable, err := g.Compile(ctx)
if err != nil {
panic(fmt.Sprintf("compile error: %v", err))
}
result, err := runnable.Invoke(ctx, "Write a function that sorts a slice of integers")
if err != nil {
panic(fmt.Sprintf("invoke error: %v", err))
}
fmt.Println("Result:", result)
}
What’s happening here
State —
reflexionStatestores the current attempt, reflections, and evaluation result. Created viaWithGenLocalStateon each graph run.Actor — embedded via
ExportGraph(). A ReAct agent with awrite_codetool. On re-entry (after reflection), it receives updated context.Evaluator — a Lambda node that “runs tests”. In a real application —
exec.Command("go", "test"). In the example — simulation: if the code is longer than 10 bytes — PASS.Reflector — a ChatModel analyzing errors. Receives test results and generates verbal reflection.
Branch — conditional branching after Evaluator: PASS → END, FAIL → Reflector → Actor (retry). Checks
Attempt < MaxAttempts.Cycle —
g.AddEdge("reflector", "actor")closes the loop.WithMaxRunSteps(20)limits the total number of steps (protection against infinite loops).
9. Engineering scenario
Code generation with tests is the ideal scenario for Reflexion. Why? Because there’s an objective external verifier: the compiler and unit tests. This isn’t a subjective LLM assessment — it’s a binary PASS/FAIL.
TDD for agents
sequenceDiagram
participant Dev as Developer
participant A as Actor Agent
participant T as go test
participant R as Reflector
participant M as Memory
Dev->>A: "Write sort([]int)"
Note over A: Attempt 1
A->>T: sort.go
T-->>A: ❌ FAIL: TestSortEmpty
expected [], got nil
A->>R: Trajectory + test failure
R->>R: "Forgot to handle
empty slice"
R->>M: Store reflection #1
Note over A: Attempt 2 (with reflection)
M-->>A: "Check empty slice first"
A->>T: sort_v2.go
T-->>A: ❌ FAIL: TestSortStable
unstable sort on equal elements
A->>R: Trajectory + test failure
R->>R: "Used unstable sort,
need stable"
R->>M: Store reflection #2
Note over A: Attempt 3 (with 2 reflections)
M-->>A: "Check empty slice first
Use stable sort"
A->>T: sort_v3.go
T-->>A: ✅ All tests passed
A-->>Dev: sort_v3.go ✅
Why this works better than Self-Refine
Self-Refine in the same scenario would give 0% improvement on Math Reasoning (Madaan et al., 2023). Why? Without tests, the model can’t distinguish sort([]int{}) from sort([]int{1}) — it lacks an objective signal. Reflexion with go test as Evaluator solves this problem: tests provide precise error diagnostics → reflection focuses on the real problem → the Actor fixes exactly what’s needed.
Production implementation
In production, the scenario expands:
| Component | Example | In production |
|---|---|---|
| Actor | react.NewAgent with write_code | + read_file, search_docs, lint_code |
| Evaluator | go test ./... | + go vet, golangci-lint, coverage ≥ 80% |
| Reflector | ChatModel “analyze failures” | Prompt with specific error patterns |
| Memory | []string in state | Redis / file with reflection history |
| Max attempts | 3 | 5 (HumanEval: 91% achieved in 12 attempts, but 3-5 usually sufficient) |
10. Practical recommendations
When to apply Reflexion
| Scenario | Applicability | Rationale |
|---|---|---|
| Code generation + tests | ✅ Excellent | Objective verifier (compiler/tests) |
| Game agents | ✅ Excellent | Environment provides clear reward signal |
| Data pipeline with validation | ✅ Good | Schema validation as verifier |
| Code review automation | ⚠️ Cautious | LLM assessment less reliable than tests |
| Creative writing | ⚠️ Cautious | No objective success criterion |
| Math / reasoning | ❌ Not recommended | Without external verifier — degrades results |
Hyperparameter tuning
Max attempts: 3-5 for tasks with a fast verifier (tests). HumanEval reached 91% in 12 attempts, but diminishing returns start after 3-5. More is more expensive, but not better.
Episodic memory size: store the last 5-10 reflections. Too many — the context grows and the model loses focus. Too few — doesn’t account for older mistakes.
Evaluator choice: a deterministic verifier (tests, compiler) is always better than LLM-based. If the verifier is unreliable — Reflexion degrades into Self-Refine with its problems.
Reflector prompt: specific > abstract. Not “analyze the error”, but “identify: (1) which test failed, (2) what input caused the failure, (3) what assumption was wrong, (4) what to change in the code”.
Cost management
Each Reflexion attempt = a full Actor + Evaluator + Reflector cycle. With 3 attempts — 3x cost. Mitigations:
- Use a cheap model for the Evaluator (deterministic checking doesn’t require GPT-4)
- Stop early: if 2 attempts didn’t help — the third probably won’t either
- Cache reflections for similar tasks
11. Summary
The Reflexion pattern is ReAct + self-assessment + episodic memory. Key takeaways:
External verifier is mandatory. Without it, self-correction degrades results (Huang et al., 2023). With it — Reflexion delivers +11% on HumanEval and +22% on AlfWorld.
Episodic memory is the key difference from Self-Refine. Not just critique-refine, but accumulating verbal lessons between attempts. This transforms a one-shot agent into a learning one.
Not a silver bullet. Reflexion doesn’t work on tasks without an objective success criterion and can degrade results when initial accuracy is already high.
In Eino — compose.Graph with a cycle. Workflow doesn’t work (no cycles). Graph +
AddEdge("reflector", "actor")+WithMaxRunSteps— the natural implementation.
What’s next? In Part 4 — Multi-Agent patterns: when one agent isn’t enough, and you need a team. And the topic of agent memory — long-term, episodic, semantic — I’ll cover in detail in Part 6.