// Feedback loops · ~13 min

Evaluator-Optimizer

Reach for it when: When a scorer can reject work until an optimizer improves it.

// 60-second mental model

How to hold it in your head

One role proposes an answer. Another scores it against a rubric. If the score is too low, an optimizer revises — and you loop until the score clears or you hit a hard stop.

Evaluator-Optimizer separates proposal from scoring: an evaluator returns a numeric or structured score; an optimizer revises until a threshold is met or maxIterations is exhausted. This is a runtime quality loop to implement — not the same thing as a CI eval harness or a fleet-wide suite.

// Mini architecture

Propose → score → revise (threshold or stop)

┌──────────────┐
│ Task + rubric│
└──────┬───────┘
       ▼
┌──────────────┐
│ Optimizer    │ → candidate vN
└──────┬───────┘
       ▼
┌──────────────┐
│ Evaluator    │ → score / reasons
└──────┬───────┘
       │ score ≥ threshold or maxIter?
       ├─ yes → Final candidate
       └─ no  → Optimizer revises → loop

The evaluator only scores. The optimizer only revises. Neither invents a new goal mid-loop.

// Mini-project

Score-gated product blurb (stub)

Optimizer + evaluator stubs with a numeric threshold

Goal: Improve a canned product blurb until a stub evaluator scores ≥8/10 on clarity and honesty, or hit maxIterations=3.

  1. Start with a hype-heavy canned blurb.
  2. Rubric axes (stub): clarity, honesty, length ≤80 words — average 0–10.
  3. Evaluator stub returns `{ score, reasons[] }` — never rewrite the blurb itself.
  4. Optimizer revises only the failing axes; log score each round.
  5. Stop at score ≥8 or maxIterations=3; emit final + history.
  6. Refuse a fourth round even if the score is still low.

Stubs are the point. Treat this as a pattern to implement in your own stack — do not read it as proof that every agent in a fleet already has a production eval suite.

// Common failure

What goes wrong

Symptom

Score thrashing — the optimizer chases nits forever and never clears the threshold.

Fix

Hard maxIterations plus a “good enough” rule: ship when remaining failures are style-only, or when iterations are exhausted.

// Self-reflection

Sit with this

Where are you rewriting by vibe instead of letting a scorer reject and a reviser improve?

Session only. Nothing is saved.