Skip to content

Skip failed evaluations in early stopping plateau count - #502

Merged
codelion merged 1 commit into
algorithmicsuperintelligence:mainfrom
akushonkamen:fix/294-early-stop-failed-evals
Oct 10, 2026
Merged

codelion merged 1 commit into
algorithmicsuperintelligence:mainfrom
akushonkamen:fix/294-early-stop-failed-evals

Conversation

@akushonkamen

Copy link
Copy Markdown
Contributor

Fixes #294

Problem

Early stopping counted failed and timed-out evaluations as if they were scores. The evaluator signals failures as metrics carrying an error/timeout key — {"error": 0.0, "timeout": True} on timeout and {"error": 0.0} on exception (openevolve/evaluator.py), or the user-side convention suggested by codelion in #294 ({"error": 1.0, "combined_score": 0.0}). In the early-stopping block of ProcessParallelController.run_evolution (openevolve/process_parallel.py), none of these shapes were distinguished from real results:

  • safe_numeric_average reduces them to 0.0, so the first failure scored 0.0 - (-inf) = +inf improvement, was recorded as best_score = 0.0, and reset the patience counter — the baseline was poisoned by a failure.
  • Every subsequent failure then scored improvement = 0 < convergence_threshold and incremented iterations_without_improvement. With early_stopping_patience: 3 and an evaluator that keeps failing (exactly theahura's report), the run stopped via the "No improvement for N iterations" path — the successful-plateau semantics — without ever succeeding once.
  • Reverse damage for negative-score tasks (e.g. loss minimization): each failure's 0.0 beat the real negative best, kept registering as improvement, and permanently pinned best_score at 0.0 so genuine negative improvements never reset the counter either.

What

  • openevolve/process_parallel.py: guarded the early-stopping block on the absence of failure markers. When child_program.metrics.get("error") is not None or child_program.metrics.get("timeout"), the iteration is skipped for early stopping entirely (debug log, no best_score update, no counter change). Both the patience-based branch and the event-based branch (negative patience) are inside the guard, so failures can no longer trigger "Task successfully solved" either. The presence check for error is is not None (not truthiness) because evaluators signal failures with both numeric ({"error": 0.0}) and string values. This follows codelion's prescription in How to handle early stopping patience with failure cases #294 ("skip iterations where 'error' or 'timeout' keys are present when counting iterations_without_improvement"); no new config option, no change to the counter semantics for successful evaluations.
  • configs/early_stopping_example.yaml: documented the contract — iterations whose metrics carry an error/timeout key do not count toward patience, and evaluators should return e.g. {"error": 1.0, "combined_score": 0.0} on failure. (There is no README early-stopping section to extend; this config file is where the feature is documented today, alongside configs/default_config.yaml's inline comments.)
  • tests/test_early_stopping_failures.py: new unittest module (no LLM, no subprocesses; reuses the _submit_iteration mock seam from tests/test_process_parallel.py::test_run_evolution_basic, mocked futures complete immediately; runs in ~0.02s):
    • test_failed_evaluations_do_not_trigger_early_stop: patience 3, 5 iterations all {"combined_score": 0.0, "error": "syntax error"} → no early stop, all 5 iterations submitted;
    • test_timeout_shaped_metrics_do_not_trigger_early_stop: same with the evaluator's current timeout shape {"error": 0.0, "timeout": True};
    • test_genuine_plateau_still_stops: 5 successful {"combined_score": 0.5} iterations → early stopping still triggers (regression guard for normal plateau behavior);
    • test_negative_best_not_poisoned_by_failures: one real {"combined_score": -0.5} best followed by 4 failures → no early stop, best_score not dragged to 0.0.

Red -> green evidence

  • Before the fix, with the repo-default convergence_threshold (0.001) and early_stopping_patience: 3. The default threshold is load-bearing for the repro: at threshold == 0 a repeated 0.0 always satisfies improvement >= threshold, so every failure counts as an "improvement", the counter never increments, and the bug cannot reproduce. The reporter's config only set early_stopping_patience, so their run effectively used the default 0.001 — which is exactly what lets consecutive failures accumulate as "no improvement":

    $ python -m unittest tests.test_early_stopping_failures -v
    ...
    FAILED (failures=3)
    AssertionError: True is not false   (x3: failed-eval, timeout-shape, negative-best tests)
    

    (test_genuine_plateau_still_stops passed before the fix too — normal plateau behavior was never broken.)

  • After the fix: Ran 4 tests in 0.023s ... OK.

Full suite: OPENAI_API_KEY=test-key-for-unit-tests python -m unittest discover tests -> Ran 578 tests in 33.128s ... OK (574 baseline + 4 new).

Notes / compatibility

  • Behavior change, intentional: runs that previously "converged" purely through repeated failures will now keep running (up to max_iterations). Anyone relying on failure-as-plateau was hitting the bug reported in How to handle early stopping patience with failure cases #294.
  • Interacts cleanly with open PR Fix timeout evaluations getting positive fitness #446 ("Fix timeout evaluations getting positive fitness", touches evaluator.py/utils/async_utils.py/utils/metrics_utils.py — no textual overlap with this change). The guard here is presence-based, so it correctly skips both the current timeout shape ({"error": 0.0, "timeout": True}, no combined_score) and Fix timeout evaluations getting positive fitness #446's proposed shape (combined_score: 0.0 + string error + timeout: True) — merge order does not matter.
  • Out of scope (kept minimal per codelion's prescription): the target-score check, the "New best" logging, and the evaluator's timeout return shape (Fix timeout evaluations getting positive fitness #446's territory). Failed programs also rarely become database.best_program_id, so the "New best" log line is unaffected in practice.
  • Known trade-off of the convention: an evaluator that legitimately uses error as a metric name (e.g. an error rate) will have those iterations excluded from early stopping. That direction is safe — it only delays stopping, never triggers it wrongly — and it is the marker key codelion prescribed in How to handle early stopping patience with failure cases #294.

Early stopping treated failed and timed-out evaluations as scores. A
failure whose metrics fall back to safe_numeric_average (e.g.
{"combined_score": 0.0, "error": ...} or the evaluator's timeout shape
{"error": 0.0, "timeout": True}) scored 0.0, so the first failure was
counted as an infinite improvement over the -inf baseline and poisoned
best_score, and every subsequent failure incremented the
iterations_without_improvement counter. A run whose evaluator never
succeeded could therefore exit through the successful-plateau path
(issue algorithmicsuperintelligence#294). For tasks with negative scores, failure results scoring
0.0 kept registering as improvements, permanently pinning best_score.

Guard the early stopping block (both the patience-based and the
event-based branches) on the absence of "error"/"timeout" markers in
the child metrics: those iterations neither update best_score nor count
toward patience, so early stopping only triggers on successful
evaluations that plateau. Evaluators should include an "error" key in
their metrics on failure per the convention suggested in algorithmicsuperintelligence#294.

Documented the convention in configs/early_stopping_example.yaml.

Fixes algorithmicsuperintelligence#294
@CLAassistant

CLAassistant commented Oct 8, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@codelion
codelion merged commit 256b55d into algorithmicsuperintelligence:main Oct 10, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

How to handle early stopping patience with failure cases

3 participants