Repository navigation
Conversation
Early stopping treated failed and timed-out evaluations as scores. A
failure whose metrics fall back to safe_numeric_average (e.g.
{"combined_score": 0.0, "error": ...} or the evaluator's timeout shape
{"error": 0.0, "timeout": True}) scored 0.0, so the first failure was
counted as an infinite improvement over the -inf baseline and poisoned
best_score, and every subsequent failure incremented the
iterations_without_improvement counter. A run whose evaluator never
succeeded could therefore exit through the successful-plateau path
(issue algorithmicsuperintelligence#294). For tasks with negative scores, failure results scoring
0.0 kept registering as improvements, permanently pinning best_score.
Guard the early stopping block (both the patience-based and the
event-based branches) on the absence of "error"/"timeout" markers in
the child metrics: those iterations neither update best_score nor count
toward patience, so early stopping only triggers on successful
evaluations that plateau. Evaluators should include an "error" key in
their metrics on failure per the convention suggested in algorithmicsuperintelligence#294.
Documented the convention in configs/early_stopping_example.yaml.
Fixes algorithmicsuperintelligence#294
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #294
Problem
Early stopping counted failed and timed-out evaluations as if they were scores. The evaluator signals failures as metrics carrying an
error/timeoutkey —{"error": 0.0, "timeout": True}on timeout and{"error": 0.0}on exception (openevolve/evaluator.py), or the user-side convention suggested by codelion in #294 ({"error": 1.0, "combined_score": 0.0}). In the early-stopping block ofProcessParallelController.run_evolution(openevolve/process_parallel.py), none of these shapes were distinguished from real results:safe_numeric_averagereduces them to0.0, so the first failure scored0.0 - (-inf) = +infimprovement, was recorded asbest_score = 0.0, and reset the patience counter — the baseline was poisoned by a failure.improvement = 0 < convergence_thresholdand incrementediterations_without_improvement. Withearly_stopping_patience: 3and an evaluator that keeps failing (exactly theahura's report), the run stopped via the "No improvement for N iterations" path — the successful-plateau semantics — without ever succeeding once.0.0beat the real negative best, kept registering as improvement, and permanently pinnedbest_scoreat 0.0 so genuine negative improvements never reset the counter either.What
openevolve/process_parallel.py: guarded the early-stopping block on the absence of failure markers. Whenchild_program.metrics.get("error") is not None or child_program.metrics.get("timeout"), the iteration is skipped for early stopping entirely (debug log, nobest_scoreupdate, no counter change). Both the patience-based branch and the event-based branch (negative patience) are inside the guard, so failures can no longer trigger "Task successfully solved" either. The presence check forerrorisis not None(not truthiness) because evaluators signal failures with both numeric ({"error": 0.0}) and string values. This follows codelion's prescription in How to handle early stopping patience with failure cases #294 ("skip iterations where 'error' or 'timeout' keys are present when countingiterations_without_improvement"); no new config option, no change to the counter semantics for successful evaluations.configs/early_stopping_example.yaml: documented the contract — iterations whose metrics carry anerror/timeoutkey do not count toward patience, and evaluators should return e.g.{"error": 1.0, "combined_score": 0.0}on failure. (There is no README early-stopping section to extend; this config file is where the feature is documented today, alongsideconfigs/default_config.yaml's inline comments.)tests/test_early_stopping_failures.py: new unittest module (no LLM, no subprocesses; reuses the_submit_iterationmock seam fromtests/test_process_parallel.py::test_run_evolution_basic, mocked futures complete immediately; runs in ~0.02s):test_failed_evaluations_do_not_trigger_early_stop: patience 3, 5 iterations all{"combined_score": 0.0, "error": "syntax error"}→ no early stop, all 5 iterations submitted;test_timeout_shaped_metrics_do_not_trigger_early_stop: same with the evaluator's current timeout shape{"error": 0.0, "timeout": True};test_genuine_plateau_still_stops: 5 successful{"combined_score": 0.5}iterations → early stopping still triggers (regression guard for normal plateau behavior);test_negative_best_not_poisoned_by_failures: one real{"combined_score": -0.5}best followed by 4 failures → no early stop,best_scorenot dragged to 0.0.Red -> green evidence
Before the fix, with the repo-default
convergence_threshold(0.001) andearly_stopping_patience: 3. The default threshold is load-bearing for the repro: atthreshold == 0a repeated0.0always satisfiesimprovement >= threshold, so every failure counts as an "improvement", the counter never increments, and the bug cannot reproduce. The reporter's config only setearly_stopping_patience, so their run effectively used the default 0.001 — which is exactly what lets consecutive failures accumulate as "no improvement":(
test_genuine_plateau_still_stopspassed before the fix too — normal plateau behavior was never broken.)After the fix:
Ran 4 tests in 0.023s ... OK.Full suite:
OPENAI_API_KEY=test-key-for-unit-tests python -m unittest discover tests->Ran 578 tests in 33.128s ... OK(574 baseline + 4 new).Notes / compatibility
max_iterations). Anyone relying on failure-as-plateau was hitting the bug reported in How to handle early stopping patience with failure cases #294.evaluator.py/utils/async_utils.py/utils/metrics_utils.py— no textual overlap with this change). The guard here is presence-based, so it correctly skips both the current timeout shape ({"error": 0.0, "timeout": True}, nocombined_score) and Fix timeout evaluations getting positive fitness #446's proposed shape (combined_score: 0.0+ stringerror+timeout: True) — merge order does not matter.database.best_program_id, so the "New best" log line is unaffected in practice.erroras a metric name (e.g. an error rate) will have those iterations excluded from early stopping. That direction is safe — it only delays stopping, never triggers it wrongly — and it is the marker key codelion prescribed in How to handle early stopping patience with failure cases #294.