Skip to content

How to handle early stopping patience with failure cases #294

Description

@theahura

Say I have an evaluator that returns a score based on some written code. If the code does not compile (e.g. a syntax error) I simply return a score of 0.

Say I also have a config with an early stopping patience set. Say that my patience score is a smallish number, like 3.

Right now what I am observing is if my evaluator cannot run a bit of iterated code 3 times, it will assume that it has succeeded and move on, even though it actually failed three times in a row.

Is there a way to distinguish between 'this was successful and that is why the score did not change (and you should stop running)' and 'this was a failure and that is why the score did not change (and you should keep running)'?

Activity

  1. codelion commented on Oct 16, 2025

    @codelion
    Member

    You can modify the evaluator to return a distinct metric key for failures (e.g., {"error": 1.0, "combined_score": 0.0}) instead of just {"combined_score": 0.0}. Then update the early stopping logic in process_parallel.py:600-642 to skip iterations where "error" or "timeout" keys are present when counting iterations_without_improvement. This way, early stopping only triggers on successful evaluations that plateau, not on repeated failures.

  2. theahura commented on Oct 17, 2025

    @theahura
    ContributorAuthor

    any chance you'd be willing to add in support for hard coded error handling logic? I can keep patching openevolve but I would much rather have these dependencies in the dependencies 😂

  3. codelion commented on Oct 17, 2025

    @codelion
    Member

    I do want to fix all such pending issues when I get some time but meanwhile feel free to send in your PRs, happy to add them in.

  4. akushonkamen commented on Oct 8, 2026

    @akushonkamen
    Contributor

    Working on this — I'm following the fix direction codelion suggested above: the early-stopping logic in process_parallel.py now skips any iteration whose metrics carry an error or timeout key when counting iterations_without_improvement, with both the patience-based branch and the negative-patience event branch wrapped in that guard, so failures can no longer reset or tick the counter (or trigger the "successfully solved" path). A PR with tests covering failed, timeout-shaped, and negative-best scenarios is coming shortly.

  5. akushonkamen commented on Oct 8, 2026

    @akushonkamen
    Contributor

    PR up: #502 (includes tests for failed, timeout-shaped, and negative-best scenarios; full suite green).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    questionFurther information is requested

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions