Replies: 5 comments 1 reply
|
Evals, or evaluations, are tests that are used to measure how well a machine learning model is performing. They are typically designed to assess the model's ability to complete a specific task or solve a specific problem. In general, it is not desirable for an evaluation to have a 0% success rate, as this means that the model is not able to perform the task at all. However, there may be cases where a 0% success rate is acceptable, depending on the specific circumstances of the evaluation. For example, if a human can easily solve the problem being tested, but the machine learning model cannot, it may still be useful to evaluate the model's performance in order to identify areas where it needs improvement. In this case, a 0% success rate would indicate that the model needs significant work in order to be able to perform the task as well as a human can. Ultimately, the decision of whether to allow evaluations with a 0% success rate will depend on the goals and objectives of the specific project or application. |
|
Evals" are like tests for robots to see how well they can understand and do things. Sometimes, a test might be too hard for the robot and it can't do it, so it gets a score of zero. It's okay if a robot gets a zero score sometimes, but if it keeps getting zero scores all the time, then we might need to change the test or help the robot get better. |
|
Just copy pasting answers from ChatGPT doesn't help here. I'd suggest to refrain from doing so in the future. |
|
ChatGPT said: Just Chill! |
|
0% success rate evals are useful as lower-bound probes — they tell you "the model definitely can't do this yet" which is meaningful information. The eval isn't broken just because no model passes it. For agent systems, 0% evals are particularly valuable for: Capability ceiling tracking — if you have an eval for "decompose this 20-step task correctly," and it starts at 0%, you can track when agents cross 10%, 50%, 80%. The trajectory matters more than the absolute score. Regression detection — if a task type is at 0% in version N and still 0% in version N+1, that's neutral. If it was 30% and dropped to 0%, that's a regression worth investigating. Aspirational benchmarks — for multi-agent coordination specifically, many of the hard problems (cycle-free delegation with budget propagation, memory consolidation without context loss) would score 0% on current agent frameworks. That's exactly why they're worth tracking. The design question is whether 0% evals slow down your CI. We separate eval suites by expected difficulty: fast / deterministic evals run on every commit; aspirational evals run nightly. The nightly suite includes 0%-category evals that track capability growth over weeks. For agent capability evals specifically — there's a write-up on what dimensions matter for multi-agent systems: https://blog.kinthai.ai/221-agents-multi-agent-coordination-lessons — covers coordination, memory, and economic decision-making as eval dimensions. What's the context — are the 0% evals for a specific capability you're trying to develop, or as a baseline for a new model? |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Are Evals with 0% success rate allowed - if I believe GPT should be able to solve it and if a human can solve it easily?
All reactions