Open source evaluation framework: accuracy + cost + latency + hallucination #1675
vignesh2027
started this conversation in
Show and tell
Replies: 1 comment
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hey OpenAI Evals community!
OpenAI Evals is fantastic for task-specific evaluation. For teams who also need production metrics alongside task accuracy, I built a complementary open source framework.
What it adds beyond task accuracy:
Key insight from running this:
GPT-4o-mini vs Gemini Flash: 78.4% vs 76.8% accuracy. But $0.0003 vs $0.0001 per 1K. For production at scale, that 2% accuracy gap rarely justifies the 3x cost difference.
Live demo (no API key needed): https://huggingface.co/spaces/vigneshwar234/llm-eval-demo
GitHub: https://github.com/vignesh2027/LLM-Evaluation-Framework
71 tests, 82% coverage, full CI/CD. Open source, free forever.
Task evaluation (OpenAI Evals) + production metrics (this) = complete evaluation stack.
All reactions