Leaderboard/Testing/evals-skills
Last commit on June 10, 2026·Created on March 1, 2026

hamelsmu/evals-skills

A framework for designing and validating the evaluation pipelines of LLM applications.
Combined rank
#194
across all skills
In Testing
#1
category rank
Stars
1.7k
7d change unavailable
Forks
170
Watchers
17
Traction scoreGitHub stars can be faked, so popularity alone can be misleading. Traction Score looks for broader signs of real attention, adoption, and active maintenance.
TL;DR

This collection provides a structured approach to measuring AI performance, focusing on the transition from raw traces to actionable insights. It guides developers through the creation of synthetic datasets, the drafting of judge prompts, and the implementation of custom annotation interfaces for human review.

By emphasizing error analysis and evaluator validation, the repository helps teams move beyond vague metrics toward a rigorous audit of how their models behave in production and RAG environments.

WHO IT'S FOR
AI Engineers
implementing rigorous LLM evaluation pipelines
AI Product Managers
defining quality benchmarks for AI features
RAG Developers
optimizing retrieval and generation quality
Data Engineers
generating synthetic datasets for testing
Repository contents

7 skill files

Compatible AgentsThe repository documents support for these agents. The skills may also work with other agents that can load SKILL.md files, but they may need some setup or small changes.
Want to install it?
Claude Marketplace · Claude Plugin · NPX Skills AddRepository instructions