Leaderboard/Testing/evals-skills
Last commit on August 16, 2026·Created on June 23, 2026

ai-evals-course/evals-skills

A framework for building and validating product-specific AI evaluation pipelines
Combined rank
#512
across all skills
In Testing
#7
category rank
Stars
295
7d change unavailable
Forks
25
Watchers
6
Traction scoreGitHub stars can be faked, so popularity alone can be misleading. Traction Score looks for broader signs of real attention, adoption, and active maintenance.
TL;DR

This collection provides a structured approach to measuring AI performance, moving from synthetic data generation to the creation of custom judge prompts. It helps engineers transition from anecdotal testing to systematic evaluation by establishing rigorous validation for the evaluators themselves.

The skills focus on the end-to-end lifecycle of quality assurance, including specialized techniques for RAG systems and the development of annotation interfaces to organize failure modes through error discovery.

WHO IT'S FOR
AI / ML Engineers
build product-specific AI evaluation pipelines
LLM Application Developers
evaluate and optimize RAG systems
AI Product Managers
perform error analysis and failure mode discovery
Data Labeling Leads
create custom human-in-the-loop annotation tools
Repository contents

8 skill files

Compatible AgentsThe repository documents support for these agents. The skills may also work with other agents that can load SKILL.md files, but they may need some setup or small changes.
Want to install it?
Claude Plugin · NPX Skills AddRepository instructions