←Leaderboard/Testing/evals-skills
Last commit on September 24, 2026·Created on June 23, 2026

ai-evals-course/evals-skills

“A structured framework for building and auditing AI evaluation pipelines”
Combined rank
#247
across all skills
In Testing
#3
category rank
Stars
1.4k
+8.4% in last 7d
Forks
97
+9.0% in last 7d
Watchers
13
0.0% in last 7d
Traction scoreGitHub stars can be faked, so popularity alone can be misleading. Traction Score looks for broader signs of recent attention, adoption, and active maintenance.
TL;DR

This collection provides the technical patterns necessary to move from basic prompting to rigorous performance measurement. It guides AI and ML engineers through the lifecycle of evaluation, from generating synthetic datasets to validating the reliability of LLM-based judges.

The skills focus on reducing uncertainty in RAG systems and generative models by implementing error discovery workflows, custom annotation interfaces, and systematic audit processes to organize failure modes.

WHO IT'S FOR
AI / ML Engineers
build product-specific AI evaluation pipelines
RAG Developers
evaluate and optimize RAG system performance
Data Scientists
perform error analysis and failure mode discovery
Frontend / Fullstack Engineers
build custom annotation and review interfaces
Repository contents

9 skill files

Compatible AgentsThe repository documents support for these agents. The skills may also work with other agents that can load SKILL.md files, but they may need some setup or small changes.

Use this skill collection

Follow the documented setup, then try a first task.

Install the complete Evals Skills bundle

Installs: Complete evals plugin bundle, including all documented skills in the repository.

  1. Run this command in a terminal to install the repository's complete skill bundle.

    npx skills add https://github.com/ai-evals-course/evals-skills
.codex-plugin/plugin.json ↗ · Checked Oct 2, 2026README.md ↗ · Checked Oct 2, 2026

Give it something to do.

Suggested first task

Start error analysis on an LLM trace dataset

Uses error-discovery

  • Path to a JSONL, CSV, or JSON file containing LLM outputs or traces
  • Availability to review samples and provide annotations interactively
Use the error-discovery skill to help me do error analysis on traces.jsonl. Treat traces.jsonl as a placeholder for my JSONL, CSV, or JSON dataset path. Begin by examining 5–10 records, inventorying their fields, and identifying the primary content and metadata before proposing the review interface and sample-selection approach.
README.md ↗ · Checked Oct 2, 2026skills/error-discovery/SKILL.md ↗ · Checked Oct 2, 2026
Suggested first task

Audit an existing LLM evaluation pipeline

Uses eval-audit

  • Local paths to available traces, evaluator configurations, judge prompts, labeled datasets, notebooks, evaluation scripts, or metrics exports
Use the eval-audit skill to inspect my existing LLM evaluation pipeline. Review the supplied traces, evaluator configurations, judge prompts, labeled data, and metrics artifacts, then give me a prioritized first-pass list of problems with concrete next steps. If I do not have every artifact, work only from the local files I provide and identify the important gaps.
skills/eval-audit/SKILL.md ↗ · Checked Oct 2, 2026README.md ↗ · Checked Oct 2, 2026
Suggested first task

Draft dimensions and tuples for synthetic eval data

Uses generate-synthetic-data

  • Description of the LLM application and what it does
  • Known failure-prone areas or failure hypotheses
  • Existing user feedback or sample traces, if available
  • Domain guidance for judging whether proposed tuples are realistic
Use the generate-synthetic-data skill for my LLM application. Based on my application description, known failure-prone areas, existing feedback, and any sample traces I provide, propose exactly three failure-focused dimensions and draft 20 diverse tuples for my review. Stop after the draft tuples so I can confirm which scenarios are realistic before any natural-language queries are generated.
README.md ↗ · Checked Oct 2, 2026skills/generate-synthetic-data/SKILL.md ↗ · Checked Oct 2, 2026