Traction scoreGitHub stars can be faked, so popularity alone can be misleading. Traction Score looks for broader signs of recent attention, adoption, and active maintenance.
This collection provides the technical patterns necessary to move from basic prompting to rigorous performance measurement. It guides AI and ML engineers through the lifecycle of evaluation, from generating synthetic datasets to validating the reliability of LLM-based judges.
The skills focus on reducing uncertainty in RAG systems and generative models by implementing error discovery workflows, custom annotation interfaces, and systematic audit processes to organize failure modes.
Compatible AgentsThe repository documents support for these agents. The skills may also work with other agents that can load SKILL.md files, but they may need some setup or small changes.
Path to a JSONL, CSV, or JSON file containing LLM outputs or traces
Availability to review samples and provide annotations interactively
Use the error-discovery skill to help me do error analysis on traces.jsonl. Treat traces.jsonl as a placeholder for my JSONL, CSV, or JSON dataset path. Begin by examining 5–10 records, inventorying their fields, and identifying the primary content and metadata before proposing the review interface and sample-selection approach.
Local paths to available traces, evaluator configurations, judge prompts, labeled datasets, notebooks, evaluation scripts, or metrics exports
Use the eval-audit skill to inspect my existing LLM evaluation pipeline. Review the supplied traces, evaluator configurations, judge prompts, labeled data, and metrics artifacts, then give me a prioritized first-pass list of problems with concrete next steps. If I do not have every artifact, work only from the local files I provide and identify the important gaps.
Draft dimensions and tuples for synthetic eval data
Uses generate-synthetic-data
Description of the LLM application and what it does
Known failure-prone areas or failure hypotheses
Existing user feedback or sample traces, if available
Domain guidance for judging whether proposed tuples are realistic
Use the generate-synthetic-data skill for my LLM application. Based on my application description, known failure-prone areas, existing feedback, and any sample traces I provide, propose exactly three failure-focused dimensions and draft 20 diverse tuples for my review. Stop after the draft tuples so I can confirm which scenarios are realistic before any natural-language queries are generated.