Leaderboard/DevOps/AI-Infra-Auto-Driven-SKILLS
Last commit on August 23, 2026·Created on April 1, 2026

BBuf/AI-Infra-Auto-Driven-SKILLS

Technical guidance for optimizing LLM serving infrastructure and analyzing model execution pipelines
Combined rank
#365
across all skills
In DevOps
#7
category rank
Stars
785
7d change unavailable
Forks
67
Watchers
8
Traction scoreGitHub stars can be faked, so popularity alone can be misleading. Traction Score looks for broader signs of real attention, adoption, and active maintenance.
TL;DR

This collection provides specialized knowledge for engineers managing high-performance LLM serving frameworks like SGLang, vLLM, and TensorRT-LLM. It focuses on the intersection of model optimization and infrastructure stability, offering structured methods to reduce regression risks and improve throughput.

The skills enable precise triage of GPU memory allocation, the analysis of torch profiler traces at the kernel level, and the use of historical pull request data to inform fusion ideas and fast-path selections.

WHO IT'S FOR
LLM infrastructure engineers
optimizing model serving performance and kernels
MLOps / Deployment engineers
planning GPU memory and request capacity
AI performance architects
benchmarking serving frameworks against latency SLAs
Deep learning researchers
estimating FLOPs and MFU for model shapes
Repository contents

12 skill files

Compatible AgentsThe repository documents support for these agents. The skills may also work with other agents that can load SKILL.md files, but they may need some setup or small changes.
Want to install it?
Claude Marketplace · Claude Plugin · Clone to Claude · Manual Copy · Symlink to AgentRepository instructions