Leaderboard/Media/claude-vision-skill
Last commit on August 7, 2026·Created on May 2, 2026

asuojun/claude-vision-skill

Converts images and clipboard captures into descriptive text for models without native vision.
Combined rank
#154
across all skills
In Media
#6
category rank
Stars
2.3k
7d change unavailable
Forks
116
Watchers
2
Traction scoreGitHub stars can be faked, so popularity alone can be misleading. Traction Score looks for broader signs of real attention, adoption, and active maintenance.
TL;DR

This skill provides a bridge for LLMs that cannot directly process visual data. By utilizing a helper script, it transforms images from local paths, URLs, or the system clipboard into text-based descriptions and analyses.

It is designed for developers who need to recognize content within images or analyze visual references while working in a text-only model environment.

WHO IT'S FOR
AI agent / automation builders
integrating image analysis into LLM workflows
Platform / DevEx teams
extending AI assistant capabilities with local tools
AI agent power users
analyzing screenshots and images via CLI tools
Backend / infrastructure engineers
configuring vision API integrations for AI tools
Repository contents

1 skill file

Compatible AgentsThe repository documents support for these agents. The skills may also work with other agents that can load SKILL.md files, but they may need some setup or small changes.
Want to install it?
Git CloneRepository instructions