←Leaderboard/Developer Tooling/modlens
Last commit on September 27, 2026·Created on February 22, 2026

liustack/modlens

“A vision bridge that provides structured image evidence for text-only models”
Combined rank
#106
across all skills
In Developer Tooling
#3
category rank
Stars
4.1k
+1.6% in last 7d
Forks
124
+0.8% in last 7d
Watchers
5
+25.0% in last 7d
Traction scoreGitHub stars can be faked, so popularity alone can be misleading. Traction Score looks for broader signs of recent attention, adoption, and active maintenance.
TL;DR

ModLens enables text-only coding agents and LLMs to process visual information by converting images into structured JSON. It extracts transcriptions, layout regions, and semantic clues from image files or URLs, allowing models without native multimodal capabilities to understand visual context.

Designed as a plugin for agent harnesses, it replaces manual OCR or byte-reading with a deterministic pipeline. It includes guardrails to detect if a model already possesses native vision, ensuring the tool is only invoked when necessary.

WHO IT'S FOR
AI agent / automation builders
adding vision capabilities to text-only LLMs
Platform / DevEx teams
integrating multimodal tools into agent harnesses
Technical leads / architects
standardizing image-to-text evidence for AI workflows
AI agent power users
processing image-based data with text-only agents
Repository contents

1 skill file

Compatible AgentsThe repository documents support for these agents. The skills may also work with other agents that can load SKILL.md files, but they may need some setup or small changes.
Want to install it?
Clone to Claude · Clone to Codex · Manual Copy · NPX Skills AddRepository instructions