Last commit on August 14, 2026·Created on February 22, 2026
liustack/modlens
“A vision bridge that converts images into structured JSON for text-only models”
Combined rank
#188
across all skills
In Data
#2
category rank
Stars
1.1k
7d change unavailable
Forks
34
Watchers
0
Traction scoreGitHub stars can be faked, so popularity alone can be misleading. Traction Score looks for broader signs of real attention, adoption, and active maintenance.
ModLens provides vision capabilities to coding agents and LLMs that lack native multimodal support. It acts as a bridge, intercepting image paths or URLs and translating them into detailed evidence including transcribed text, layout regions, and visual semantics.
Designed for automation builders and infrastructure engineers, the tool replaces manual OCR or basic image processing with a structured CLI pipeline. It supports multiple providers including Gemini, Claude, and OpenAI-compatible endpoints to ensure reliable image analysis.
WHO IT'S FOR
AI agent / automation builders
adding vision capabilities to text-only LLMs
DeepSeek Harness users
processing images via structured JSON evidence
Claude Code users
configuring vision providers for coding assistants