给纯文本 LLM agent 装上眼睛:图片问答、OCR、截图分析、视觉定位等一套视觉工具箱 + skill,并可无缝接入 Codex、Claude Code、OpenCode、Pi | Give text-only LLM agents vision: image Q&A, OCR, screenshot understanding, visual grounding, image-to-SVG - a vision toolkit & skill, with drop-in integration for Codex, Claude Code, OpenCode, Pi
313stars14forksPython
agentagent-skillsclaude-codeclicodexcomputer-usedeepseekdeepseek-v4glmharness-engineeringkimillmmultimodalocropencodevisionvision-language-model
Real data pulled from GitHub this week. The author's original repo lives upstream.
View on GitHub