Anionex/agent-vision-toolkit
A text-only coding agent can add vision through shell tools, a skill, and optional local integrations instead of changing the model.
The repo provides CLIs for image Q&A, OCR, grounding, detection, cropping, tracing, and related visual workflows. Its included `vision-skills` package teaches agents when to use those tools and how to verify results. Optional integrations let pasted images and built-in image tools route through a local proxy or native plugin for Codex, Claude Code, Pi, Oh My Pi, and OpenCode. The project says its pipeline has been verified in real sessions across those agents. GitHub · LLM repos' note
The repo provides CLIs for image Q&A, OCR, grounding, detection, cropping, tracing, and related visual workflows. Its included `vision-skills` package teaches agents when to use those tools and how to verify results. Optional integrations let pasted images and built-in image tools route through a local proxy or native plugin for Codex, Claude Code, Pi, Oh My Pi, and OpenCode. The project says its pipeline has been verified in real sessions across those agents. GitHub · LLM repos' note
score 4