Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
- GitHub stars
- 797
- Version
- v0.1.39
- License
- MIT
- Added
- 2026-08-13
GitHub info
- GitHub stars
- 797
- License
- MIT
- Primary language
- TypeScript
- Last push
- Aug 20, 2026, 8:50 PM
- Maintainer
- Anionex
- Added
- 2026-08-13
Preview from README
Screenshots and GIFs extracted from the repository README (badges and avatars filtered out).

QR code for the agent-vision-toolkit community group
https://raw.githubusercontent.com/Anionex/dsh-vision-toolkit/main/assets/community-group-qr.png

WeChat reward code
https://raw.githubusercontent.com/Anionex/dsh-vision-toolkit/main/assets/wechat-reward.png

DSH Vision Toolkit helps text-only DeepSeek Harness agents understand images and complete visual tasks
https://raw.githubusercontent.com/Anionex/dsh-vision-toolkit/main/assets/hero-v2.png

A text-only DeepSeek model answering a question about a pasted image through Vision Toolkit in DSH Web
https://raw.githubusercontent.com/Anionex/dsh-vision-toolkit/main/assets/dsh-view-example.png

Reference infographic screenshot used for restoration
https://raw.githubusercontent.com/Anionex/dsh-vision-toolkit/main/assets/upstream/infographic-reference.webp

Editable HTML and CSS reconstruction created from the reference screenshot
https://raw.githubusercontent.com/Anionex/dsh-vision-toolkit/main/assets/upstream/infographic-result.webp
Install
dsh plugin --profile web add @anionex/dsh-vision-toolkitTags
README badge
Add this Markdown to your plugin README to link back to its listing.
[](https://dshget.com/plugins/Anionex/dsh-vision-toolkit)Related plugins
Vision & Multimodaldsh-computer-use
★ 34Accessibility-first macOS computer use: fresh observations, stale-state rejection, scoped permissions, and safe input.
modlens
★ 3768Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
dsh-vision-router
★ 1030Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.