Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
- GitHub stars
- 3,445
- License
- MIT
- Added
- 2026-08-14
GitHub info
- GitHub stars
- 3,445
- License
- MIT
- Primary language
- TypeScript
- Last push
- Aug 20, 2026, 8:05 PM
- Maintainer
- liustack
- Added
- 2026-08-14
Preview from README
Screenshots and GIFs extracted from the repository README (badges and avatars filtered out).

The skill triggering on its own in a DeepSeek Claude Code session and reading a pasted slide
https://raw.githubusercontent.com/liustack/modlens/main/assets/demo-claude-paste-recovery.jpg

Text-only DeepSeek reading a tweet screenshot in full detail via ModLens
https://raw.githubusercontent.com/liustack/modlens/main/assets/demo-codex-app.jpg

Three images dropped together, read one by one
https://raw.githubusercontent.com/liustack/modlens/main/assets/demo-codex-batch.jpg

The 128-model scatter plot read in full: axes, log scale, and highlighted region
https://raw.githubusercontent.com/liustack/modlens/main/assets/demo-codex-chart.jpg

Pasting an image straight into DeepSeek Harness, read through the modlens vision plugin
https://raw.githubusercontent.com/liustack/modlens/main/assets/demo-dsh-paste.jpg

The ModLens vision-engine card in the dsh settings page, shown in Chinese: switch the engine, tick which local CLIs auto mode reuses
https://raw.githubusercontent.com/liustack/modlens/main/assets/demo-dsh-settings-card.jpg
Install
dsh plugin --profile web add @liustack/modlensTags
README badge
Add this Markdown to your plugin README to link back to its listing.
[](https://dshget.com/plugins/liustack/modlens)Related plugins
Vision & Multimodalmodsearch
★ 321Web search bridge for text-only agents: ask the web or X, get structured JSON evidence (search, fetch, citations).
agent-vision-toolkit
★ 1121为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
dsh-vision-toolkit
★ 842Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.