面向纯文本模型的免费视觉桥与生图:粘贴读图、GLM-4V-Flash 与 Gemini 引擎故障转移、modlens 同款结构化证据输出,并自动播种免费视觉模型路由。
- GitHub 星标
- 16
- 许可证
- MIT
- 收录日期
- 2026-08-18
GitHub 信息
- GitHub 星标
- 16
- 许可证
- MIT
- 主要语言
- Python
- 最近推送
- 2026年8月21日 05:30
- 维护者
- MJorgin
- 收录日期
- 2026-08-18
README 演示图
从仓库 README 提取的截图/GIF(已过滤徽章、头像等装饰图)。

Demo: paste an image into a text-only DeepSeek Harness session, the vision model reads it, and the model answers; the same bundle can also generate images
https://raw.githubusercontent.com/MJorgin/dsh-media-skills/main/docs/screenshots/demo-paste.png

How paste-image reading works: paste → vision model describes → text description arrives at the current model
https://raw.githubusercontent.com/MJorgin/dsh-media-skills/main/docs/screenshots/how-it-works.png

dsh-media-skills — free image reading & generation for DeepSeek Harness
https://raw.githubusercontent.com/MJorgin/dsh-media-skills/main/docs/social-preview.png
安装
dsh plugin --profile web add github:MJorgin/dsh-media-skillsREADME 徽章
将这段 Markdown 添加到插件 README,链接到对应的详情页。
[](https://dshget.com/plugins/MJorgin/dsh-media-skills)相关插件
视觉与多模态modlens
★ 3768为纯文本模型架起视觉桥梁:粘贴图片,输出结构化 JSON 证据(OCR、版面、语义)。
dsh-vision-router
★ 1030为纯文本 Agent 提供视觉能力:内置免 Key 视觉链 + 像素级视觉工具(看图问答、定位、裁剪、像素对比、取色、OCR、矢量化、抠图、截图);粘贴图片即可用。
dsh-vision-toolkit
★ 842让纯文本模型处理视觉任务:粘贴图片后自动切换到 Vision Toolkit 变体,支持图片问答、多图比较、长截图 OCR、截图还原前端 UI、元素定位与像素对比。默认无需 API Key——图片经作者自建的免费服务处理,每台机器每天 100 张;也可改为指向自己的服务商。