DSH GetDSH Get

Vision & Multimodal

Browse Vision & Multimodal plugins for DeepSeek Harness on dshget.com.

92 plugins

Sort

modlens

3768

Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).

Vision & Multimodalliustack

dsh-vision-router

1030

Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.

Vision & Multimodalysr666

dsh-vision-toolkit

842

Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.

Vision & MultimodalAnionex

dsh-imagegen

40

AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.

Vision & Multimodaldickpy

dsh-comfyui

36

Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.

Vision & Multimodalfandc520

picturereader

35

Image "reading" for text-only models: downscale + reduce color depth + structure/color fingerprints into text grids fed back to the conversation, letting the model zoom, sample and OCR autonomously like a multimodal model; fully local with zero external model dependency, ships an image-reading methodology skill and optional PaddleOCR.

Vision & Multimodaljing-hy

dsh-media-skills

19

Free vision bridge and image generation for text-only models: paste-image reading, GLM-4V-Flash and Gemini engine failover, ModLens-style structured evidence, and a seeded free vision model route.

Vision & MultimodalMJorgin

dsh-vision-proxy

15

DeepSeek brain + automatic image transcription: attach images in the GUI and each one is transcribed via the official deepseek-v4-flash-vision-exp by default (a pure-text V4-Pro brain can see images), with any OpenAI-compatible VLM or local Ollama as alternatives.

Vision & MultimodalFlyvhidbwo

dsh-vision-opencode

13

Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.

Vision & Multimodalpoiuyjie

dsh-visual-plugin

12

Gives text-only models vision: forwards user images to an OpenAI-compatible vision model and shows the descriptions in a Web UI right panel.

Vision & Multimodaljyh20030112

dsh-chat-imagine

11

Automatically generates and displays images in the DSH chat via API channels or local CLIs (mmx / codex / agy), and can also recognize images using the corresponding CLI.

Vision & Multimodalcorrinehu

dsh-vision

11

External vision plugin for DeepSeek Harness: whale-button config panel, image recognition with auto-reply, and agent screenshot/recognize tools.

Vision & Multimodallinenxi-ctrl

Gemini-Eyes

11

MCP bridge to gemini.google.com: vision analysis of images and videos, Imagen image and Veo video generation, and conversation management using the logged-in browser session with no API key.

Vision & MultimodalConsoleSun

dsh-web-ui#dsh-tool-describe-image

10

Gives a text-only model image understanding via a vision-language model, exposed as a `describe_image` tool.

Vision & MultimodalDamonKoy

dsh-deepseek-vision

8

A vision-language gateway provider route: pasted images are described by a configurable VL model (Qwen-VL by default) before the DeepSeek wire.

Vision & Multimodalsiegfly

dsh-free-vision

8

Free vision bridge for text-only models: image understanding, OCR, UI and debug analysis via free-tier providers (Qwen3-VL-Flash, Doubao, DeepSeek-OCR) with a settings GUI.

Vision & MultimodalFuzzySoul

dsh-vision

8

Vision for text-only DeepSeek via Doubao Web by default (zero-cost, no API key — drives your logged-in Chrome through a Windows CDP bridge), with Antigravity IDE quota (flash/pro) or Gemini fallback; auto detail escalation, vision evidence memory with compaction rehydration, content-hash cache, and a bilingual client panel.

Vision & Multimodal54xkeee

dsh-highres-vision

7

For the DeepSeek Harness native vision model deepseek-v4-flash-vision-exp, raises image admission limits to 32 MiB / 8192 px / 600 images and adds a highres_read tool that tiles large images, then returns the whole image plus 800x800 tiles through the host read_image tool.

Vision & Multimodalazwosile

dsh-windows-ocr

7

Local OCR for attached images via the built-in Windows engine (Windows.Media.Ocr): only the recognized text is sent to the model, never the image bytes; vision passthrough is opt-in.

Vision & Multimodalmaxwell-feng

dsh-plugin-appshot

6

Codex Appshots for DSH: capture the frontmost active window via global shortcut and seamlessly mount it into the composer for agent queries.

Vision & MultimodalTaurusWood

deepseek-vision#plugins/dsh-plugin-deepseek-vision

5

Vision MCP and DSH bundle for text-only DeepSeek: analyze_image, analyze_clipboard, compare_images and vision_status tools, a visual settings page, free GLM-4.6V-Flash by default, result caching and rate-limit tolerance; keys stay out of logs.

Vision & MultimodalGOU-GEE

dsh-guide-dog

5

MiniMax-powered multimodal plugin: real-time voice call mode (streaming conversation, floating dock UI), voice mode and mic voice input, plus image/video/music/speech generation and vision inspection tools.

Vision & MultimodalAtropinolTT

dsh-image-pathify

5

Lets text-only models handle pasted chat images, with a native vision experience, batch image viewing, and a built-in OpenAI-compatible analyze_image tool; vision-capable models are unaffected.

Vision & Multimodaldami9527

dsh-vision-bridge

5

Bridges session images to configurable vision providers and returns text-only analysis to eligible DeepSeek Harness model routes.

Vision & MultimodalGXX182

dsh-vision-hub#tool-vision

5

Enhanced vision toolbox: 14 pixel-level vision tools (describe, ground, detect, crop, pixel-diff, OCR, long-screenshot OCR, vectorize, colors, cutout, screenshot, present, materialize, html-screenshot) driven by one OpenAI-compatible endpoint, with clean \[图片: path] bridge markers, content-safety classification and rate-limit auto-retry.

Vision & Multimodalxing666173

free-vision-skill

5

Fully-local image understanding & OCR via macOS Vision Framework: `ocr_image` (text, table layout + coordinates) and `view_image` (scene, faces, QR) — paste multiple images into the web input box or pass path/URL/base64; images never leave your Mac.

Vision & Multimodalniyongsheng

dsh-llm-vision-bridge

4

Native LLM-provider vision bridge: images pasted in the chat are described by a vision model (Qwen3-VL via pi-ai/llama.cpp) and the text description is fed to text-only DeepSeek for the reply — image admission, routing and compaction all run through harness-native mechanisms, with an LRU description cache and 503 retry.

Vision & MultimodalEinskyle

dsh-plugin-multimodal

4

Advertise image paste on text-only DeepSeek routes, describe attachments with a vision sidecar, and leave native vision models untouched.

Vision & Multimodalshinjiyu

dsh-screenshot

4

Zero-dependency screen capture for DSH: instant full-screen shots, or stage your desktop (move/resize any window) and box-select a region — agent-callable capture, path-only delivery. Works with any vision consumer; pair with modlens (optional) for one-call structured evidence.

Vision & Multimodalpaicat1

dsh-vision-mix

4

Combine text, vision, and image-generation APIs into one Mix model with automatic routing: text-only requests go to the chat model, user images and agent screenshots go to the vision model, follow-ups keep using the same session image, and agents can generate or edit images with session-scoped call history.

Vision & Multimodalhaiziyao

dsh-vision-tools

4

Full vision-capability bundle for DeepSeek Harness: a vision_understand tool (OpenAI-compatible vision APIs, free Zhipu GLM-4V-Flash by default) plus paste/drag-and-drop/button entry points for image recognition.

Vision & Multimodalmoon09300731

Deepseek-Continuity

3

Local image, voice, music and SFX generation plus transcription, with pinned identity: characters, animals, objects and actor voices are defined once and reused on every later call, degenerate output (a near-flat image, silent audio) is rejected instead of returned as success, and a generated line can be read back as text so a clone that swallowed its ending becomes visible. The engines unload when idle, and image generation and transcription can each be pointed at an OpenAI-shaped API instead of the local Vulkan backend.

Vision & Multimodallinxuhao

dsh-draw-router

3

Unified image generation router for DeepSeek Harness (DSH): auto-discovers image models from any OpenAI-compatible endpoint, provides draw_image and draw_list_sources tools, supports SenseNova, StepFun, Agnes, Qwen, Flux, SD, Imagen and more.

Vision & Multimodalxiaozhe7772222

DSH-dseyes

3

Native image attachments for text-only DeepSeek in the Web GUI: pasted or dropped images appear as thumbnails in the session, and before dispatch the host reads them with the free Zhipu GLM-4V-Flash vision API (glm-4v-flash fallback chain) and substitutes the description, so DeepSeek answers about the image while the original is kept in history.

Vision & MultimodalOkkay712

dsh-image-gen

3

GPT Image 2 `image_gen` with Codex subscription OAuth by default or explicit API-key mode: developing card, up to three live API partials, durable attachment replay/lightbox/download, text-only model output, and bounded credential-safe requests.

Vision & MultimodalLeemanCheung

dsh-image-vision

3

Image understanding for any DSH model: vision, OCR, grounding, and crop tools with domain presets for histopathology, cell biology, anatomy, clinical images and scientific figures.

Vision & Multimodalxiaoyuink

View all in category (92) →