Skip to main content
kRouter
All posts
Fix an error

Screenshots and text-only models like GLM-5.3: what works

GLM-5.3 can't read a pasted screenshot. Three fixes -- a model that reads images, a vision MCP server, a per-turn switch -- and where each falls short.

Kodelyth · The team behind kRouter
· Updated
9 min read

You run Claude Code on GLM-5.3 because it is good at code and cheap. You paste a screenshot of a broken layout, ask what is wrong, and one of two things happens. The request fails with a 400 that complains about the image block, worded differently by every text-only backend. Or it succeeds, and you get a confident answer about a screenshot the model never saw. Somewhere between your client and the model the image was dropped, and the model answered the text that was left.

Neither is a bug in GLM. Z.ai's model page says GLM-5.3 "currently supports text-only inputs". Check the model card of any cheap model before you build screenshots into a workflow.

There are three ways to get screenshots working again, and each one fails somewhere different.

1. Use a model that reads images

The simplest fix is to paste into a model that takes images. On the GLM Coding Plan that is GLM-5.3-Flash, which Z.ai lists with video, image, text and file input and a 1M context window. Claude, GPT-5 and Gemini models read images too.

The cost is the model choice itself. If you picked GLM-5.3 for how it handles your repository, moving every turn elsewhere for the occasional screenshot is a poor trade.

2. A vision MCP server: describe, then reason

An MCP server can give a text-only agent an image tool. The agent calls the tool with a file path, the server sends the image to a vision model, and your text model works from the text that comes back. Z.ai publishes one for GLM Coding Plan subscribers, with tools such as extract_text_from_screenshot, diagnose_error_screenshot and ui_diff_check. Community servers do the same with other vision models.

It works in any MCP client, with any text model, and nothing changes in the request path. Where it falls short:

  • The agent has to decide to call it. Z.ai's own page says that, outside Claude Code, pasting an image sends it straight to the model rather than to the server, so you save the image and give its path.
  • The text model reasons from a description. If the describer leaves out the detail that matters, a 2px misalignment or one wrong digit in an error code, the text model never learns it was there.
  • It is one more tool call per image.

3. Switch that one request to a vision model

The third option sits in the request path. When the newest user message holds an image and the target model cannot read images, a router between the agent and the model can send that one request to a model that can. You paste the screenshot as usual and change nothing in your prompt.

kRouter calls this the capacity adapter; the dashboard shows it as Vision Adapter. Here is how it decides:

  • It reads only the newest user turn, looking for Anthropic image blocks, OpenAI image_url parts and Responses input_image parts. Those are what the Messages, Chat Completions and Responses endpoints carry, so it works for Claude Code, Codex, OpenCode, Cline or a script.
  • It checks the target against its capability table. GLM-5.x models are listed as text-only there, so an image turn to glm/glm-5.3 qualifies.
  • It tries the Vision pool first. Pool models go in order, and your original model stays at the end as the last fallback. If the target is a combo that already contains an image-capable model, kRouter skips the pool and moves that member to the front for this turn.

When it fires, the server log shows a line like this:

[CHAT] Capacity adapter for [vision] on "glm/glm-5.3" → trying <your-vision-model>, glm/glm-5.3

For OpenCode, kRouter's OpenCode card writes each model into opencode.json with image input declared, so OpenCode treats those models as image-capable and passes pasted images through.

Setting it up in kRouter

The Vision pool is on by default in kRouter 0.5.163, but out of the box it does not help. An empty pool falls back to mmf/mimo-auto, Xiaomi's free MiMo Auto channel, and Xiaomi ended MiMo Code's free phase on 26 July 2026 (report). The endpoint now answers Unsupported model mimo-auto. So with the default pool, a screenshot turn tries MiMo, fails, and falls back to your text model with the image still attached: the original problem plus a wasted request.

Add your own model. Install or update kRouter and start it in the foreground, with logs:

npm install -g @sifxprime/krouter
krouter -l
  1. Open http://localhost:20128/dashboard, go to Providers and connect an account that serves a model that reads images.
  2. Open Combos and scroll to Vision Adapter. Leave the Vision row switched on, click Add Model and pick that model. If the row lists mmf/mimo-auto, remove it once yours is in. Add a second model for a fallback chain, or turn on Round to spread image turns across them.
  3. Paste a screenshot into your agent and check the log. Your model should be the first one after "trying".

The Add Model list shows only models that kRouter's capability table marks as image-capable. In 0.5.163 the table misses some obvious picks: GLM-5.3-Flash and OpenCode Go's deepseek-v4-flash-vision-exp both read images but are listed as text-only, so neither appears. A Claude or Gemini model works, or ocg/kimi-k2.6 if you are on OpenCode Go. The table can be wrong in the other direction too, so test the pool model before you rely on it.

To test the route without an agent, send a screenshot from disk. Put your own text model in place of glm/glm-5.3. Local requests need no key unless Require API key is on (on the Endpoint page) or REQUIRE_API_KEY=true is set; then add -H "Authorization: Bearer <your-key>" to the curl command:

node -e '
const data = require("fs").readFileSync(process.argv[1]).toString("base64");
process.stdout.write(JSON.stringify({ model: "glm/glm-5.3", max_tokens: 300, stream: false,
  messages: [{ role: "user", content: [
    { type: "image", source: { type: "base64", media_type: "image/png", data } },
    { type: "text", text: "Describe this screenshot in one sentence." } ] }] }));
' screenshot.png | curl -s http://localhost:20128/v1/messages -H "Content-Type: application/json" -d @-

A description that matches the picture means the switch worked. It is a one-message test, so it cannot show the long-session problem described below; for that, paste a screenshot deep into a real agent session and read the log.

Two variations. To keep the switch per combo instead of global, put a recognized vision model in the combo after your text model. kRouter moves it to the front only on image turns; on other turns it is an ordinary fallback. The reverse case is a main model that reads images while the table says it does not, GLM-5.3-Flash today. The adapter still sends its image turns to the pool, so switch the Vision row off to keep them on GLM-5.3-Flash.

What the vision model sees, and what happens next

Five details decide whether the switch helps in a long session.

The pool model sees a trimmed conversation. For the switched request, kRouter keeps the system prompt, the first six messages of the session and your screenshot turn. Everything in between is dropped, whether or not it would fit. After an hour of work, the vision model has not seen that hour. Make the screenshot turn stand alone: name the file, say what you expected and what you see.

The trim can break the request. If the sixth message is a tool call, its result is one of the messages dropped, and in an agent session the opening messages usually alternate tool calls and results. Anthropic's API rejects a request in which a tool call is not followed by its result, so a Claude model served by Anthropic fails there. kRouter moves on to the next pool model and, at the end, to your text model with the image attached. The log shows the pool model failing and kRouter trying the next one. A vision model placed in a combo is not trimmed: it gets the full history, which makes the combo variation the sturdier setup for agent sessions.

Only that request switches. When the agent runs a tool and sends the result back, the newest user message is the tool result, not your screenshot, so the next request goes to your text model. The vision model reads the image and picks the first step, and your text model does the rest, knowing the image only through what the vision model wrote.

Images that arrive through a tool are not caught. Give Claude Code a file path instead of pasting, and it opens the file with its Read tool. The image comes back inside a tool result, which kRouter does not inspect, so that request goes to your text model with the image in it. Paste or drag the screenshot instead.

The old image stays in the history. kRouter does not remove it before later requests. If the text model's provider rejects a request with an error kRouter recognizes as "no image support", kRouter strips the images, retries once on the same account, and keeps stripping them for that model for an hour. An error it does not recognize goes back to your agent.

What does not work

Pasting into a text-only model and trusting the answer. If the image was dropped on the way, the reply describes a picture the model never saw, and the model cannot tell you so.

Leaving the Vision pool at its default. It looks configured, but the only model in it no longer answers.

One more thing to weigh: a switched screenshot goes to the pool model's provider, not your text model's. If your screenshots can show secrets or customer data, pick the pool model with that in mind.

Choosing

Your situationBest fit
Your main model reads imagesNothing to add. Behind kRouter, if the log shows a switch anyway (GLM-5.3-Flash in 0.5.163), turn the Vision row off
Text-only model, occasional screenshots, any MCP clientA vision MCP server, given file paths
Text-only model behind kRouter, short sessions or one-off questionsAn image-capable model in the Vision pool
Long agent sessions, or the vision model needs the whole conversationA combo with the vision model after your text model
The agent opens image files itselfA vision MCP server, or paste instead

Common questions

Can GLM-5.3 read screenshots?

No. Z.ai documents GLM-5.3 as text-only input. GLM-5.3-Flash, on the same Coding Plan, accepts images, video and files.

Why did my agent answer as if it saw the image?

The image was dropped before it reached the model, by the client, a translation step or the provider, and the model answered the text that was left. A text-only model gets no signal that a picture was removed, so it rarely says so.

Does kRouter's vision switch work without setup?

Not in 0.5.163. The Vision pool is on, but its built-in fallback is Xiaomi's free MiMo Auto channel, whose free phase ended in July 2026; it now returns "Unsupported model". Add an image-capable model from one of your own providers under Combos, then Vision Adapter.

Does it work with Codex, OpenCode or Cline, or only Claude Code?

Any client that sends the image in the request to the Anthropic Messages, OpenAI Chat Completions or Responses endpoint. Two exceptions: images an agent loads through a tool call are not inspected, and kRouter's Gemini-format endpoint forwards only the text of each message in 0.5.163, so an image sent that way never reaches any model.

Is a vision MCP server or the router switch better?

The router switch needs no new habit, and the vision model answers from the pixels. Through the Vision pool it sees a trimmed history, which can fail in long agent sessions; through a combo it sees the full history. Either way it handles only the request you pasted into. An MCP server works with any client and can be called at any point, but the agent has to choose to call it, and your text model works from a description. You can use both.

Kodelyth · The team behind kRouter

Published by Kodelyth, the team that builds kRouter. Posts are drafted with AI assistance and reviewed by a person before they go out. kRouter is free and MIT licensed.

Install kRouter