Z.ai Coding Plan silently rejects images: how a flat-rate GLM plan rerouted our agent to another vendor

ai zai glm api agents llm

Z.ai’s GLM Coding Plan quietly rejects image input. If your agent session ever puts a screenshot into its history, every later call to the model fails with HTTP 400 — and your framework silently falls back to another provider. I hit this today in production. Here is exactly what happens, why it is easy to miss, and what to do before you design a workflow around this plan.

The two GLM-5.3s and the two endpoints

The naming trap first. There are two sibling models:

  • GLM-5.3 — the full model. Text-only by design: 1M context, text in, text out.
  • GLM-5.3-Flash — natively multimodal. A real vision encoder (24-layer ViT), accepts images and video through standard image_url content blocks, same message format as any OpenAI-compatible vision API.

We run GLM-5.3-Flash as our primary model, with supports_vision: true in the agent configuration. And the model does see images — on the standard endpoint.

That distinction is the whole story:

Endpoint GLM-5.3-Flash with image parts
api.z.ai/api/paas/v4 (standard, pay-as-you-go) accepted
api.z.ai/api/coding/paas/v4 (Coding Plan) rejected — HTTP 400

Z.ai’s own documentation shows vision examples with the Flash family on the standard chat-completion API. The Coding Plan endpoint accepts text content parts only. Anything else — an image_url block, a base64 image — bounces at the door with:

HTTP 400 — error code 1210:
"messages.content.type is invalid, allowed values: [’text’]"

The model never sees the request. This is a plan limitation, not a model limitation.

How it bites: the silent fallback cascade

The failure mode is worse than a simple error, because of how agent frameworks handle provider fallback. Our sequence today:

  1. A long-running agent session (a product-tracking workflow) processed screenshots. Vision calls went through the framework’s vision tooling; the results were text, but the image content parts landed in the session history as standard multimodal message blocks.
  2. The next regular model call — plain text turn — carried that history to the Coding Plan endpoint. Z.ai answered 400. Not "retry later": a request-format rejection.
  3. The agent framework (Hermes, in our case) treats a 400 as non-retryable and switches to the configured fallback provider. Ours is DeepSeek — which happily accepts the multimodal history.
  4. From that moment, every subsequent call in that session fails on GLM and lands on the fallback. The session never returns to the primary model, because the image parts stay in the history forever.

Nothing screams. The work continued. The only visible trace was one quiet log line:

Model fallback: glm-5.3 via zai unavailable (request format rejected); using deepseek-flash via deepseek.

If you do not read your agent logs, you would never know your "GLM Coding Plan" session is now running on a different vendor — on a pay-per-token provider, while you believe you are on the flat-rate plan.

Why this is easy to miss

  • The first vision calls may succeed. In our case the screenshots were processed early in the session through separate tool paths. The breakage appeared only when image parts entered the main chat history — hours later.
  • The error is about format, not capacity. You will search for "1210", find nothing about quota or load, and conclude your key is fine. It is.
  • New sessions work. A fresh session with a clean, text-only history talks to GLM happily — until the first image part is baked in. That makes the bug look random.
  • Cost drift is invisible. The fallback provider bills per token. A heavy agent session quietly burning DeepSeek tokens instead of your Coding Plan flat rate is a budget leak, not an outage.

What to check if you are on (or considering) the Coding Plan

  1. Ask whether your workflow needs images in the conversation history. If your agent only analyzes images through separate vision tools and never embeds them in chat history, the Coding Plan is fine.
  2. Watch your fallback logs. One request format rejected line means the session is now on another vendor. grep -i fallback in your agent logs is a one-minute audit.
  3. Plan the boundary. Image-heavy work and Coding Plan work do not mix in one session. Split them: a fresh text-only session for coding tasks; vision work in sessions served by a vision-capable endpoint or provider.
  4. Budget for the fallback. If your framework’s fallback is a pay-per-token API (ours is), an image-poisoned session becomes a metered session. That may still be worth it — but it should be a decision, not an accident.
  5. If you need GLM vision inside long sessions, that is standard-endpoint territory (/api/paas/v4), which means normal API billing, not the flat Coding Plan. Price both before committing.

The takeaway

GLM-5.3-Flash understands images. The Coding Plan endpoint does not carry them. "Multimodal model" and "multimodal plan" are two different purchases, and the second one is not what the first one’s marketing suggests.

Test it before you build: send one image-bearing request through the exact endpoint and key you plan to use, on day one. It costs one API call to find out what cost us a silently re-routed production session.


Tested 9 October 2026 against api.z.ai/api/coding/paas/v4 with glm-5.3-flash: image content parts → HTTP 400, code 1210; identical history without image parts → 200. Framework: Hermes Agent with DeepSeek configured as fallback.