Cinematic radar, a side quest — what generative-video models and tool-bounded agents diverge on
Asking Gemini Veo to make a radar widget cinematic produced one destruction-arc video after another. Here's why generative-video priors do that, and what tool-bounded agents do differently.
Cinematic radar, a side quest — what generative-video models and tool-bounded agents diverge on
The side quest
You can open the public Gemini share here and watch it for yourself. The setup was deliberately small: a still screenshot of the live radar widget, and a prompt asking the model to expand the static graphic into something more cinematic. No editorial direction beyond "cinematic." No request for a destruction arc, no request for drama, no request for a narrative. Just cinematic.
The pattern across the generated videos was consistent: each one started faithful-ish to the static frame — a recognizable radar sweep, recognizable cloud cover — and then drifted, over four to eight seconds, toward damage. Flooding streets. Lightning fork onto a power line. Wind tearing roofing material loose. Buildings buckling. The model couldn't seem to extend a weather radar graphic into "cinematic" without ending the world.
This is funny, and the share is worth opening for the comedy alone, but the architectural reason it happens is what the post is actually about.
Why "cinematic weather" is a destruction arc
Generative video models like Veo learn from a corpus where the phrase cinematic weather footage and the phrase destruction footage live very close together in latent space. Storm chaser videos, news broadcast B-roll, climate-disaster trailers, insurance-claim reels, blockbuster opening credits — that's the training data for "weather is dramatic." If you're optimizing a 4-to-8-second video against that prior, the gradient pulls hard toward damage frames. The narrative arc is baked in.
The image grounding mitigates this only weakly. The model honors the input frame for the first second or two — radar tile shapes, color gradient, time-of-day lighting — and then the generator's job becomes "what should happen next?" The "next" prior is not "tile-faithful continuation"; it's whatever the model has learned makes a 4-second clip cinematic. And cinematic-weather-prior-given-radar-input ends in destruction.
To be very specific about the claim: I'm not saying Veo always generates destruction. I'm saying that in this specific experiment, with this specific prompt and this specific input, the prior dominated. Run it with a different prompt ("a calm time-lapse of a clear sky," say) and the prior changes. The lesson isn't "the model is broken" — it's "the prompt is the steering wheel and the prior is the road grade. You're going to slide downhill no matter how you hold the wheel."
The contrast — the radar-v2 build, same week
The radar-v2 build happened in parallel. Two days. Fifteen TDD tasks. About thirty commits. Eight master states, four per-frame states. A scrubber that knows which frames have failed. A diagnostic surface that draws itself.
Across that whole arc, no part of the agent producing the work tried to "make it cinematic." The temptation didn't exist, because the agent's tools didn't permit it. The reducer is (state, event) → state. The test runner cares about whether expect(reducer(s, e)).toEqual(s') passes. The mermaid generator cares about whether the diagram parses. The build runner cares about whether runner_build() returns success. None of these tools have a "make the output more dramatic" channel, and none of them have a destruction prior buried inside them.
The system prompt isn't enforcing faithfulness. The tool set is.
What's actually different
Both Gemini-generating-Veo-output and Claude-driving-the-radar-v2-build are agents, and both are "AI" in the loose marketing sense. But they sit in profoundly different positions on the prior-vs-tools axis:
| Property | Gemini → Veo cinematic-radar | xl-dev-agent on radar-v2 |
|---|---|---|
| Output channel | One giant pixel-emitting function | Narrow tools (Bash, Edit, Read, Grep, MCP build/deploy/secrets) — each tightly typed |
| Verification | None: the output is opaque to the model itself | Constant: tests, type checks, build status, eyeball-the-snapshot |
| What dominates behavior | The training prior on "cinematic weather" | The tool set's accept/reject signals |
| Failure mode | Plausible-looking output that drifts from input | Loud, specific errors at the tool boundary |
| Source of "faithfulness" | Politeness from the prompt, mostly | Tools that refuse misuse |
The interesting line is the bottom one. Faithfulness in a tool-bounded agent isn't earned by being well-prompted; it's earned by being unable to do the unfaithful thing. The reducer can't lie about its previous state. The test can't pass while the bug is present. The MCP build can't deploy a broken image past the canary health probe (well, in theory — but the kind of failure mode is "the deploy didn't happen" not "the deploy happened and stayed quiet"). Each tool is a place where bad outputs get caught.
A generative video model, by contrast, has no such places. There is one tool — emit pixels — and it is allowed to emit any pixels. The prior is the only steering signal, and the prior was trained on a corpus that says cinematic weather is destructive weather.
What this means for builders
The takeaway isn't Veo bad, Claude good. Veo is excellent at the job it has — emitting plausible video against a prior. The takeaway is don't ask a single-channel generative model to extend a dashboard. Dashboards encode constraints (the radar gradient means something; the tile timestamps mean something), and a generative model has no tool for honoring constraints it can't observe.
If you need an agent to extend a dashboard faithfully, you need:
- A tool that can read the dashboard's underlying data, not just its rendered pixels. (For radar-v2, that's the RainViewer JSON API. For Grafana, it's the Prometheus query. For an industrial HMI, it's the OPC UA tag list.)
- A tool that can verify the extension against the data — a test, an assertion, a constraint solver, even an eyeball.
- A tool to emit the extension in the right format — markup, code, parameter changes — not a giant raster blob.
When all three exist, the agent's prior gets constrained at every step and the output stays faithful. When they don't — when the only tool is emit pixels — the prior wins and your dashboard ends in flames.
What's been worth saying out loud
- Agent behavior = system prompt × training prior × tool set, with the tool set's weight scaling fast as the tool set narrows. A one-tool agent is a prior in a costume.
- "Cinematic" is a steering wheel; the training corpus is the road grade. If the road tilts toward destruction in your genre, the prompt won't stop the slide — only different tools can.
- Faithfulness in a tool-bounded agent is a property of the tools, not of the prompt. That's a much sturdier place to build from than "ask nicely."
- A funny side quest can teach as much as a 15-task plan. The radar-v2 plan taught us how the build pattern works. The Gemini side quest taught us why the build pattern is necessary.
What's next
I'm going to keep the radar widget non-cinematic. The radar at weather.kraftware.dev is rendered, not generated — every tile a real fetch from a real provider, every frame timestamped, every animation a function over the timeline. None of it is dramatic. That's the point.
If you want to verify the pattern yourself, open the share. Try it with your own dashboards if you want to see the prior-vs-tools axis in motion. The fastest way to learn the difference between generating and extending is to ask a generative model to extend something where the constraints actually matter.