Skip to content

Why Prompt-to-Video Has Bigger Problems than Spaghetti

Sam Lester

7 min read

Short on time?

Get a quick summary in your LLM.

We sell AI video. So take this for what it is: an argument against a chunk of our own category, made because we think it's true, not because it's comfortable to say.

Why Prompt-to-Video Has Bigger Problems than Spaghetti

Fully generative prompt-to-video, the kind where you type a description and a model hands back a finished clip, is the wrong tool for regulated, reputation-sensitive organisations, law firms and global accountancy firms we work with among them. Not because the technology is bad. It's improving fast, and it'll keep improving. The problem is structural. Prompt-to-video removes the one thing a regulated brand actually needs from its video: a human who understands that brand, in the loop, before anything goes out with the firm's name on it.

That's the whole argument. Everything below is why the usual counterpoints miss it.

The meeting that convinced me this needed writing down

I was making this exact case to a senior stakeholder at a large firm, someone sharp and genuinely well-read on the space. Every few minutes she came back to the Will Smith spaghetti test, the benchmark clip people regenerate each year to show how far the models have come.

She was right about the progress. It's real and it's remarkable. But she never engaged with the actual point, which is that a model's raw generation ability and whether generation serves the business outcome are two different questions. The pasta renders better every year. The question of who checks the output before it carries your firm's name hasn't moved at all.

The failures are obvious once you've tested for them

Anyone who's spent real time with these tools has their own list. Ours starts with hands. Extra fingers, wrong counts, limbs that resolve strangely on a close look, the kind of thing that's easy to miss on a first watch and impossible to unsee after that. On the current leading models this still isn't solved. Any close shot of hands doing something fine-motor regularly produces artefacts, which is exactly the sort of detail a client notices in a video with your name on it.

Text is worse, because it fails twice. On-screen signs and labels come back garbled or illegible more often than they should. And when you regenerate a clip to fix some unrelated detail, the text frequently changes again, sometimes into something worse than where you started. At that point you've stopped editing a video and started gambling on one.

Then the one brand teams care about most: logos. A brand mark has to be exactly right, exact colours, exact proportions, every single time, and generative models don't hold that reliably across frames.

None of this is anecdotal grumbling. A 2024 survey of AI-generated video evaluation sorts these failures into distinct categories: consistency errors, where an object drifts unexpectedly from frame to frame, quality errors that cover unreadable fine detail like text, and physical errors where the footage ignores gravity. These are overlapping ways for a clip to go wrong, not one bug waiting on a patch. A regulated brand only needs one of them to surface once, in one client-facing video, for the whole exercise to backfire.

The market has noticed. IAS and YouGov's 2026 Industry Pulse Report found 83% of media experts see the rise of AI-generated content as a significant concern that needs careful monitoring. Audiences and buyers are both getting warier of visibly synthetic content at the exact moment more of it is being produced. The lesson isn't to avoid AI video. It's to be deliberate about which parts of the process you let it touch.

Get the next piece by email.

One email when we publish. Unsubscribe anytime.

The cost maths doesn't help either

Generative video looks cheap until you price in how it's actually used. At the premium end, Google's Veo 3.1 runs around $0.40 per second for full-quality output with audio, which puts a finished two-minute video near $48 in raw generation cost. That's before a single retry.

And retries are the default, not the exception, for anything that has to look professional. Get the logo wrong, get the text wrong, get a stray hand wrong in the background of a shot you weren't even watching, and you regenerate. A cheap two-minute video quietly becomes an expensive one, with no guarantee the next attempt is clean.

The supply side looks shaky too. When OpenAI wound down Sora this year, reporting put its compute cost near $1 million a day. That's a reported figure, not an audited account, so apply the usual caution. Even so, when one of the best-funded AI companies in the world can't make consumer video generation pay for itself, "raw generation is basically free now" deserves more scepticism than it tends to get.

We won't pretend Hyperframe does an identical job for less. It solves a different problem: turning footage a brand has already licensed and approved into something new, rather than generating pixels from nothing.

What's actually missing is governance, not polish

Strip away the artefacts and the pricing, and the real gap is who's accountable. NIST's AI Risk Management Framework, the closest thing the US has to a standard for managing AI risk, organises the work into four functions: govern, map, measure and manage. Govern sits at the centre on purpose, because the framework treats it as the thing that has to inform everything else.

Prompt-to-video walks straight past it. There's no step between "type a prompt" and "get a rendered clip" where anyone checks whether the output reflects the brand's values, matches its guidelines, or even tells a coherent story.

At the large enterprises we've worked with, that gap gets patched by hand. A brand or marketing function generates stock-style footage internally, reviews it clip by clip, and colour-grades the lot afterward to make it feel consistent. It's a real, working process, and it looks nothing like what a consumer prompt-to-video tool hands to someone who has no marketing background, has never seen the brand guidelines, and doesn't know how a piece of content reflects on the firm.

Closing that gap is the point of Hyperframe. Rather than asking someone to describe a video and hope, we draft the script from what the brand actually wants said, match scenes against a footage library the brand has already licensed and approved, and apply its own styling on top. AI does the parts it's genuinely good at, and writing is one of them, while a person, or a process shaped by one, still decides what the final thing says and how. Speed and cost come down. The parts that create risk for a regulated brand don't get automated away.

Where this leaves the argument

We don't think fully generative AI video is permanently unworkable. Model quality is moving fast enough that some of today's failures, the hands, the text, the drifting logos, will likely improve. What won't fix itself is the governance gap. Someone still has to decide what a video should say before it says it, and prompt-to-video, by design, hands that decision to whoever happens to be typing.

For a bank, a law firm, a partner at an industry-leading accounting firm, or any regulated, reputation-sensitive organisation putting their name behind a piece of content, that isn't a workflow. It's an unmanaged risk with a rendering engine attached.

See the product build a video, live.

20 minutes with James, Hyperframe's co-founder. Add your brief when you book, and the demo runs on it instead of a showreel.

More from Hyperframe

Your Budget Doesn't Believe You're a Trusted Advisor

Calling yourself a trusted advisor is close to a contradiction. By the Trust Equation's own logic, saying it out loud reads as a self-orientation signal.

Sam Lester
Why Your Best Expert Gives Your Worst Pitch

Why Your Best Expert Gives Your Worst Pitch

The person who knows the most often gives the worst pitch. It's a cognitive bias no presentation training fixes. The fix is a format.

James Keal
Your Video's First Ten Seconds Are Wasted on a Logo

Your Video's First Ten Seconds Are Wasted on a Logo

Psychology research says viewers judge a video in about two seconds. Most corporate videos spend that time on a logo animation, the worst use of it.

Sam Lester