Skip to content

Nobody at a Bank Asks How Photorealistic Your Video Is

Sam Lester

9 min read

Short on time?

Get a quick summary in your LLM.

We launched Hyperframe at The MarTech Summit in London last November, and I spent two days on the conference floor talking to people from professional services firms, banks, brokers, insurers and other vendors working the same room. I went in braced for questions about how good the video looks. Almost none came. What they wanted to know was who owns the output, what data trained the model, what the review step looks like, and who signs it off.

Nobody at a Bank Asks How Photorealistic Your Video Is

That gap is the whole story of the category right now. Generative AI video demos are scored on one thing: whether a single generated clip looks and moves like real footage. It's the axis enterprise buyers care about least. The technology is overhyped and underrated at once, usually by the same people, depending which corner of it they mean.

This is what we keep seeing selling into these firms, and it lines up with the named data rather than the vendor reels. So here's the honest scorecard for enterprise teams in mid-2026, in three parts.

The verdict, up front

Overhyped: the prompt-to-video demo as the enterprise story. The clip that dominates AI video marketing, one striking shot from a text prompt, is built to impress on the exact axis enterprise buyers care about least. It says almost nothing about consistency across a campaign, rights and training-data provenance, or whether the output survives a brand and legal review. Judging the category's enterprise readiness by demo-reel quality is grading the wrong exam. We made the longer version of that argument in our piece on why prompt-to-video won't work for regulated brands.

Underrated: asset-led, governed workflows, and the instructional-content data nobody quotes. Tools that keep generation inside approved brand and stock assets, hold a human review step in the loop by default, and build around scripting and scene-matching rather than open-ended generation solve the actual bottleneck: the legal, IT and brand questions that stall a pilot before it scales. And the most actionable finding in the whole set isn't a model benchmark. Specific, useful content dramatically outperforms generic content, and the outcome is largely decided in the first few seconds. That's a scripting discipline, not a model capability, and it's solvable now.

Still genuinely unresolved: what "good enough" consistency looks like at scale. No model in this generation has fully cracked character, logo and product consistency across long, multi-video use, and no amount of prompt engineering closes the gap yet. This is the honest technical limitation still worth tracking, separate from governance, and the one place where the next wave of model improvements could actually change the maths for enterprise buyers.

The rest of the piece is the evidence behind each of those three calls.

The demo measures the wrong thing

Start with the overhyped part, because it's where the models have improved most and mattered least. Model quality has jumped, but the leaders have split into specialisms rather than converging on one general-purpose winner, and the split is documented in the vendors' own material. Google's Veo 3.1 is built around natively synchronised audio and 4K output, tuned for high temporal consistency inside a shot. Kuaishou's Kling 3.0 is built around multi-shot storyboarding, up to six camera cuts in a single generation, with element-binding to hold a character steady as the camera moves. Different tools now win on different axes.

Here's the catch they share: strong inside a single clip, still fragile across a longer sequence. Holding a character, a logo or a product steady over many scenes is hard enough that whole secondary tools exist just to patch it. Higgsfield's Soul ID, one of the better-known consistency tools, recommends 20 or more reference photos of a face and accepts up to 80, plus a per-character training pass, just to keep one person looking like themselves across generations. That's the tell. If consistency were solved, you wouldn't need a separate model trained on eighty photos to enforce it.

None of this shows up in a curated demo reel, which is why it's easy to misjudge a vendor from marketing rather than real output. The generative demo answers one question: does a single clip look real. Enterprise buyers are asking a different one. Does the output hold up across dozens or hundreds of videos, made by different people, over months. On that question the model you pick matters far less than the process around it.

Prices are falling. The economics still aren't settled.

Cost and speed improved alongside quality, which is easy to miss under the realism headlines. Per-clip generation has dropped far enough that experimentation is cheap, even when the finished output still needs a human pass. That doesn't solve consistency or governance, and it says nothing about whether the economics work for the companies running the models. OpenAI shut down its Sora app in 2026 after reported compute costs it could not come close to covering with revenue. Falling per-clip prices don't prove the category's economics are settled. The most visible test of that so far failed.

Two different things get called "AI video"

Most category coverage bundles two very different things under one label, and the difference decides almost everything for an enterprise buyer. It's also the hinge between what's overhyped and what's underrated.

Get the next piece by email.

One email when we publish. Unsubscribe anytime.

The first is fully generative, prompt-to-video output: a model producing footage from scratch. The second is AI-assisted assembly, where AI handles scripting, structure, narration and matching a message to footage the organisation already approved, while the visuals come from a brand or stock library it already trusts. The model comparisons above, and nearly all the coverage, are about the first. Most enterprise work actually shipping in front of clients today is the second.

The two carry almost opposite risk profiles. Generate-from-scratch inherits every training-data, consistency and rights question in this piece. Asset-led assembly inherits almost none of them, because the footage was approved before any AI touched it. That distinction decides everything.

Enterprise buyers aren't grading your shots

This is the underrated half, and the published evidence points the same way the room did. Superside's AI-readiness work with Vimeo's creative organisation ran structured interviews across creative, legal and IT rather than a model bake-off, and located more than 20% efficiency potential in specific workflows like ideation and early concepting, not in wholesale content generation.

The real bottleneck it surfaced was organisational, getting legal, IT and creative aligned on an approved toolkit and a responsible-use policy before adoption could scale past a pilot group.

That matches what legal and security teams consistently ask before signing off on any AI content tool: who owns the output, what data trained the model, whether there are commercial-use restrictions, and how the tool handles data privacy. Wolters Kluwer's analysis of generative AI risk names the two that come up most, data leakage and hallucination, the second a serious problem in regulated legal, financial and medical work. Its recommended mitigations (restrict inputs, classify approved sources, require human verification before high-risk output ships) map closely onto NIST's AI Risk Management Framework and its govern, map, measure and manage structure.

The audience data explains why this concern rises rather than fades as models improve. IAS and YouGov's 2026 Industry Pulse Report, surveying nearly 300 US media professionals, found 61% excited about AI's potential in digital media, so this isn't reflexive technophobia. The wariness sits alongside the enthusiasm. The 2025 Edelman Trust Barometer AI poll found US adults more likely to reject AI outright than embrace it, roughly 49% against 17%, with a sharp generational split in the UK, where far more under-35s trust it than over-55s. So the market is excited about AI and uneasy about it in front of clients at the same time. The buyer's question has moved on from whether a model is good to whether using it visibly, with their name on it, is a reputational risk.

That's why the teams who build the governance answers into the tool and the workflow from the start (approved asset sourcing, a defined human review step, a documented data policy) clear that same review meaningfully faster. Not because their model is more advanced, but because the questions legal and IT ask have already been answered by design. Involving those stakeholders during the pilot, even lightly, beats treating their sign-off as a later phase. That's a process change, available to any team today whatever model or vendor they're evaluating.

The most useful finding isn't a model benchmark

The other underrated thing has nothing to do with model quality, and it's the most actionable finding in the set. Companies want more video and aren't funding it to match. Wistia's 2026 State of Video Report, drawn from a survey of 900+ professionals plus analysis of over 13 million videos and 79 million hours of viewing on its own platform, found teams producing more while roughly 46% hold budgets flat year over year. That's the condition that turns faster, cheaper production from a nice-to-have into the point.

Where the video lands has shifted too. LinkedIn is now B2B's leading video channel for these teams, ahead of YouTube, and social engagement has climbed up the list of metrics marketers actually track.

The real prize is buried underneath. Specific, instructional content wins attention fast. Wistia puts average engagement for a three-to-five-minute video at 43%, but instructional content of the same length runs closer to 74%. Engagement holds roughly flat from one minute to five, then drops. Say something genuinely useful in the opening seconds, or the rest of the runtime is academic.

None of that waits on better models. It's a scripting and discipline problem, solvable today with tools that already exist.

Still unresolved: consistency at scale

One line in the verdict isn't a marketing problem, it's a real technical limit. No model in this generation has fully cracked character, logo and product consistency across long, multi-video use, and no amount of prompt engineering closes the gap yet. This sits separate from governance, and it's the one place where the next wave of model improvements could actually change the maths for enterprise buyers. Everything else here is available to any team today. This isn't, at least not yet.

Which is really just what two days on a conference floor already told me. Nobody at a bank has ever asked me how photorealistic our videos are.

See the product build a video, live.

20 minutes with James, Hyperframe's co-founder. Add your brief when you book, and the demo runs on it instead of a showreel.

More from Hyperframe

Your Budget Doesn't Believe You're a Trusted Advisor

Calling yourself a trusted advisor is close to a contradiction. By the Trust Equation's own logic, saying it out loud reads as a self-orientation signal.

Sam Lester
Why Your Best Expert Gives Your Worst Pitch

Why Your Best Expert Gives Your Worst Pitch

The person who knows the most often gives the worst pitch. It's a cognitive bias no presentation training fixes. The fix is a format.

James Keal
Your Video's First Ten Seconds Are Wasted on a Logo

Your Video's First Ten Seconds Are Wasted on a Logo

Psychology research says viewers judge a video in about two seconds. Most corporate videos spend that time on a logo animation, the worst use of it.

Sam Lester