Skip to content

Enterprise AI Pilots Don't Stall on the Model

Sam Lester

7 min read

Short on time?

Get a quick summary in your LLM, or watch this article as a Hyperframe.

The demo goes well. The pilot team is impressed, the champion is sold, the numbers look good enough to write up. Six months later the tool still hasn't moved past the handful of people who ran the trial, and nobody can quite tell you why. Spend any time inside a large professional services firm and you've watched this happen, probably more than once. The model is almost never the reason.

Enterprise AI Pilots Don't Stall on the Model

The gap between a pilot and a rollout is organisational, not technical

What stalls is agreement, not capability. Superside's AI readiness work with Vimeo is a clean look at where the friction actually sits. Rather than ask whether the models were good enough, they ran 15 one-on-one interviews across the creative teams and pulled legal and IT into the same view. The blocker wasn't model quality. It was getting those functions to agree on a single approved toolkit and a responsible-use policy before anyone scaled past the pilot. The work ended with seven vetted tools, not a verdict on whether AI was ready.

The technical proof of concept had already worked. The agreement on how to use it hadn't.

That pattern holds well beyond one case. A pilot runs inside a single team, with a champion who wants it to succeed and enough autonomy to route quietly around the questions a wider rollout can't dodge. Scaling past that team means surviving contact with three groups who weren't in the room.

Legal wants to know who owns the output and what data trained it. IT needs to know whether it fits the existing security posture or triggers a fresh vendor review. Brand or compliance want to know whether it produces consistent, on-brand results across teams that never saw the pilot.

Most pilots don't arrive with those answers ready, because the whole point of a pilot was to move fast and prove value before dealing with any of it.

I know what that gap looks like from the other side of the table, because filling it in is most of what selling a tool into a large firm involves. Before anything gets near production there's the security questionnaire, a spreadsheet that can run to a couple of hundred rows, asking where the data is hosted, whether anything you're handed trains a model, how long it's kept, who at our end can see it and what happens to all of it when the contract ends. Then the documents behind the questionnaire: how the AI actually works, the data-processing terms, the information-security policy, the list of every sub-processor we touch. We've filled these in, more than once, for firms that hadn't yet let a single video out of the pilot team. None of it asks whether the video looks good. All of it has to be answered before anyone will let the tool grow.

The two questions that freeze a review

Reviewers stall on two specific risks, the same two almost every time. Wolters Kluwer's breakdown of generative AI risk names them plainly. Data leakage, where someone feeds confidential or proprietary material into a public model that may retain it and train on it. And hallucination, where the model produces something plausible and wrong, which stops being an annoyance and becomes a liability in legal, financial or medical work.

Their recommended mitigations are worth reading, because they're the exact shape of thing a reviewer wants to see before signing off: restrict what data can go in, classify which sources are approved, and require a human to verify anything high-risk before it's used. That's a checkable process, not a promise. A pilot that didn't build those in has to stop and retrofit them at review, and that pause is usually where the stall actually happens.

Get the next piece by email.

One email when we publish. Unsubscribe anytime.

There's a third reviewer that case studies tend to underweight, and in my experience with large firms it's often the decisive one. Brand review can sit on a piece of work indefinitely, even when the guidelines were followed to the letter, because the reviewer's job is to catch the thing nobody anticipated, not to wave work through. The people pushing to use video are almost always on the ground in sales and marketing, the ones who feel the pain. The people who can clear it for the whole firm sit somewhere else entirely, and don't personally use the tool. A pilot that only convinced the enthusiasts hasn't touched the part of the org that actually says yes.

Why this gets blamed on the technology

When a pilot stalls, "the tool wasn't good enough" is the easiest story to reach for, because it's simpler than admitting IT, legal and brand were never aligned on a shared policy before anyone started. It's also more comfortable. It points at a vendor choice rather than an internal process gap, and nobody has to own it.

The trouble is that swapping tools almost never fixes a stall caused by review friction. The next tool walks into the same unanswered questions, just later in the cycle. You can watch it happen on a delay. A team runs a second pilot with a different vendor eighteen months on, hits the same legal and IT questions the first one never resolved, and lands in roughly the same place. From the outside that reads as two failed technology bets. From the inside it's one unanswered question surfacing twice, because switching vendors did nothing to answer it.

Design for the review, not just the demo

The way out is to build the review answers into the tool before a pilot starts, rather than treating them as a cleanup job for after it succeeds. That's a design decision, and it's mostly about narrowing what the tool is allowed to do.

Default to approved, licensed assets instead of open-ended generation, and the IP question has a clean answer before legal thinks to ask it. Keep a human review step in the workflow by default, and the hallucination and brand-consistency worries have a built-in answer rather than an improvised one. Neither makes for a flashier demo. Both make for a much shorter review.

None of this comes from getting it right every time. We've run pilots that fizzled out, same as anyone. The demo lands, everyone's keen, and then the thing quietly stops, and six months on nobody's using it and you can't honestly pin it on the tool. Sitting in that a few times taught us more about where pilots actually die than any demo ever did.

So the tool we build now leans the other way on purpose. The approved library and the human sign-off aren't procurement features we bolted on late, they're how the thing works, and because every scene is a clip a person chose from that approved library, the record of what went in exists without anyone having to build one. That's the whole trade we're making: efficiency with a bit less risk, fewer of the questions that stall a pilot still open by the time it reaches review.

This is the same split we wrote about in the realism arms race piece: the criteria that win a demo and the criteria that clear a review have almost nothing to do with each other. Tools built for the second set from the start tend to clear enterprise review faster, and not because their model is more advanced. It's because they were designed for how sign-off actually works inside a large organisation, instead of assuming a good pilot would carry them through the rest of the building. It won't. It never has.

See this article become a video.

Write the brief behind this article and preview the video free, no card needed. Or run it past James on a call.

More from Hyperframe

Made for the Schedule, Watched by Nobody

Made for the Schedule, Watched by Nobody

Video made because the comms calendar said one was due rarely gets watched. Wistia's data shows instructional content holds nearly double the attention.

Sam Lester
The Leak Is Already Inside the Building

The Leak Is Already Inside the Building

Enterprise AI video pitches stall on questions a demo can't answer, like where the model trained and who's liable if it copies something.

Sam Lester
The Trust Test a Synthetic Presenter Can't Pass

The Trust Test a Synthetic Presenter Can't Pass

Synthetic presenters keep getting more realistic. That isn't the same as more trusted, and for client-facing video the two are pulling apart.

Sam Lester