The Leak Is Already Inside the Building
Sam Lester
Short on time?
Get a quick summary in your LLM, or watch this article as a Hyperframe.
Every enterprise AI video pitch I've sat in eventually reaches the same turn. The picture looks great, everyone agrees the picture looks great, and then someone from legal or security asks a question the demo has no way to answer. Where was this model trained. What happens to the brief we just typed into it. Who pays if the output borrows something it shouldn't have.

None of those are questions about video quality. They're questions about rights and data, and they're the ones that actually decide whether a tool gets bought. A demo is built to settle how the video looks. Everything else, it leaves for someone else to worry about later.
The leak is already inside the building
The data question isn't hypothetical, and it isn't waiting on some future rollout. It's already happening. WRITER's 2026 enterprise survey, run with Workplace Intelligence across about 2,400 C-suite leaders and employees, found 67% of executives believe their company has already had a data leak or breach from unapproved AI tools, and 79% say they're struggling to adopt AI at all.
The mechanism behind that first number is simple. When an approved tool is slow to arrive, people reach for whatever's fastest. They paste a client brief, a draft contract, a set of unpublished figures into a consumer AI tool whose terms may quietly train on everything it's handed. Every upload is an exposure, and it happens below the level any procurement process can see. The cost of not having an approved path isn't that people avoid AI. It's that they use the wrong one, with your data.
A demo can't touch this. The exposure isn't in the output. It's in what your staff feed the thing to get the output.
Nobody can tell you what the model learned from
Provenance is the question legal asks first, and most generative video tools can't answer it. A model trained on an open scrape of the web has, by definition, no clean account of what went into it. Ask who owns the output, or whether anything in the training set was licensed, and the honest answer is that nobody can fully say. For a global professional services firm putting its name on a client-facing video, "we can't fully say" is where the conversation ends.
And the hidden version of this isn't hypothetical. In the ongoing copyright case against Meta, court records reported by Tom's Hardware allege staff torrented nearly 82TB of pirated books to help train the company's Llama models. Picture how ordinary that is: an employee, a work laptop, a torrent client left running in the background. Meta disputes the claim and the case is still open. But the reason we know about it at all is the tell. It surfaced because litigation forced the documents into daylight, not because anyone could inspect the training set from outside. You can't audit what a vendor won't show you.
Get the next piece by email.
One email when we publish. Unsubscribe anytime.
This is the part that doesn't improve when the model gets better. A more convincing clip generated from an unaccountable training set is still generated from an unaccountable training set. The deeper version of that problem is that prompt-to-video removes the human who would catch a bad output before it ships. But even upstream of that, it removes the ability to say where the material on screen came from in the first place.
Sitting right behind provenance is liability. If the output does reproduce something it shouldn't, who eats the cost. The demo has no answer, because indemnification isn't a feature you can show on a screen. It's a clause in a contract, and most tools don't offer one.
What an answer you can actually check looks like
A real answer to the rights questions is specific, and survives a lawyer reading it closely. Adobe Firefly is the clearest example, worth studying even if you never touch it. Firefly is trained only on content Adobe has the rights to (Adobe Stock, openly licensed and public-domain material) rather than an open scrape of the web, and Adobe offers IP indemnification for generated content on qualifying enterprise plans.
That's two concrete things: a named training-data policy you can point to, and a contractual backstop for when the policy fails. Neither is a "brand-safe" label. Both are the kind of answer legal and security are asking for. The lesson isn't that Firefly is the only tool that clears review. It's that clearing review takes this shape, a policy plus a backstop, and any tool that can't produce both is asking your legal team to sign off on trust.
The provenance problem you don't have
The cleanest way to answer the rights questions is to never create them. Build video from footage a brand has already licensed and approved, and provenance is settled before any model runs, because nothing new is being scraped or generated from an unknown training set. The clips were cleared long ago. AI handles the scripting, the structure and the matching of a message to the right approved shot, and the rights status of what appears on screen never changes.
That's a narrower approach than generating pixels from a prompt, and it gives up the party trick of conjuring a scene that never existed. What it buys is the one thing that matters at the point of sale: a legal reviewer can trace every frame to something the company already owns the right to use. Most enterprises are already sitting on exactly that, a DAM full of approved, licensed footage they've paid for and barely use. The rights problem the rest of the category is trying to indemnify its way out of is one they solved years ago, just by owning their material.
The demo will always look impressive. It was built to. The questions that decide the purchase were never about how it looks.


