AI & Production

Why Brands Break in AI-Generated Video — and How We Keep Them Locked

There’s a specific moment I’ve watched happen in a dozen client meetings now.

The studio plays an AI-generated product video. It’s beautiful — the lighting is cinematic, the motion is fluid, the whole thing looks like it cost four times what it did. The room is impressed. Then the brand manager asks to see it again, leans in, and says: “wait, go back… is the logo moving?”

It is. And once you’ve seen it, you can’t unsee it.

This is the single biggest unsolved problem in using generative video for commercial work, and it’s the one studios talk about least, because it’s the part that makes AI look less magical. So let’s talk about it properly: why it happens, why it matters more in FMCG than almost any other category, and what we actually do about it.

What “breaking” looks like in practice

When people say AI video isn’t brand-safe yet, they usually can’t name what goes wrong. Here’s the actual list, in rough order of how often we see it:

Text degradation. Your label copy turns into a convincing-looking alphabet that isn’t your alphabet. Arabic copy suffers worse than Latin — letterforms connect wrong, diacritics drift, and the result reads as gibberish to anyone who can actually read it. This happens even when the first frame is perfect.

Logo drift. The lockup subtly changes proportion across the shot. The icon rotates a few degrees relative to the wordmark. A registered trademark symbol appears, disappears, or relocates.

Silhouette instability. Your bottle neck gets 4% longer. The cap changes from a flip-top to a screw-top somewhere around frame 60. A carton’s corner radius softens. For a product with a distinctive, protected shape, this is a genuine problem and not only an aesthetic one.

Color slip. Brand red becomes a slightly different red, and because it shifts gradually, nobody catches it until it’s next to the correct version on the same page.

Detail invention. The model adds texture, embossing, or a highlight that doesn’t exist on the real pack — usually because it’s seen ten thousand similar products that had one.

Individually these are small. Cumulatively they mean the asset can’t be the primary brand asset, which is exactly what a product film usually needs to be.

Why FMCG suffers most

Most creative categories can absorb some drift. A fashion film, a travel piece, an abstract brand mood film — small inconsistencies read as style.

FMCG can’t, for three reasons.

First, the pack is the brand. In a category where a shopper makes a decision in under two seconds at a shelf, packaging recognition is the entire game. A pack that’s 95% right is a pack that doesn’t trigger recognition.

Second, the copy is often regulated. Nutritional claims, net weight, certifications, ingredient callouts. A model that reinvents the small print isn’t producing a stylistic variation; it’s producing a compliance problem.

Third, these assets get reused everywhere. The same hero frame ends up on packaging mockups, trade presentations, e-commerce listings, and printed POS. An inconsistency that was invisible in an eight-second Reel becomes obvious at A1.

The mistake studios make

The instinctive fix is to fight the model. Write a longer prompt. Add “consistent logo, accurate text, do not alter packaging” and hope. Add negative prompts. Generate thirty versions and pick the one where the label happened to hold.

We went through that phase. It doesn’t work, and more importantly it’s the wrong shape of solution. You’re asking a probabilistic system to behave deterministically, and then paying for the variance in generation credits and revision rounds.

The fix isn’t a better prompt. It’s a different division of labor.

How we actually keep a brand locked

Four things, and they’re not clever — they’re just disciplined.

1. The pack is never generated. Ever.

The product itself is CGI. Modeled to the real dielines, with your print-ready artwork wrapped onto the geometry. That’s not nostalgia for the old pipeline; it’s the only method that gives you a pack that is definitionally correct, frame after frame, because it’s the actual artwork file rendered in 3D space rather than a model’s impression of it.

Everything a brand owns and controls stays on this side of the line: geometry, label, logo, color, copy.

2. AI handles everything the brand doesn’t own

This is where generative tools genuinely outperform what we could afford to simulate before: splashes, smoke, drifting particles, atmospheric haze, ingredients tumbling through frame, fabric and foliage movement, background environments, organic chaos of every kind.

These elements have no correct version. Nobody is going to compare your splash against a brand guideline. So the tolerance for variance is high, and the payoff — motion that looks genuinely physical rather than simulated on a budget — is real.

3. Composite, don’t choose

The two halves meet in compositing. The CGI product is rendered on its own layers with mattes; the generated elements are graded, tracked, and assembled around it.

The craft here is in making them belong in the same frame: matching the light direction so the generated splash is lit from where the product’s key light is, matching grain and depth of field so one layer doesn’t read sharper than the other, and adding the optical imperfections — slight bloom, chromatic edges, a little atmospheric falloff — that make a composite look photographed instead of assembled.

When it’s done right, a viewer can’t tell you which parts of the frame came from where. That’s the whole objective.

4. Reference locking for everything upstream

For previsualization and stills, we don’t prompt a product from scratch. We feed the actual product render or pack artwork in as a reference image and let the model work around it, with the product identity anchored rather than described.

We also keep a per-brand reference library — approved renders, the locked lighting setup, the exact material definitions — so that campaign two, six months later, matches campaign one without anyone re-deriving the look from memory. This is the least exciting item on the list and the one that saves the most money over time.

What this means if you’re the client

Three practical takeaways.

Ask which parts of a proposed film are generated. Not as a gotcha. The answer tells you what’s controllable in revisions. “The splash is AI, the pack is CGI” means a splash note is cheap and a pack note is cheap. “It’s all AI” means every note is a re-roll, and the thing you approved may not be the thing you receive.

Judge AI work on the last frame, not the first. Everyone’s first frame is good. Scrub to the end of the shot and look at the label, the logo edges, and the silhouette. Then look at the small print. That fifteen-second check will tell you more about a studio’s process than a whole showreel.

Send real files, and the pack will never be the problem. Print-ready artwork with dielines and correct color values. The entire method above depends on having a source of truth. Without it, even a CGI pipeline is approximating.

The honest summary

Generative video hasn’t solved brand consistency, and I don’t think a prompt will ever solve it — you can’t instruct your way to determinism. What has happened is more useful than that: we now have two tools with genuinely complementary properties, and the skill worth paying for is knowing exactly where the line between them falls.

Anyone selling you a fully generative product film for a brand with a regulated label and a recognizable pack is either not looking closely at their own output, or is confident you won’t.