Back to Work
Craft · Quality

9 out of 10 is the floor, not the target.

AI x3 is our own house campaign: forty-six sarcastic fifteen-second spots about AI saturating daily life. On day one we graded our own output honestly. Three to five out of ten. This is what we wrote down, and what we changed.

Loom walkthroughThe AI x3 quality journey: the honest grade, the anchor-frame pivot, and a before-and-after shot.
46
Concepts in the campaign
3-5
Honest starting grade / 10
9+
The floor to ship

Grading our own homework, honestly

A studio that sells taste should be embarrassed to write this sentence, so we wrote it into the project kickoff on day one: current honest state, our outputs are 3 to 5 out of 10.

The rule that followed was just as blunt. Nine out of ten is the floor, not the target. We would rather burn extra credits and reroll ten times than ship a single eight. Naming the gap in writing came first. Closing it was the mission.

Fix 01

The realism reset

The failure mode had a name: outputs that look organized from far away but read as fully AI-generated up close. The root cause was the workflow itself. We were feeding the video model polished, labeled, studio-lit reference grids, so it faithfully produced what we asked for: a polished AI reference sheet in motion, not a scene.

The fix flipped the order. Build one genuinely photoreal anchor frame first, with real light, real skin texture, correct grain. Get a human sign-off on that single frame. Then animate that exact frame forward as one continuous take. The anchor pins identity, room, grade, and grain from frame one, so nothing drifts. The proof shot passed on its first generation after two silent failures under the old method.

Anchor candidates

Takes compete for the anchor. One gets locked and signed off, and every frame of motion is generated from it.

The reset, in one pair

Old input: a studio reference grid

New input: one photoreal anchor frame

Next
Fix 02

Physics, not adjectives

Our old prompts stacked director names, lens jargon, hex codes, and a paragraph of things to avoid. That pile was itself a cause of the AI look: a model juggling twenty contradictory descriptors averages them into generic AI cinema. Worse, naming the defect can summon it.

The replacement rule: describe real light and real skin in plain sentences. One bare tungsten bulb overhead, hard top light falling into damp shadow, sweat sheen and visible pores. One grounding reference, not five. A lean prompt around a hundred and fifty words. The anchor frame carries the look; the prompt carries the motion.

The manifest cannot overrule the frames.

Fix 03

Frames outrank paperwork

Our tracking records used to say PASS while the footage said otherwise. So we made a rule with teeth: before anything is marked final, extract a one-frame-per-second contact sheet from the actual render and check it against the promised beats. If the contact sheet does not show the hook, the shot fails, no matter how good the prompt read.

And when a shot fails, fix upstream before rerolling. Tighten the joke into two or three visible physical beats, regenerate clean references, simplify the storyboard, recompile a lean prompt, generate once. Rerolling the same broken inputs is just paying the same tax twice.

The traps
Fix 04

Silent failures cost real money

Video model moderation fails silently. No error message, no diagnosis, just a dead render and the credits gone, at roughly sixty credits per fifteen-second generation.

We reverse-engineered the traps the hard way. Naming a well-known AI product in a prompt trips a trademark filter, even when the brand is the joke. Injury and fight language trips a safety filter, even in comedy, so a punch became a near-miss and bruised became clean. And a child-safety classifier hates close-up framing on a baby even in a wholesome family scene, so the framing pulled back and the model changed. Now every prompt gets pre-flighted against the trigger list before submission. Blind retries are for people with unlimited credits.

The point

Sycophancy is a quality risk

One soft does-this-look-good check passes everything, which is why we replaced it with layers: an image-stage audit that catches most problems cheaply, an adversarial pass that hunts defects with a checklist, and a human on the contact sheet at the end.

That is what moved the floor. Not a better model, the same models under a process that refuses to flatter itself. Taste stayed the bottleneck, on purpose. Everything else got systematized so it could be.

Want output with a floor, not a ceiling?

The difference between AI slop and shippable work is the process that grades it honestly. Fifteen minutes and we will show you ours.

Book a 15 minute call