Teaching a video model to speak in shots.
AI video models do not know film language on their own. Ask for cinematic and you get generic, weightless motion. So we stopped describing shots in adjectives and encoded the actual craft: a 186-entry bank of real techniques, each with the exact words that make a model reproduce it.
Adjectives are not direction
Film language is operational. A shot needs a size, a lens, a camera move, and a lighting setup. Without those decisions, our generation engine fills the gaps with an average.
A director does not say cinematic. A director says slow push-in on an 85, noir key, hold the close-up. We built that specificity into the pipeline, indexed by technique and translated into language our generation engine can follow.
From technique to prompt
Every entry carries the same four things, always ending in a ready-to-paste phrase.
Definition
What the technique is
Effect
Why a director reaches for it
Film reference
Where it was used
AI phrasing
The words that reproduce it
The categories cover a real DP's decisions
Where the camera sits, how it moves, what lens is on it, how the scene is lit, how bodies are staged in the frame, how shots cut together. Nine shot sizes, twenty-three camera movements, thirteen lens behaviors, ten lighting signatures, editing patterns down to the Kuleshov effect, plus recipe categories for famous shots and directorial styles.
It was built from three tiers of sources merged together: industry shot taxonomy, the academic film-language canon from film school, and working cinematographer craft references. Film knowledge, formally indexed. Not vibes.
One example: the dolly zoom
The entry defines it: dolly one way while zooming the opposite way, so the subject stays constant while the background warps. It names the effect: internal psychological rupture, the gut-drop. It cites where you know it from: Vertigo, Jaws.
Then it hands the model the phrasing: background compresses and stretches while the subject stays fixed size, disorienting spatial distortion. That last line is the whole point. The craft is real, and the translation is exact.
Perfection reads as CGI. Imperfection reads as capture.
Name the artifacts
The single biggest trick for footage that reads as real is to name the flaws a physical camera would produce. Heat-haze swim on a long lens. Focus hunting before it settles. Gimbal micro-jitter. Sensor bloom, gate weave, grain, motion smear.
A model told to be flawless produces something waxy and dead. A model told the lens trembles at 800mm and the focus takes a beat to find its subject produces something that looks shot. The bank writes imperfection in on purpose.
Recipes, stacked by code
Entries are indexed by code and stacked into recipes. A tense revelation is a close-up plus a slow push-in plus an 85mm plus a noir key. An epic reveal is an extreme wide plus a crane up plus anamorphic glass plus golden hour. Interrogation menace is pooled top light, a low angle, and a push-in too slow to notice.
The discipline that holds it together: one shot size, one camera move, one lens per shot, camera motion kept separate from subject motion. The grammar is strict because film grammar is strict.
Taste, built into the plumbing
Taste chooses the shot. The bank makes that choice reusable, so the automated half of the pipeline can follow it without pretending to replace it.
This is why footage from a film team looks different. The model is not guessing what cinematic means. It is being told, in the language of the craft, one technique at a time.
Want footage with actual film craft under it?
The difference between generic AI video and ours is the grammar underneath. Fifteen minutes and we will show you what that means for your brief.
Book a 15 minute call