One character, a stack of real office renders, and a tight set of rules — turned into a two-minute cinematic short, one scene at a time inside an AI video tool. This walks through exactly how I did it, and why each choice mattered.
Made for a corporate workplace change programme. The tools move every few months, so nothing here names one: the method is the part that keeps working.
There's no magic to it. Most shots follow the same path — get it right once and the rest of the day mostly builds itself. Here's the whole thing end to end; I'll open up each step below.
The day, beat by beat
Maya, built as a reusable reference
One page, every scene sketched
Designers' real space renders as refs
Stylise + lock the frame
Animate, iterate, pick the best
Scenebuilder → the final cut
Consistency is the whole game in AI video, so I start with the person rather than the plot. I write Maya: the Sociable Engineer as a proper character brief — how she moves, how she listens, how she works. Everything after this gets measured against it.
The model can only keep a character consistent if I'm consistent first. The brief is the contract: same face, same hair, same soft blue sweater, same calm energy, in every single scene.
Overall presence — Maya exudes a warm, approachable energy grounded by sharp technical focus. Her default expression is a genuine, easy smile. She moves with calm, deliberate confidence — never frantic.
How she works — deep concentration at a standing desk or tucked into a phone booth; leans in and gestures fluidly when explaining technical concepts; methodical and precise in every action.
Movement — a fluid, unhurried cadence. Relaxed but attentive posture, entirely at home in the space.
The bible is pasted into the tool's character profile so the agent can craft scenes that stay true to her.
Next I turn Maya into something reusable. In these tools a consistent character is an ingredient — a reference you build once from an image (or generate one) and reuse everywhere. I built her out as a full reference sheet: front, three-quarter and profile, headshots, and her accessories — tote bag, pen, jewellery — so the model always has the same Maya to lock onto.
Build the reference sheet — multiple angles in one image — before you touch any scene. A handful of consistent views gives the video model far more to anchor to than a single headshot ever will.
Before I touch video, I sketch the entire day as one concept page — arrival, desk, focus, coffee, meeting, collaboration, pack-up. It's the cheapest place to fix the pacing, and it becomes the shot list everything else follows. I leaned on the look of traditional film boards, then reimagined the whole thing in Maya's world.
For the hero beats I push further: each board carries the actual camera direction — lens, aperture, the move, the foreground, the light. The shot is half-decided before the model ever runs.



The film sells a real workplace — the KLS office design — so I don't let the model invent rooms. Every scene starts from an actual 3D design render I treat as unbreakable ground truth — reception, collaboration corners, the town square, the social hubs. Its only job is to restyle what's there, never redesign it.
Real architecture in, stylised architecture out. Walls, pillars, windows and furniture stay exactly where the designers put them.




Four of the actual KLS design renders — the raw material every scene is built from.
This is the part that does the heavy lifting. Rather than re-explaining the look in every prompt, I wrote one set of Universal System Guardrails and loaded them as the agent's standing instructions. Every scene inherits them automatically — which is why twelve separate shots still feel like one film.
Treat the provided render as unalterable ground truth. Do not add, move or remove walls, pillars, furniture or windows. You are a style-transfer filter — apply the sketch aesthetic and lighting over the exact existing composition. Nothing more.
Uniform detailed pencil / fine-pen linework with varied weights over every element. Hybrid rendering: clean colour planes + soft painterly gradients. No photorealism, no flat vector look.
Deep forest greens, cool indigos, teals, soft baby blues, warm wood and beige, crisp whites and greys — layered for dimensional depth.
Slow, smooth, stabilised, deliberate motion — like a high-end slider. Very shallow depth of field, a heavily blurred foreground edge, no jerky pans or sudden drops.
Vibrant, diffused morning daylight flooding through glass, with soft volumetric god-rays cutting through the space for atmosphere and depth.
Strict hot-desking / ABW look. Absolutely no under-desk pedestals, drawers or personal clutter. Surfaces stay clean and open.
Predominantly Asian professionals (Hong Kong context), in calm small groups of 2–3. Present to add life — never a chaotic crowd.
Most of what makes this read as expensive isn't the model — it's the camera language I write into every prompt. I brief each shot the way I'd brief a DP: a real camera, a real lens, a real move. That's the difference between "an AI clip" and something that feels shot.
Every prompt opens with the same camera DNA. These are the numbers and choices I repeat on purpose:


Two things I add to almost every frame: a slider move for slow, deliberate motion, and volumetric god rays for depth and atmosphere.
The trick I lean on hardest: put something between the camera and Maya and throw it right out of focus. A plant leaf, a pillar, a monitor, the backs of two colleagues. It builds instant depth and stops every shot looking flat.
"Cinematic commercial aesthetic shot on Arri Alexa, extremely shallow depth of field. A slow lateral slide from left to right. The shot begins with a lush green indoor plant entirely blurred in the extreme foreground, revealing Maya in sharp focus, with soft volumetric god-rays through the glazing."
Negative: photorealism, 3D render, fast pans, jerky motion, under-desk clutter.
Now I feed the render plus my guardrails into the image model and it hands back the same room as a living sketch. Drag the handle — left is the source render, right is the frame I lock and use as the seed for the video.
Render
Stylised
I generate a few variations, then lock the single best image. That frame — not the text — is the real instruction for the video step. Every line, colour and light position is already decided, so the motion model has nothing left to invent.
A locked start frame is the difference between rolling the dice on a prompt and animating a picture I've already signed off.
Every room runs through the exact same step, each one anchored to its own render:








These tools are built around a few zones, and once you know them the rest is just repetition. Here's my actual project — the sidebar down the left, and every generated take sitting in the scenes grid.
My real workspace, untouched — the sidebar holds the project's assets; every tile in the grid is a generated scene, all sharing the one sketch look.
The project's assets — characters, scenes and saved Ingredients. This is where Maya lives, ready to drop into any shot.
Every take I've generated, laid out as tiles. I scrub through, compare, and pull the keepers straight into the cut.
Where I describe each shot, switch the model, attach reference frames — and set the standing instructions that run on every scene.
This is the part people miss. The prompt bar runs an agent, and the agent has standing instructions — so my guardrails from step 05 aren't retyped each time, they're attached once and applied to every scene automatically.
Plain language in. Pick the model — one for stills, one for video — attach a reference frame, hit send.
Universal System GuardrailsCore directives — applied to every shot
Scene 4–5 · Desk arrivalReference image + camera + negative prompt
Scene 6 · Work caféSlow drift past a blurred espresso machine
Scene 7 · Confidential meetingShoot through glass, blurred frame foreground
A faithful recreation of my agent-instructions panel — each rule toggled on per scene, the guardrails always live.
These tools split the work across models and let you choose which — you're never locked to one. There are two jobs:
One model handles the stills, another handles the motion. Video engines usually come in a fast tier and a quality tier, trading speed against fidelity: generate your takes on the fast one, and re-run only the keeper on the slow one.
The locked frame and the scene prompt go to the video model. Every current engine works the same way. I almost never get the hero shot first try, so I generate several takes and cut between them like an editor — watching hands, faces and the camera move. Click the take you'd keep.
Scene 01 — arrival. Take A is the full taxi pull-up; Take B catches the door-exit beat. Which one sells it?
Scene 04 — desk height adjust. Subtle motion is hard; iterate until it's smooth.
Scene 03 — greeting colleagues. Watch hands and faces — the usual giveaways.
With a keeper for every scene, the scene builder becomes the storyboard I actually edit — I drop the best takes in order, use Extend to let a beat breathe, Jump To to carry Maya between rooms, and trim everything to rhythm.
These are pulled directly from the final film below — the same takes, now timed and edited together:
Twelve scenes, one character who never breaks, every shot grown from a real design render and the same page of rules I wrote at the start. This is the finished cut.
A Day in the Life — Maya · the finished two-minute short.
▶ press playA guided, floor-by-floor walk-through — move through arrive, focus, collaborate and wind-down.
Step through the day →Pick a character and write them properly. Storyboard a day. Gather your real reference imagery. Write your guardrails once. Then stylise → lock → animate → iterate → stitch. The same seven steps scale to any space and any story — this one just happened to be Maya's.