By now you've got the vocabulary.
Realism triggers, focus control, skin texture, film grain, light direction, negatives - a dictionary of hundreds of prompt terms that each do a specific job.
And here's where I fell apart for months: I dumped all of it into one long run-on sentence, hit generate, and got back mush. The model can't tell the important words from the filler, so it treats the whole thing as noise and guesses.
Vocabulary was never the problem. Assembly is.
This is the seven-slot formula I run on every AI ad prompt - the same order every time, so the model reads structure instead of soup. You'll get the formula, the three rules that keep it from breaking, and fill-in templates for three different ad types.
Let's assemble the thing.
Why Prompt Order Beats Prompt Vocabulary
Quick truth that took me too long to accept:
The same words in a different order produce a different image.
That's not a quirk - it's how these models read. Earlier words get weighted more heavily than later ones. So the first thing in your prompt sets the category, and everything after it gets interpreted through that category.
Which means structure is a lever completely separate from vocabulary. You can have every right word and still get a bad image if they're in the wrong order.
So we fix the order once and never think about it again.
The Seven-Slot Formula
Every prompt from here follows the same structure. Fill each slot in order. Never reverse it.
The Seven-Slot Formula: Intent → Subject → Action → Environment → Light → Technical → Negatives. A fixed prompt skeleton where each slot does one job and sits in priority order, so the model reads a structured photograph brief instead of a pile of adjectives.
Here's each slot, its job, and an example fill:
| Slot | Its one job | Example fill |
|---|---|---|
| 1 - Intent | The realism trigger. Sets the category before anything else is read. | Ultra-realistic documentary-style photography |
| 2 - Subject | Who or what - most important physical details first. | of a woman in her early 30s, natural skin tone variation, visible pores, light freckles |
| 3 - Action | What they're doing. Specific and physical, never vague. | applying serum with fingertips to her cheek, eyes softly closed |
| 4 - Environment | Where it happens. Enough to place it, not to compete. | in a bright minimal bathroom, marble vanity |
| 5 - Light | Source, direction, quality. One primary, one optional secondary. | soft morning window light from the left |
| 6 - Technical | Camera, lens, grain, color, dynamic range. Always last in the positive section. | 85mm lens, shallow depth of field, subtle 35mm film grain |
| 7 - Negatives | The universal baseline plus ad-specific exclusions. | no CGI, no plastic skin, no text in image |
1 - Intent
Its one job
The realism trigger. Sets the category before anything else is read.Example fill
Ultra-realistic documentary-style photography2 - Subject
Its one job
Who or what - most important physical details first.Example fill
of a woman in her early 30s, natural skin tone variation, visible pores, light freckles3 - Action
Its one job
What they're doing. Specific and physical, never vague.Example fill
applying serum with fingertips to her cheek, eyes softly closed4 - Environment
Its one job
Where it happens. Enough to place it, not to compete.Example fill
in a bright minimal bathroom, marble vanity5 - Light
Its one job
Source, direction, quality. One primary, one optional secondary.Example fill
soft morning window light from the left6 - Technical
Its one job
Camera, lens, grain, color, dynamic range. Always last in the positive section.Example fill
85mm lens, shallow depth of field, subtle 35mm film grain7 - Negatives
Its one job
The universal baseline plus ad-specific exclusions.Example fill
no CGI, no plastic skin, no text in image
Two slots I used to fumble, worth a note each:
Slot 1 (Intent) is the realism trigger - it decides which dialect of "real" the model speaks. If that phrase means nothing to you yet, it's the whole of the photorealism triggers guide.
Slot 6 (Technical) is where the quality layers live - and where over-stacking wrecks the image. That's the four-layer fix in why AI ads look fake.
And here's the principle that makes the formula flexible instead of rigid. I call it Slot Depth:
Slot Depth: the formula never changes - but which slots you fill deep and which you keep short does. A close-up beauty shot loads the Subject and Light slots and keeps Environment to a few words. A lifestyle shot loads Environment and Action and keeps Subject brief. Match each slot's depth to how much that element matters in this creative, and the prompt writes itself.
That one idea is why the same seven slots build a skincare macro and a fitness action shot without any new vocabulary.
The 3 Stacking Rules
Three rules stop the most common ways a prompt collapses.
Rule 1: One conflict kills the whole prompt
Write shallow depth of field and everything in sharp focus in the same prompt and the model splits the difference - you get neither.
Before you submit anything, read it once looking only for contradictions. Delete the weaker instruction. This single pass fixes more bad generations than any new keyword you could add.
Rule 2: Five detail layers, maximum
You can stack skin texture, hair detail, fabric texture, environmental grounding, and grain. That's five. It's enough.
Add a sixth, seventh, eighth layer and you get diminishing returns, then noise. Use Slot Depth to choose which five matter for this ad instead of cramming in all of them.
Rule 3: Negatives always go last
Negatives dropped in the middle of a prompt interfere with the positive instructions that come after them.
So every "no -" term goes together, at the very end, after all your positives are done.
The Formula in Practice: 3 Ad Types
Same seven slots. Three completely different ads. Watch the structure stay identical while Slot Depth shifts.
Skincare UGC close-up
Subject and Light run deep. Environment stays short.
Ultra-realistic documentary-style photography of a woman in her early 30s,
average body type, natural skin tone variation, visible pores, light
freckles, slight under-eye circles, loose hair with flyaway strands,
applying face serum with fingertips to her cheek, eyes softly closed,
relaxed candid expression, in a bright minimal bathroom with a marble
vanity, soft morning window light from the left diffused through sheer
curtains, shallow depth of field, subtle 35mm film grain, natural muted
color grading, no CGI, no plastic skin, no airbrushing, no extra fingers,
no text in image, no fantasy lighting, no overpolished result
Fitness lifestyle action
Now Action and Environment run deep. The skin description shrinks to a few words.
Cinematic realism, true-to-life photography of a woman with an athletic
build, toned physique, sweat on forehead and wet hairline, mid-stride
running on an urban street, arms pumping naturally, hair moving with motion
blur, modern city street with buildings receding in background, harsh
midday sun overhead casting short hard shadows, 35mm lens, motion blur on
legs, high contrast natural color grading, subtle film grain, no CGI, no
stock photo athlete, no fake effort, no 3D render, no plastic skin, no
fantasy lighting
Premium product e-commerce
No human at all. Subject and Technical carry the weight; Action becomes "product sitting still."
Ultra-realistic product photography of a glass skincare serum bottle, on a
white Carrara marble surface with natural grey veining, product sitting
still with a realistic cast shadow falling to the right, single softbox key
light from upper left, contact shadow at bottle base, condensation droplets
on glass exterior, 85mm macro lens, sharp focus on label and glass detail,
micro-detail visible, natural color grading, highlights rolled off, no CGI,
no 3D render, no floating product, no label text artifacts, no background
artifacts, no fantasy lighting
Three ads, one skeleton. The only thing that moved was Slot Depth.
The 30-Second Checklist
Before you hit generate, run this. It takes under thirty seconds and it catches the expensive mistakes:
- Intent is the first thing written
- Subject details ordered most important → least
- Action is specific and physical, not vague
- Environment supports the subject, doesn't compete
- One named light source, with direction
- Camera and lens specified
- Negatives all at the end
- No two instructions contradict each other
Every box checked? Submit. One box empty? Fix that gap first.
Pro Tip
Note
Conclusion: One Skeleton, Infinite Ads
Four things to lock in:
1. Order beats vocabulary. Early slots get weighted more, so the sequence controls the image as much as the words. Fix the order once.
2. Seven slots, always the same: Intent, Subject, Action, Environment, Light, Technical, Negatives.
3. Slot Depth is the flexibility. The formula never changes - you just load the slots that matter for this specific creative and keep the rest brief.
4. Read for contradictions before every submit. One conflict produces a muddy in-between and quietly wastes your generation.
So here's the move: take the last prompt you wrote and reorder it into the seven slots. My bet is you'll spot a contradiction and a missing light source in the first ten seconds - and the reordered version will beat the original on the first try.
Steal my ad concept & hook library
Every prompt, archetype, and hook I use in my own accounts. Updated weekly.
Join the free Skool community