How to Get the Photo You Actually Wanted

Eight things worth describing when you prompt an AI image, so you stop getting stock-photo averages and start getting something you’d actually use.

A woman laughing in warm late-afternoon backlight, shot with a shallow depth of field

Most AI images disappoint for the same reason: the prompt described a category, not a photograph. Ask for “a woman cooking in a kitchen” and you get the average of every stock photo ever captioned that way. Clean counters, even light, nobody in mid-motion.

The fix is not a longer prompt. It is knowing which eight things are worth describing, and what each one buys you.

The same model, the same day, two prompts

One line

A woman cooking in a kitchen

Described

A woman in her early 30s at a kitchen counter, mid-motion, tipping chopped herbs from a board into a pan. Relaxed, unposed, half-smiling at something off-frame. Worn linen shirt, sleeves pushed up, muted olive tones. A lived-in home kitchen, late afternoon, dishes still out behind her. Low side light through a window to camera left, warm and soft, falling off quickly into shadow on the far wall. Eye level, medium shot, shallow depth of field so the background goes soft. Candid and documentary, intimate rather than staged. Warm and slightly muted. Shot on Portra 400, 35mm, f/1.4.

compare-oneline
compare-described

The two prompts above, run the same day on the same model.

The second one is not cleverer. It just answers the questions the model was otherwise going to answer for you.

Change one variable at a time

This is the part that actually transfers. When you are learning what a detail does, hold everything else still and change only that one thing. Same subject, same wardrobe, same room.

You end up with a mental library of what each word buys you. That is what lets you brief an image in one pass later, instead of rerolling twenty times hoping something lands.

If you change three things and the image improves, you have learned nothing about which of the three did it.

01 · Reference and inspiration

Start with images you like. You probably already have a screenshot folder or a Pinterest board somewhere. Look at what your favourites share. It is usually one or two things, the light, the angle, the colour palette, not everything.

Name a photographer or a style. A named visual style does more work in three words than a paragraph of adjectives. “In the style of Petra Collins” or “like a Kinfolk editorial” gives the model a whole universe to pull from.

Say what you are not after. Ruling out the obvious default is often faster than describing the alternative. “Not stock photography. Not overhead. Not looking at camera.”

02 · Subject

Who or what is in frame, and what they are doing. How many people, roughly what age, and anything that actually matters about them.

Then the choice most people skip: action or still. Mid-motion reads as candid. Held poses read as staged. This single decision sets the register of the whole image, so make it deliberately rather than inheriting it.

Exhibit · Subject
subject-1
subject-2
subject-3

03 · Wardrobe and styling

Clothing register — casual, formal, athletic, workwear. Consistency with the environment matters more than the specific choice.

Visual aesthetic — retro, vintage, minimalist, maximalist. This one tends to leak into the whole frame rather than staying on the clothes. Useful when you want it, worth watching when you do not.

Colour and texture. Naming two or three colours keeps a palette coherent. Texture words do a lot of quiet work for realism: linen, denim, worn.

04 · Environment

Indoor or outdoor sets the lighting expectations before you have described the light at all.

Be specific about the place. Kitchen, backyard, campsite, studio, cafe. Specific places carry their own props and their own light for free.

Say how much background you want. A busy, lived-in background reads documentary. A bare one reads editorial. Say which, or you inherit the model’s default.

05 · Lighting

This is where most of the believability lives.

What is the source — sun, overcast sky, window, flash, practical lamps. One dominant source is usually more convincing than several.

Where is it coming from — behind, side, above, front. Direction does more for shape and mood than any adjective you could add.

What is bouncing it — a white wall, a wooden floor, sand, grass. Bounce is what fills the shadows, and naming it is the difference between an image that looks lit and one that looks photographed.

A handful of named conditions are worth memorising, because each one is a complete lighting setup in two words: rim light, overcast, golden hour, blue hour, on-camera flash.

Exhibit · Lighting
Rim light
Rim lightLight shining at an angle from behind
Overcast
OvercastSoft, even, no visible shadow edge
Evening ambient
Evening ambientNatural light
Golden hour
Golden hourLow, warm, long shadows
Blue hour
Blue hourCool, dim, after the sun has gone
On-camera flash
On-camera flashFlat, bright, hard falloff behind the subject

06 · Composition and camera

Angle — eye level, low, overhead, over-the-shoulder. Low angles add stature. Overhead flattens and abstracts.

Depth of field is just how blurry the background is. Shallow means a soft background and a subject that pops. Deep means everything sharp and more context visible. Saying “shot at f/1.4” is shorthand for “very blurry background.”

07 · Mood and style

Tone — candid, editorial, documentary, staged. This is the one people most often leave out, and most often end up unhappy about.

Feeling — joyful, intimate, dramatic, peaceful. Vaguer than the other slots, but it steers expression and light together.

Colour treatment — warm, cool, muted, vibrant. Say it explicitly or you inherit whatever your reference implied.

08 · Technical look, optional

Lens. Focal length changes the image more than people expect. 16mm and 135mm of the same subject are barely the same photograph.

Post-processing — film grain, vintage fade, sharp digital. Easy to overdo. One term is usually enough.

This section is optional for a reason. If the first seven are vague, no amount of “shot on Portra 400” will save the image.

Exhibit · Technical look
16mm
16mm
24mm
24mm
50mm
50mm
135mm
135mm
iPhone 4
iPhone 4
Point and shoot with flash
Point and shoot with flash
Instax
Instax
Medium format film
Medium format film

When this is worth the effort

The underlying skill is not prompting. It is art direction. These are the same eight things you would tell a photographer on a shoot, and the reason the descriptive prompt works is that it briefs the model the way you would brief a person.

The template

Delete the lines you do not need. An empty slot is a decision handed to the model.

REFERENCE:   in the style of ____ / like ____
SUBJECT:     ____, aged ____, ____ (expression), ____ (doing what)
WARDROBE:    ____ (register), ____ (aesthetic), ____ (colours, textures)
ENVIRONMENT: ____ (indoor/outdoor), ____ (specific place), ____ background
LIGHTING:    ____ (source) from ____ (direction), ____ (hard/soft,
             warm/cool), bouncing off ____
COMPOSITION: ____ (angle), ____ (framing), ____ depth of field
MOOD:        ____ (tone), ____ (feeling), ____ (colour treatment)
TECHNICAL:   ____ (film stock), ____ (camera/lens), ____ (processing)

Where to start

Pick one slot and run a grid on it, holding everything else still. Lighting is the one worth doing first. It changes the most, and it transfers to every image you brief afterwards.

Take it with you

I packaged the whole framework as an agent skill: all eight slots, the vocabulary for each, and both worked examples. Drop it in ~/.claude/skills/ai-image-prompts/ and it will build a prompt from a one-line brief, or take apart one you already wrote.

Download the skill
All Thoughts