Prompting for media: the transferable skill
The durable prompting craft that works across every image and video tool.
This is the most transferable, longest-lasting skill in the course: how to prompt a media model well. The tools change; this craft doesn't. And it's different from prompting a chatbot — you're describing a scene, layering visual information the model composes.
A tool-agnostic structure that works everywhere — think of it as stacking layers (coverage matters more than exact wording):
- Subject — concrete detail: materials, textures, clothing, expression, age. "A weathered fisherman" beats "a man."
- Action / pose — what they're doing.
- Environment — where, and what's around them.
- Composition / framing — shot type and angle: close-up, wide shot, low angle, overhead, shallow depth of field.
- Lighting — the single biggest quality lever: source, direction, quality, time of day. "Soft golden-hour side light" transforms an image.
- Camera / lens language — "85mm portrait lens," "shot on film," "shallow depth of field." This steers photorealism far more reliably than tags like "8K, ultra-detailed."
- Style & medium — photograph, oil painting, 3D render, anime.
- Mood & color — the emotional and palette direction.
The important 2026 split: classic diffusion tools (Midjourney, Stable Diffusion) reward dense, comma-separated keyword stacks; LLM-native tools (gpt-image-1, Gemini) reward clear natural-language descriptions and explicit constraints stated in plain English ("make sure there are exactly three apples"). Learn to feel which kind of tool you're in.
Key mechanics you'll use constantly:
- Aspect ratio — set it deliberately; it changes composition, not just crop.
- Negative prompts — list what to exclude (classic tools). Newer LLM-native tools have no negative field — you just say "without a hat" in the prompt.
- Iterate deliberately — your first prompt is a draft. Change one thing at a time (hold the seed) so you learn what each word does, rather than rewriting everything and starting over.
The mindset: you're a director giving specific, visual direction — not typing keywords and hoping. Master this layered, iterative approach and every tool in your stack produces dramatically better output, because they all run on the same underlying craft.
Write a prompt for one image using all eight layers. Generate it, then improve it by changing only the lighting layer. Compare — lighting alone usually transforms the result.
Enjoying the free lessons? Get an email when we publish new courses and updates — no spam, unsubscribe anytime.
Discussion (0)
Ask a question or share what worked for you. Comments are reviewed before they appear.
No comments yet. Be the first to start the discussion!