Getting one good image is luck management. Getting twelve that look like they came from the same shoot is a process. The difference is whether you are rerolling prompts or controlling the things that actually make output repeatable.
Why the same prompt gives different pictures
Image models start from random noise and shape it according to your prompt. Change the starting noise and you get a different picture from identical words. That randomness is why exploration is easy and consistency is hard.
Everything below is a way of removing randomness from one part of the process so the model only improvises where you want it to.
References are the strongest control
Words describe; references show. If the subject exists, supplying it through image to image removes the model's freedom to reinvent it. This is far more reliable than any amount of adjective-stacking.
How many references a model accepts varies a lot. GPT Image 2 and GPT Image 2.5 accept up to sixteen. Nano Banana 2 takes fourteen, Seedream 5 Pro ten, Nano Banana Pro and several others eight. A few models accept exactly one.
What to put in each slot
A set that works: two or three angles of the subject, one image that carries the lighting you want, one that carries the setting or surface. Five references chosen for different reasons beat twelve near-duplicates, which mostly teach the model to average.
Keep the same reference set across a series. Swapping one image between renders is the most common cause of a set that almost matches.
Seeds: the repeat button
A seed fixes the starting noise. Same seed, same prompt, same settings, same model gives you the same picture again. Change one word and you get a controlled variation rather than a fresh roll of the dice.
Not every model exposes it. Among those that do, Qwen, Midjourney Diffusion, Ideogram and the Wan models let you set it directly. When it is available, the workflow is: explore without a seed until something is close, note the seed of that render, then iterate with it fixed while you adjust wording.
The mistake is fixing a seed too early. A seed locks in a composition you have not chosen yet.
Guidance and steps, in plain terms
| Setting | What it changes | Sensible range |
|---|---|---|
| Guidance scale | How literally the model follows the prompt | Middle of the range; high values look stiff |
| Inference steps | How long it refines before stopping | Default, unless output looks unfinished |
| Strength (image to image) | How far it may drift from your reference | Low to keep the subject, high to restyle |
| Negative prompt | What to avoid | Short and concrete, not a wishlist |
Guidance is the one people over-tune. Pushing it high forces literal compliance and usually costs you naturalism — hard edges, stiff poses, over-saturated light. If the model is ignoring something important, the fix is almost always a clearer prompt or a reference image, not a higher number.
Strength is the one people under-use. In image to image it is the dial between "touch this up" and "reinterpret this". If your product keeps changing shape, strength is too high.
Writing a prompt that survives a rerun
A repeatable prompt names four things: the subject, the light, the framing, and what must not be there. Anything vaguer gets filled in differently each time.
Compare "a nice product photo of a bottle, professional" with "a matte ceramic bottle on a dark slate surface, single soft key light from the left, shallow depth of field, no text, no reflections on the label". The second one is boring to read and comes back consistent.
Keep the structure fixed across a series and change only the variable part. If the subject changes but lighting and framing should not, only those words should move.
Consistency across a set
For a series that has to hang together: fix the model, fix the aspect ratio, fix the reference set, fix the lighting sentence, and vary only the subject or angle. Run them in one sitting rather than over days, because your own wording drifts more than the model does.
Resolution should also stay constant within a set. Mixing a 4K hero with 1K supporting shots is fine; mixing 1K and 2K for images shown side by side is visible in the detail. Our guide to frames and resolution covers which tier to choose for each slot.
When output still will not settle
If a subject keeps mutating despite references, the model may simply not be the right one — reference limits and instruction-following differ sharply across the catalogue. If detail looks smeared, resolution is more likely the issue than the prompt. If text in the image comes out garbled, that is a model choice: Ideogram is the one built for typography, and no amount of prompting fixes it elsewhere.
And if the picture is right except for one element, stop generating. Remove the element with object removal or fix the backdrop with the background tools. We covered that boundary in cleaning up images.
A loop that converges
Explore cheaply on a 2-credit model with no seed until the composition is right. Note the seed if the model offers one. Move to text to image or image to image on the model that suits the destination, carrying the same wording. Add references for anything that must stay accurate. Change one variable per run, never three.
Most jobs land in three or four renders that way. The workspace keeps prompt, references and settings in one place while you switch engines, and plans shows what each of those runs draws from your balance. If you are still choosing between engines, this comparison is the shortcut.
A prompt template worth reusing
Keep the skeleton fixed and swap only the variable clause. Something like: subject and material, then surface or setting, then light direction and quality, then depth of field, then the exclusions.
Written out: "a matte ceramic bottle on dark slate, single soft key light from the left, shallow depth of field, no text, no harsh reflections". To photograph a second product in the same set, only the first clause changes. Everything the model uses to build the look stays word-for-word identical, which is the point.
Exclusions belong at the end and should be concrete. "No text" works. "Nothing ugly" does not describe anything the model can act on.
Troubleshooting by symptom
| Symptom | Likely cause | Fix |
|---|---|---|
| Subject changes between runs | No reference, or strength too high | Add references; lower strength |
| Output looks stiff or over-lit | Guidance pushed too high | Return to the middle of the range |
| Detail smeared at full size | Resolution too low for the use | Re-run at 2K or 4K |
| Text in the image is garbled | Wrong model for typography | Use Ideogram |
| Set does not match | Prompt or references drifted | Fix the wording and the reference set; re-run together |
| One element ruins a good frame | Trying to solve it by regenerating | Remove it instead |
Change one thing at a time
The discipline that makes all of the above work is boring: one variable per run. Change the seed or the wording or the model, never two together, or you will not know which change produced the improvement.
It feels slower and finishes sooner. Three deliberate runs usually beat a dozen hopeful ones, and on the cheap engines those three runs cost six credits.
