Image prompts grow longer as people try to pin down what they want, and past a certain length the additions stop helping. A reference image does the same job more directly.
What a description has to encode
A text prompt must specify subject, composition, lighting, colour, texture, lens behaviour and style, all through words that were never precise about any of them.
Terms like moody, cinematic or minimal each cover an enormous range, and the model resolves them towards whatever was most common in training.
The requester usually has a specific image in mind, and the vocabulary available cannot address that specificity. Describing it accurately enough in words would take longer than producing it another way.
Why extra adjectives compete
Each term in a prompt influences the result, and influence is shared, so adding words dilutes the ones already there.
Long prompts also contain unnoticed contradictions, since soft natural light and high contrast pull in opposite directions.
The visible symptom is a prompt where adding a word changes something unrelated, which indicates the terms are interacting rather than accumulating. At that point removing words usually helps more than adding them.
How a reference collapses ambiguity
An image supplied as a reference carries composition, palette, contrast and treatment simultaneously and without ambiguity.
The model extracts those properties directly rather than reconstructing them from vocabulary, which removes the interpretation step where meaning was lost.
Two or three sentences of text alongside a reference typically outperform a paragraph of description alone. The text then only has to carry what the reference does not already show.
Why influence settings matter
Reference features have a strength control, and its position determines whether the output resembles the reference or copies it.
High influence reproduces the reference closely, which is useful for consistency across a set and unhelpful when a new composition is wanted.
Working from a low setting upwards, rather than starting high and reducing, gives a clearer sense of what the reference is contributing. It also avoids the trap of accepting an early close copy because it looked correct.
Where description still wins
Text remains the only way to specify what does not yet exist in any image you have.
Unusual combinations, specific subjects and narrative content all have to be described, and no reference supplies them.
The arrangement that works is a reference for treatment and words for content, which uses each channel for what it encodes efficiently rather than asking either to carry the whole request.