Running one prompt at a square ratio and then at a wide one produces two different pictures rather than two framings of one. The reason lies in how training material is composed.
Frame shape is correlated with subject matter
Photographs are not composed at random ratios. Portraits are overwhelmingly vertical, cinematic stills are wide, product photography is square, and social formats have their own conventions.
A model trained on that material learns the association. Asking for a wide frame supplies a signal about what kind of image this is, before the prompt text is considered.
So the ratio functions as an additional instruction, and it can override parts of the written prompt when the two disagree.
Subject scale shifts with the frame
The same subject occupies a different proportion of a vertical frame than a wide one, because photographers frame differently in each.
A person requested in a vertical image tends to arrive closer, cropped nearer the waist or shoulders. The same request in a wide frame usually produces a fuller figure with visible surroundings.
This means changing the ratio changes how much environment the model has to invent, which is where new and unrequested content enters.
Wide frames generate context that was never specified
A wide image has space that must be filled. If the prompt describes only a subject, the model supplies a setting consistent with its guess about the genre.
That invented context frequently carries implications the prompt did not intend, such as a period, a location or a mood implied by the surroundings.
Practitioners describe this as the ratio adding words to the prompt, and it is the most common source of surprise when the same prompt is reused at a new size.
Unusual ratios sit outside the training distribution
Standard shapes are well represented. Extreme panoramas and very tall frames appear far less often, so the model has weaker guidance about how to compose them.
Results at these ratios commonly show duplicated elements, repeated subjects, or a composition that reads as two images joined, because the model is extending a pattern it has less evidence for.
Generating at a conventional ratio and extending afterward usually produces a more coherent result than requesting the unusual shape outright.
Choose the ratio before writing the prompt
Because the ratio influences content rather than only framing, it belongs in the initial decision rather than as an adjustment afterward.
Prompts refined at one shape frequently need retuning at another, and treating that as a fault in the prompt leads to changes that were never necessary.
The efficient sequence is to fix the output size the deliverable requires, then develop the prompt inside that constraint.