A generated shot of a busy sidewalk or a stadium reads correctly at first viewing and breaks down on a second pass. The failure is specific to background human figures and follows from how detail is allocated.
Fidelity is distributed by prominence
A generation system does not model each person and then render them. It produces a frame in which large, central and high-contrast elements receive the most consistent treatment.
Background figures are small, partially occluded and frequently in motion, so they receive whatever the surrounding context implies rather than individual attention.
What appears is something crowd-shaped: correct silhouette, plausible density, correct clothing distribution, without the internal consistency a real person would have.
Motion masks the errors in the first viewing
Human vision samples a moving scene coarsely. Attention follows the subject, and the periphery is accepted as long as it moves in a plausible way.
A background walker with an unstable gait or a limb that changes length passes unnoticed at normal speed, because nothing directed the eye there.
Pausing removes that protection, and the same frame that looked fine becomes obviously wrong, which is why these failures surface in review rather than in playback.
Persistence across frames is the harder constraint
An individual frame can contain a plausible crowd. The difficulty is that the same person must remain the same person across a sequence.
Background figures frequently change jacket color, swap direction or merge with a neighbor between frames, because nothing enforces their identity over time.
The result is a crowd that is internally coherent at every instant and incoherent as a group across a shot lasting a few seconds.
Depth boundaries produce the most visible artifacts
Where a background figure passes behind a foreground object, the system must decide what is occluded and restore it correctly on the other side.
That restoration frequently fails, producing a person who emerges from behind a car slightly different from the one who entered, or who does not emerge at all.
These transitions draw the eye because motion across an edge is exactly what peripheral vision is built to detect.
Production practice adapts by hiding the problem
Editors working with generated footage keep crowd shots short, hold them wide, and avoid the slow push that gives a viewer time to inspect the background.
Depth of field is used aggressively, since a defocused crowd is both cinematically ordinary and forgiving of exactly these errors.
The constraint shows up as a stylistic tendency in generated video: shallow focus, brief cuts, and subjects isolated from the people behind them.