Generated backing tracks pass unnoticed in a way that generated vocals do not. The difference is in how closely listeners attend to voices compared with everything else in a mix.

What a listener hears in a voice

People spend their lives listening to voices and extracting information from them: age, effort, emotion, health, sincerity and distance.

That perceptual machinery is far more finely tuned than anything applied to a guitar tone, and it operates involuntarily on singing as well as speech.

Small irregularities that would pass unnoticed on an instrument are registered on a voice as something being wrong, even when the listener cannot identify what.

Why breath and effort are the tell

Singing is physical. Breath is taken at particular points, phrases end where air runs out, and strain appears predictably at the top of a range.

These constraints shape phrasing so consistently that listeners expect them, and generated vocals frequently omit them because nothing in the model has lungs.

A phrase held longer than a human could sustain, or a leap into a high register with no change in tone quality, reads as synthetic immediately.

How consonants expose the model

Vowels carry pitch and are comparatively easy to generate convincingly, since they are sustained and harmonically regular.

Consonants are brief, noisy transients that vary enormously by singer, language and microphone placement, and they are where generated vocals most often smear.

The effect is a voice that sounds acceptable on sustained notes and indistinct on rapid lyrics, which is why generated singing tends to be presented at slow tempos.

Why instrumental parts get away with more

An instrument has a narrower expressive range and a more mechanical relationship between action and sound.

Production also conceals a great deal, since compression, reverb, layering and effects are applied to instruments as a matter of routine and mask small irregularities.

Vocals sit forward in a mix by convention, receive the most attention, and are the least protected by that same production treatment.

Where the work is concentrated now

Systems that transform an existing sung performance into another voice outperform systems generating singing from nothing, because the human performance supplies the breath, timing and effort.

That approach preserves the physical realism and changes only the timbre, which is the part models handle well.

It also relocates the difficulty from acoustics to permission, since the performance being transformed and the voice being imitated both belong to someone, and the technique works whether or not either of them agreed.