The founding gap
Voiceover is one of the few areas of modern software where the shortcut still gives itself away. A text field pours into a speaker, and the listener knows within two seconds that no human said this.
The reason is not that the underlying models are weak. It is that the layer above them - the one that decides where a breath goes, which noun to stress, whether a clause lifts or lands - has been treated as an afterthought for years.
Koythu started there. Not with a bigger model, but with the layer that decides how a line is spoken before any waveform is generated. That is what makes the output stop sounding read.
