Language Tags, Voices, and Denoising Steps
Supertonic accepts Unicode text together with an explicit language selection. The integration wraps the text in the corresponding language tag before inference, so choosing the matching language is part of the model input rather than a display-only filter. Its ten named choices map to the same M1–M5 and F1–F5 built-in style files; the fictional names are catalog labels and do not identify real speakers.
Denoising steps trade computation for another refinement pass, but more steps do not guarantee that every sentence sounds better. Begin around the interface default, compare a fixed sample, and increase steps only when the audible result justifies the extra time. Speed and steps interact with device performance, text length, punctuation, and backend, so retain those settings with an important export.