Controllable Stress in Text-to-Speech via Weakly-Supervised Prosody Detection
Abstract
Spoken communication is shaped not only by what words are said but by how they are realized: through stress, rhythm, and intonation that signal emphasis, information structure, and emotion. Modern text-to-speech (TTS) and speech-to-speech translation systems translate and synthesize words faithfully but discard the prominence patterns that make speech expressive and intelligible. This shortfall is most acute for low-resource Indian languages, where annotated prosodic resources are almost non-existent, and it directly limits applications such as automatic dubbing, and expressive speech translation. This research develops a unified framework that detects, controls, and transfers prosodic stress across the speech pipeline, with a focus on this under-served setting.