HuPER: A Human-Inspired Framework for Phonetic Perception
Abstract
We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic--phonetic evidence and linguistic knowledge. HuPER first learns an acoustic-grounded phone recognizer from limited human-annotated data and transcript-only speech: canonical G2P transcriptions are treated as auxiliary linguistic cues rather than ground truth, and are corrected into realized phone proxies for self-training. At inference time, HuPER supports multiple perceptual routes: it can rely on bottom-up phone evidence when the signal is clear, incorporate explicit expectations when a reference is available, or use lexical constraints when acoustic evidence is weak. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic feature error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER also improves robustness on weak-evidence and disordered speech, demonstrating the benefit of adaptive, multi-path phonetic perception under diverse acoustic and task conditions. All training data, models, and code are open-sourced. Code and demo are available at https://github.com/HuPER29/HuPER.