Laryngeal Structure Segmentation in High-Speed Videoendoscopy Using Deep Learning
Teaching computers to automatically map the moving parts of your voice box
Researchers trained artificial intelligence systems to automatically identify and outline five key structures in high-speed videos of the larynx—the vocal folds, arytenoid cartilages, epiglottis, aryepiglottic folds, and the opening between the vocal folds—achieving over 95% accuracy. The breakthrough was applying this method to natural connected speech rather than just sustained vowels, where tissue movement and video quality make the task far harder.
Voice doctors currently spend significant time manually analyzing frame-by-frame videos to diagnose hoarseness, paralysis, and other laryngeal disorders. This automated approach could speed up diagnosis, make measurements more consistent across patients, and enable early detection of voice problems by spotting abnormal laryngeal patterns that doctors might miss by eye. It particularly matters for detecting voice disorders that only appear during natural speech, not during simple sustained vowel tests.