Hanan Aldarmaki
Hanan Aldarmaki is an Assistant Professor in the Computing and Mathematical Sciences division at Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI). Her research broadly covers speech and language technologies for accessibility, education, and culture, including dialectal Arabic speech and language processing, clinical speech processing for communication disorders, sign language processing, speech privacy, and multimodal language understanding. She is also a co-founder of Potion AI, a UAE-based technology startup building games for educational, clinical, and cultural purposes.
Beyond Benchmarks: What Real Applications Reveal About Speech Recognition
Strong benchmark performance does not guarantee that automatic speech recognition (ASR) meets the needs of real-world applications. In accessibility tools and educational applications, seemingly minor recognition errors can undermine the application’s purpose. Even the definition of a correct transcription could depend on whom the system serves. This talk draws on our work on various applications (namely, clinical speech and language assessment and educational reading games) to examine recurring failure modes in contemporary ASR systems. These applications expose various challenges, such as limited linguistic context, unfamiliar accents, and the distinction between verbatim and intended speech. For instance, context can improve recognition, but it can also encourage models to “correct” the disfluencies or pronunciation errors an application needs to observe. Practical requirements for efficiency and privacy further complicate model selection. Motivated by these observations, we developed a dataset designed to quantify key challenges involving isolated-word recognition, context, and accent variation. I will present experiments examining how recognition performance changes across these conditions and conclude with directions for developing and evaluating speech systems around the needs of target speakers and applications.
Her talk takes place on Tuesday, November 3, 2026 at 13:00 in room TBD.