Hi! My name is Junkai Wu. I’m a third year Ph.D. student in the Department of Electrical & Computer Engineering at the University of Washington, advised by Prof. Mari Ostendorf. My research is on audio-centric multimodal AI, and especially on text and multimodal LLMs that understand and generate speech, music, and sound, either end to end or as components in larger systems.
Before coming to UW, I got my B.S. in Computer Engineering from the University of Illinois Urbana-Champaign, where I worked on audio processing with Prof. Paris Smaragdis and speech processing with Prof. Mark Hasegawa-Johnson.
news
- AVMeme Exam was accepted to COLM 2026!
- Started my research internship with the Music AI Group at Adobe.
- Bridging Ears and Eyes won the Best Paper Award🥇 at WASPAA 2025!
- Website lauched
!
selected publications
- Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal DistillationIn IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2025
- Just ASR + LLM? A Study on Speech Large Language Models’ Ability to Identify and Understand Speaker in Spoken DialogueIn IEEE Spoken Language Technology Workshop (SLT), 2024
experience
- Summer 2026Research Scientist/Engineer InternImproving audio–text alignment in music generation models, with Zhepei Wang and Nicholas J. Bryan.
education
- 2023 – presentPh.D. in Electrical & Computer EngineeringAdvised by Prof. Mari Ostendorf.
- 2019 – 2023B.S. in Computer EngineeringAdvised by Prof. Paris Smaragdis and Prof. Mark Hasegawa-Johnson.
teaching
- Spring 2026EE P 598 · Introduction to Digital AudioUniversity of Washington·Seattle, WATA for the course, covering audio synthesis and sound programming in SuperCollider.