publications
Papers and preprints, newest first.
Also listed on Google Scholar.
2026
- Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPOarXiv:2607.27756, 2026
2025
- Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal DistillationIn IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2025
2024
- Just ASR + LLM? A Study on Speech Large Language Models’ Ability to Identify and Understand Speaker in Spoken DialogueIn IEEE Spoken Language Technology Workshop (SLT), 2024
- Meta-AF Echo Cancellation for Improved Keyword SpottingIn IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024
2023
- Listen, Decipher and Sign: Toward Unsupervised Speech-to-Sign Language RecognitionIn Findings of the Association for Computational Linguistics: ACL 2023, Jul 2023