Zhaoheng Ni

Research Scientist

Meta Reality Labs · New York, US

I am a research scientist at Meta Reality Labs working on generative models for audio, text, and video.

Previously, I was a maintainer of TorchAudio, the official audio library of PyTorch. Before Meta, I was a PhD student advised by Michael I. Mandel and an undergraduate student advised by Yan Xu.

Research Interests

News

Aug 2026
I am co-organizing the ICASSP 2027 special session Speaker Grounding in Speech LLMs with Dr. Hao Shi, Dr. Hexin Liu, and colleagues. Paper submissions are welcome.
Jun 2026
One paper was accepted by Interspeech 2026.
Jan 2026
Three papers were accepted by ICASSP 2026.
Aug 2025
Marvin Sach, Robin Scheibler, and I are organizing the ICASSP 2026 special session Promise and Perils of Generative AI for the Evaluation of Speech and Audio.
Aug 2025
Two main papers and one demo paper were accepted by ASRU 2025.
May 2025
Two papers were accepted by Interspeech 2025.
Dec 2024
One paper was accepted by ICASSP 2025.
Nov 2024
We are organizing the URGENT 2025 Challenge at Interspeech 2025.
Sep 2024
Our MelodyFlow demo supports text-guided music editing and generation at a 48 kHz sample rate.
Earlier news
Sep 2024
Three papers were accepted by IEEE SLT 2024.
Jun 2024
We organized the Audio Imagination Workshop at NeurIPS 2024.
May 2024
We organized the URGENT Challenge at the NeurIPS 2024 Competition Track.
Apr 2024
Our MMS paper was accepted by the Journal of Machine Learning Research.
Feb 2024
We released demo videos and the paper for FoleyGen.

Selected Publications

For a complete and current list, see Google Scholar. My name is shown in bold.

2025
  1. Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding

    Jiahui Zhao, Hao Shi, Chenrui Cui, Tianrui Wang, Hexin Liu, Zhaoheng Ni, Lingxuan Ye, Longbiao Wang

    IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) [paper]

2024
  1. SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

    Haohe Liu, Gael Le Lan, Xinhao Mei, Zhaoheng Ni, Anurag Kumar, Varun Nagaraja, Wenwu Wang, Mark D. Plumbley, Yangyang Shi, Vikas Chandra

    arXiv preprint

  2. High Fidelity Text-Guided Music Generation and Editing via Single-Stage Flow Matching

    Gael Le Lan, Bowen Shi, Zhaoheng Ni, Sidd Srinivasan, Anurag Kumar, Brian Ellis, David Kant, Varun Nagaraja, Ernie Chang, Wei-Ning Hsu, Yangyang Shi, Vikas Chandra

    arXiv preprint [paper] [demo]

  3. FoleyGen: Visually-Guided Audio Generation

    Xinhao Mei, Varun Nagaraja, Gael Le Lan, Zhaoheng Ni, Ernie Chang, Yangyang Shi, Vikas Chandra

    IEEE Workshop on Machine Learning for Signal Processing (MLSP) [paper] [demo]

  4. Data Efficient Reflow for Few Step Audio Generation

    Lemeng Wu, Zhaoheng Ni, Bowen Shi, Gael Le Lan, Anurag Kumar, Varun Nagaraja, Xinhao Mei, Yunyang Xiong, Bilge Soran, Raghuraman Krishnamoorthi, Wei-Ning Hsu, Yangyang Shi, Vikas Chandra

    IEEE Spoken Language Technology Workshop (SLT)

  5. Massively Multilingual Forced Aligner Leveraging Self-Supervised Discrete Units

    Hirofumi Inaguma, Ilia Kulikov, Zhaoheng Ni, Sravya Popuri, Paden Tomasello

    IEEE Spoken Language Technology Workshop (SLT)

  6. URGENT Challenge: Universality, Robustness, and Generalizability for Speech Enhancement

    Wangyou Zhang, Robin Scheibler, Kohei Saijo, Samuele Cornell, Chenda Li, Zhaoheng Ni, Anurag Kumar, Jan Pirklbauer, Marvin Sach, Shinji Watanabe, Tim Fingscheidt, Yanmin Qian

    Interspeech [paper]

  7. Scaling Speech Technology to 1,000+ Languages

    Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, et al.

    Journal of Machine Learning Research [paper] [code]

2023
  1. TorchAudio 2.1: Advancing Speech Recognition, Self-Supervised Learning, and Audio Processing Components for PyTorch

    Jeff Hwang, Moto Hira, Caroline Chen, Xiaohui Zhang, Zhaoheng Ni, Guangzhi Sun, Pingchuan Ma, Ruizhe Huang, Vineel Pratap, Yuekai Zhang, et al.

    IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) [paper] [code]

  2. TorchAudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in TorchAudio

    Anurag Kumar, Ke Tan, Zhaoheng Ni, Pranay Manocha, Xiaohui Zhang, Ethan Henderson, Buye Xu

    IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) [paper]

  3. Reducing Barriers to Self-Supervised Learning: HuBERT Pre-Training with Academic Compute

    William Chen, Xuankai Chang, Yifan Peng, Zhaoheng Ni, Soumi Maiti, Shinji Watanabe

    Interspeech

Earlier selected publications
2022–2019
  1. TorchAudio: Building Blocks for Audio and Speech Processing

    Yao-Yuan Yang, Moto Hira, Zhaoheng Ni, Anjali Chourdia, Artyom Astafurov, Caroline Chen, et al.

    ICASSP 2022 [paper] [code]

  2. WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation

    Zhaoheng Ni, Yong Xu, Meng Yu, Bo Wu, Shixiong Zhang, Dong Yu, Michael I. Mandel

    IEEE Spoken Language Technology Workshop (SLT), 2021 [paper]

  3. Mask-Dependent Phase Estimation for Monaural Speaker Separation

    Zhaoheng Ni, Michael I. Mandel

    ICASSP 2020

  4. ONSSEN: An Open-Source Speech Separation and Enhancement Library

    Zhaoheng Ni, Michael I. Mandel

    arXiv, 2019 [paper]

Professional Service

Open Source

I was a maintainer of TorchAudio, the official PyTorch library for audio and signal processing. My open-source work has supported speech recognition, self-supervised learning, audio processing, and speech quality evaluation.