Zhaoheng Ni
Research Scientist
Meta Reality Labs · New York, US
I am a research scientist at Meta Reality Labs working on generative models for audio, text, and video.
Previously, I was a maintainer of TorchAudio, the official audio library of PyTorch. Before Meta, I was a PhD student advised by Michael I. Mandel and an undergraduate student advised by Yan Xu.
Research Interests
- Generative audio and music: high-fidelity generation, editing, flow matching, and controllable synthesis.
- Multimodal generation: temporally aligned audio-video generation and visually guided audio creation.
- Speech processing: recognition, enhancement, separation, forced alignment, and quality evaluation.
- Efficient and open machine learning: data-efficient training, on-device inference, and production audio tools.
News
- Aug 2026
- I am co-organizing the ICASSP 2027 special session Speaker Grounding in Speech LLMs with Dr. Hao Shi, Dr. Hexin Liu, and colleagues. Paper submissions are welcome.
- Jun 2026
- One paper was accepted by Interspeech 2026.
- Jan 2026
- Three papers were accepted by ICASSP 2026.
- Aug 2025
- Marvin Sach, Robin Scheibler, and I are organizing the ICASSP 2026 special session Promise and Perils of Generative AI for the Evaluation of Speech and Audio.
- Aug 2025
- Two main papers and one demo paper were accepted by ASRU 2025.
- May 2025
- Two papers were accepted by Interspeech 2025.
- Dec 2024
- One paper was accepted by ICASSP 2025.
- Nov 2024
- We are organizing the URGENT 2025 Challenge at Interspeech 2025.
- Sep 2024
- Our MelodyFlow demo supports text-guided music editing and generation at a 48 kHz sample rate.
Earlier news
- Sep 2024
- Three papers were accepted by IEEE SLT 2024.
- Jun 2024
- We organized the Audio Imagination Workshop at NeurIPS 2024.
- May 2024
- We organized the URGENT Challenge at the NeurIPS 2024 Competition Track.
- Apr 2024
- Our MMS paper was accepted by the Journal of Machine Learning Research.
- Feb 2024
- We released demo videos and the paper for FoleyGen.
Selected Publications
For a complete and current list, see Google Scholar. My name is shown in bold.
-
Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) [paper]
-
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
arXiv preprint
-
High Fidelity Text-Guided Music Generation and Editing via Single-Stage Flow Matching
-
FoleyGen: Visually-Guided Audio Generation
IEEE Workshop on Machine Learning for Signal Processing (MLSP) [paper] [demo]
-
Data Efficient Reflow for Few Step Audio Generation
IEEE Spoken Language Technology Workshop (SLT)
-
Massively Multilingual Forced Aligner Leveraging Self-Supervised Discrete Units
IEEE Spoken Language Technology Workshop (SLT)
-
URGENT Challenge: Universality, Robustness, and Generalizability for Speech Enhancement
Interspeech [paper]
-
Scaling Speech Technology to 1,000+ Languages
-
TorchAudio 2.1: Advancing Speech Recognition, Self-Supervised Learning, and Audio Processing Components for PyTorch
IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) [paper] [code]
-
TorchAudio-Squim: Reference-Less Speech Quality and Intelligibility Measures in TorchAudio
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) [paper]
-
Reducing Barriers to Self-Supervised Learning: HuBERT Pre-Training with Academic Compute
Interspeech
Earlier selected publications
-
TorchAudio: Building Blocks for Audio and Speech Processing
-
WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation
IEEE Spoken Language Technology Workshop (SLT), 2021 [paper]
-
Mask-Dependent Phase Estimation for Monaural Speaker Separation
ICASSP 2020
-
ONSSEN: An Open-Source Speech Separation and Enhancement Library
arXiv, 2019 [paper]
Professional Service
- Co-organizer, Promise and Perils of Generative AI for the Evaluation of Speech and Audio, ICASSP 2026 special session.
- Organizer, URGENT 2025 Challenge, Interspeech 2025.
- Organizer, Audio Imagination Workshop, NeurIPS 2024.
- Organizer, URGENT Challenge, NeurIPS 2024 Competition Track.
Open Source
I was a maintainer of TorchAudio, the official PyTorch library for audio and signal processing. My open-source work has supported speech recognition, self-supervised learning, audio processing, and speech quality evaluation.