Jiaqi Li's Homepage

iclr2026.jpg
jiaqili3[at]link.cuhk.edu.cn

I am a third-year Ph.D. student at the Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), SDS, supervised by Professor Zhizheng Wu. Before that, I received B.S. degree at CUHK-Shenzhen.

My research interest includes speech language models, text-to-speech synthesis, and neural audio coding. I am one of the main contributors and leaders of the open-source Amphion toolkit. My recent work includes FlexiSLM, a spoken language model with dynamic and controllable frame rates, as well as DualCodec and FlexiCodec for low-frame-rate speech generation.

news

Aug 21, 2026 Our SimulS2ST-Omni paper was accepted to EMNLP 2026 Main Conference!
Aug 21, 2026 Our FlexiSLM paper was accepted to EMNLP 2026 Main Conference!
Apr 26, 2026 I presented our FlexiCodec paper at ICLR 2026 in Brazil. ICLR 2026
May 17, 2025 Our DualCodec paper was accepted to InterSpeech 2025!
Feb 01, 2025 We released the Amphion v0.2 technical report, summarizing our development of Amphion in 2024.
Dec 03, 2024 I presented our new paper, Investigating neural audio codecs for speech language model-based speech generation in SLT 2024. SLT 2024
Aug 25, 2024 🎉 Our papers, Amphion and Emila, got accepted by IEEE SLT 2024!
Jul 28, 2024 🔥 We released Emila: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation, with 101k hours of speech in six languages and features diverse speech with varied speaking styles.
Apr 19, 2024 I presented our paper, An initial investigation of neural replay simulator for over-the-air adversarial perturbations to automatic speaker verification in ICASSP 2024 in Korea! SLT 2024
Nov 26, 2023 🔥 We released Amphion v0.1 GitHub stars, which is an open-source toolkit for audio, music, and speech generation.

selected publications

  1. EMNLP 2026
    FlexiSLM: A Spoken Language Model with Dynamic and Controllable Frame Rates
    Jiaqi Li, Chaoren Wang, Xiaohai Tian, Mingjie Chen, Xinyu Liang, Xu Li, Yufan Lin, Junwen Qiu, Jun Zhang, Lu Lu, Haizhou Li, and Zhizheng Wu
    In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing, 2026
    TL;DR: We introduce the first spoken language model with dynamic and controllable frame rates for speech input and output.
  2. EMNLP 2026
    SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
    Rongshen He, Xinyu Liang, Dekun Chen, Jiaqi Li, Mingjie Chen, and Zhizheng Wu
    In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing, 2026
    TL;DR: We enable data-efficient long-form streaming speech-to-speech translation with joint text-code trajectory supervision.
  3. ICLR 2026
    FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
    Jiaqi Li, Yao Qian, Yuxuan Hu, Leying Zhang, Xiaofei Wang, Heng Lu, Manthan Thakker, Jinyu Li, Sheng Zhao, and Zhizheng Wu
    In International Conference on Learning Representations (ICLR), 2026
    TL;DR: We develop a dynamic neural audio codec with controllable low frame rates for efficient speech generation.
  4. Interspeech 2025
    DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
    Jiaqi Li, Xiaolong Lin, Zhekai Li, Shixi Huang, Yuancheng Wang, Chaoren Wang, Zhenpeng Zhan, and Zhizheng Wu
    In Proceedings of Interspeech 2025, 2025
    TL;DR: We propose a low-frame-rate neural audio codec that strengthens semantic information for speech generation.
  5. Tech Report
    Overview of the Amphion Toolkit (v0. 2)
    Jiaqi Li, Xueyao Zhang, Yuancheng Wang, Haorui He, Chaoren Wang, Li Wang, Huan Liao, Junyi Ao, Zeyu Xie, Yiqiao Huang, and  others
    arXiv preprint arXiv:2501.15442, 2025
    TL;DR: This is the technical report for the second version of the Amphion toolkit.
  6. SLT 2024
    Investigating neural audio codecs for speech language model-based speech generation
    Jiaqi Li, Dongmei Wang, Xiaofei Wang, Yao Qian, Long Zhou, Shujie Liu, Midia Yousefi, Canrun Li, Chung-Hsien Tsai, Zhen Xiao, and  others
    In 2024 IEEE Spoken Language Technology Workshop (SLT), 2024
  7. SLT 2024
    Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
    Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu, Chen Yang, Jiaqi Li, Peiyang Shi, Yuancheng Wang, Kai Chen, Pengyuan Zhang, and Zhizheng Wu
    In 2024 IEEE Spoken Language Technology Workshop (SLT), 2024
    TL;DR: We collect a 100k hours in-the-wild speech dataset for speech generation.
  8. SLT 2024
    Amphion: an Open-Source Audio, Music, and Speech Generation Toolkit
    Xueyao Zhang*, Liumeng Xue*, Yicheng Gu*, Yuancheng Wang*, Jiaqi Li, Haorui He, Chaoren Wang, Songting Liu, Xi Chen, Junan Zhang, Tze Ying Tang, Lexiao Zou, Mingxuan Wang, Jun Han, Kai Chen, Haizhou Li, and Zhizheng Wu
    In 2024 IEEE Spoken Language Technology Workshop (SLT), 2024
    TL;DR: We develop a unified toolkit for audio, music, and speech generation.
  9. ICASSP 2024
    An initial investigation of neural replay simulator for over-the-air adversarial perturbations to automatic speaker verification
    Jiaqi Li, Li Wang, Liumeng Xue, Lei Wang, and Zhizheng Wu
    In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024
  10. ICASSP 2024
    Advsv: An over-the-air adversarial attack dataset for speaker verification
    Li Wang, Jiaqi Li, Yuhao Luo, Jiahao Zheng, Lei Wang, Hao Li, Ke Xu, Chengfang Fang, Jie Shi, and Zhizheng Wu
    In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024