Audio and Speech Processing

Authors and titles for recent submissions

See today's new changes

Total of 62 entries : 1-50 51-62

Showing up to 50 entries per page: fewer | more | all

[1] arXiv:2501.16761 [pdf, html, other]: Title: CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions

Xinfa Zhu, Wenjie Tian, Xinsheng Wang, Lei He, Xi Wang, Sheng Zhao, Lei Xie

Comments: 12 pages, 5 figures, 7 tables

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[2] arXiv:2501.16641 [pdf, html, other]: Title: SCDiar: a streaming diarization system based on speaker change detection and speech recognition

Naijun Zheng, Xucheng Wan, Kai Liu, Zhou Huan

Comments: Accepted at ICASSP 2025

Subjects: Audio and Speech Processing (eess.AS); Human-Computer Interaction (cs.HC); Sound (cs.SD)
[3] arXiv:2501.16542 [pdf, html, other]: Title: UniPET-SPK: A Unified Framework for Parameter-Efficient Tuning of Pre-trained Speech Models for Robust Speaker Verification

Mufan Sang, John H. L. Hansen

Comments: Accepted to IEEE/ACM Transactions on Audio, Speech, and Language Processing

Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[4] arXiv:2501.16367 [pdf, html, other]: Title: Neural Kalman Filters for Acoustic Echo Cancellation

Ernst Seidel, Gerald Enzner, Pejman Mowlaee, Tim Fingscheidt

Comments: Published in IEEE Signal Processing Magazine: Special Issue On Model-Based and Data-Driven Audio Signal Processing

Journal-ref: in IEEE Signal Processing Magazine, vol. 41, no. 6, pp. 24-38, Nov. 2024

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[5] arXiv:2501.16344 [pdf, html, other]: Title: WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning

Rajath Rao, Adithya Ganesan, Oscar Kjell, Jonah Luby, Akshay Raghavan, Scott Feltman, Whitney Ringwald, Ryan L. Boyd, Benjamin Luft, Camilo Ruggero, Neville Ryant, Roman Kotov, H. Andrew Schwartz

Comments: 13 pages, 6 figures, ACL ARR 2024

Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)
[6] arXiv:2501.16341 [pdf, other]: Title: Developing Enhanced Conversational Agents for Social Virtual Worlds

D. Griol, A. Sanchis, J. M. Molina, Z. Callejas

Comments: Neurocomputing 2019

Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[7] arXiv:2501.17048 (cross-list from q-bio.NC) [pdf, other]: Title: Cortical Temporal Mismatch Compensation in Bimodal Cochlear Implant Users: Selective Attention Decoding and Pupillometry Study

Hanna Dolhopiatenko, Waldo Nogueira

Comments: 28 pages, 15 figures

Subjects: Neurons and Cognition (q-bio.NC); Audio and Speech Processing (eess.AS)
[8] arXiv:2501.17011 (cross-list from cs.SD) [pdf, html, other]: Title: MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music Composition

Philippe Pasquier, Jeff Ens, Nathan Fradet, Paul Triana, Davide Rizzotti, Jean-Baptiste Rolland, Maryam Safi

Comments: AAAI 25

Subjects: Sound (cs.SD); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[9] arXiv:2501.16813 (cross-list from cs.CL) [pdf, other]: Title: Whispers of Sound-Enhancing Information Extraction from Depression Patients' Unstructured Data through Audio and Text Emotion Recognition and Llama Fine-tuning

Lindy Gan, Yifan Huang, Xiaoyang Gao, Jiaming Tan, Fujun Zhao, Tao Yang

Comments: 21 pages,7 figures.1 table

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[10] arXiv:2501.16780 (cross-list from cs.SD) [pdf, html, other]: Title: AVE Speech Dataset: A Comprehensive Benchmark for Multi-Modal Speech Recognition Integrating Audio, Visual, and Electromyographic Signals

Dongliang Zhou, Yakun Zhang, Jinghan Wu, Xingyu Zhang, Liang Xie, Erwei Yin

Subjects: Sound (cs.SD); Human-Computer Interaction (cs.HC); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[11] arXiv:2501.16643 (cross-list from cs.CL) [pdf, html, other]: Title: An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue

Koji Inoue, Divesh Lala, Mikey Elmers, Keiko Ochi, Tatsuya Kawahara

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[12] arXiv:2501.16471 (cross-list from cs.LG) [pdf, html, other]: Title: SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching Experiments

Simon Dahan, Gabriel Bénédict, Logan Z. J. Williams, Yourong Guo, Daniel Rueckert, Robert Leech, Emma C. Robinson

Comments: 27 pages, accepted to ICLR 2025

Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV); Neurons and Cognition (q-bio.NC)

[13] arXiv:2501.16201 [pdf, html, other]: Title: Enhancing and Exploring Mild Cognitive Impairment Detection with W2V-BERT-2.0

Yueguan Wang, Tatsunari Matsushima, Soichiro Matsushima, Toshimitsu Sakai

Comments: Submitted to ICASSP-SPADE workshop 2025

Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[14] arXiv:2501.16171 [pdf, html, other]: Title: Separate This, and All of these Things Around It: Music Source Separation via Hyperellipsoidal Queries

Karn N. Watcharasupat, Alexander Lerch

Comments: Submitted to the 2025 International Joint Conference on Artificial Intelligence

Subjects: Audio and Speech Processing (eess.AS); Information Retrieval (cs.IR); Machine Learning (cs.LG); Sound (cs.SD)
[15] arXiv:2501.15965 [pdf, html, other]: Title: EDSep: An Effective Diffusion-Based Method for Speech Source Separation

Jinwei Dong, Xinsheng Wang, Qirong Mao

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[16] arXiv:2501.15764 [pdf, html, other]: Title: Introducing RIFT: A Hierarchical Entropic Filtering Scheme for Ideal Time-Frequency Reconstruction

James M. Cozens, Simon J. Godsill

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Signal Processing (eess.SP)
[17] arXiv:2501.15744 [pdf, html, other]: Title: Noise disturbance and lack of privacy: Modeling acoustic dissatisfaction in open-plan offices

Manuj Yadav, Jungsoo Kim, Valtteri Hongisto, Densil Cabrera, Richard de Dear

Comments: The following article has been submitted to The Journal of the Acoustical Society of America. After it is published, it will be found at this https URL

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[18] arXiv:2501.15496 [pdf, html, other]: Title: Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer

Hu Hu, Sabato Marco Siniscalchi, Chao-Han Huck Yang, Chin-Hui Lee

Comments: Accepted to TASLP

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[19] arXiv:2501.15466 [pdf, html, other]: Title: End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario

Mohsen Ghane, Mohammad Sadegh Safari

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[20] arXiv:2501.16327 (cross-list from cs.CL) [pdf, html, other]: Title: LUCY: Linguistic Understanding and Control Yielding Early Stage of Her

Heting Gao, Hang Shao, Xiong Wang, Chaofan Qiu, Yunhang Shen, Siqi Cai, Yuchen Shi, Zihan Xu, Zuwei Long, Yike Zhang, Shaoqi Dong, Chaoyou Fu, Ke Li, Long Ma, Xing Sun

Comments: Demo Link: this https URL

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[21] arXiv:2501.15907 (cross-list from cs.SD) [pdf, html, other]: Title: Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation

Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu, Chen Yang, Jiaqi Li, Peiyang Shi, Yuancheng Wang, Kai Chen, Pengyuan Zhang, Zhizheng Wu

Comments: Extended version of arXiv:2407.05361, submitted to TASLP, dataset is available at: this https URL

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[22] arXiv:2501.15858 (cross-list from cs.CL) [pdf, other]: Title: Potential Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech

Eunjung Yeo, Julie Liss, Visar Berisha, David Mortensen

Comments: 10 pages, 1 figure, 2 tables

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[23] arXiv:2501.15613 (cross-list from cs.SD) [pdf, html, other]: Title: Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning

Qian Yang, Calbert Graham

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[24] arXiv:2501.15442 (cross-list from cs.SD) [pdf, html, other]: Title: Overview of the Amphion Toolkit (v0.2)

Jiaqi Li, Xueyao Zhang, Yuancheng Wang, Haorui He, Chaoren Wang, Li Wang, Huan Liao, Junyi Ao, Zeyu Xie, Yiqiao Huang, Junan Zhang, Zhizheng Wu

Comments: Github: this https URL

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[25] arXiv:2501.15417 (cross-list from cs.SD) [pdf, html, other]: Title: AnyEnhance: A Unified Generative Model with Prompt-Guidance and Self-Critic for Voice Enhancement

Junan Zhang, Jing Yang, Zihao Fang, Yuancheng Wang, Zehua Zhang, Zhuo Wang, Fan Fan, Zhizheng Wu

Comments: 12 pages, 4 figures

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[26] arXiv:2501.15368 (cross-list from cs.CL) [pdf, html, other]: Title: Baichuan-Omni-1.5 Technical Report

Yadong Li, Jun Liu, Tao Zhang, Tao Zhang, Song Chen, Tianpeng Li, Zehuan Li, Lijun Liu, Lingfeng Ming, Guosheng Dong, Da Pan, Chong Li, Yuanbo Fang, Dongdong Kuang, Mingrui Wang, Chenglin Zhu, Youwei Zhang, Hongyu Guo, Fengyu Zhang, Yuran Wang, Bowen Ding, Wei Song, Xu Li, Yuqi Huo, Zheng Liang, Shusen Zhang, Xin Wu, Shuai Zhao, Linchu Xiong, Yozhen Wu, Jiahui Ye, Wenhao Lu, Bowen Li, Yan Zhang, Yaqi Zhou, Xin Chen, Lei Su, Hongda Zhang, Fuzhong Chen, Xuezhen Dong, Na Nie, Zhiying Wu, Bin Xiao, Ting Li, Shunya Dang, Ping Zhang, Yijia Sun, Jincheng Wu, Jinjie Yang, Xionghai Lin, Zhi Ma, Kegeng Wu, Jia li, Aiyuan Yang, Hui Liu, Jianqiang Zhang, Xiaoxi Chen, Guangwei Ai, Wentao Zhang, Yicong Chen, Xiaoqin Huang, Kun Li, Wenjing Luo, Yifei Duan, Lingling Zhu, Ran Xiao, Zhe Su, Jiani Pu, Dian Wang, Xu Jia, Tianyu Zhang, Mengyu Ai, Mang Wang, Yujing Qiao, Lei Zhang, Yanjun Shen, Fan Yang, Miao Zhen, Yijie Zhou, Mingyang Chen, Fei Li, Chenzheng Zhu, Keer Lu, Yaqi Zhao, Hao Liang, Youquan Li, Yanzhao Qin, Linzhuang Sun, Jianhua Xu, Haoze Sun, Mingan Lin, Zenan Zhou, Weipeng Chen

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[27] arXiv:2501.15310 (cross-list from cs.CL) [pdf, html, other]: Title: The Multicultural Medical Assistant: Can LLMs Improve Medical ASR Errors Across Borders?

Ayo Adedeji, Mardhiyah Sanni, Emmanuel Ayodele, Sarita Joshi, Tobi Olatunji

Comments: 15 pages, 8 figures

Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[28] arXiv:2501.15304 (cross-list from cs.SD) [pdf, other]: Title: Music Generation using Human-In-The-Loop Reinforcement Learning

Aju Ani Justus

Comments: This is a preprint of a paper presented at the 2023 IEEE International Conference on Big Data (BigData). It has been made public for the benefit of the community and should be considered a preprint rather than a formally reviewed paper

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[29] arXiv:2501.15302 (cross-list from cs.SD) [pdf, html, other]: Title: The ICME 2025 Audio Encoder Capability Challenge

Junbo Zhang, Heinrich Dinkel, Qiong Song, Helen Wang, Yadong Niu, Si Cheng, Xiaofeng Xin, Ke Li, Wenwu Wang, Yujun Wang, Jian Luan

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[30] arXiv:2501.15177 (cross-list from cs.SD) [pdf, html, other]: Title: Audio-Language Models for Audio-Centric Tasks: A survey

Yi Su, Jisheng Bai, Qisheng Xu, Kele Xu, Yong Dou

Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[31] arXiv:2501.15032 (cross-list from cs.SD) [pdf, html, other]: Title: Stealthy Voice Eavesdropping with Acoustic Metamaterials: Unraveling a New Privacy Threat

Zhiyuan Ning, Zhanyong Tang, Juan He, Weizhi Meng, Yuntian Chen

Subjects: Sound (cs.SD); Cryptography and Security (cs.CR); Audio and Speech Processing (eess.AS)
[32] arXiv:2501.14994 (cross-list from cs.SD) [pdf, html, other]: Title: Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition

Satwinder Singh, Qianli Wang, Zihan Zhong, Clarion Mendes, Mark Hasegawa-Johnson, Waleed Abdulla, Seyed Reza Shahamiri

Comments: Accepted to ICASSP 2025

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[33] arXiv:2501.14790 (cross-list from q-bio.NC) [pdf, other]: Title: Towards Dynamic Neural Communication and Speech Neuroprosthesis Based on Viseme Decoding

Ji-Ha Park, Seo-Hyun Lee, Soowon Kim, Seong-Whan Lee

Comments: 5 pages, 5 figures, 1 table, Name of Conference: 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing

Subjects: Neurons and Cognition (q-bio.NC); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[34] arXiv:2501.14788 (cross-list from cs.SD) [pdf, html, other]: Title: Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages

Alexan Ayrapetyan, Sofia Kostandian, Ara Yeroyan, Mher Yerznkanyan, Nikolay Karpov, Nune Tadevosyan, Vitaly Lavrukhin, Boris Ginsburg

Comments: The first four authors contributed equally

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

[35] arXiv:2501.14680 [pdf, html, other]: Title: Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning

Jisi Zhang, Pablo Peso Parada, Md Asif Jalal, Karthikeyan Saravanan

Comments: Accepted at ICASSP 2025

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[36] arXiv:2501.14477 [pdf, html, other]: Title: Enhancing Intelligibility for Generative Target Speech Extraction via Joint Optimization with Target Speaker ASR

Hao Ma, Rujin Chen, Ruihao Jing, Xiao-Lei Zhang, Ju Liu, Xuelong Li

Comments: demo: this https URL

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[37] arXiv:2501.14350 [pdf, html, other]: Title: FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration

Kai-Tuo Xu, Feng-Long Xie, Xu Tang, Yao Hu

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[38] arXiv:2501.14273 [pdf, html, other]: Title: Characteristic-Specific Partial Fine-Tuning for Efficient Emotion and Speaker Adaptation in Codec Language Text-to-Speech Models

Tianrui Wang, Meng Ge, Cheng Gong, Chunyu Qiang, Haoyu Wang, Zikang Huang, Yu Jiang, Xiaobao Wang, Xie Chen, Longbiao Wang, Jianwu Dang

Comments: 13 pages

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[39] arXiv:2501.14240 [pdf, html, other]: Title: Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation

Wen Huang, Yanmei Gu, Zhiming Wang, Huijia Zhu, Yanmin Qian

Comments: Accepted to ICASSP 2025

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[40] arXiv:2501.14610 (cross-list from cs.SD) [pdf, html, other]: Title: Leveraging Spatial Cues from Cochlear Implant Microphones to Efficiently Enhance Speech Separation in Real-World Listening Scenes

Feyisayo Olalere, Kiki van der Heijden, Christiaan H. Stronks, Jeroen Briaire, Johan HM Frijns, Marcel van Gerven

Comments: 10 pages, 5 figures

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)

[41] arXiv:2501.13884 [pdf, html, other]: Title: Exploring Finetuned Audio-LLM on Heart Murmur Features

Adrian Florea, Xilin Jiang, Nima Mesgarani, Xiaofan Jiang

Comments: 5 pages, 1 figure, and 3 tables. Submitted to IEEE/ACM Conference on Connected Health: Applications, Systems , and Engineering Technologies

Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Sound (cs.SD)
[42] arXiv:2501.13642 [pdf, html, other]: Title: Learning-based A Posteriori Speech Presence Probability Estimation and Applications

Shuai Tao, Jesper Rindom Jensen, Yang Xiang, Himavanth Reddy, Qingzheng Zhang, Mads Græsbøll Christensen

Subjects: Audio and Speech Processing (eess.AS)
[43] arXiv:2501.13372 [pdf, html, other]: Title: Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement

Jae-Sung Bae, Anastasia Kuznetsova, Dinesh Manocha, John Hershey, Trausti Kristjansson, Minje Kim

Comments: Accepted to ICASSP 2025 Satellite Workshop: Generative Data Augmentation for Real-World Signal Processing Applications

Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI)
[44] arXiv:2501.13250 [pdf, html, other]: Title: Generative Data Augmentation Challenge: Synthesis of Room Acoustics for Speaker Distance Estimation

Jackie Lin, Georg Götz, Hermes Sampedro Llopis, Haukur Hafsteinsson, Steinar Guðjónsson, Daniel Gert Nielsen, Finnur Pind, Paris Smaragdis, Dinesh Manocha, John Hershey, Trausti Kristjansson, Minje Kim

Comments: Accepted to the Workshop on Generative Data Augmentation at ICASSP 2025. Challenge website: this https URL

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[45] arXiv:2501.13887 (cross-list from cs.LG) [pdf, html, other]: Title: What Does an Audio Deepfake Detector Focus on? A Study in the Time Domain

Petr Grinberg, Ankur Kumar, Surya Koppisetti, Gaurav Bharaj

Comments: Accepted to ICASSP 2025

Subjects: Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[46] arXiv:2501.13870 (cross-list from cs.SD) [pdf, html, other]: Title: Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference

Shuqi Dai, Yunyun Wang, Roger B. Dannenberg, Zeyu Jin

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[47] arXiv:2501.13772 (cross-list from cs.SD) [pdf, html, other]: Title: Tune In, Act Up: Exploring the Impact of Audio Modality-Specific Edits on Large Audio Language Models in Jailbreak

Erjia Xiao, Hao Cheng, Jing Shao, Jinhao Duan, Kaidi Xu, Le Yang, Jindong Gu, Renjing Xu

Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[48] arXiv:2501.13720 (cross-list from cs.CL) [pdf, html, other]: Title: Musical ethnocentrism in Large Language Models

Anna Kruspe

Journal-ref: Proceedings of the 3rd Workshop on NLP for Music and Audio (NLP4MusA) 2024

Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[49] arXiv:2501.13497 (cross-list from cs.SD) [pdf, html, other]: Title: DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition

Qijie Shao, Linhao Dong, Kun Wei, Sining Sun, Lei Xie

Comments: Submitted to the IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)

Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[50] arXiv:2501.13465 (cross-list from cs.SD) [pdf, html, other]: Title: Neural Vocoders as Speech Enhancers

Andong Li, Zhihang Sun, Fengyuan Hao, Xiaodong Li, Chengshi Zheng

Comments: 6 pages, 3 figures

Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)

Total of 62 entries : 1-50 51-62

Showing up to 50 entries per page: fewer | more | all

Audio and Speech Processing

Authors and titles for recent submissions

Wed, 29 Jan 2025 (showing 12 of 12 entries )

Tue, 28 Jan 2025 (showing 22 of 22 entries )

Mon, 27 Jan 2025 (showing 6 of 6 entries )

Fri, 24 Jan 2025 (showing first 10 of 12 entries )