Area of research
Artificial Intelligence · Signal Processing
Research interest
Research interests include Speech Recognition and Synthesis, Speech and Audio Processing, Topic Modeling, and Natural Language Processing Techniques.
A Survey on Speech Large Language Models for Understanding
Recent Advances in Discrete Speech Tokens: A Review.
SelfSE: Self-Supervised Speech Enhancement via Noisy Speech Refinement
Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
Developing ChemDFM as a large language foundation model for chemistry
MOS-GAN: Mean Opinion Score GAN for Unsupervised Speech Enhancement
MFA-KWS: Effective Keyword Spotting With Multi-Head Frame-Asynchronous Decoding
DFM: Dialogue foundation model for universal large-scale dialogue-oriented task learning
EveMRC: Two-Stage Bidirectional Evidence Modeling for Multi-Choice Machine Reading Comprehension
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
Beyond the Status Quo: A Contemporary Survey of Advances and Challenges in Audio Captioning
Spatial-temporal distribution of labeled set bias remote sensing estimation: An implication for supervised machine learning in water quality monitoring
Towards Weakly Supervised Text-to-Audio Grounding
Genre: generative multi-turn question answering with contrastive learning for entity–relation extraction
Unsupervised Speech Enhancement Using Optimal Transport and Speech Presence Probability
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
E$^{3}$TTS: End-to-End Text-Based Speech Editing TTS System and Its Applications
DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
Speech Enhancement With Integration of Neural Homomorphic Synthesis and Spectral Masking
Speaker Adaptive Text-to-Speech With Timbre-Normalized Vector-Quantized Feature
A Heterogeneous Graph to Abstract Syntax Tree Framework for Text-to-SQL.
Semantic Enhancement Framework for Robust Speech Recognition
Phone-Level Prosody Modelling With GMM-Based MDN for Diverse and Controllable Speech Synthesis
Neural Fusion for Voice Cloning
Data augmentation based non-parallel voice conversion with frame-level speaker disentangler
Towards Duration Robust Weakly Supervised Sound Event Detection
Voice Activity Detection in the Wild: A Data-Driven Approach Using Teacher-Student Training
Towards a new generation of artificial intelligence in China