Area of research
Signal Processing · Artificial Intelligence
Research interest
Research interests include Computer science, Speech recognition, Artificial intelligence, Feature (linguistics), Encoder, and Sound (geography).
A Dual Consistency Training (DCT) strategy for polyphonic sound event detection
A Singing Melody Extraction Network Via Self-Distillation and Multi-Level Supervision
MN-Net: Speech Enhancement Network via Modeling the Noise
HANet: A Harmonic Attention-Based Network for Singing Melody Extraction from Polyphonic Music
SAT-SED: Semi-Supervised Sound Event Detection by Pseudo-Labeling with Self-Adaptive Threshold Strategy
TDE-VC: Timbre Disentanglement and Extraction Via Consistency for Zero-Shot Voice Conversion
GLFER-Net: a polyphonic sound source localization and detection network based on global-local feature extraction and recalibration
SMMA-Net: An Audio Clue-Based Target Speaker Extraction Network with Spectrogram Matching and Mutual Attention
Improving Speaker Verification With Noise-Aware Label Ensembling and Sample Selection: Learning and Correcting Noisy Speaker Labels
Introducing Multilingual Phonetic Information to Speaker Embedding for Speaker Verification
Speaker Recognition Based on Pre-Trained Model and Deep Clustering
Multi-branch Network with Cross-Domain Feature Fusion for Anomalous Sound Detection
A Lightweight Music Source Separation Model with Graph Convolution Network
SESNet: A Speech Enhancement and Separation Network in Noisy Reverberant Environments
ASD-Diff: Unsupervised Anomalous Sound Detection with Masked Diffusion Model
IIFC-Net: A Monaural Speech Enhancement Network With High-Order Information Interaction and Feature Calibration
W2VC: WavLM representation based one-shot voice conversion with gradient reversal distillation and CTC supervision
How to Boost Anti-Spoofing with X-Vectors
Speakeraugment: Data Augmentation for Generalizable Source Separation via Speaker Parameter Manipulation
CRA-DIFFUSE: Improved Cross-Domain Speech Enhancement Based on Diffusion Model with T-F Domain Pre-Denoising
Speech Topic Classification Based on Pre-trained and Graph Networks
A Polyphonic SELD Network Based on Attentive Feature Fusion and Multi-stage Training Strategy
Improved Self-Consistency Training with Selective Feature Fusion for Sound Event Detection
Hierarchic Temporal Convolutional Network With Cross-Domain Encoder for Music Source Separation
A Multi-grained based Attention Network for Semi-supervised Sound Event Detection
Dual-Path Hybrid Attention Network for Monaural Speech Separation
D<sup>2</sup>Net: A Denoising and Dereverberation Network Based on Two-branch Encoder and Dual-path Transformer
Multi-stage music separation network with dual-branch attention and hybrid convolution
Mining Hard Samples Locally And Globally For Improved Speech Separation
Self-Consistency Training with Hierarchical Temporal Aggregation for Sound Event Detection