Area of research
Computer Vision and Pattern Recognition · Artificial Intelligence
Research interest
Research interests include Multimodal Machine Learning Applications, Domain Adaptation and Few-Shot Learning, Human Pose and Action Recognition, and Advanced Image and Video Retrieval Techniques.
DreamJourney: Perpetual View Generation With Video Diffusion Models
HiFi3D: Improving Text-to-3D With High-Fidelity Multi-View Diffusion
HiDream-I1: An Open-Source High-Efficient Image Generative Foundation Model
Stream-ViT: Learning Streamlined Convolutions in Vision Transformer
Kernel Masked Image Modeling Through the Lens of Theoretical Understanding
Exploring Vision-Language Foundation Model for Novel Object Captioning
Creatively Upscaling Images with Global-Regional Priors
Identity-Preserving Video Generation Challenge
Talk, Imagine, Evolve: A Unified Multimodal Agent for Seamless Visual Generation and Editing
Edit-by-Example: Adaptive Exemplar-Based Image Editing
HIRI-ViT: Scaling Vision Transformer With High Resolution Inputs
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
End-to-End Video Scene Graph Generation With Temporal Propagation Transformer
Improving Virtual Try-On with Garment-Focused Diffusion Models
Contextual Transformer Networks for Visual Recognition
Semantic-Conditional Diffusion Networks for Image Captioning*
Control3D: Towards Controllable Text-to-3D Generation
Control3D: Towards Controllable Text-to-3D Generation
ControlStyle: Text-Driven Stylized Image Generation Using Diffusion Priors
Retrieval Augmented Convolutional Encoder-decoder Networks for Video Captioning
ControlStyle: Text-Driven Stylized Image Generation Using Diffusion Priors
3DStyle-Diffusion: Pursuing Fine-grained Text-driven 3D Stylization with 2D Diffusion Models
A Low Rank Promoting Prior for Unsupervised Contrastive Learning
3D Creation at Your Fingertips: From Text or Image to 3D Assets
Boosting Relationship Detection in Images with Multi-Granular Self-Supervised Learning
Wave-ViT: Unifying Wavelet and Transformers for Visual Representation Learning
Retrieval Augmented Convolutional Encoder-decoder Networks for Video Captioning
Unpaired Image Captioning With semantic-Constrained Self-Learning
Bottom-up and Top-down Object Inference Networks for Image Captioning