← Search

Yun Cao

8 accepted papers

2026

MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention

AAAI 2026technical

Vision large language models (VLLMs) are focusing primarily on handling complex and fine-grained visual information by incorporating advanced vision encoders and scaling up visual models. However, these approaches face high training and inference costs, as well as challenges in extracting visual det

Cited by 0SourcePDFScholar
2026

SwiftVideo: A Unified Framework for Few-Step Video Generation Through Trajectory-Distribution Alignment

AAAI 2026technical

Diffusion-based or flow-based models have achieved significant progress in video synthesis but require multiple iterative sampling steps, which incurs substantial computational overhead. While many distillation methods that are solely based on trajectory-preserving or distribution-matching have been

Cited by 0SourcePDFScholar
2025

Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning

ICASSP 2025accepted

Although Large Language Models (LLMs) excel in reasoning and generation for language tasks, they are not specifically designed for multimodal challenges. Training Multimodal Large Language Models (MLLMs), however, is resource-intensive and constrained by various training limitations. In this paper,…

Cited by 0SourceScholar
2025

Object-Based Video Tampering Localization via Trace Consistency Analysis

ICASSP 2025accepted

With the rapid advancement of object-based video inpainting and splicing tampering techniques, the dissemination of malicious videos on the internet poses significant risks. Existing localization methods, however, exhibit limitations such as restriction to specific datasets, limited performance in d…

Cited by 0SourceScholar
2022

Learning To Memorize Feature Hallucination for One-Shot Image Generation

CVPR 2022poster

This paper studies the task of One-Shot image Generation (OSG), where generation network learned on base dataset should be generalizable to synthesize images of novel categories with only one available sample per novel category. Most existing methods for feature transfer in one-shot image generation…

Cited by 10PDFScholar
2022

SeedFormer: Patch Seeds Based Point Cloud Completion with Upsample Transformer

ECCV 2022poster

"Point cloud completion has become increasingly popular among generation tasks of 3D point clouds, as it is a challenging yet indispensable problem to recover the complete shape of a 3D object from its partial observation. In this paper, we propose a novel SeedFormer to improve the ability of detail…

2021

Frequency Consistent Adaptation for Real World Super Resolution

AAAI 2021technical

Recent deep-learning based Super-Resolution (SR) methods have achieved remarkable performance on images with known degradation. However, these methods always fail in real-world scene, since the Low-Resolution (LR) images after the ideal degradation (e.g., bicubic down-sampling) deviate from real sou…

Cited by 12SourcePDFScholar
2021

Spectrum-to-Kernel Translation for Accurate Blind Image Super-Resolution

NeurIPS 2021poster

Deep-learning based Super-Resolution (SR) methods have exhibited promising performance under non-blind setting where blur kernel is known; however, blur kernels of Low-Resolution (LR) images in different practical applications are usually unknown. It may lead to a significant performance drop when…

Cited by 27SourcePDFScholar