← Search

Quan Chen

21 accepted papers

2026

Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step Generation

CVPR 2026

Video generation has recently emerged as a central task in the field of generative AI. However, the substantial computational cost inherent in video synthesis makes model distillation a critical technique for efficient deployment. Despite its significance, there is a scarcity of methods specifically

Cited by 0SourceScholar
2026

AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation

CVPR 2026

Short-form videos have become a primary medium for digital advertising, requiring scalable and efficient content creation. However, current workflows and AI tools remain disjoint and modality-specific, leading to high production costs and low overall efficiency. To address this issue, we propose Aut

Cited by 0SourcecodeScholar
2026

ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling

ICLR 2026poster

Process reward models (PRMs) have demonstrated significant efficacy in enhancing the mathematical reasoning capabilities of large language models (LLMs) by leveraging test-time scaling (TTS). However, while most PRMs exhibit substantial gains in mathematical domains, the scarcity of domain-specific…

Cited by 0SourceScholar
2026

Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning

CVPR 2026

We present Narrative Weaver, a novel framework that addresses a fundamental challenge in generative AI: achieving controllable, long-range, and consistent visual content generation. While existing models excel at generating high-fidelity short-form visual content, they struggle to maintain narrative

Cited by 0SourcecodeScholar
2026

Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders

ICML 2026poster

Recent progress in text-to-image (T2I) diffusion models (DMs) has enabled high-quality visual synthesis from diverse textual prompts. Yet, most existing T2I DMs, even those equipped with large language model (LLM)-based text encoders, remain text-pixel mappers -- they employ LLMs merely as text enco…

Cited by 0SourceScholar
2025

D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matching

AAAI 2025technical

Videos showcasing specific products are increasingly important for E-commerce. Key moments naturally exist as the first appearance of a specific product, presentation of its distinctive features, the presence of a buying link, etc. Adding proper sound effects (SFX) to such moments, or video decorati…

Cited by 0SourcePDFScholar
2025

Improving Preference Alignment of LLM with Inference-Free Self-Refinement

EMNLP 2025

Large language models (LLMs) develop the in-context learning capability through pretraining and instruction tuning, enabling task adaptation without parameter updates. Self-refinement is a manifestation of this capability, which allows LLMs to iteratively refine the output using self-generated feedb

2025

LEARN: Knowledge Adaptation from Large Language Model to Recommendation for Practical Industrial Application

AAAI 2025technical

Contemporary recommendation systems predominantly rely on ID embedding to capture latent associations among users and items. However, this approach overlooks the wealth of semantic information embedded within textual descriptions of items, leading to suboptimal performance and poor generalizations.…

2025

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

ICML 2025poster

We introduce Orthus, a unified multimodal model that excels in generating interleaved images and text from mixed-modality inputs by simultaneously handling discrete text tokens and continuous image features under the \textbf{AR} modeling principle. The continuous treatment of visual signals minimize…

Cited by 8SourcePDFScholar
2025

SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization

ICCV 2025poster

This paper presents the Semantic-aWarE spatial-tEmporal Tokenizer (SweetTok), a novel video tokenizer to overcome the limitations in current video tokenization methods for compacted yet effective discretization. Unlike previous approaches that process flattened local visual patches via direct discre…

Cited by 0SourcePDFScholar
2024

Quad Bayer Joint Demosaicing and Denoising Based on Dual Encoder Network with Joint Residual Learning

AAAI 2024technical

The recent imaging technology Quad Bayer CFA brings better imaging PSNR and higher visual quality compared to traditional Bayer CFA, but also serious challenges for demosaicing and denoising during the ISP pipeline. In this paper, we propose a novel dual encoder network, namely DRNet, to achieve joi…

Cited by 12SourcePDFScholar
2024

Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning

AAAI 2024technical

In recent years, text-to-video retrieval methods based on CLIP have experienced rapid development. The primary direction of evolution is to exploit the much wider gamut of visual and textual cues to achieve alignment. Concretely, those methods with impressive performance often design a heavy fusion…

2023

Cross-Domain Product Representation Learning for Rich-Content E-Commerce

ICCV 2023poster

The proliferation of short video and live-streaming platforms has revolutionized how consumers engage in online shopping. Instead of browsing product pages, consumers are now turning to rich-content e-commerce, where they can purchase products through dynamic and interactive media like short videos…

Cited by 3PDFcodeScholar
2023

Cross-view Semantic Alignment for Livestreaming Product Recognition

ICCV 2023poster

Live commerce is the act of selling products online through livestreaming. The customer's diverse demands for online products introduces more challenges to Livestreaming Product Recognition. Previous works are either focus on fashion clothing data or subject to single-modal input, thus inconsistent…

Cited by 4PDFcodeScholar
2023

Improving Dynamic HDR Imaging with Fusion Transformer

AAAI 2023technical

Reconstructing a High Dynamic Range (HDR) image from several Low Dynamic Range (LDR) images with different exposures is a challenging task, especially in the presence of camera and object motion. Though existing models using convolutional neural networks (CNNs) have made great progress, challenges s…

2022

DCCF: Deep Comprehensible Color Filter Learning Framework for High-Resolution Image Harmonization

ECCV 2022poster

"Image color harmonization algorithm aims to automatically match the color distribution of foreground and background images captured in different conditions. Previous deep learning based models neglect two issues that are critical for practical applications, namely high resolution (HR) image process…

2021

Cognitive Memory Constrained Human Decision Making based on Multi-source Information

ICASSP 2021accepted

Unlike decision making systems made up of physical sensors where the system parameters are known a priori and can be controlled at will, human behavior in decision making is complex and uncertain. The objective of this work is to study how humans make decisions based on internal and external sources…

Cited by 0SourceScholar
2020

How Far Does BERT Look At: Distance-based Clustering and Analysis of BERT’s Attention

COLING 2020main

Recent research on the multi-head attention mechanism, especially that in pre-trained models such as BERT, has shown us heuristics and clues in analyzing various aspects of the mechanism. As most of the research focus on probing tasks or hidden states, previous works have found some primitive patter…

Cited by 25SourcePDFScholar
2019

Adversarial Defense Through Network Profiling Based Path Extraction

CVPR 2019poster

Recently, researchers have started decomposing deep neural network models according to their semantics or functions. Recent work has shown the effectiveness of decomposed functional blocks for defending adversarial attacks, which add small input perturbation to the input image to fool the DNN models…

Cited by 65PDFScholar