← Search

Keyu Chen

17 accepted papers

2026

Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation

ICML 2026poster

Text-to-Image (T2I) models and Unified Multimodal Models (UMMs) have achieved remarkable progress in visual generation. However, their reliance on a single-pass generation paradigm limits their ability to handle complex prompts requiring iterative refinement. To enable multi-round Reflective Visual …

Cited by 0SourceScholar
2026

ExpertAD: Enhancing Autonomous Driving Systems with Mixture of Experts

AAAI 2026technical

Recent advancements in end-to-end autonomous driving systems (ADSs) underscore their potential for perception and planning capabilities. However, challenges remain. Complex driving scenarios contain rich semantic information, yet ambiguous or noisy semantics can compromise decision reliability, whil

Cited by 0SourcePDFScholar
2026

FIRM-MoE:Fine-GrainedExpert Decomposition for Resource-Adaptive MoE Inference

AAAI 2026technical

Mixture-of-Experts (MoE) is a sparse neural architecture that significantly increases model capacity while maintaining low computational complexity. However, deploying MoE-based large language models (LLMs) on memory-constrained edge devices remains challenging due to their substantial memory requir

Cited by 0SourcePDFScholar
2026

FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging

ICLR 2026oral

Although Video Large Language Models (VLLMs) have shown remarkable capabilities in video understanding, they are required to process high volumes of visual tokens, causing significant computational inefficiency. Existing VLLMs acceleration frameworks usually compress spatial and temporal redundancy…

Cited by 0SourcecodeScholar
2026

ForSim: Stepwise Forward Simulation for Traffic Policy Fine-Tuning

ICRA 2026poster

As the foundation of closed-loop training and evaluation in autonomous driving, traffic simulation still faces two fundamental challenges: covariate shift introduced by open-loop imitation learning and limited capacity to reflect the multimodal behaviors observed in real-world traffic. Although rece…

2026

Training-Free Hashing-Based Attention via Binary Principal Components

ICML 2026poster

Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-value (KV) caches. Existing sparse attention reduce…

Cited by 0SourceScholar
2024

FREA: Feasibility-Guided Generation of Safety-Critical Scenarios with Reasonable Adversariality

CoRL 2024poster

Generating safety-critical scenarios, which are essential yet difficult to collect at scale, offers an effective method to evaluate the robustness of autonomous vehicles (AVs). Existing methods focus on optimizing adversariality while preserving the naturalness of scenarios, aiming to achieve a bala…

Cited by 3SourceScholar
2024

iMCB-PGO: Incremental Minimum Cycle Basis Construction and Application to Online Pose Graph Optimization

RA-L 2024

Pose graph optimization (PGO) is a fundamental technique for robot localization. It is typically encoded with a sparse graph. The recent work on the cycle-based PGO reveals the merits of solving PGOs in the graph cycle space, which brings the computation of the minimum cycle basis (MCB) into the rob

Cited by 1SourceScholar
2023

CO-NET: Classification-Oriented Point Cloud Sampling via Informative Feature Learning and Non-Overlapped Local Adjustment

ICASSP 2023accepted

Recent studies have proven the strength of task-oriented point cloud sampling methods over traditional non-learned ones. However, previous task-oriented samplers are not adequate to extract local details and spatial patterns of point clouds, limiting the quality of synthesized points. In this paper,…

Cited by 0SourceScholar
2023

Collaborative Noisy Label Cleaner: Learning Scene-Aware Trailers for Multi-Modal Highlight Detection in Movies

CVPR 2023poster

Movie highlights stand out of the screenplay for efficient browsing and play a crucial role on social media platforms. Based on existing efforts, this work has two observations: (1) For different annotators, labeling highlight has uncertainty, which leads to inaccurate and time-consuming annotations…

2023

DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-Centric Rendering

ICCV 2023poster

Realistic human-centric rendering plays a key role in both computer vision and computer graphics. Rapid progress has been made in the algorithm aspect over the years, yet existing human-centric rendering datasets and benchmarks are rather impoverished in terms of diversity (e.g., outfit's fabric/mat…

Cited by 61PDFcodeScholar
2023

NewsNet: A Novel Dataset for Hierarchical Temporal Segmentation

CVPR 2023poster

Temporal video segmentation is the get-to-go automatic video analysis, which decomposes a long-form video into smaller components for the following-up understanding tasks. Recent works have studied several levels of granularity to segment a video, such as shot, event, and scene. Those segmentations…

2022

Scene Consistency Representation Learning for Video Scene Segmentation

CVPR 2022poster

A long-term video, such as a movie or TV show, is composed of various scenes, each of which represents a series of shots sharing the same semantic story. Spotting the correct scene boundary from the long-term video is a challenging task, since a model must understand the storyline of the video to fi…

Cited by 20PDFcodeScholar
2021

AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis

ICCV 2021poster

Generating high-fidelity talking head video by fitting with the input audio sequence is a challenging problem that receives considerable attentions recently. In this paper, we address this problem with the aid of neural scene representation networks. Our method is completely different from existing…

Cited by 452PDFcodeScholar
2021

Learning a Facial Expression Embedding Disentangled From Identity

CVPR 2021poster

The facial expression analysis requires a compact and identity-ignored expression representation. In this paper, we model the expression as the deviation from the identity by a subtraction operation, extracting a continuous and identity-invariant expression embedding. We propose a Deviation Learning…

Cited by 92PDFScholar