← Search

Zhiwen Chen

8 accepted papers

2026

FHAvatar: Fast and High-Fidelity Reconstruction of Face-and-Hair Composable 3D Head Avatar from Few Casual Captures

CVPR 2026

We present FHAvatar, a novel framework for reconstructing 3D Gaussian avatars with composable face and hair components from an arbitrary number of views. Unlike previous approaches that couple facial and hair representations within a unified modeling process, we explicitly decouple two components in

Cited by 0SourceScholar
2026

Prism-MoE: Efficient Dense-to-MoE Conversion for Visual Autoregressive Generation

ICML 2026poster

Scaling up visual autoregressive models improves generation quality but incurs substantial inference costs. Mixture-of-Experts (MoE) architectures mitigate this issue through sparse activation and have proven effective in large language models. However, training MoE models from scratch remains prohi…

Cited by 0SourceScholar
2026

Scaling Dense Event-Stream Pretraining from Visual Foundation Models

CVPR 2026

Learning versatile, fine-grained representations from irregular event streams is pivotal yet nontrivial, primarily due to the heavy annotation that hinders scalability in dataset size, semantic richness, and application scope. To mitigate this dilemma, we launch a novel self-supervised pretraining m

Cited by 0SourcecodeScholar
2025

HCRMP: An LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving

NeurIPS 2025poster

Integrating the understanding and reasoning capabilities of Large Language Models (LLM) with the self-learning capabilities of Reinforcement Learning (RL) enables more reliable driving performance under complex driving conditions. There has been a lot of work exploring LLM-Dominated RL methods in th…

Cited by 0SourceScholar
2025

SaMer: A Scenario-aware Multi-dimensional Evaluator for Large Language Models

ICLR 2025poster

Evaluating the response quality of large language models (LLMs) for open-ended questions poses a significant challenge, especially given the subjectivity and multi-dimensionality of "quality" in natural language generation. Existing LLM evaluators often neglect that different scenarios require disti…

Cited by 0SourcePDFScholar
2025

TaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian Splatting

CVPR 2025highlight

Realistic 3D full-body talking avatars hold great potential in AR, with applications ranging from e-commerce live streaming to holographic communication. Despite advances in 3D Gaussian Splatting (3DGS) for lifelike avatar creation, existing methods struggle with fine-grained control of facial expre…

2024

Segment Any Event Streams via Weighted Adaptation of Pivotal Tokens

CVPR 2024poster

In this paper we delve into the nuanced challenge of tailoring the Segment Anything Models (SAMs) for integration with event data with the overarching objective of attaining robust and universal object segmentation within the event-centric domain. One pivotal issue at the heart of this endeavor is t…