← Search

Yongjie Zhu

8 accepted papers

2026

Beyond the Golden Data: Resolving the Motion-Vision Quality Dilemma via Timestep Selective Training

CVPR 2026

Recent advances in video generation models have achieved impressive results. However, these models heavily rely on the use of high-quality data that combines both high visual quality and high motion quality. In this paper, we identify a key challenge in video data curation: the Motion-Vision Quality

Cited by 0SourceScholar
2025

MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding

ICML 2025spotlight

Multimodal large language models (MLLMs) recently showed strong capacity in integrating data among multiple modalities, empowered by generalizable attention architecture. Advanced methods predominantly focus on language-centric tuning while less exploring multimodal tokens mixed through attention, p…

Cited by 0SourcePDFScholar
2025

VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models

NeurIPS 2025poster

Understanding and predicting emotions from videos has gathered significant attention in recent studies, driven by advancements in video large language models (VideoLLMs). While advanced methods have made progress in video emotion analysis, the intrinsic nature of emotions—characterized by their open…

Cited by 0SourceScholar
2023

Complementary Intrinsics From Neural Radiance Fields and CNNs for Outdoor Scene Relighting

CVPR 2023poster

Relighting an outdoor scene is challenging due to the diverse illuminations and salient cast shadows. Intrinsic image decomposition on outdoor photo collections could partly solve this problem by weakly supervised labels with albedo and normal consistency from multi-view stereo. With neural radiance…

Cited by 9SourcePDFScholar
2022

Estimating Spatially-Varying Lighting in Urban Scenes with Disentangled Representation

ECCV 2022poster

"We present an end-to-end network for spatially-varying outdoor lighting estimation in urban scenes given a single limited field-of-view LDR image and any assigned 2D pixel position. We use three disentangled latent spaces learned by our network to represent sky light, sun light, and lighting-indepe…

Cited by 15SourcePDFScholar
2021

Disentangling Identifiable Features from Noisy Data with Structured Nonlinear ICA

NeurIPS 2021poster

We introduce a new general identifiable framework for principled disentanglement referred to as Structured Nonlinear Independent Component Analysis (SNICA). Our contribution is to extend the identifiability theory of deep generative models for a very broad class of structured models. While previous…

2019

Measuring the Task Induced Oscillatory Brain Activity Using Tensor Decomposition

ICASSP 2019accepted

The characterization of dynamic electrophysiological brain activity, which form and dissolve in order to support ongoing cognitive function, is one of the most important goals in neuroscience. Here, we introduce a method with tensor decomposition for measuring the task-induced oscillations in the hu…

Cited by 0SourceScholar