← Search

Zixuan Wang

37 accepted papers

2026

Chain of Event-Centric Causal Thought for Physically Plausible Video Generation

CVPR 2026

Physically Plausible Video Generation (PPVG) has emerged as a promising avenue for modeling real-world physical phenomena. PPVG requires an understanding of commonsense knowledge, which remains a challenge for video diffusion models. Current approaches leverage commonsense reasoning capability of la

Cited by 0SourcecodeScholar
2026

FACESLEUTH-R: ADAPTIVE ORIENTATION-AWARE ATTENTION FOR ROBUST MICRO-EXPRESSION RECOGNITION

ICASSP 2026oral

Micro-expression recognition (MER) has achieved impressive accuracy in controlled laboratory settings. However, its real-world applicability faces a significant generalization cliff, severely hindering practical deployment due to poor performance on unseen data and susceptibility to domain shifts. E…

Cited by 0SourcePDFScholar
2026

GaussianMatch: Semi-Supervised Regression with Pseudo-Label Filtering via Multi-View Gaussian Consistency

CVPR 2026

Semi-Supervised Regression (SSR) is essential in domains like sentiment analysis and healthcare where labeled data is limited but unlabeled data is plentiful. Despite its practical importance, SSR remains underexplored due to the lack of effective pseudo-labeling strategies for continuous outputs. U

Cited by 0SourcecodeScholar
2026

Hydra-Nav: Object Navigation via Adaptive Dual-Process Reasoning

ICML 2026poster

While large vision-language models (VLMs) show promise for object goal navigation, current methods still struggle with low success rates and inefficient localization of unseen objects—failures primarily attributed to weak temporal-spatial reasoning. Meanwhile, recent attempts to inject reasoning int…

Cited by 0SourceScholar
2026

Imitating the Truth: Attention-aware Truth-Guided Enhancement for Hallucination Mitigation in Large Vision-Language Models

ICLR 2026poster

Large Vision-Language Models (LVLMs) achieve impressive multimodal reasoning but remain prone to hallucinations, generating content inconsistent with visual evidence. Existing mitigation methods often rely on auxiliary modules or coarse decoding-time adjustments, overlooking the fine-grained dynamic…

Cited by 0SourceScholar
2026

Is Data Shapley Not Better than Random in Data Selection? Ask NASH

ICML 2026spotlight

Data selection studies the problem of identifying high-quality subsets of training data. While some existing works have considered selecting the subset of data with top-$m$ Data Shapley or other semivalues as they account for the interaction among every subset of data, other works argue that Data Sh…

Cited by 0SourceScholar
2026

Milestone-Guided Policy Learning for Long-Horizon Language Agents

ICML 2026poster

While long-horizon agentic tasks require language agents to perform dozens of sequential decisions, training such agents with reinforcement learning remains challenging. We identify two root causes: credit misattribution, where correct early actions are penalized due to terminal failures, and sample…

Cited by 0SourceScholar
2026

One Flow Fits All! A Scale-Aware Generative Framework for Diverse Data

IJCAI 2026

Real-world systems increasingly require coherent reasoning and generation over diverse data modalities simultaneously. Current generative frameworks rely on complex, multi-stage training, resulting in low efficiency due to iterative inference and high computational cost. They also struggle with unif

Cited by 0Scholar
2026

SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models

ICLR 2026poster

Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We identify that this limitation stems from a critical gap: existing methods attempt to learn spatial reasoning directly with…

Cited by 0SourcecodeScholar
2026

Training-free Motion Factorization for Compositional Video Generation

CVPR 2026

Compositional video generation aims to synthesize multiple instances with diverse appearance and motion. However, current approaches mainly focus on binding semantics, neglecting to understand diverse motion categories specified in prompts. In this paper, we propose a motion factorization framework

Cited by 0SourcecodeScholar
2026

WET: Mitigating World-Conditioned Knowledge Conflicts via World Entropy Tethering

ICML 2026poster

Large language models (LLMs) face a "loyalty dilemma" when correctness is conditioned on an active world-of-discourse. We identify a systemic failure mode---world misattribution---where models implicitly ground generation in an incompatible regime and drift from the target world. We propose World En…

Cited by 0SourceScholar
2025

Beyond Words: Augmenting Discriminative Richness via Diffusions in Unsupervised Prompt Learning

CVPR 2025poster

Fine-tuning vision-language models (VLMs) with large amounts of unlabeled data has recently garnered significant interest. However, a key challenge remains the lack of high-quality pseudo-labeled data. Current pseudo-labeling strategies often struggle with mismatches between semantic and visual info…

2025

COLA: Characterizing and Optimizing the Tail Latency for Safe Level-4 Autonomous Vehicle Systems

ICRA 2025

Autonomous vehicles (AVs) systems are envisioned to revolutionize our life by providing safe, relaxing, and convenient ground transportation. To ensure safety, AV systems need to make timely driving decisions in response to complicated and highly dynamic real-world driving environments. We present a

Cited by 4SourceScholar
2025

CoPRA: Bridging Cross-domain Pretrained Sequence Models with Complex Structures for Protein-RNA Binding Affinity Prediction

AAAI 2025technical

Accurately measuring protein-RNA binding affinity is crucial in many biological processes and drug design. Previous computational methods for protein-RNA binding affinity prediction rely on either sequence or structure features, unable to capture the binding mechanisms comprehensively. The recent em…

2025

Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Prediction

EMNLP 2025

Natural Language Inference (NLI) is a fundamental task in natural language processing. While NLI has developed many subdirections such as sentence-level NLI, document-level NLI and cross-lingual NLI, Cross-Document Cross-Lingual NLI (CDCL-NLI) remains largely unexplored. In this paper, we propose a

Cited by 0SourcePDFScholar
2025

FocusLLM: Precise Understanding of Long Context by Dynamic Condensing

ACL 2025long

Empowering LLMs with the ability to precisely understand long contexts is crucial for many downstream applications. However, handling long contexts with conventional transformer architecture requires substantial training and inference resources. Existing context condensing methods cannot accurately…

2025

FontAnimate: High Quality Few-shot Font Generation via Animating Font Transfer Process

ICCV 2025poster

Few-shot font generation (FFG) aims to create new font images by imitating the style from a limited set of reference images, while maintaining the content from the source images. Although this task has achieved significant progress, most existing methods still suffer from the incorrect generation of…

2025

KARMA: Augmenting Embodied AI Agents with Long-and-Short Term Memory Systems

ICRA 2025

Embodied AI agents responsible for executing interconnected, long-sequence household tasks often face difficulties with in-context memory, leading to inefficiencies and errors in task execution. To address this issue, we introduce KARMA, an innovative memory system that integrates longterm and short

Cited by 23SourcecodeScholar
2025

LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?

NeurIPS 2025poster

Recent reports claim that large language models (LLMs) now outperform elite humans in competitive programming. Drawing on knowledge from a group of medalists in international algorithmic contests, we revisit this claim, examining how LLMs differ from human experts and where limitations still remain.…

Cited by 0SourceScholar
2025

MeGA: Hybrid Mesh-Gaussian Head Avatar for High-Fidelity Rendering and Head Editing

CVPR 2025poster

Creating high-fidelity head avatars from multi-view videos is essential for many AR/VR applications. However, current methods often struggle to achieve high-quality renderings across all head components (e.g., skin vs. hair) due to the limitations of using one single representation for elements with…

2025

Minimal Impact ControlNet: Advancing Multi-ControlNet Integration

ICLR 2025poster

With the advancement of diffusion models, there is a growing demand for high-quality, controllable image generation, particularly through methods that utilize one or multiple control signals based on ControlNet. However, in current ControlNet training, each control is designed to influence all areas…

Cited by 0SourcePDFScholar
2025

Neural Encoding and Decoding at Scale

ICML 2025spotlight

Recent work has demonstrated that large-scale, multi-animal models are powerful tools for characterizing the relationship between neural activity and behavior. Current large-scale approaches, however, focus exclusively on either predicting neural activity from behavior (encoding) or predicting behav…

Cited by 1SourcePDFScholar
2025

Training-free Dense-Aligned Diffusion Guidance for Modular Conditional Image Synthesis

CVPR 2025poster

Conditional image synthesis is a crucial task with broad applications, such as artistic creation and virtual reality. However, current generative methods are often task-oriented with a narrow scope, handling a restricted condition with constrained applicability. In this paper, we propose a novel app…

2025

Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought

ICLR 2025spotlight

Chain of Thought (CoT) prompting has been shown to significantly improve the performance of large language models (LLMs), particularly in arithmetic and reasoning tasks, by instructing the model to produce intermediate reasoning steps. Despite the remarkable empirical success of CoT and its theoreti…

Cited by 0SourcePDFScholar
2025

What Makes a Reward Model a Good Teacher? An Optimization Perspective

NeurIPS 2025spotlight

The success of Reinforcement Learning from Human Feedback (RLHF) critically depends on the quality of the reward model. However, while this quality is primarily evaluated through accuracy, it remains unclear whether accuracy fully captures what makes a reward model an effective teacher. We address t…

Cited by 0SourcecodeScholar
2024

DanceCamera3D: 3D Camera Movement Synthesis with Music and Dance

CVPR 2024poster

Choreographers determine what the dances look like while cameramen determine the final presentation of dances. Recently various methods and datasets have showcased the feasibility of dance synthesis. However camera movement synthesis with music and dance remains an unsolved challenging problem due t…

2024

Generate Like Experts: Multi-Stage Font Generation by Incorporating Font Transfer Process into Diffusion Models

CVPR 2024poster

Few-shot font generation (FFG) produces stylized font images with a limited number of reference samples which can significantly reduce labor costs in manual font designs. Most existing FFG methods follow the style-content disentanglement paradigm and employ the Generative Adversarial Network (GAN) t…

2024

Inner Classifier-Free Guidance and Its Taylor Expansion for Diffusion Models

ICLR 2024poster

Classifier-free guidance (CFG) is a pivotal technique for balancing the diversity and fidelity of samples in conditional diffusion models. This approach involves utilizing a single model to jointly optimize the conditional score predictor and unconditional score predictor, eliminating the need for a…

Cited by 2SourcePDFScholar
2024

Learning and Transferring Sparse Contextual Bigrams with Linear Transformers

NeurIPS 2024poster

Transformers have achieved significant success in natural language modeling because of their exceptional capabilities to combine contextual information and global knowledge, yet their theoretical basis remains unclear. In this paper, we first propose Sparse Contextual Bigram (SCB), a natural extensi…

Cited by 2SourcePDFScholar
2024

Towards a "Universal Translator" for Neural Dynamics at Single-Cell, Single-Spike Resolution

NeurIPS 2024poster

Neuroscience research has made immense progress over the last decade, but our understanding of the brain remains fragmented and piecemeal: the dream of probing an arbitrary brain region and automatically reading out the information encoded in its neural activity remains out of reach. In this work, w…

Cited by 6SourcePDFScholar
2024

Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot

ICML 2024poster

The transformer architecture has prevailed in various deep learning settings due to its exceptional capabilities to select and compose structural information. Motivated by these capabilities, Sanford et al. (2023) proposed the *sparse token selection* task, in which transformers excel while fully-co…

Cited by 14SourcePDFScholar
2023

Understanding Edge-of-Stability Training Dynamics with a Minimalist Example

ICLR 2023poster

Recently, researchers observed that gradient descent for deep neural networks operates in an ``edge-of-stability'' (EoS) regime: the sharpness (maximum eigenvalue of the Hessian) is often larger than stability threshold $2/\eta$ (where $\eta$ is the step size). Despite this, the loss oscillates and…

Cited by 45SourcePDFScholar
2022

Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of Stability

NeurIPS 2022accept

Recent findings demonstrate that modern neural networks trained by full-batch gradient descent typically enter a regime called Edge of Stability (EOS). In this regime, the sharpness, i.e., the maximum Hessian eigenvalue, first increases to the value 2/(step size) (the progressive sharpening phase) a…

Cited by 26SourcePDFScholar
2022

Residual-Guided Personalized Speech Synthesis based on Face Image

ICASSP 2022accepted

Previous works derive personalized speech features by training the model on a large dataset composed of his/her audio sounds. It was reported that face information has a strong link with the speech sound. Thus in this work, we innovatively extract personalized speech features from human faces to syn…

Cited by 0SourceScholar