← Search

Yifan Shen

13 accepted papers

2026

Controllable Video Generation with Provable Disentanglement

ICLR 2026poster

Controllable video generation remains a significant challenge, despite recent advances in generating high-quality and consistent videos. Most existing methods for controlling video generation treat the video as a whole, neglecting intricate fine-grained spatiotemporal relationships, which limits bot…

Cited by 0SourceScholar
2026

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

CVPR 2026

Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify task-relevant interaction cues or track progress within a su

Cited by 0SourceScholar
2025

CausalVerse: Benchmarking Causal Representation Learning with Configurable High-Fidelity Simulations

NeurIPS 2025spotlight

Causal Representation Learning (CRL) aims to uncover the data-generating process and identify the underlying causal variables and relations, whose evaluation remains inherently challenging due to the requirement of known ground-truth causal variables and causal structure. Existing evaluations often…

Cited by 0SourcecodeScholar
2025

Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs

NeurIPS 2025poster

Current Vision-Language Models (VLMs) struggle with fine-grained spatial reasoning, particularly when multi-step logic and precise spatial alignment are required. In this work, we introduce SpatialReasoner-R1, a vision-language reasoning model designed to address these limitations. To construct high…

Cited by 0SourceScholar
2025

Flow: Modularized Agentic Workflow Automation

ICLR 2025poster

Multi-agent frameworks powered by large language models (LLMs) have demonstrated great success in automated planning and task execution. However, the effective adjustment of agentic workflows during execution has not been well studied. An effective workflow adjustment is crucial in real-world scenar…

2025

On the Identification of Temporal Causal Representation with Instantaneous Dependence

ICLR 2025oral

Temporally causal representation learning aims to identify the latent causal process from time series observations, but most methods require the assumption that the latent causal processes do not have instantaneous relations. Although some recent methods achieve identifiability in the instantaneous…

Cited by 6SourcePDFScholar
2025

Reflection-Window Decoding: Text Generation with Selective Refinement

ICML 2025poster

The autoregressive decoding for text generation in large language models (LLMs), while widely used, is inherently suboptimal due to the lack of a built-in mechanism to perform refinement and/or correction of the generated content. In this paper, we consider optimality in terms of the joint probabili…

Cited by 2SourcePDFScholar
2025

Rescorla-Wagner Steering of LLMs for Undesired Behaviors over Disproportionate Inappropriate Context

EMNLP 2025

Incorporating external context can significantly enhance the response quality of Large Language Models (LLMs). However, real-world contexts often mix relevant information with disproportionate inappropriate content, posing reliability risks. How do LLMs process and prioritize mixed context? To study

2025

Towards Adversarially Robust Dataset Distillation by Curvature Regularization

AAAI 2025technical

Dataset distillation (DD) allows datasets to be distilled to fractions of their original size while preserving the rich distributional information so that models trained on the distilled datasets can achieve a comparable accuracy while saving significant computational loads. Recent research in this…

2025

Towards Identifiability of Hierarchical Temporal Causal Representation Learning

NeurIPS 2025poster

Modeling hierarchical latent dynamics behind time series data is critical for capturing temporal dependencies across multiple levels of abstraction in real-world tasks. However, existing temporal causal representation learning methods fail to capture such dynamics, as they fail to recover the joint…

Cited by 0SourceScholar
2024

CaRiNG: Learning Temporal Causal Representation under Non-Invertible Generation Process

ICML 2024poster

Identifying the underlying time-delayed latent causal processes in sequential data is vital for grasping temporal dynamics and making downstream reasoning. While some recent methods can robustly identify these latent causal variables, they rely on strict assumptions about the invertible generation p…

2022

Explainable Survival Analysis with Convolution-Involved Vision Transformer

AAAI 2022technical

Image-based survival prediction models can facilitate doctors in diagnosing and treating cancer patients. With the advance of digital pathology technologies, the big whole slide images (WSIs) provide increasing resolution and more details for diagnosis. However, the gigabyte-size WSIs would make mos…