← Search

Siyi Chen

6 accepted papers

2026

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

CVPR 2026

Vision Language Models (VLMs) demonstrate strong qualitative visual understanding, but struggle with metrically precise spatial reasoning required for embodied applications. The agentic paradigm promises that VLMs can use a wide variety of tools that could augment these capabilities, such as depth e

Cited by 0SourcecodeScholar
2025

FlowDAS: A Stochastic Interpolant-based Framework for Data Assimilation

NeurIPS 2025poster

Data assimilation (DA) integrates observations with a dynamical model to estimate states of PDE-governed systems. Model-driven methods (e.g., Kalman Filter, Particle Filter) presuppose full knowledge of the true dynamics, which is not always satisfied in practice, while purely data-driven solvers le…

Cited by 0SourcecodeScholar
2025

Learning Diffusion Model from Noisy Measurement using Principled Expectation-Maximization Method

ICASSP 2025accepted

Diffusion models have demonstrated exceptional ability in modeling complex image distributions, making them versatile plug-and-play priors for solving imaging inverse problems. However, their reliance on large-scale clean datasets for training limits their applicability in scenarios where acquiring…

Cited by 0SourceScholar
2025

Understanding Representation Dynamics of Diffusion Models via Low-Dimensional Modeling

NeurIPS 2025poster

Diffusion models, though originally designed for generative tasks, have demonstrated impressive self-supervised representation learning capabilities. A particularly intriguing phenomenon in these models is the emergence of unimodal representation dynamics, where the quality of learned features peaks…

Cited by 0SourceScholar
2024

Exploring Low-Dimensional Subspace in Diffusion Models for Controllable Image Editing

NeurIPS 2024poster

Recently, diffusion models have emerged as a powerful class of generative models. Despite their success, there is still limited understanding of their semantic spaces. This makes it challenging to achieve precise and disentangled image generation without additional training, especially in an unsupe…

2022

Understanding 3D Object Articulation in Internet Videos

CVPR 2022poster

We propose to investigate detecting and characterizing the 3D planar articulation of objects from ordinary RGB videos. While seemingly easy for humans, this problem poses many challenges for computers. Our approach is based on a top-down detection system that finds planes that can be articulated. Th…

Cited by 24PDFcodeScholar