← Search

Jingxi Chen

10 accepted papers

2026

First Frame Is the Place to Go for Video Content Customization

CVPR 2026

What role does the first frame play in video generation models? Traditionally, it's viewed as the spatial-temporal starting point of a video, merely a seed for subsequent animation. In this work, we reveal a fundamentally different perspective: video models implicitly treat the first frame as a conc

Cited by 0SourcecodeScholar
2026

From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition

CVPR 2026

Images can be viewed as layered compositions, foreground objects over background, with potential occlusions. This layered representation enables independent editing of elements, offering greater flexibility for content creation. Despite the progress in large generative models, decomposing a single i

Cited by 0SourceScholar
2026

Self-Rewarding Vision-Language Model via Reasoning Decomposition and Multi-Reward Policy Optimization

ICLR 2026poster

Vision-Language Models (VLMs) often suffer from visual hallucinations – generating things that are not consistent with visual inputs – and language shortcuts, where they skip the visual part and just rely on text priors. These issues arise because most post-training methods for VLMs rely on simple v…

Cited by 0SourceScholar
2025

Learning Normal Flow Directly From Events

ICCV 2025poster

Event-based motion field estimation is an important task. However, current optical flow methods face challenges: learning-based approaches, often frame-based and relying on CNNs, lack cross-domain transferability, while model-based methods, though more robust, are less accurate. To address the limit…

2025

Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation

CVPR 2025poster

Video Frame Interpolation aims to recover realistic missing frames between observed frames, generating a high-frame-rate video from a low-frame-rate video. However, without additional guidance, large motion between frames makes this problem ill-posed. Event-based Video Frame Interpolation (EVFI) add…

Cited by 4SourcePDFScholar
2024

Active Human Pose Estimation via an Autonomous UAV Agent

IROS 2024poster

One of the core activities of an active observer involves moving to secure a "better" view of the scene, where the definition of "better" is task-dependent. This paper focuses on the task of human pose estimation from videos capturing a person’s activity. Self-occlusions within the scene can complic…

Cited by 2SourceScholar
2024

CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event Cameras

CVPR 2024poster

Point-spread-function (PSF) engineering is a well-established computational imaging technique that uses phase masks and other optical elements to embed extra information (e.g. depth) into the images captured by conventional CMOS image sensors. To date however PSF-engineering has not been applied to…

Cited by 2SourcePDFScholar
2024

Temporally Consistent Atmospheric Turbulence Mitigation with Neural Representations

NeurIPS 2024poster

Atmospheric turbulence, caused by random fluctuations in the atmosphere's refractive index, introduces complex spatio-temporal distortions in imagery captured at long range. Video Atmospheric Turbulence Mitigation (ATM) aims to restore videos affected by these distortions. However, existing video AT…

2023

ProxMaP: Proximal Occupancy Map Prediction for Efficient Indoor Robot Navigation

IROS 2023poster

Planning a path for a mobile robot typically requires building a map (e.g., an occupancy grid) of the environment as the robot moves around. While navigating in an unknown environment, the map built by the robot online may have many as-yet-unknown regions. A conservative planner may avoid such regio…

Cited by 9SourceScholar
2021

Multi-Agent Reinforcement Learning for Visibility-based Persistent Monitoring

IROS 2021poster

The Visibility-based Persistent Monitoring (VPM) problem seeks to find a set of trajectories (or controllers) for robots to persistently monitor a changing environment. Each robot has a sensor, such as a camera, with a limited field-of-view that is obstructed by obstacles in the environment. The rob…

Cited by 22SourcecodeScholar