← Search

Peixi Peng

34 accepted papers

2026

A Training-Free Framework for Long Video Understanding via Video-Query-Options Similarity

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved remarkable success in image and short video understanding tasks, but their performance on hour-long videos remains limited due to constraint of input token capacity. Existing approaches often require costly training procedures, hindering their a…

Cited by 0SourceScholar
2026

BulletTime4D: Towards High Spatio-Temporal Resolution Dynamic Scene Rendering via Spike-Guided Stereo Vision

AAAI 2026technical

High spatio‑temporal resolution novel‑view scene rendering is crucial for applications such as sports analysis and scientific experiments. However, existing Dynamic Scene Rendering (DSR) approaches typically rely on conventional RGB cameras with limited frame rates, making it difficult to achieve hi

Cited by 0SourcePDFScholar
2026

COVR: Collaborative Optimization of VLMs and RL Agent for Visual-Based Control

AAAI 2026technical

Visual reinforcement learning (RL) suffers from poor sample efficiency due to high-dimensional observations in complex tasks. While existing works have shown that vision-language models (VLMs) can assist RL, they often focus on knowledge distillation from the VLM to RL, overlooking the potential of

Cited by 0SourcePDFScholar
2026

Fine-flow Distilling Coarse-flow Video Generation for Long-Term Driving World Model

AAAI 2026technical

Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error accumulations when predicting the long-term future, which limits practical applications. Recent studies utilize the Diffu

Cited by 0SourcePDFScholar
2026

Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement Learning

CVPR 2026

Multimodal Large Language Models (MLLMs) have made great progress in video understanding tasks. However, when it comes to understanding complex or lengthy videos, MLLMs tend to overlook details or produce hallucinations. To alleviate these issues, recent work has attempted to leverage reinforcement

Cited by 0SourcecodeScholar
2026

LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding

AAAI 2026technical

The development of multimodal large language models (MLLMs) has advanced general video understanding. However, existing video evaluation benchmarks primarily focus on non-interactive videos, such as movies and recordings. To fill this gap, this paper proposes the first omnimodal benchmark for intera

Cited by 0SourcePDFScholar
2026

MER-Tracker: Towards High-Speed 3D Point Tracking via Multi-View Event-RGB Hybrid Cameras

CVPR 2026

This paper proposes the first task for high-speed 3D point tracking using multi-view Event-RGB hybrid cameras. We design a cuboid observation device comprising 4 RGB cameras (30fps) and 2 Event cameras to synchronously capture high-speed motions, and propose MER-Tracker, a high-frame-rate 3D point-t

Cited by 0SourceScholar
2026

Multi-timescale Reinforcement Learning by Value Reconstruction

ICML 2026poster

Most reinforcement learning (RL) baselines maximize future cumulative rewards with a fixed single discount factor, which limits their performance in complex sequential decision-making tasks due to a failure to balance short-term objectives and long-term planning. To address this issue, this paper fo…

Cited by 0SourceScholar
2026

Perceiving the Knowledge Boundary: Uncertainty-Guided Exploration and Imagination for World Models

AAAI 2026technical

World-model-based reinforcement learning achieves high sample efficiency by learning from imagined rollouts. However, its success critically depends on the accuracy of the learned world model, which is prone to producing unrealistic or hallucinated rollouts when queried beyond its domain of competen

Cited by 0SourcePDFScholar
2026

Resolving the Stability-Plasticity Dilemma in Reinforcement Learning via Complementary Continual Critics

CVPR 2026

This paper proposes the Continual Dual-Critic with Cross-Attention (CD-CCA) framework for visual reinforcement learning to address the plasticity-stability conflict. Our method introduces continual learning techniques into the visual RL architecture, constructing two complementary critics using Cont

Cited by 0SourcecodeScholar
2026

Structured Expert Routing with Multi-View Task Priors for Offline Meta-Reinforcement Learning

ICML 2026poster

Offline meta-reinforcement learning requires agents to generalize to unseen tasks from fixed datasets, yet existing sequence-based and MoE-based methods rely on implicit or token-level routing signals that fail to capture task-level structure. We propose the **Task-Guided Router (TGR)**, a structure…

Cited by 0SourceScholar
2025

Exploiting Continuous Motion Clues for Vision-Based Occupancy Prediction

AAAI 2025technical

Occupancy networks aim to reconstruct the surroundings with occupied semantic voxels. However, frequent object occlusions often occur in dynamic real-world scenarios, which cannot be captured by independent frames. Most existing occupancy networks generate results without explicitly considering past…

2025

Spike4DGS: Towards High-Speed Dynamic Scene Rendering with 4D Gaussian Splatting via a Spike Camera Array

NeurIPS 2025poster

Spike camera with high temporal resolution offers a new perspective on high-speed dynamic scene rendering. Most existing rendering methods rely on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) for static scenes using a monocular spike camera. However, these methods struggle with dyna…

Cited by 0SourcecodeScholar
2025

Towards Building Human-like Smart Agents in Modern 3D Video Games (Student Abstract)

AAAI 2025technical

In recent years, reinforcement learning has been widely applied in the field of games. However, most studies focus on assisting agents to achieve victory, with less attention paid to whether the agents exhibit human-like characteristics. In order to build human-like agents with high performance, we…

Cited by 0SourcePDFScholar
2025

VLMs-Guided Representation Distillation for Efficient Vision-Based Reinforcement Learning

CVPR 2025poster

Vision-based Reinforcement Learning (VRL) attempts to establish associations between visual inputs and optimal actions through interactions with the environment. Given the high-dimensional and complex nature of visual data, it becomes essential to learn policy upon high-quality state representation.…

Cited by 0SourcePDFScholar
2025

When Every Millisecond Counts: Real-Time Anomaly Detection via the Multimodal Asynchronous Hybrid Network

ICML 2025spotlight

Anomaly detection is essential for the safety and reliability of autonomous driving systems. Current methods often focus on detection accuracy but neglect response time, which is critical in time-sensitive driving scenarios. In this paper, we introduce real-time anomaly detection for autonomous driv…

Cited by 0SourcePDFScholar
2024

Adaptive Discovering and Merging for Incremental Novel Class Discovery

AAAI 2024technical

One important desideratum of lifelong learning aims to discover novel classes from unlabelled data in a continuous manner. The central challenge is twofold: discovering and learning novel classes while mitigating the issue of catastrophic forgetting of established knowledge. To this end, we introduc…

Cited by 12SourcePDFScholar
2024

DMR: Decomposed Multi-Modality Representations for Frames and Events Fusion in Visual Reinforcement Learning

CVPR 2024poster

We explore visual reinforcement learning (RL) using two complementary visual modalities: frame-based RGB camera and event-based Dynamic Vision Sensor (DVS). Existing multi-modality visual RL methods often encounter challenges in effectively extracting task-relevant information from multiple modaliti…

2024

Density-Adaptive Model Based on Motif Matrix for Multi-Agent Trajectory Prediction

CVPR 2024poster

Multi-agent trajectory prediction is essential in autonomous driving risk avoidance and traffic flow control. However the heterogeneous traffic density on interactions which caused by physical laws social norms and so on is often overlooked in existing methods. When the density varies the number of…

Cited by 1SourcePDFScholar
2024

Seek Commonality but Preserve Differences: Dissected Dynamics Modeling for Multi-modal Visual RL

NeurIPS 2024poster

Accurate environment dynamics modeling is crucial for obtaining effective state representations in visual reinforcement learning (RL) applications. However, when facing multiple input modalities, existing dynamics modeling methods (e.g., DeepMDP) usually stumble in addressing the complex and volatil…

Cited by 0SourcePDFScholar
2023

Dynamic Belief for Decentralized Multi-Agent Cooperative Learning

IJCAI 2023poster

Decentralized multi-agent cooperative learning is a practical task due to the partially observed setting both in training and execution. Every agent learns to cooperate without access to the observations and policies of others. However, the decentralized training of multi-agent is of great difficult…

Cited by 2SourcePDFScholar
2023

Hierarchical Adaptive Value Estimation for Multi-modal Visual Reinforcement Learning

NeurIPS 2023poster

Integrating RGB frames with alternative modality inputs is gaining increasing traction in many vision-based reinforcement learning (RL) applications. Existing multi-modal vision-based RL methods usually follow a Global Value Estimation (GVE) pipeline, which uses a fused modality feature to obtain a…

2023

Learning With Fantasy: Semantic-Aware Virtual Contrastive Constraint for Few-Shot Class-Incremental Learning

CVPR 2023poster

Few-shot class-incremental learning (FSCIL) aims at learning to classify new classes continually from limited samples without forgetting the old classes. The mainstream framework tackling FSCIL is first to adopt the cross-entropy (CE) loss for training at the base session, then freeze the feature ex…

2023

Simoun: Synergizing Interactive Motion-appearance Understanding for Vision-based Reinforcement Learning

ICCV 2023accepted

Efficient motion and appearance modeling are critical for vision-based Reinforcement Learning (RL). However, existing methods struggle to reconcile motion and appearance information within the state representations learned from a single observation encoder. To address the problem, we present Synergi…

Cited by 1SourcePDFScholar
2023

Stabilizing Visual Reinforcement Learning via Asymmetric Interactive Cooperation

ICCV 2023poster

Vision-based reinforcement learning (RL) depends on discriminative representation encoders to abstract the observation states. Despite the great success of increasing CNN parameters for many supervised computer vision tasks, reinforcement learning with temporal-difference (TD) losses cannot benefit…

Cited by 4PDFScholar
2022

Spectrum Random Masking for Generalization in Image-based Reinforcement Learning

NeurIPS 2022accept

Generalization in image-based reinforcement learning (RL) aims to learn a robust policy that could be applied directly on unseen visual environments, which is a challenging task since agents usually tend to overfit to their training environment. To handle this problem, a natural approach is to incre…

Cited by 19SourcePDFScholar
2021

Allocating DNN Layers Computation Between Front-End Devices and The Cloud Server for Video Big Data Processing

ICASSP 2021accepted

With the development of intelligent hardware, front-end devices can also perform DNN computation. Moreover, the deep neural network can be divided into several layers. In this way, part of the computation of DNN models can be migrated to the front-end devices, which can alleviate the cloud burden an…

Cited by 0SourceScholar
2021

Amplitude-Phase Recombination: Rethinking Robustness of Convolutional Neural Networks in Frequency Domain

ICCV 2021poster

Recently, the generalization behavior of Convolutional Neural Networks (CNN) is gradually transparent through explanation techniques with the frequency components decomposition. However, the importance of the phase spectrum of the image for a robust vision system is still ignored. In this paper, we…

Cited by 127PDFcodeScholar
2020

Learning Open Set Network with Discriminative Reciprocal Points

ECCV 2020poster

Open set recognition is an emerging research area that aims to simultaneously classify samples from predefined classes and identify the rest as 'unknown'. In this process, one of the key challenges is to reduce the risk of generalizing the inherent characteristics of numerous unknown samples learned…

Cited by 266SourcePDFScholar
2018

Visual Tracking via Spatially Aligned Correlation Filters Network

ECCV 2018poster

Correlation filters based trackers rely on a periodic assumption of the search sample to efficiently distinguish the target from the background. This assumption however yields undesired boundary effects and restricts aspect ratios of search samples. To handle these issues, an end-to-end deep archite…

2016

Unsupervised Cross-Dataset Transfer Learning for Person Re-Identification

CVPR 2016poster

Most existing person re-identification (Re-ID) approaches follow a supervised learning framework, in which a large number of labelled matching pairs are required for training. This severely limits their scalability in real-world applications. To overcome this limitation, we develop a novel cross-dat…

Cited by 457PDFScholar