← Search

Zhipeng Zhang

43 accepted papers

2026

AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation

ICML 2026poster

We present \textbf{AAD-1}, an \textbf{A}symmetric \textbf{A}dversarial \textbf{D}istillation framework for \textbf{O}ne-step autoregressive image-to-video generation. State-of-the-art methods adopt adversarial distillation but suffer from motion collapse and training instability, resulting in static…

Cited by 1SourceScholar
2026

AdaHC: Accelerating Multi-Token Prediction with Adaptive Head Chunking with Pipeline Parallelism

ICML 2026poster

Multi-token prediction (MTP) architecture is widely adopted in LLMs. MTP blocks can be appended to the tail of model to predict additional tokens. However, when training with pipeline parallel, MTP leads to more pipeline bubbles and deteriorates the pipeline efficiency. Based on in-depth analysis of…

Cited by 0SourceScholar
2026

AutoQVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization

ICLR 2026poster

The advent of Vision-Language-Action (VLA) models represents a significant leap for embodied intelligence, yet their immense computational demands critically hinder deployment on resource-constrained robotic platforms. Intuitively, low-bit quantization is a prevalent and preferred technique for larg…

Cited by 0SourcecodeScholar
2026

Beyond Isolation: A Unified Benchmark for General-Purpose Navigation

RSS 2026poster

The pursuit of general-purpose embodied agents is currently hindered by fragmented evaluation protocols that isolate navigation skills and fixate on specific robot morphologies. This disconnect fails to reflect real-world scenarios where agents must orchestrate diverse behaviors across varying physi…

Cited by 0SourceScholar
2026

FastVGGT: Fast Visual Geometry Transformer

ICLR 2026poster

Scaling visual geometry transformers for long image sequences poses a significant computational and memory challenge. In this work, we diagnose this issue in the state-of-the-art model VGGT, and trace the primary bottleneck to its Global Attention layer. Our analysis reveals a ``token collapse'' phe…

Cited by 0SourcecodeScholar
2026

FlowAD: Ego-Scene Interactive Modeling for Autonomous Driving

ICLR 2026poster

Effective environment modeling is the foundation for autonomous driving, underpinning tasks from perception to planning. However, current paradigms often inadequately consider the feedback of ego motion to the observation, which leads to an incomplete understanding of the driving process and consequ…

Cited by 0SourcecodeScholar
2026

HDGS: Hierarchical Dynamic Gaussian Splatting for Urban Driving Scenes

AAAI 2026technical

This paper tackles the challenging task of achieving storage-efficient yet high-fidelity motion representation in large-scale dynamic 3D Gaussian Splatting. Our motivation stems from the truth that existing urban-scale methods, which rely on massive and unstructured individual Gaussians for scene mo

Cited by 0SourcePDFScholar
2026

Integrating Diverse Assignment Strategies into DETRs

AAAI 2026technical

Label assignment is a critical component in object detectors, particularly within DETR-style frameworks where the one-to-one matching strategy, despite its end-to-end elegance, suffers from slow convergence due to sparse supervision. While recent works have explored one-to-many assignments to enrich

Cited by 0SourcePDFScholar
2026

OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding

ICLR 2026poster

We introduce spatio-temporal omni-object video grounding, dubbed $\textbf{OmniSTVG}$, a new STVG task aiming to localize spatially and temporally all targets mentioned in the textual query within videos. Compared to classic STVG locating only a single target, OmniSTVG enables localization of not onl…

Cited by 0SourcecodeScholar
2026

RPE-PAD: Relative Pose Estimation for Pose-agnostic Anomaly Detection

AAAI 2026technical

Pose-agnostic Anomaly Detection (PAD) aims to detect anomalies when the poses of query images are unknown and differ from those in the training set. Therefore, accurately estimating the camera poses for the query images in the test set is critical for this task. Existing query-specific framework met

Cited by 0SourcePDFScholar
2026

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation

CVPR 2026

Efficient streaming video generation is critical for simulating interactive and dynamic worlds. Existing methods distill few-step video diffusion models with sliding window attention, using initial frames as sink tokens to maintain attention performance and reduce error accumulation. However, video

Cited by 0SourcecodeScholar
2026

SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker

CVPR 2026

Parameter-efficient fine-tuning (PEFT) in multimodal tracking reveals a concerning trend where recent performance gains are often achieved at the cost of inflated parameter budgets, which fundamentally erodes PEFT's efficiency promise. In this work, we introduce SEATrack, a Simple, Efficient, and Ad

Cited by 0SourcecodeScholar
2026

Sparse Annotation, Dense Supervision: Unleashing Self-Training Power for Occupancy Prediction With 2D Labels

RA-L 2026

Serving as a fundamental task in robotic navigation and autonomous driving, occupancy prediction is gaining increasing attention for its fine-grained perception of the 3D environment. Most existing methods rely on dense 3D annotations, which are expensive, labor-intensive, and difficult to scale in

Cited by 1SourceScholar
2026

Structured Labeling Enables Faster Vision-Language Models for End-To-End Autonomous Driving

ICRA 2026poster

Vision-Language Models (VLMs) offer a promising approach to end-to-end autonomous driving due to their human-like reasoning capabilities. However, troublesome gaps remains between current VLMs and real-world autonomous driving applications. One major limitation is that existing datasets with loosely…

2026

The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models

CVPR 2026

The integration of Vision-Language Models (VLMs) into autonomous driving promises to solve long-tail scenarios, but this paradigm faces the critical and unaddressed challenge of catastrophic forgetting. The very fine-tuning process used to adapt these models to driving-specific data simultaneously e

Cited by 0SourceScholar
2026

Towards Visual Query Localization in the 3D World

CVPR 2026

Visual query localization (VQL) aims to predict a spatial-temporal response of the most recent occurrence from a sequence given a query. Currently, most research focuses on visual query localization from 2D videos, while its counterpart in 3D space has received little attention. In this paper, we ma

Cited by 0SourcecodeScholar
2025

CorrBEV: Multi-View 3D Object Detection by Correlation Learning with Multi-modal Prototypes

CVPR 2025poster

Camera-only multi-view 3D object detection in autonomous driving has witnessed encouraging developments in recent years, largely attributed to the revolution of fundamental architectures in modeling bird's eye view (BEV). Despite the growing overall average performance, we contend that the explorati…

Cited by 0SourcePDFScholar
2025

Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning

ICCV 2025poster

Training visual reinforcement learning (RL) in practical scenarios presents a significant challenge, i.e., RL agents suffer from low sample efficiency in environments with variations. While various approaches have attempted to alleviate this issue by disentangled representation learning, these metho…

Cited by 0SourcePDFScholar
2025

DreamTrack: Dreaming the Future for Multimodal Visual Object Tracking

CVPR 2025poster

Aiming to achieve class-agnostic perception in visual object tracking, current trackers commonly formulate tracking as a one-shot detection problem with the template-matching architecture. Despite the success, severe environmental variations in long-term tracking raise challenges to generalizing the…

Cited by 0SourcePDFScholar
2025

Each Complexity Deserves a Pruning Policy

NeurIPS 2025poster

The established redundancy in visual tokens within large vision–language models (LVLMs) allows for pruning to effectively reduce their substantial computational demands. Empirical evidence from previous works indicates that visual tokens in later decoder stages receive less attention than shallow la…

Cited by 0SourcecodeScholar
2025

EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models

NeurIPS 2025poster

Vision-Language-Action (VLA) models, particularly diffusion-based architectures, demonstrate transformative potential for embodied intelligence but are severely hampered by high computational and memory demands stemming from extensive inherent and inference-time redundancies. While existing accelera…

Cited by 0SourceScholar
2025

Evolving High-Quality Rendering and Reconstruction in a Unified Framework with Contribution-Adaptive Regularization

CVPR 2025poster

Representing 3D scenes from multiview images is a core challenge in computer vision and graphics, which requires both precise rendering and accurate reconstruction. Recently, 3D Gaussian Splatting (3DGS) has garnered significant attention for its high-quality rendering and fast inference speed. Yet,…

Cited by 2SourcePDFScholar
2025

Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection

ICCV 2025poster

We present the first work demonstrating that a pure Mamba block can achieve efficient Dense Global Fusion, meanwhile guaranteeing top performance for camera-LiDAR multi-modal 3D object detection. Our motivation stems from the observation that existing fusion strategies are constrained by their inabi…

2025

LoRATv2: Enabling Low-Cost Temporal Modeling in One-Stream Trackers

NeurIPS 2025spotlight

Transformer-based algorithms, such as LoRAT, have significantly enhanced object-tracking performance. However, these approaches rely on a standard attention mechanism, which incurs quadratic token complexity, making real-time inference computationally expensive. In this paper, we introduce LoRATv2,…

Cited by 0SourcecodeScholar
2025

Mamba-3VL: Taming State Space Model for 3D Vision Language Learning

ICCV 2025poster

3D vision-language (3D-VL) reasoning, connecting natural language with 3D physical world, represents a milestone in advancing spatial intelligence. While transformer-based methods dominate 3D-VL research, their quadratic complexity and simplistic positional embedding mechanisms severely limits effec…

2025

Online Segment Any 3D Thing as Instance Tracking

NeurIPS 2025poster

Online, real-time, and fine-grained 3D segmentation constitutes a fundamental capability for embodied intelligent agents to perceive and comprehend their operational environments. Recent advancements employ predefined object queries to aggregate semantic information from Vision Foundation Models (VF…

Cited by 0SourcecodeScholar
2025

ScoreNet: Consistency-driven Framework with Multi-side Information Fusion for Session-based Recommendation

AAAI 2025technical

Fusing side information in session-based recommendation is crucial for improving the performance of next-item prediction by providing additional context. Recent methods optimize attention weights by combining item and side information embeddings. However, semantic heterogeneity between item IDs and…

2025

The Devil is in the Quality: Exploring Informative Samples for Semi-Supervised Monocular 3D Object Detection

ICRA 2025

This paper tackles the challenging problem of semi-supervised monocular 3D object detection with a general framework. In specific, having observed that the bottleneck of this task lies in lacking reliable and informative samples from unlabeled data for detector learning, we introduce a novel simple

Cited by 0SourceScholar
2024

"Tracking Meets LoRA: Faster Training, Larger Model, Stronger Performance"

ECCV 2024poster

"Motivated by the Parameter-Efficient Fine-Tuning (PEFT) in large language models, we propose LoRAT, a method that unveils the power of larger Vision Transformers (ViT) for tracking within laboratory-level resources. The essence of our work lies in adapting LoRA, a technique that fine-tunes a small…

2024

A-Teacher: Asymmetric Network for 3D Semi-Supervised Object Detection

CVPR 2024poster

This work proposes the first online asymmetric semi-supervised framework namely A-Teacher for LiDAR-based 3D object detection. Our motivation stems from the observation that 1) existing symmetric teacher-student methods for semi-supervised 3D object detection have characterized simplicity but impede…

Cited by 2SourcePDFScholar
2024

Image Fusion via Vision-Language Model

ICML 2024poster

Image fusion integrates essential information from multiple images into a single composite, enhancing structures, textures, and refining imperfections. Existing methods predominantly focus on pixel-level and semantic visual features for recognition, but often overlook the deeper text-level semantic…

2024

Multi-View Point Cloud Registration Based on Improved NDT Algorithm and ODM Optimization Method

RA-L 2024

The acquisition of targets' complete point cloud model is crucial for tasks such as 3D reconstruction and disordered grasping. Shooting targets from multiple perspectives and registering point clouds from different perspectives can obtain a relatively complete point cloud model. However, small scene

Cited by 8SourceScholar
2024

VastTrack: Vast Category Visual Object Tracking

NeurIPS 2024poster

In this paper, we propose a novel benchmark, named VastTrack, aiming to facilitate the development of general visual tracking via encompassing abundant classes and videos. VastTrack consists of a few attractive properties: (1) Vast Object Category. In particular, it covers targets from 2,115 categor…

2023

AUNet: Learning Relations Between Action Units for Face Forgery Detection

CVPR 2023poster

Face forgery detection becomes increasingly crucial due to the serious security issues caused by face manipulation techniques. Recent studies in deepfake detection have yielded promising results when the training and testing face forgeries are from the same domain. However, the problem remains chall…

Cited by 56SourcePDFScholar
2022

Learning Target-aware Representation for Visual Tracking via Informative Interactions

IJCAI 2022poster

We introduce a novel backbone architecture to improve target-perception ability of feature representation for tracking. Having observed de facto frameworks perform feature matching simply using the backbone outputs for target localization, there is no direct feedback from the matching module to the…

Cited by 61SourcePDFScholar
2022

One More Check: Making “Fake Background” Be Tracked Again

AAAI 2022technical

The one-shot multi-object tracking, which integrates object detection and ID embedding extraction into a unified network, has achieved groundbreaking results in recent years. However, current one-shot trackers solely rely on single-frame detections to predict candidate bounding boxes, which may be u…

2022

SwinTrack: A Simple and Strong Baseline for Transformer Tracking

NeurIPS 2022accept

Recently Transformer has been largely explored in tracking and shown state-of-the-art (SOTA) performance. However, existing efforts mainly focus on fusing and enhancing features generated by convolutional neural networks (CNNs). The potential of Transformer in representation learning remains under-e…

2021

Learn To Match: Automatic Matching Network Design for Visual Tracking

ICCV 2021poster

Siamese tracking has achieved groundbreaking performance in recent years, where the essence is the efficient matching operator cross-correlation and its variants. Besides the remarkable success, it is important to note that the heuristic matching network design relies heavily on expert experience. M…

Cited by 239PDFcodeScholar
2021

Reality Transform Adversarial Generators for Image Splicing Forgery Detection and Localization

ICCV 2021poster

When many forged images become more and more realistic with the help of image editing tools and deep learning techniques, authenticators need to improve their ability to verify these forged images. The process of generating and detecting forged images is thus similar to the principle of Generative A…

Cited by 33PDFScholar
2021

Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and Benchmark

CVPR 2021poster

Tracking by natural language specification is a new rising research topic that aims at locating the target object in the video sequence based on its language description. Compared with traditional bounding box (BBox) based tracking, this setting guides object tracking with high-level semantic inform…

Cited by 220PDFScholar