← Search

Lei Ke

31 accepted papers

2026

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling

ICML 2026poster

Preference optimization for diffusion models relies on reward functions that are both discriminative and computationally efficient. Vision-Language Models (VLMs) have emerged as powerful reward providers. However, their computation and memory cost can be substantial, and optimizing a latent diffusio…

Cited by 0SourceScholar
2026

FlowSteer: Guiding Few-Step Image Synthesis with Authentic Trajectories

CVPR 2026

With the success of flow matching in visual generation, sampling efficiency remains a critical bottleneck for its practical application. Among flow models' accelerating methods, ReFlow has been somehow overlooked although it has theoretical consistency with flow matching. This is primarily due to it

Cited by 0SourceScholar
2026

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

ICML 2026poster

Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundamental dimensionality mismatch that precludes reliable source localization and spatial reasoning in complex 3D environment…

Cited by 0SourceScholar
2026

RobotArena $\infty$: Unlimited Robot Benchmarking via Real-to-Sim Translation

ICLR 2026poster

The pursuit of robot generalists—instructable agents capable of performing diverse tasks across diverse environments—demands rigorous and scalable evaluation. Yet real-world testing of robot policies remains fundamentally constrained: it is labor-intensive, slow, unsafe at scale, and difficult to re…

Cited by 0SourcecodeScholar
2026

Stable and Efficient Single-Rollout RL for Multimodal Reasoning

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) has become a key paradigm to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, prevalent group-based algorithms such as GRPO require multi-rollout sampling for each prompt. While more efficient single-rollou

Cited by 0SourceScholar
2025

M^3PC: Test-time Model Predictive Control using Pretrained Masked Trajectory Model

ICLR 2025poster

Recent work in Offline Reinforcement Learning (RL) has shown that a unified transformer trained under a masked auto-encoding objective can effectively capture the relationships between different modalities (e.g., states, actions, rewards) within given trajectory datasets. However, this information…

2025

Multi-View 3D Point Tracking

ICCV 2025poster

We introduce the first data-driven multi-view 3D point tracker, designed to track arbitrary points in dynamic scenes using multiple camera views. Unlike existing monocular trackers, which struggle with depth ambiguities and occlusion, or prior multi-camera methods that require over 20 cameras and te…

2025

ProReflow: Progressive Reflow with Decomposed Velocity

CVPR 2025poster

Diffusion models have achieved significant progress in both image and video generation while still suffering from huge computation costs. As an effective solution, rectified flow aims to rectify the diffusion process of diffusion models into a straight line for few-step and even one-step generation.…

Cited by 1SourcePDFScholar
2025

Robust Multi-Object 4D Generation for In-the-wild Videos

CVPR 2025poster

We address the challenge of generating dynamic 4D scenes from monocular multi-object videos with heavy occlusions and introduce Robust4DGen, a novel approach that integrates rendering-based deformable 3D Gaussian optimization with generative priors for view synthesis. While existing view-synthesis m…

Cited by 0SourcePDFScholar
2025

TAPIP3D: Tracking Any Point in Persistent 3D Geometry

NeurIPS 2025poster

We introduce TAPIP3D, a novel approach for long-term 3D point tracking in monocular RGB and RGB-D videos. TAPIP3D represents videos as camera-stabilized spatio-temporal feature clouds, leveraging depth and camera motion information to lift 2D video features into a 3D world space where camera moveme…

Cited by 0SourcecodeScholar
2024

"SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking"

ECCV 2024poster

"Open-vocabulary Multiple Object Tracking (MOT) aims to generalize trackers to novel categories not in the training set. Currently, the best-performing methods are mainly based on pure appearance matching. Due to the complexity of motion patterns in the large-vocabulary scenarios and unstable classi…

2024

DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos

NeurIPS 2024poster

View-predictive generative models provide strong priors for lifting object-centric images and videos into 3D and 4D through rendering and score distillation objectives. A question then remains: what about lifting complete multi-object dynamic scenes? There are two challenges in this direction: First…

2024

Matching Anything by Segmenting Anything

CVPR 2024highlight

The robust association of the same objects across video frames in complex scenes is crucial for many applications especially object tracking. Current methods predominantly rely on labeled domain-specific video datasets which limits cross-domain generalization of learned similarity embeddings. We pro…

2023

BiMatting: Efficient Video Matting via Binarization

NeurIPS 2023poster

Real-time video matting on edge devices faces significant computational resource constraints, limiting the widespread use of video matting in applications such as online conferences and short-form video production. Binarization is a powerful compression approach that greatly reduces computation and…

2023

Cascade-DETR: Delving into High-Quality Universal Object Detection

ICCV 2023poster

Object localization in general environments is a fundamental part of vision systems. While dominating on the COCO benchmark, recent Transformer-based detection methods are not competitive in diverse domains. Moreover, these methods still struggle to very accurately estimate the object bounding boxes…

Cited by 39PDFcodeScholar
2023

Mask-Free Video Instance Segmentation

CVPR 2023poster

The recent advancement in Video Instance Segmentation (VIS) has largely been driven by the use of deeper and increasingly data-hungry transformer-based models. However, video masks are tedious and expensive to annotate, limiting the scale and diversity of existing VIS datasets. In this work, we aim…

2023

OVTrack: Open-Vocabulary Multiple Object Tracking

CVPR 2023poster

The ability to recognize, localize and track dynamic objects in a scene is fundamental to many real-world applications, such as self-driving and robotic systems. Yet, traditional multiple object tracking (MOT) benchmarks rely only on a few object categories that hardly represent the multitude of pos…

Cited by 66SourcePDFScholar
2023

Segment Anything in High Quality

NeurIPS 2023poster

The recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being trained with 1.1 billion masks, SAM's mask prediction quality falls short in many cases, particularly when dealing with…

2022

Mask Transfiner for High-Quality Instance Segmentation

CVPR 2022poster

Two-stage and query-based instance segmentation methods have achieved remarkable results. However, their segmented masks are still very coarse. In this paper, we present Mask Transfiner for high-quality and efficient instance segmentation. Instead of operating on regular dense tensors, our Mask Tran…

Cited by 154PDFcodeScholar
2022

Video Mask Transfiner for High-Quality Video Instance Segmentation

ECCV 2022poster

"While Video Instance Segmentation (VIS) has seen rapid progress, current approaches struggle to predict high-quality masks with accurate boundary details. Moreover, the predicted segmentations often fluctuate over time, suggesting that temporal consistency cues are neglected or not fully utilized.…

Cited by 38SourcePDFScholar
2021

Prototypical Cross-Attention Networks for Multiple Object Tracking and Segmentation

NeurIPS 2021spotlight

Multiple object tracking and segmentation requires detecting, tracking, and segmenting objects belonging to a set of given classes. Most approaches only exploit the temporal dimension to address the association problem, while relying on single frame predictions for the segmentation mask itself. We p…

2020

Cascaded Deep Monocular 3D Human Pose Estimation With Evolutionary Training Data

CVPR 2020oral

End-to-end deep representation learning has achieved remarkable accuracy for monocular 3D human pose estimation, yet these models may fail for unseen poses with limited and fixed training data. This paper proposes a novel data augmentation method that: (1) is scalable for synthesizing massive amount…

Cited by 226PDFcodeScholar
2020

Commonality-Parsing Network across Shape and Appearance for Partially Supervised Instance Segmentation

ECCV 2020poster

Partially supervised instance segmentation aims to perform learning on limited mask-annotated categories of data thus eliminating expensive and exhaustive mask annotation. The learned models are expected to be generalizable to novel categories. Existing methods either learn a transfer function from…

2020

GSNet: Joint Vehicle Pose and Shape Reconstruction with Geometrical and Scene-aware Supervision

ECCV 2020poster

We present a novel end-to-end framework named as GSNet ( extbf{\underline{G}}eometric and extbf{\underline{S}}cene-aware \underline{ extbf{Net}}work), which jointly estimates 6DoF poses and reconstructs detailed 3D car shapes from single urban street view. GSNet utilizes a unique four-way feature ex…

2019

Memory-Attended Recurrent Network for Video Captioning

CVPR 2019poster

Typical techniques for video captioning follow the encoder-decoder framework, which can only focus on one source video being processed. A potential disadvantage of such design is that it cannot capture the multiple visual context information of a word appearing in more than one relevant videos in tr…

Cited by 294PDFScholar