← Search

Hengshuang Zhao

118 accepted papers

2026

AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in Unified Multimodal Models via Decompositional Verifiable Reward

ICML 2026poster

In this paper, we propose **AlphaGRPO**, a novel framework that applies Group Relative Policy Optimization (GRPO) to AR-Diffusion Unified Multimodal Models (UMMs) to enhance multimodal generation capabilities without relying on external knowledge injection. Our approach unlocks the model's intrinsic…

Cited by 0SourceScholar
2026

Anime-Ready: Controllable 3D Anime Character Generation with Body-Aligned Component-Wise Garment Modeling

ICLR 2026poster

3D anime character generation has become increasingly important in digital entertainment, including animation production, virtual reality, gaming, and virtual influencers. Unlike realistic human modeling, anime-style characters require exaggerated proportions, stylized surface details, and artistica…

Cited by 0SourceScholar
2026

Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

ICML 2026poster

Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. How can we incorporate 3D information to enhance VLA capabilities? We conduct a pilot study across different observation spaces and visual representation…

Cited by 0SourceScholar
2026

DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning

CVPR 2026

Although multi-modal large language models (MLLMs) have shown strong capabilities across diverse domains, their application in generating fine-grained 3D perception and prediction outputs in autonomous driving remains underexplored. In this paper, we propose DrivePI, a novel spatial-aware 4D MLLM th

Cited by 0SourcecodeScholar
2026

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

ICLR 2026poster

In recent years, 2D Vision-Language Models (VLMs) have made significant strides in image-text understanding tasks. However, their performance in 3D spatial comprehension, which is critical for embodied intelligence, remains limited. Recent advances have leveraged 3D point clouds and multi-view image…

Cited by 0SourcecodeScholar
2026

Game Ground Bench: Probing the Limits of LVLMs in Complex Semantic Grounding Across Game Universes

AAAI 2026technical

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities, yet their ability to ground language in complex, interactive environments such as video games remains a critical frontier. Existing benchmarks are inadequate for this purpose: real-world datasets like RefCOCO introduce a

Cited by 0SourcePDFScholar
2026

GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation

CVPR 2026

Physics-aware driving world model is essential for drive planning, out-of-distribution data synthesis, and closed-loop evaluation. However, existing methods often rely on a single diffusion model to directly map driving actions to videos, which makes learning difficult and leads to physically incons

Cited by 0SourceScholar
2026

Hint2Gen: Bridging Understanding and Generation via Code-structured Hints

CVPR 2026

Recent unified models have made remarkable strides in generating high-quality images, yet they consistently fail on reasoning-intensive tasks, i.e., solving mazes, assembling tangrams. Intriguingly, we find that vision-language models (VLMs) and large language models (LLMs) can accurately solve thes

Cited by 0SourceScholar
2026

In Pursuit of Pixel Supervision for Visual Pre-training

CVPR 2026

Data matters. In computer vision, data (or pixels) are the primary source of information containing signals that span from low-level attributes to high-level concepts. At scale, the success of modern vision systems has been closely tied to how data is curated for semantic understanding (e.g., ImageN

Cited by 0SourcecodeScholar
2026

Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting

ICLR 2026poster

Existing feed-forward 3D Gaussian Splatting methods typically rely on pixel-aligned primitives, which makes scaling to higher resolutions (e.g., 4K) prohibitive as the number of Gaussians grows quadratically with image resolution. We introduce LGTM (Less Gaussians, Texture More), a feed-forward and…

Cited by 0SourcecodeScholar
2026

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search

ICLR 2026poster

Recent advances in large multimodal models have leveraged image-based tools with reinforcement learning to tackle visual problems. However, existing open-source approaches often exhibit monotonous reasoning patterns and allow only a limited number of interaction turns, making them inadequate for dif…

Cited by 0SourcecodeScholar
2026

RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving Scenes

CVPR 2026

High-quality video editing and processing are crucial in domains such as filmmaking and autonomous driving, where accurate visual refinement and data preparation are essential. However, it is challenging to achieve precise control over dynamic objects while maintaining spatiotemporal consistency. Cu

Cited by 0SourcecodeScholar
2026

SpatialHand: Generative Object Manipulation from 3D Prespective

ICLR 2026poster

We introduce SpatialHand, a novel framework for generative object insertion with precise 3D control. Current generative object manipulation methods primarily operate within the 2D image plane, but often fail to grasp 3D scene complexities, leading to ambiguities in an object's 3D position, orientati…

Cited by 0SourceScholar
2026

Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents

ICML 2026poster

Large language model (LLM) agents increasingly rely on external tools such as search engines to solve complex, multi-step problems, yet their rollouts are structurally heterogeneous: variations in tool-call number, placement, and outcomes induce distinct behaviors and reward distributions. As a resu…

Cited by 5SourceScholar
2026

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation

CVPR 2026

Diffusion-based video motion customization facilitates the acquisition of human motion representations from a few video samples, while achieving arbitrary subjects transfer through precise textual conditioning. Existing approaches often rely on semantic-level alignment, expecting the model to learn

Cited by 0SourceScholar
2026

Temporal Equilibrium MeanFlow: Bridging the Scale Gap for One-Step Generation

CVPR 2026

MeanFlow is a powerful few-step generative framework that can be trained from scratch, but its performance degrades significantly when the one-step loss uses a large portion of training data. This stems from a temporal scale imbalance: gradients from different stages of generation contribute unevenl

Cited by 0SourceScholar
2026

Utonia: Toward One Encoder for All Point Clouds

ICML 2026poster

We dream of a future where point clouds from all domains can come together to shape a single model that benefits them all. Toward this goal, we present Utonia, a first step toward training a single self-supervised point transformer encoder across heterogeneous domains, spanning remote sensing, outdo…

Cited by 0SourceScholar
2026

WorldCompass: Reinforcement Learning for Long-Horizon World Models

ICML 2026poster

This work presents WorldCompass, a novel Reinforcement Learning (RL) post-training framework for the long-horizon, interactive video-based world models, enabling them to explore the world more accurately and consistently based on interaction signals. To effectively "steer" the world model's explorat…

Cited by 0SourceScholar
2025

BOOD: Boundary-based Out-Of-Distribution Data Generation

ICML 2025poster

Harnessing the power of diffusion models to synthesize auxiliary training data based on latent space features has proven effective in enhancing out-of-distribution (OOD) detection performance. However, extracting effective features outside the in-distribution (ID) boundary in latent space remains ch…

Cited by 0SourcePDFScholar
2025

Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

NeurIPS 2025poster

Humans learn abstract concepts through multisensory synergy, and once formed, such representations can often be recalled from a single modality. Inspired by this principle, we introduce Concerto, a minimalist simulation of human concept learning for spatial cognition, combining 3D intra-modal self-d…

Cited by 0SourceScholar
2025

DiffDoctor: Diagnosing Image Diffusion Models Before Treating

ICCV 2025poster

In spite of recent progress, image diffusion models still produce artifacts. A common solution is to leverage the feedback provided by quality assessment systems or human annotators to optimize the model, where images are generally rated in their entirety. In this work, we believe problem-solving st…

Cited by 0SourcePDFScholar
2025

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

ICCV 2025poster

In video Multimodal Large Language Models (video MLLMs), the visual encapsulation process plays a pivotal role in converting video contents into representative tokens for LLM input. While linear projectors are widely employed for encapsulation, they introduce semantic indistinctness and temporal inc…

2025

DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving

CVPR 2025highlight

Multimodal large language models (MLLMs) possess the ability to comprehend visual images or videos, and show impressive reasoning ability thanks to the vast amounts of pretrained knowledge, making them highly suitable for autonomous driving applications. Unlike the previous work, DriveGPT4-V1, which…

Cited by 0SourcePDFScholar
2025

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

CVPR 2025poster

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging…

Cited by 23SourcePDFScholar
2025

Empowering Large Language Models with 3D Situation Awareness

CVPR 2025poster

Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer in 3D scenes can change, resulting in different descriptions…

Cited by 0SourcePDFScholar
2025

Enhancing LLM Knowledge Learning through Generalization

EMNLP 2025

As Large language models (LLMs) are increasingly deployed in diverse applications, faithfully integrating evolving factual knowledge into these models remains a critical challenge. Continued pre-training on paraphrased data has shown empirical promise for enhancing knowledge acquisition. However, th

2025

GenSpace: Benchmarking Spatially-Aware Image Generation

NeurIPS 2025poster

Humans can intuitively compose and arrange scenes in the 3D space for photography. However, can advanced AI image generators plan scenes with similar 3D spatial awareness when creating images from text or image prompts? We present GenSpace, a novel benchmark and evaluation pipeline to comprehensivel…

Cited by 0SourceScholar
2025

HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation

ICCV 2025poster

Driving World Models (DWMs) have become essential for autonomous driving by enabling future scene prediction. However, existing DWMs are limited to scene generation and fail to incorporate scene understanding, which involves interpreting and reasoning about the driving environment. In this paper, we…

2025

HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding

ICML 2025poster

Recent advancements in large language models (LLMs) have significantly propelled the development of large multi-modal models (LMMs), highlighting the potential for general and intelligent assistants. However, most LMMs model visual and textual modalities separately, leading to recent efforts to deve…

Cited by 0SourcePDFScholar
2025

HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models

CVPR 2025poster

High-resolution image inputs allow Large Vision-Language Models (LVLMs) to capture finer visual details, improving comprehension. However, the increased training and computational costs associated with such inputs pose significant challenges. A common approach to mitigate these costs involves slicin…

Cited by 8SourcePDFScholar
2025

LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence

ICML 2025poster

Recent embodied agents are primarily built based on reinforcement learning (RL) or large language models (LLMs). Among them, RL agents are efficient for deployment but only perform very few tasks. By contrast, giant LLM agents (often more than 1000B parameters) present strong generalization while de…

Cited by 1SourcePDFScholar
2025

LiteReality: Graphic-Ready 3D Scene Reconstruction from RGB-D Scans

NeurIPS 2025poster

We propose LiteReality, a novel pipeline that converts RGB-D scans of indoor environments into compact, realistic, and interactive 3D virtual replicas. LiteReality not only reconstructs scenes that visually resemble reality but also supports key features essential for graphics pipelines, such as obj…

Cited by 0SourceScholar
2025

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

NeurIPS 2025poster

This work explores enabling Chain-of-Thought (CoT) reasoning to link visual cues across multiple images. A straightforward solution is to adapt rule-based reinforcement learning for Vision-Language Models (VLMs). However, such methods typically rely on manually curated question-answer pairs, which c…

Cited by 0SourceScholar
2025

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

ICLR 2025poster

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Meanwhile, multimodal representation models have emerged as the foundation for these versatile multimodal understanding and generation pipeline. Models like CLIP, CLAP and ImageBind…

Cited by 11SourcePDFScholar
2025

Orient Anything V2: Unifying Orientation and Rotation Understanding

NeurIPS 2025spotlight

This work presents Orient Anything V2, an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. Building upon Orient Anything V1, which defines orientation via a single unique front face, V2 extends this capability to handle objects w…

Cited by 0SourceScholar
2025

Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models

ICML 2025poster

Orientation is a fundamental attribute of objects, essential for understanding their spatial pose and arrangement. However, practical solutions for estimating the orientation of open-world objects in monocular images remain underexplored. In this work, we introduce Orient Anything, the first foundat…

2025

PanDA: Towards Panoramic Depth Anything with Unlabeled Panoramas and Mobius Spatial Augmentation

CVPR 2025poster

Recently, Depth Anything Models (DAMs) - a type of depth foundation models - have demonstrated impressive zero-shot capabilities across diverse perspective images. Despite its success, it remains an open question regarding DAMs' performance on panorama images that enjoy a large field-of-view (180x36…

Cited by 0SourcePDFScholar
2025

ROSE: Remove Objects with Side Effects in Videos

NeurIPS 2025poster

Video object removal has achieved advanced performance due to the recent success of video generative models. However, when addressing the side effects of objects, \textit{e.g.,} their shadows and reflections, existing works struggle to eliminate these effects for the scarcity of paired video data as…

Cited by 0SourceScholar
2025

Seg-VAR:Image Segmentation with Visual Autoregressive Modeling

NeurIPS 2025poster

While visual autoregressive modeling (VAR) strategies have shed light on image generation with the autoregressive models, their potential for segmentation, a task that requires precise low-level spatial perception, remains unexplored. Inspired by the multi-scale modeling of classic Mask2Former-based…

Cited by 0SourceScholar
2025

Sonata: Self-Supervised Learning of Reliable Point Representations

CVPR 2025highlight

In this paper, we question whether we have a reliable self-supervised point cloud model that can be used for diverse 3D tasks via simple linear probing, even with limited data and minimal computation. We find that existing 3D self-supervised learning approaches fall short when evaluated on represent…

2025

SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language

CVPR 2025poster

Contrastive Language-Image Pre-training (CLIP) learns robust visual models through language supervision, making it a crucial visual encoding technique for various applications. However, CLIP struggles with comprehending spatial concepts in images, potentially restricting the spatial intelligence of…

2025

StableDepth: Scene-Consistent and Scale-Invariant Monocular Depth

ICCV 2025poster

Recent advances in monocular depth estimation significantly improve robustness and accuracy. However, relative depth models exhibit flickering and 3D inconsistency in video data, limiting 3D reconstruction applications. We introduce StableDepth, a scene-consistent and scale-invariant depth estimatio…

Cited by 0SourcePDFScholar
2025

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

ICML 2025poster

Recent advancements in reinforcement learning from human feedback have shown that utilizing fine-grained token-level reward models can substantially enhance the performance of Proximal Policy Optimization (PPO) in aligning large language models. However, it is challenging to leverage such token-leve…

2025

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

CVPR 2025highlight

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation…

2025

VIP: Vision Instructed Pre-training for Robotic Manipulation

ICML 2025poster

The effectiveness of scaling up training data in robotic manipulation is still limited. A primary challenge in manipulation is the tasks are diverse, and the trained policy would be confused if the task targets are not specified clearly. Existing works primarily rely on text instruction to describe…

Cited by 0SourcePDFScholar
2025

ViLLa: Video Reasoning Segmentation with Large Language Model

ICCV 2025poster

Recent efforts in video reasoning segmentation (VRS) integrate large language models (LLMs) with perception models to localize and track objects via textual instructions, achieving barely satisfactory results in simple scenarios. However, they struggled to discriminate and deduce the objects from us…

2025

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

NeurIPS 2025poster

Recent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens. However, we observe that most real-world scenarios do not require such an extensive number of visual tokens. While the perf…

Cited by 0SourceScholar
2025

Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance

NeurIPS 2025poster

We present Wan-Move, a simple and scalable framework that brings motion control to video generative models. Existing motion-controllable methods typically suffer from coarse control granularity and limited scalability, leaving their outputs insufficient for practical use. We narrow this gap by achie…

Cited by 0SourceScholar
2024

AnyDoor: Zero-shot Object-level Image Customization

CVPR 2024poster

This work presents AnyDoor a diffusion-based image generator with the power to teleport target objects to new scenes at user-specified locations with desired shapes. Instead of tuning parameters for each object our model is trained only once and effortlessly generalizes to diverse object-scene combi…

Cited by 268SourcePDFScholar
2024

Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

CVPR 2024poster

This work presents Depth Anything a highly practical solution for robust monocular depth estimation. Without pursuing novel technical modules we aim to build a simple yet powerful foundation model dealing with any images under any circumstances. To this end we scale up the dataset by designing a dat…

2024

DreamComposer: Controllable 3D Object Generation via Multi-View Conditions

CVPR 2024poster

Utilizing pre-trained 2D large-scale generative models recent works are capable of generating high-quality novel views from a single in-the-wild image. However due to the lack of information from multiple views these works encounter difficulties in generating controllable novel views. In this paper…

2024

DriveGPT4: Interpretable End-to-End Autonomous Driving Via Large Language Model

RA-L 2024

Multimodallarge language models (MLLMs) have emerged as a prominent area of interest within the research community, given their proficiency in handling and reasoning with non-textual data, including images and videos. This study seeks to extend the application of MLLMs to the realm of autonomous dri

Cited by 603SourceScholar
2024

GPT4Point: A Unified Framework for Point-Language Understanding and Generation

CVPR 2024highlight

Multimodal Large Language Models (MLLMs) have excelled in 2D image-text comprehension and image generation but their understanding of the 3D world is notably deficient limiting progress in 3D language understanding and generation. To solve this problem we introduce GPT4Point an innovative groundbrea…

Cited by 43SourcePDFScholar
2024

GroupContrast: Semantic-aware Self-supervised Representation Learning for 3D Understanding

CVPR 2024poster

Self-supervised 3D representation learning aims to learn effective representations from large-scale unlabeled point clouds. Most existing approaches adopt point discrimination as the pretext task which assigns matched points in two distinct views as positive pairs and unmatched points as negative pa…

2024

GroupLane: End-to-End 3D Lane Detection With Channel-Wise Grouping

RA-L 2024

Efficiency is quite important for 3D lane detection while previous detectors are either computationally expensive or difficult for optimization. To bridge this gap, we propose a fully convolutional detector named GroupLane, which is simple, fast, and still maintains high detection precision. Specifi

Cited by 20SourceScholar
2024

Influencer Backdoor Attack on Semantic Segmentation

ICLR 2024spotlight

When a small number of poisoned samples are injected into the training dataset of a deep neural network, the network can be induced to exhibit malicious behavior during inferences, which poses potential threats to real-world applications. While they have been intensively studied in classification, b…

2024

InsMapper: Exploring Inner-instance Information for Vectorized HD Mapping

ECCV 2024poster

"Vectorized high-definition (HD) maps contain detailed information about surrounding road elements, which are crucial for various downstream tasks in modern autonomous vehicles, such as motion planning and vehicle control. Recent works attempt to directly detect the vectorized HD map as a point set…

Cited by 7SourcePDFScholar
2024

LION: Linear Group RNN for 3D Object Detection in Point Clouds

NeurIPS 2024poster

The benefit of transformers in large-scale 3D point cloud perception tasks, such as 3D object detection, is limited by their quadratic computation cost when modeling long-range relationships. In contrast, linear RNNs have low computational complexity and are suitable for long-range modeling. Toward…

2024

LiT: Unifying LiDAR "Languages" with LiDAR Translator

NeurIPS 2024poster

LiDAR data exhibits significant domain gaps due to variations in sensors, vehicles, and driving environments, creating “language barriers” that limit the effective use of data across domains and the scalability of LiDAR perception models. To address these challenges, we introduce the LiDAR Translato…

2024

Mind the Interference: Retaining Pre-trained Knowledge in Parameter Efficient Continual Learning of Vision-Language Models

ECCV 2024poster

"This study addresses the Domain-Class Incremental Learning problem, a realistic but challenging continual learning scenario where both the domain distribution and target classes vary across tasks. To handle these diverse tasks, pre-trained Vision-Language Models (VLMs) are introduced for their stro…

2024

OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation

CVPR 2024poster

The booming of 3D recognition in the 2020s began with the introduction of point cloud transformers. They quickly overwhelmed sparse CNNs and became state-of-the-art models especially in 3D semantic segmentation. However sparse CNNs are still valuable networks due to their efficiency treasure and eas…

2024

OV-Uni3DETR: Towards Unified Open-Vocabulary 3D Object Detection via Cycle-Modality Propagation

ECCV 2024poster

"In the current state of 3D object detection research, the severe scarcity of annotated 3D data, substantial disparities across different data modalities, and the absence of a unified architecture, have impeded the progress towards the goal of universality. In this paper, we propose OV-Uni3DETR, a u…

2024

One for All: Multi-Domain Joint Training for Point Cloud Based 3D Object Detection

NeurIPS 2024poster

The current trend in computer vision is to utilize one universal model to address all various tasks. Achieving such a universal model inevitably requires incorporating multi-domain data for joint training to learn across multiple problem scenarios. In point cloud based 3D object detection, however,…

Cited by 2SourcePDFScholar
2024

Point Transformer V3: Simpler Faster Stronger

CVPR 2024poster

This paper is not motivated to seek innovation within the attention mechanism. Instead it focuses on overcoming the existing trade-offs between accuracy and efficiency within the context of point cloud processing leveraging the power of scale. Drawing inspiration from recent advances in 3D large-sca…

Cited by 981SourcePDFScholar
2024

SyncVIS: Synchronized Video Instance Segmentation

NeurIPS 2024poster

Recent DETR-based methods have advanced the development of Video Instance Segmentation (VIS) through transformers' efficiency and capability in modeling spatial and temporal information. Despite harvesting remarkable progress, existing works follow asynchronous designs, which model video sequences v…

2024

Towards Large-scale 3D Representation Learning with Multi-dataset Point Prompt Training

CVPR 2024poster

The rapid advancement of deep learning models is often attributed to their ability to leverage massive training data. In contrast such privilege has not yet fully benefited 3D deep learning mainly due to the limited availability of large-scale 3D datasets. Merging multiple available data sources and…

2024

UniPAD: A Universal Pre-training Paradigm for Autonomous Driving

CVPR 2024poster

In the context of autonomous driving the significance of effective feature learning is widely acknowledged. While conventional 3D self-supervised pre-training methods have shown widespread success most methods follow the ideas originally designed for 2D images. In this paper we present UniPAD a nove…

2024

Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding

CVPR 2024poster

3D Visual Grounding (3DVG) aims at localizing 3D object based on textual descriptions. Conventional supervised methods for 3DVG often necessitate extensive annotations and a predefined vocabulary which can be restrictive. To address this issue we propose a novel visual programming approach for zero-…

2024

Zero-shot Image Editing with Reference Imitation

NeurIPS 2024poster

Image editing serves as a practical yet challenging task considering the diverse demands from users, where one of the hardest parts is to precisely describe how the edited image should look like. In this work, we present a new form of editing, termed imitative editing, to help users exercise their c…

Cited by 24SourcePDFScholar
2023

BT^2: Backward-compatible Training with Basis Transformation

ICCV 2023poster

Modern retrieval system often requires recomputing the representation of every piece of data in the gallery when updating to a better representation model. This process is known as backfilling and can be especially costly in the real world where the gallery often contains billions of samples. Recent…

Cited by 6PDFcodeScholar
2023

CorresNeRF: Image Correspondence Priors for Neural Radiance Fields

NeurIPS 2023poster

Neural Radiance Fields (NeRFs) have achieved impressive results in novel view synthesis and surface reconstruction tasks. However, their performance suffers under challenging scenarios with sparse input views. We present CorresNeRF, a novel method that leverages image correspondence priors computed…

2023

Detecting Everything in the Open World: Towards Universal Object Detection

CVPR 2023poster

In this paper, we formally address universal object detection, which aims to detect every scene and predict every category. The dependence on human annotations, the limited visual information, and the novel categories in the open world severely restrict the universality of traditional detectors. We…

2023

FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation Models

NeurIPS 2023poster

Semantic segmentation has witnessed tremendous progress due to the proposal of various advanced network architectures. However, they are extremely hungry for delicate annotations to train, and the acquisition is laborious and unaffordable. Therefore, we present FreeMask in this work, which resorts t…

2023

Masked Scene Contrast: A Scalable Framework for Unsupervised 3D Representation Learning

CVPR 2023poster

As a pioneering work, PointContrast conducts unsupervised 3D representation learning via leveraging contrastive learning over raw RGB-D frames and proves its effectiveness on various downstream tasks. However, the trend of large-scale unsupervised learning in 3D has yet to emerge due to two stumblin…

2023

Mod-Squad: Designing Mixtures of Experts As Modular Multi-Task Learners

CVPR 2023poster

Optimization in multi-task learning (MTL) is more challenging than single-task learning (STL), as the gradient from different tasks can be contradictory. When tasks are related, it can be beneficial to share some parameters among them (cooperation). However, some tasks require additional parameters…

Cited by 107SourcePDFScholar
2023

Open-vocabulary Panoptic Segmentation with Embedding Modulation

ICCV 2023poster

Open-vocabulary segmentation is attracting increasing attention due to its critical applications in the real world. Traditional closed-vocabulary segmentation methods are not able to characterize novel objects, whereas several recent open-vocabulary attempts obtain unsatisfactory results, i.e., nota…

Cited by 34PDFScholar
2023

Semantics-Aware Dynamic Localization and Refinement for Referring Image Segmentation

AAAI 2023technical

Referring image segmentation segments an image from a language expression. With the aim of producing high-quality masks, existing methods often adopt iterative learning approaches that rely on RNNs or stacked attention layers to refine vision-language features. Despite their complexity, RNN-based me…

Cited by 28SourcePDFScholar
2023

Shrinking Class Space for Enhanced Certainty in Semi-Supervised Learning

ICCV 2023poster

Semi-supervised learning is attracting blooming attention, due to its success in combining unlabeled data. To mitigate potentially incorrect pseudo labels, recent frameworks mostly set a fixed confidence threshold to discard uncertain samples. This practice ensures high-quality pseudo labels, but in…

Cited by 24PDFcodeScholar
2023

TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation

NeurIPS 2023poster

Training on large-scale datasets can boost the performance of video instance segmentation while the annotated datasets for VIS are hard to scale up due to the high labor cost. What we possess are numerous isolated filed-specific datasets, thus, it is appealing to jointly train models across the aggr…

2023

Uni3DETR: Unified 3D Detection Transformer

NeurIPS 2023poster

Existing point cloud based 3D detectors are designed for the particular scene, either indoor or outdoor ones. Because of the substantial differences in object distribution and point density within point clouds collected from various environments, coupled with the intricate nature of 3D metrics, ther…

2022

DecoupleNet: Decoupled Network for Domain Adaptive Semantic Segmentation

ECCV 2022poster

"Unsupervised domain adaptation in semantic segmentation alleviates the reliance on expensive pixel-wise annotation. It uses a labeled source domain dataset as well as unlabeled target domain images to learn a segmentation network. In this paper, we observe two main issues of existing domain-invaria…

2022

FocalClick: Towards Practical Interactive Image Segmentation

CVPR 2022poster

Interactive segmentation allows users to extract target masks by making positive/negative clicks. Although explored by many previous works, there is still a gap between academic approaches and industrial needs: first, existing models are not efficient enough to work on low power devices; second, the…

Cited by 185PDFcodeScholar
2022

Generalized Few-Shot Semantic Segmentation

CVPR 2022poster

Training semantic segmentation models requires a large amount of finely annotated data, making it hard to quickly adapt to novel classes not satisfying this condition. Few-Shot Segmentation (FS-Seg) tackles this problem with many constraints. In this paper, we introduce a new benchmark, called Gener…

Cited by 109PDFcodeScholar
2022

LAVT: Language-Aware Vision Transformer for Referring Image Segmentation

CVPR 2022poster

Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key challenges behind this task is leveraging the referring expression for highlighting relevant positions in the image. A para…

Cited by 385PDFcodeScholar
2022

MTFormer: Multi-task Learning via Transformer and Cross-Task Reasoning

ECCV 2022poster

"In this paper, we explore the advantages of utilizing transformer structures for addressing multi-task learning (MTL). Specifically, we demonstrate that models with transformer structures are more appropriate for MTL than convolutional neural networks (CNNs), and we propose a novel transformer-base…

Cited by 67SourcePDFScholar
2022

PhysFormer: Facial Video-Based Physiological Measurement With Temporal Difference Transformer

CVPR 2022poster

Remote photoplethysmography (rPPG), which aims at measuring heart activities and physiological signals from facial video without any contact, has great potential in many applications. Recent deep learning approaches focus on mining subtle rPPG clues using convolutional neural networks with limited s…

Cited by 248PDFcodeScholar
2022

Point Transformer V2: Grouped Vector Attention and Partition-based Pooling

NeurIPS 2022accept

As a pioneering work exploring transformer architecture for 3D point cloud understanding, Point Transformer achieves impressive results on multiple highly competitive benchmarks. In this work, we analyze the limitations of the Point Transformer and propose our powerful and efficient Point Transforme…

2022

Prototype-Voxel Contrastive Learning for LiDAR Point Cloud Panoptic Segmentation

ICRA 2022poster

LiDAR point cloud panoptic segmentation, including both semantic and instance segmentation, plays a critical role in meticulous scene understanding for autonomous driving. Existing 3D voxelized approaches either utilize 3D sparse convolution that only focuses on local scene understanding, or add ext…

Cited by 20SourceScholar
2022

SegPGD: An Effective and Efficient Adversarial Attack for Evaluating and Boosting Segmentation Robustness

ECCV 2022poster

"Deep neural network-based image classifications are vulnerable to adversarial perturbations. The image classifications can be easily fooled by adding artificial small and imperceptible perturbations to input images. As one of the most effective defense strategies, adversarial training was proposed…

Cited by 97SourcePDFScholar
2022

Stratified Transformer for 3D Point Cloud Segmentation

CVPR 2022poster

3D point cloud segmentation has made tremendous progress in recent years. Most current methods focus on aggregating local features, but fail to directly model long-range dependencies. In this paper, we propose Stratified Transformer that is able to capture long-range contexts and demonstrates strong…

Cited by 520PDFcodeScholar
2021

Bidirectional Projection Network for Cross Dimension Scene Understanding

CVPR 2021poster

2D image representations are in regular grids and can be processed efficiently, whereas 3D point clouds are unordered and scattered in 3D space. The information inside these two visual domains is well complementary, e.g., 2D images have fine-grained texture while 3D point clouds contain plentiful ge…

Cited by 147PDFcodeScholar
2021

Do Different Tracking Tasks Require Different Appearance Models?

NeurIPS 2021poster

Tracking objects of interest in a video is one of the most popular and widely applicable problems in computer vision. However, with the years, a Cambrian explosion of use cases and benchmarks has fragmented the problem in a multitude of different experimental setups. As a consequence, the literature…

2021

Dual-Cross Central Difference Network for Face Anti-Spoofing

IJCAI 2021poster

Face anti-spoofing (FAS) plays a vital role in securing face recognition systems. Recently, central difference convolution (CDC) has shown its excellent representation capacity for the FAS task via leveraging local gradient features. However, aggregating central difference clues from all neighbors/d…

2021

Dynamic Divide-and-Conquer Adversarial Training for Robust Semantic Segmentation

ICCV 2021poster

Adversarial training is promising for improving robustness of deep neural networks towards adversarial perturbations, especially on the classification task. The effect of this type of training on semantic segmentation, contrarily, just commences. We make the initial attempt to explore the defense st…

Cited by 47PDFcodeScholar
2021

Fully Convolutional Networks for Panoptic Segmentation

CVPR 2021poster

In this paper, we present a conceptually simple, strong, and efficient framework for panoptic segmentation, called Panoptic FCN. Our approach aims to represent and predict foreground things and background stuff in a unified fully convolutional pipeline. In particular, Panoptic FCN encodes each objec…

Cited by 223PDFcodeScholar
2021

PAConv: Position Adaptive Convolution With Dynamic Kernel Assembling on Point Clouds

CVPR 2021poster

We introduce Position Adaptive Convolution (PAConv), a generic convolution operation for 3D point cloud processing. The key of PAConv is to construct the convolution kernel by dynamically assembling basic weight matrices stored in Weight Bank, where the coefficients of these weight matrices are self…

Cited by 555PDFcodeScholar
2021

Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With Transformers

CVPR 2021poster

Most recent semantic segmentation methods adopt a fully-convolutional network (FCN) with an encoder-decoder architecture. The encoder progressively reduces the spatial resolution and learns more abstract/semantic visual concepts with larger receptive fields. Since context modeling is critical for se…

Cited by 4009PDFcodeScholar
2021

Semi-Supervised Semantic Segmentation With Directional Context-Aware Consistency

CVPR 2021poster

Semantic segmentation has made tremendous progress in recent years. However, satisfying performance highly depends on a large number of pixel-level annotations. Therefore, in this paper, we focus on the semi-supervised segmentation problem where only a small set of labeled data is provided with a mu…

Cited by 277PDFcodeScholar
2020

PointGroup: Dual-Set Point Grouping for 3D Instance Segmentation

CVPR 2020oral

Instance segmentation is an important task for scene understanding. Compared to the fully-developed 2D, 3D instance segmentation for point clouds have much room to improve. In this paper, we present PointGroup, a new end-to-end bottom-up architecture, specifically focused on better grouping the poin…

Cited by 519PDFScholar
2019

Hierarchical Point-Edge Interaction Network for Point Cloud Semantic Segmentation

ICCV 2019poster

We achieve 3D semantic scene labeling by exploring semantic relation between each point and its contextual neighbors through edges. Besides an encoder-decoder branch for predicting point labels, we construct an edge branch to hierarchically integrate point features and generate edge features. To inc…

Cited by 245PDFScholar
2019

PointWeb: Enhancing Local Neighborhood Features for Point Cloud Processing

CVPR 2019poster

This paper presents PointWeb, a new approach to extract contextual features from local neighborhood in a point cloud. Unlike previous work, we densely connect each point with every other in a local neighborhood, aiming to specify feature of each point based on the local region characteristics for be…

Cited by 953PDFcodeScholar
2019

UPSNet: A Unified Panoptic Segmentation Network

CVPR 2019oral

In this paper, we propose a unified panoptic segmentation network (UPSNet) for tackling the newly proposed panoptic segmentation task. On top of a single backbone residual network, we first design a deformable convolution based semantic segmentation head and a Mask R-CNN style instance segmentation…

Cited by 548PDFcodeScholar
2018

Compositing-aware Image Search

ECCV 2018poster

We present a new image search technique that, given a background image, returns compatible foreground objects for image compositing tasks. The compatibility of a foreground object and a background scene depends on various aspects such as semantics, surrounding context, geometry, style and color. How…

Cited by 21SourcePDFScholar
2018

ICNet for Real-Time Semantic Segmentation on High-Resolution Images

ECCV 2018poster

We focus on the challenging task of real-time semantic segmentation in this paper. It finds many practical applications and yet is with fundamental difficulty of reducing a large portion of computation for pixel-wise label inference. We propose an image cascade network (ICNet) that incorporates mult…

2018

PSANet: Point-wise Spatial Attention Network for Scene Parsing

ECCV 2018poster

We notice information flow in convolutional neural networks is restricted inside local neighborhood regions due to the physical design of convolutional filters, which limits the overall understanding of complex scenes. In this paper, we propose the point-wise spatial attention network (PSANet) to re…

2018

SegStereo: Exploiting Semantic Information for Disparity Estimation

ECCV 2018poster

Disparity estimation for binocular stereo images finds a wide range of applications. Traditional algorithms may fail on featureless regions, which could be handled by high-level clues such as semantic segments. In this paper, we suggest that appropriate incorporation of semantic cues can greatly rec…

Cited by 429SourcePDFScholar