← Search

WEI-SHI ZHENG

131 accepted papers

2026

A Closed-Loop Multi-Agent Framework for Robust Multi-Robot Manipulation

RSS 2026poster

Multi-robot systems provide the parallelism and redundancy necessary for long-horizon tasks, while Large Language Models (LLMs) offer the reasoning capabilities to decompose these objectives into actionable plans. However, effectively grounding this high-level reasoning in physical multi-agent execu…

Cited by 0SourceScholar
2026

Beyond Mimicry: Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations

CVPR 2026

Enabling humanoid robots to physically interact with humans is a critical frontier, but progress is hindered by the scarcity of high-quality Human-Humanoid Interaction (HHoI) data. While leveraging abundant Human-Human Interaction (HHI) data presents a scalable alternative, we first demonstrate that

Cited by 0SourceScholar
2026

CycleManip: Enabling Cycle-based Manipulation via Effective History Perception and Understanding

CVPR 2026

In this paper, we explore an important yet underexplored task in robot manipulation: cycle-based manipulation, where robots need to perform cyclic or repetitive actions with an expected terminal time. These tasks are crucial in daily life, such as shaking a bottle or knocking a nail. However, few pr

Cited by 0SourceScholar
2026

DexGrasp-Zero: A Morphology-Aligned Policy for Zero-Shot Cross-Embodiment Dexterous Grasping

RSS 2026poster

To meet the demands of increasingly diverse dexterous hand hardware, it is crucial to develop a policy that enables zero-shot cross-embodiment grasping without redundant re-learning. Cross-embodiment alignment is challenging due to heterogeneous hand kinematics and physical constraints. Existing app…

Cited by 0SourceScholar
2026

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

ICML 2026oral

Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varying environments. While existing approaches attempt to construct behavior representations through action-centric latent v…

Cited by 0SourceScholar
2026

Learning Explicit Continuous Motion Representation for Dynamic Gaussian Splatting from Monocular Videos

CVPR 2026

We present an approach for high-quality dynamic Gaussian Splatting from monocular videos. To this end, we in this work go one step further beyond previous methods to explicitly model continuous position and orientation deformation of dynamic Gaussians, using an SE(3) B-spline motion bases with a com

Cited by 0SourcecodeScholar
2026

Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

CVPR 2026

Referring video object segmentation (RVOS) aims to identify, track and segment the objects in a video based on language descriptions, which has received great attention in recent years. However, existing datasets remain focus on short video clips within several seconds, with salient objects visible

Cited by 0SourcecodeScholar
2026

ObjEmbed: Towards Universal Multimodal Object Embeddings

ICML 2026poster

Aligning objects with corresponding textual descriptions is a fundamental challenge and a realistic requirement in vision-language understanding. While recent multimodal embedding models excel at global image-text alignment, they often struggle with fine-grained alignment between image regions and s…

Cited by 0SourceScholar
2026

OmniDexGrasp: Generalizable Dexterous Grasping Via Foundation Model and Force Feedback

ICRA 2026poster

Enabling robots to dexterously grasp and manipulate objects based on human commands is a promising direction in robotics. However, existing approaches are challenging to generalize across diverse objects or tasks due to the limited scale of semantic dexterous grasp datasets. Foundation models offer …

2026

PhysiGen: Integrating Collision-Aware Physical Constraints for High-Fidelity Human-Human Interaction Generation

ICASSP 2026poster

Despite substantial progress in text-driven 3D human motion synthesis, generating realistic multi-person interaction sequences remains challenging. Notably, body inter-penetration is a pervasive issue from both data acquisition to the generated results, which significantly undermines the realism and…

Cited by 0SourcePDFScholar
2026

Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation

CVPR 2026

Referring Video Object Segmentation (RVOS) aims to segment objects in videos based on textual queries. Current methods mainly rely on large-scale supervised fine-tuning (SFT) of Multi-modal Large Language Models (MLLMs). However, this paradigm suffers from heavy data dependence and limited scalabili

Cited by 0SourcecodeScholar
2026

RehearseVLA: Simulated Post-Training for VLAs with Physically-Consistent World Model

CVPR 2026

Vision-Language-Action (VLA) models trained via imitation learning suffer from significant performance degradation in data-scarce scenarios due to their reliance on large-scale demonstration datasets. Although reinforcement learning (RL)-based post-training has proven effective in addressing data sc

Cited by 0SourcecodeScholar
2026

SGS-Intrinsic: Semantic-Invariant Gaussian Splatting for Sparse-View Indoor Inverse Rendering

CVPR 2026

We present SGS-Intrinsic, an indoor inverse rendering framework that works well for sparse-view images. Unlike existing 3D Gaussian Splatting (3DGS) based methods that focus on object-centric reconstruction and fail to work under sparse view settings, our method allows to achieve high-quality geomet

Cited by 0SourcecodeScholar
2026

Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search

ICML 2026poster

Segmentation based on language has been a popular topic in computer vision. While recent advances in multimodal large language models (MLLMs) have endowed segmentation systems with reasoning capabilities, these efforts remain confined by the frozen internal knowledge of MLLMs, which limits their pot…

Cited by 0SourceScholar
2026

VLANeXt: Recipes for Building Strong VLA Models

ICML 2026poster

Following the rise of large foundation models, Vision–Language–Action models (VLAs) emerged, leveraging strong visual and language understanding for general-purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA mo…

Cited by 0SourceScholar
2026

WeDetect: Fast Open-Vocabulary Object Detection as Retrieval

CVPR 2026

Open-vocabulary object detection aims to detect arbitrary classes via text prompts. Methods without cross-modal fusion layers (non-fusion) offer faster inference by treating recognition as a retrieval problem, i.e., matching regions to text queries in a shared embedding space. In this work, we fully

Cited by 0SourcecodeScholar
2026

You Only Erase Once: Erasing Anything without Bringing Unexpected Content

CVPR 2026

We present YOEO, an approach for object erasure. Unlike recent diffusion-based methods which struggle to erase target objects without generating unexpected content within the masked regions due to lack of sufficient paired training data and explicit constraint on content generation, our method allow

Cited by 0SourcecodeScholar
2025

AffordDexGrasp: Open-set Language-guided Dexterous Grasp with Generalizable-Instructive Affordance

ICCV 2025poster

Language-guided robot dexterous generation enables robots to grasp and manipulate objects based on human commands. However, previous data-driven methods are hard to understand intention and execute grasping with unseen categories in the open set. In this work, we explore a new task, Open-set Languag…

Cited by 0SourcePDFScholar
2025

CLIP-RestoreX: Restore Image Structure and Perception in Exposure Correction

AAAI 2025technical

Exposure correction aims to adjust the exposure of an under- and over-exposed image to enhance its overall visual quality. The core challenge of this task lies in that it requires to faithfully restore both the structure and perception information. In this work, we present a novel exposure correctio…

2025

Chain of Methodologies: Scaling Test Time Computation without Training

ACL 2025finding

Large Language Models (LLMs) often struggle with complex reasoning tasks due to insufficient in-depth insights in their training data, which are frequently absent in publicly available documents. This paper introduces the Chain of Methodologies (CoM), a simple and innovative iterative prompting fram…

Cited by 0SourcePDFScholar
2025

ChainHOI: Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation

CVPR 2025poster

We propose ChainHOI, a novel approach for text-driven human-object interaction (HOI) generation that explicitly models interactions at both the joint and kinetic chain levels. Unlike existing methods that implicitly model interactions using full-body poses as tokens, we argue that explicitly mode…

Cited by 2SourcePDFScholar
2025

DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse Rendering

ICCV 2025poster

Recent methods have shown that pre-trained diffusion models can be fine-tuned to enable generative inverse rendering by learning image-conditioned noise-to-intrinsic mapping. Despite their remarkable progress, they struggle to robustly produce high-quality results as the noise-to-intrinsic paradigm…

2025

Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric Tasks

CVPR 2025highlight

In this work, we present DEcoupLEd Distillation To Erase (DELETE), a general and strong unlearning method for any class-centric tasks. To derive this, we first propose a theoretical framework to analyze the general form of unlearning loss and decompose it into forgetting and retention terms. Through…

Cited by 2SourcePDFScholar
2025

Distilling LLM Prior to Flow Model for Generalizable Agent’s Imagination in Object Goal Navigation

NeurIPS 2025poster

The Object Goal Navigation (ObjectNav) task challenges agents to locate a specified object in an unseen environment by imagining unobserved regions of the scene. Prior approaches rely on deterministic and discriminative models to complete semantic maps, overlooking the inherent uncertainty in indoor…

Cited by 0SourceScholar
2025

EntityErasure: Erasing Entity Cleanly via Amodal Entity Segmentation and Completion

CVPR 2025poster

This paper presents EntityErasure, a novel diffusion-based inpainting method that can effectively erase entities without inducing unwanted sundries. To this end, we propose to address this problem by dividing it into amodal entity segmentation and completion, such that the region to inpaint takes on…

2025

FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution Detection

ICCV 2025poster

Pre-trained vision-language models (VLMs) have advanced out-of-distribution (OOD) detection recently. However, existing CLIP-based methods often focus on learning OOD-related knowledge to improve OOD detection, showing limited generalization or reliance on external large-scale auxiliary datasets. In…

2025

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

CVPR 2025highlight

Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed captions for each image can further improve performance. To achi…

2025

Learning Implicit Features with Flow-Infused Transformations for Realistic Virtual Try-On

ICCV 2025poster

Diffusion-based virtual try-on aims to synthesize a realistic image that seamlessly integrating the specific garment into a target model. The primary challenge lies in effectively guiding the warping process of the latent diffusion model. However, previous methods either lack direct guidance or expl…

Cited by 0SourcePDFScholar
2025

Less Static, More Private: Towards Transferable Privacy-Preserving Action Recognition by Generative Decoupled Learning

ICCV 2025poster

This work focuses on the task of privacy-preserving action recognition (PPAR), which aims to protect individual privacy in action videos without compromising recognition performance. Despite recent advancements, existing PPAR models still struggle with video domain shifts. To address this challenge,…

Cited by 0SourcePDFScholar
2025

Light-T2M: A Lightweight and Fast Model for Text-to-motion Generation

AAAI 2025technical

Despite the significant role text-to-motion (T2M) generation plays across various applications, current methods involve a large number of parameters and suffer from slow inference speeds, leading to high usage costs. To address this, we aim to design a lightweight model to reduce usage costs. First,…

2025

MaintaAvatar: A Maintainable Avatar Based on Neural Radiance Fields by Continual Learning

AAAI 2025technical

The generation of a virtual digital avatar is a crucial research topic in the field of computer vision. Many existing works utilize Neural Radiance Fields (NeRF) to address this issue and have achieved impressive results. However, previous works assume the images of the training person are available…

Cited by 0SourcePDFScholar
2025

Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks

CVPR 2025poster

Error detection in procedural activities is essential for consistent and correct outcomes in AR-assisted and robotic systems. Existing methods often focus on temporal ordering errors or rely on static prototypes to represent normal actions. However, these approaches typically overlook the common sce…

2025

MotionGrasp: Long-Term Grasp Motion Tracking for Dynamic Grasping

RA-L 2025

Dynamic grasping, which aims to grasp moving objects in unstructured environment, is crucial for robotics community. Previous methods propose to track the initial grasps or objects by matching between the latest two frames. However, this neighbour-frame matching strategy ignores the long-term histor

Cited by 6SourceScholar
2025

Panorama Generation From NFoV Image Done Right

CVPR 2025highlight

Generating 360-degree panoramas from narrow field of view (NFoV) image is a promising computer vision task for Virtual Reality (VR) applications. Existing methods mostly assess the generated panoramas with InceptionNet or CLIP based metrics, which tend to perceive the image quality and is not suitab…

2025

ParGo: Bridging Vision-Language with Partial and Global Views

AAAI 2025technical

This work presents ParGo, a novel Partial-Global projector designed to connect the vision and language modalities for Multimodal Large Language Models (MLLMs). Unlike previous works that rely on global attention-based projectors, our ParGo bridges the representation gap between the separately pre-tr…

2025

Person De-reidentification: A Variation-guided Identity Shift Modeling

CVPR 2025poster

Person re-identification (ReID) is to associate images of individuals from different camera views against cross-view variations. Like other surveillance technologies, Re-ID faces serious privacy challenges, particularly the potential for unauthorized tracking. Although various tasks (e.g., face reco…

Cited by 0SourcePDFScholar
2025

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

ICCV 2025poster

Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language understanding, pixel-level dense prediction and spatiotemporal reasoning. Despite notable progress in recent years, existi…

Cited by 0SourcePDFScholar
2025

Rethinking Bimanual Robotic Manipulation: Learning with Decoupled Interaction Framework

ICCV 2025poster

Bimanual robotic manipulation is an emerging and critical topic in the robotics community. Previous works primarily rely on integrated control models that take the perceptions and states of both arms as inputs to directly predict their actions. However, we think bimanual manipulation involves not on…

Cited by 0SourcePDFScholar
2025

RoGSplat: Learning Robust Generalizable Human Gaussian Splatting from Sparse Multi-View Images

CVPR 2025poster

This paper presents RoGSplat, a novel approach for synthesizing high-fidelity novel views of unseen human from sparse multi-view images, while requiring no cumbersome per-subject optimization. Unlike previous methods that typically struggle with sparse views with few overlappings and are less effect…

2025

Structure-Guided Diffusion Models for High-Fidelity Portrait Shadow Removal

ICCV 2025poster

We present a diffusion-based portrait shadow removal approach that can robustly produce high-fidelity results. Unlike previous methods, we cast shadow removal as diffusion-based inpainting. To this end, we first train a shadow-independent structure extraction network on a real-world portrait dataset…

2025

TacCap: A Wearable FBG-Based Tactile Sensor for Efficient Human-to-Robot Skill Transfer

IROS 2025

Tactile sensing is essential for dexterous manipulation, yet large-scale human demonstration datasets lack tactile feedback, limiting their effectiveness in skill transfer to robots. To address this, we introduce TacCap, a wearable Fiber Bragg Grating (FBG)-based tactile sensor designed for seamless

Cited by 0SourceScholar
2025

TypeTele: Releasing Dexterity in Teleoperation by Dexterous Manipulation Types

CoRL 2025poster

Dexterous teleoperation plays a crucial role in robotic manipulation for real-world data collection and remote robot control. Previous dexterous teleoperation mostly relies on hand retargeting to closely mimic human hand postures. However, these approaches may fail to fully leverage the inherent dex…

Cited by 0SourceScholar
2025

VIPerson: Flexibly Generating Virtual Identity for Person Re-Identification

ICCV 2025poster

Person re-identification (ReID) is to match the person images under different camera views. Training ReID models necessitates a substantial amount of labeled real-world data, leading to high labeling costs and privacy issues. Although several ReID data synthetic methods are proposed to address these…

2025

ViSpeak: Visual Instruction Feedback in Streaming Videos

ICCV 2025poster

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive characteristics. In this work, we aim to extend the streaming…

2025

When Shadow Removal Meets Intrinsic Image Decomposition: A Joint Learning Framework Using Unpaired Data

AAAI 2025technical

We present a framework that achieves shadow removal by learning intrinsic image decomposition (IID) from unpaired shadow and shadow-free images. Although it is well-known that intrinsic images, \ie, illumination and reflectance, are highly beneficial to shadow removal, IID is rarely adopted by previ…

Cited by 0SourcePDFScholar
2025

iManip: Skill-Incremental Learning for Robotic Manipulation

ICCV 2025poster

The development of a generalist agent with adaptive multiple manipulation skills has been a long-standing goal in the robotics community.In this paper, we explore a crucial task, skill-incremental learning, in robotic manipulation, which is to endow the robots with the ability to learn new manipulat…

Cited by 0SourcePDFScholar
2025

monoVLN: Bridging the Observation Gap between Monocular and Panoramic Vision and Language Navigation

ICCV 2025poster

Vision and Language Navigation(VLN) requires agents to navigate 3D environments by following natural language instructions. While existing methods predominantly assume access to panoramic observations, many practical robotics are equipped with monocular RGBD cameras, creating a significant configura…

Cited by 0SourcePDFScholar
2024

Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary Alignment

CVPR 2024poster

Weakly-supervised action segmentation is a task of learning to partition a long video into several action segments where training videos are only accompanied by transcripts (ordered list of actions). Most of existing methods need to infer pseudo segmentation for training by serial alignment between…

2024

Exploiting Discrepancy in Feature Statistic for Out-of-Distribution Detection

AAAI 2024technical

Recent studies on out-of-distribution (OOD) detection focus on designing models or scoring functions that can effectively distinguish between unseen OOD data and in-distribution (ID) data. In this paper, we propose a simple yet novel ap- proach to OOD detection by leveraging the phenomenon that the…

2024

Factorized Diffusion Autoencoder for Unsupervised Disentangled Representation Learning

AAAI 2024technical

Unsupervised disentangled representation learning aims to recover semantically meaningful factors from real-world data without supervision, which is significant for model generalization and interpretability. Current methods mainly rely on assumptions of independence or informativeness of factors, re…

2024

FeatWalk: Enhancing Few-Shot Classification through Local View Leveraging

AAAI 2024technical

Few-shot learning is a challenging task due to the limited availability of training samples. Recent few-shot learning studies with meta-learning and simple transfer learning methods have achieved promising performance. However, the feature extractor pre-trained with the upstream dataset may neglect…

2024

Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models

NeurIPS 2024poster

Recent vision foundation models can extract universal representations and show impressive abilities in various tasks. However, their application on object detection is largely overlooked, especially without fine-tuning them. In this work, we show that frozen foundation models can be a versatile feat…

Cited by 4SourcePDFScholar
2024

Grasp as You Say: Language-guided Dexterous Grasp Generation

NeurIPS 2024poster

This paper explores a novel task "Dexterous Grasp as You Say'' (DexGYS), enabling robots to perform dexterous grasping based on human commands expressed in natural language. However, the development of this field is hindered by the lack of datasets with natural human guidance; thus, we propose a lan…

2024

PRET: Planning with Directed Fidelity Trajectory for Vision and Language Navigation

ECCV 2024poster

"Vision and language navigation is a task that requires an agent to navigate according to a natural language instruction. Recent methods predict sub-goals on constructed topology map at each step to enable long-term action planning. However, they suffer from high computational cost when attempting t…

2024

Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels

CVPR 2024poster

This paper focuses on open-ended video question answering which aims to find the correct answers from a large answer set in response to a video-related question. This is essentially a multi-label classification task since a question may have multiple answers. However due to annotation costs the labe…

Cited by 2SourcePDFScholar
2024

Real-to-Sim Grasp: Rethinking the Gap between Simulation and Real World in Grasp Detection

CoRL 2024poster

For 6-DoF grasp detection, simulated data is expandable to train more powerful model, but it faces the challenge of the large gap between simulation and real world. Previous works bridge this gap with a sim-to-real way. However, this way explicitly or implicitly forces the simulated data to adapt to…

Cited by 4SourcecodeScholar
2024

Rethinking Few-shot Class-incremental Learning: Learning from Yourself

ECCV 2024poster

"Few-shot class-incremental learning (FSCIL) aims to learn sequential classes with limited samples in a few-shot fashion. Inherited from the classical class-incremental learning setting, the popular benchmark of FSCIL uses averaged accuracy (aAcc) and last-task averaged accuracy (lAcc) as the evalua…

2024

Revealing Distribution Discrepancy by Sampling Transfer in Unlabeled Data

NeurIPS 2024poster

There are increasing cases where the class labels of test samples are unavailable, creating a significant need and challenge in measuring the discrepancy between training and test distributions. This distribution discrepancy complicates the assessment of whether the hypothesis selected by an algorit…

Cited by 0SourcePDFScholar
2024

Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training

CVPR 2024poster

Contrastive learning has emerged as a promising paradigm for 3D open-world understanding i.e. aligning point cloud representation to image and text embedding space individually. In this paper we introduce MixCon3D a simple yet effective method aiming to sculpt holistic 3D representation in contrasti…

2024

Selective Hourglass Mapping for Universal Image Restoration Based on Diffusion Model

CVPR 2024poster

Universal image restoration is a practical and potential computer vision task for real-world applications. The main challenge of this task is handling the different degradation distributions at once. Existing methods mainly utilize task-specific conditions (e.g. prompt) to guide the model to learn d…

2024

Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding

CVPR 2024poster

Video Paragraph Grounding (VPG) is an emerging task in video-language understanding which aims at localizing multiple sentences with semantic relations and temporal order from an untrimmed video. However existing VPG approaches are heavily reliant on a considerable number of temporal labels that are…

Cited by 4SourcePDFScholar
2024

Single-View Scene Point Cloud Human Grasp Generation

CVPR 2024poster

In this work we explore a novel task of generating human grasps based on single-view scene point clouds which more accurately mirrors the typical real-world situation of observing objects from a single viewpoint. Due to the incompleteness of object point clouds and the presence of numerous scene poi…

2024

TagFog: Textual Anchor Guidance and Fake Outlier Generation for Visual Out-of-Distribution Detection

AAAI 2024technical

Out-of-distribution (OOD) detection is crucial in many real-world applications. However, intelligent models are often trained solely on in-distribution (ID) data, leading to overconfidence when misclassifying OOD data as ID classes. In this study, we propose a new learning framework which leverage…

2023

ASAG: Building Strong One-Decoder-Layer Sparse Detectors via Adaptive Sparse Anchor Generation

ICCV 2023poster

Recent sparse detectors with multiple, e.g. six, decoder layers achieve promising performance but much inference time due to complex heads. Previous works have explored using dense priors as initialization and built one-decoder-layer detectors. Although they gain remarkable acceleration, their perfo…

Cited by 8PDFcodeScholar
2023

AsyFOD: An Asymmetric Adaptation Paradigm for Few-Shot Domain Adaptive Object Detection

CVPR 2023poster

In this work, we study few-shot domain adaptive object detection (FSDAOD), where only a few target labeled images are available for training in addition to sufficient source labeled images. Critically, in FSDAOD, the data-scarcity in the target domain leads to an extreme data imbalance between the s…

2023

Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video Grounding

CVPR 2023poster

Spatio-Temporal Video Grounding (STVG) aims to localize the target object spatially and temporally according to the given language query. It is a challenging task in which the model should well understand dynamic visual cues (e.g., motions) and static visual cues (e.g., object appearances) in the la…

2023

Diversifying Spatial-Temporal Perception for Video Domain Generalization

NeurIPS 2023poster

Video domain generalization aims to learn generalizable video classification models for unseen target domains by training in a source domain. A critical challenge of video domain generalization is to defend against the heavy reliance on domain-specific cues extracted from the source domain when reco…

2023

Estimator Meets Equilibrium Perspective: A Rectified Straight Through Estimator for Binary Neural Networks Training

ICCV 2023poster

Binarization of neural networks is a dominant paradigm in neural networks compression. The pioneering work BinaryConnect uses Straight Through Estimator (STE) to mimic the gradients of the sign function, but it also causes the crucial inconsistency problem. Most of the previous methods design differ…

Cited by 19PDFcodeScholar
2023

Event-Guided Procedure Planning from Instructional Videos with Text Supervision

ICCV 2023poster

In this work, we focus on the task of procedure planning from instructional videos with text supervision, where a model aims to predict an action sequence to transform the initial visual state into the goal visual state. A critical challenge of this task is the large semantic gap between observed vi…

Cited by 20PDFScholar
2023

Generating Anomalies for Video Anomaly Detection With Prompt-Based Feature Mapping

CVPR 2023poster

Anomaly detection in surveillance videos is a challenging computer vision task where only normal videos are available during training. Recent work released the first virtual anomaly detection dataset to assist real-world detection. However, an anomaly gap exists because the anomalies are bounded in…

Cited by 42SourcePDFScholar
2023

Grasp Region Exploration for 7-DoF Robotic Grasping in Cluttered Scenes

IROS 2023poster

Robotic grasping is a fundamental skill for robots, but it is quite challenging in cluttered scenes. In cluttered scenes, the precise prediction of high-quality grasp configurations such as rotation and grasping width while avoiding collisions is essential. To accomplish this, the grasp detection mo…

Cited by 7SourceScholar
2023

Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding

CVPR 2023poster

Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of the critical challenges in addressing this problem is to comprehend the complex sem…

Cited by 24SourcePDFScholar
2023

Inner-Outer Aware Reconstruction Model for Monocular 3D Scene Reconstruction

NeurIPS 2023poster

Monocular 3D scene reconstruction aims to reconstruct the 3D structure of scenes based on posed images. Recent volumetric-based methods directly predict the truncated signed distance function (TSDF) volume and have achieved promising results. The memory cost of volumetric-based methods will grow cub…

2023

Revisit PCA-based Technique for Out-of-Distribution Detection

ICCV 2023poster

Out-of-distribution (OOD) detection is a desired ability to ensure the reliability and safety of intelligent systems. A scoring function is often designed to measure the degree of any new data being an OOD sample. While most designed scoring functions are based on a single source of information (e.g…

Cited by 7PDFcodeScholar
2023

Shape-Erased Feature Learning for Visible-Infrared Person Re-Identification

CVPR 2023poster

Due to the modality gap between visible and infrared images with high visual ambiguity, learning diverse modality-shared semantic concepts for visible-infrared person re-identification (VI-ReID) remains a challenging problem. Body shape is one of the significant modality-shared cues for VI-ReID. To…

2023

Temporal Continual Learning with Prior Compensation for Human Motion Prediction

NeurIPS 2023poster

Human Motion Prediction (HMP) aims to predict future poses at different moments according to past motion sequences. Previous approaches have treated the prediction of various moments equally, resulting in two main limitations: the learning of short-term predictions is hindered by the focus on long-t…

2022

AcroFOD: An Adaptive Method for Cross-Domain Few-Shot Object Detection

ECCV 2022poster

"Under the domain shift, cross-domain few-shot object detection aims to adapt object detectors in the target domain with a few annotated target data. There exists two significant challenges: (1) Highly insufficient target domain data; (2) Potential over-adaptation and misleading caused by inappropri…

2022

Learning To Imagine: Diversify Memory for Incremental Learning Using Unlabeled Data

CVPR 2022poster

Deep neural network (DNN) suffers from catastrophic forgetting when learning incrementally, which greatly limits its applications. Although maintaining a handful of samples (called "exemplars") of each task could alleviate forgetting to some extent, existing methods are still limited by the small nu…

Cited by 38PDFcodeScholar
2022

Lifelong Person Re-identification by Pseudo Task Knowledge Preservation

AAAI 2022technical

In real world, training data for person re-identification (Re-ID) is collected discretely with spatial and temporal variations, which requires a model to incrementally learn new knowledge without forgetting old knowledge. This problem is called lifelong person re-identification (LReID). Variations o…

2022

SIOD: Single Instance Annotated per Category per Image for Object Detection

CVPR 2022poster

Object detection under imperfect data receives great attention recently. Weakly supervised object detection (WSOD) suffers from severe localization issues due to the lack of instance-level annotation, while semi-supervised object detection (SSOD) remains challenging led by the inter-image discrepanc…

Cited by 32PDFcodeScholar
2022

Text-Adaptive Multiple Visual Prototype Matching for Video-Text Retrieval

NeurIPS 2022accept

Cross-modal retrieval between videos and texts has gained increasing interest because of the rapid emergence of videos on the web. Generally, a video contains rich instance and event information and the query text only describes a part of the information. Thus, a video can have multiple different…

Cited by 32SourcePDFScholar
2021

Action-guided 3D Human Motion Prediction

NeurIPS 2021poster

The ability of forecasting future human motion is important for human-machine interaction systems to understand human behaviors and make interaction. In this work, we focus on developing models to predict future human motion from past observed video frames. Motivated by the observation that human mo…

Cited by 10SourcePDFScholar
2021

Fine-Grained Shape-Appearance Mutual Learning for Cloth-Changing Person Re-Identification

CVPR 2021poster

Recently, person re-identification (Re-ID) has achieved great progress. However, current methods largely depend on color appearance, which is not reliable when a person changes the clothes. Cloth-changing Re-ID is challenging since pedestrian images with clothes change exhibit large intra-class vari…

Cited by 206PDFScholar
2021

Graph-Based High-Order Relation Modeling for Long-Term Action Recognition

CVPR 2021poster

Long-term actions involve many important visual concepts, e.g., objects, motions, and sub-actions, and there are various relations among these concepts, which we call basic relations. These basic relations will jointly affect each other during the temporal evolution of long-term actions, which forms…

Cited by 67PDFScholar
2021

Learning 3D Shape Feature for Texture-Insensitive Person Re-Identification

CVPR 2021poster

It is well acknowledged that person re-identification (person ReID) highly relies on visual texture information like clothing. Despite significant progress has been made in recent years, texture-confusing situations like clothing changing and persons wearing the same clothes receive little attention…

Cited by 147PDFScholar
2021

Learning To Know Where To See: A Visibility-Aware Approach for Occluded Person Re-Identification

ICCV 2021poster

Person re-identification (ReID) has gained an impressive progress in recent years. However, the occlusion is still a common and challenging problem for recent ReID methods. Several mainstream methods utilize extra cues (e.g., human pose information) to distinguish human parts from obstacles to allev…

Cited by 86PDFScholar
2021

MIST: Multiple Instance Self-Training Framework for Video Anomaly Detection

CVPR 2021poster

Weakly supervised video anomaly detection (WS-VAD) is to distinguish anomalies from normal events based on discriminative representations. Most existing works are limited in insufficient video representations. In this work, we develop a multiple instance self-training framework (MIST) to efficiently…

Cited by 342PDFcodeScholar
2021

Predictive Feature Learning for Future Segmentation Prediction

ICCV 2021poster

Future segmentation prediction aims to predict the segmentation masks for unobserved future frames. Most existing works addressed it by directly predicting the intermediate features extracted by existing segmentation models. However, these segmentation features are learned to be local discriminative…

Cited by 20PDFScholar
2021

Weakly Supervised Text-Based Person Re-Identification

ICCV 2021poster

The conventional text-based person re-identification methods heavily rely on identity annotations. However, this labeling process is costly and time-consuming. In this paper, we consider a more practical setting called weakly supervised text-based person re-identification, where only the text-image…

Cited by 40PDFcodeScholar
2020

Adaptive Interaction Modeling via Graph Operations Search

CVPR 2020poster

Interaction modeling is important for video action analysis. Recently, several works design specific structures to model interactions in videos. However, their structures are manually designed and non-adaptive, which require structures design efforts and more importantly could not model interactions…

Cited by 7PDFcodeScholar
2020

An Asymmetric Modeling for Action Assessment

ECCV 2020poster

Action assessment is a task of assessing the performance of an action. It is widely applicable to many real-world scenarios such as medical treatment and sporting events. However, existing methods for action assessment are mostly limited to individual actions, especially lacking modeling of the asym…

Cited by 58SourcePDFScholar
2020

Contextual Heterogeneous Graph Network for Human-Object Interaction Detection

ECCV 2020poster

Human-object interaction (HOI) detection is an important task for understanding human activity. Graph structure is appropriate to denote the HOIs in the scene. Since there is an subordination between human and object---human play subjective role and object play objective role in HOI, the relations b…

Cited by 114SourcePDFScholar
2020

Do Not Disturb Me: Person Re-identification Under the Interference of Other Pedestrians

ECCV 2020poster

In the conventional person Re-ID setting, it is assumed that cropped images are the person images within the bounding box for each individual. However, in a crowded scene, off-shelf-detectors may generate bounding boxes involving multiple people, where the large proportion of background pedestrians…

2020

Learning to Detect Important People in Unlabelled Images for Semi-Supervised Important People Detection

CVPR 2020poster

Important people detection is to automatically detect the individuals who play the most important roles in a social event image, which requires the designed model to understand a high-level pattern. However, existing methods rely heavily on supervised learning using large quantities of annotated ima…

Cited by 21PDFScholar
2020

MINI-Net: Multiple Instance Ranking Network for Video Highlight Detection

ECCV 2020poster

We address the weakly supervised video highlight detection problem for learning to detect segments that are more attractive in training videos given their video event label but without expensive supervision of manually annotating highlight segments. While manually averting localizing highlight segme…

Cited by 87SourcePDFScholar
2020

Spatial-Temporal Graph Convolutional Network for Video-Based Person Re-Identification

CVPR 2020poster

While video-based person re-identification (Re-ID) has drawn increasing attention and made great progress in recent years, it is still very challenging to effectively overcome the occlusion problem and the visual ambiguity problem for visually similar negative samples. On the other hand, we observe…

Cited by 266PDFScholar
2020

Squeeze-and-Attention Networks for Semantic Segmentation

CVPR 2020poster

The recent integration of attention mechanisms into segmentation networks improves their representational capabilities through a great emphasis on more informative features. However, these attention mechanisms ignore an implicit sub-task of semantic segmentation and are constrained by the grid struc…

Cited by 288PDFScholar
2020

Weakly Supervised Discriminative Feature Learning With State Information for Person Identification

CVPR 2020poster

Unsupervised learning of identity-discriminative visual feature is appealing in real-world tasks where manual labelling is costly. However, the images of an identity can be visually discrepant when images are taken under different states, e.g. different camera views and poses. This visual discrepanc…

Cited by 32PDFcodeScholar
2019

Patch-Based Discriminative Feature Learning for Unsupervised Person Re-Identification

CVPR 2019poster

While discriminative local features have been shown effective in solving the person re-identification problem, they are limited to be trained on fully pairwise labelled data which is expensive to obtain. In this work, we overcome this problem by proposing a patch-based unsupervised learning framewor…

Cited by 268PDFcodeScholar
2019

Progressive Teacher-Student Learning for Early Action Prediction

CVPR 2019poster

The goal of early action prediction is to recognize actions from partially observed videos with incomplete action executions, which is quite different from action recognition. Predicting early actions is very challenging since the partially observed videos do not contain enough action information fo…

Cited by 166PDFcodeScholar
2019

Underexposed Photo Enhancement Using Deep Illumination Estimation

CVPR 2019oral

This paper presents a new neural network for enhancing underexposed photos. Instead of directly learning an image-to-image mapping as previous work, we introduce intermediate illumination in our network to associate the input with expected enhancement result, which augments the network's capability…

Cited by 1084PDFcodeScholar
2019

Unsupervised Person Re-Identification by Camera-Aware Similarity Consistency Learning

ICCV 2019poster

For matching pedestrians across disjoint camera views in surveillance, person re-identification (Re-ID) has made great progress in supervised learning. However, it is infeasible to label data in a number of new scenes when extending a Re-ID system. Thus, studying unsupervised learning for Re-ID is i…

Cited by 144PDFScholar
2019

Unsupervised Person Re-Identification by Soft Multilabel Learning

CVPR 2019oral

Although unsupervised person re-identification (RE-ID) has drawn increasing research attentions due to its potential to address the scalability problem of supervised RE-ID models, it is very challenging to learn discriminative information in the absence of pairwise labels across disjoint camera view…

Cited by 487PDFcodeScholar
2018

Deep Bilinear Learning for RGB-D Action Recognition

ECCV 2018poster

In this paper, we focus on exploring modality-temporal mutual information for RGB-D action recognition. In order to learn time-varying information and multi-modal features jointly, we propose a novel deep bilinear learning framework. In the framework, we propose bilinear blocks that consist of two l…

Cited by 116SourcePDFScholar
2017

Cross-View Asymmetric Metric Learning for Unsupervised Person Re-Identification

ICCV 2017poster

While metric learning is important for Person re-identification (RE-ID), a significant problem in visual surveillance for cross-view pedestrian matching, existing metric models for RE-ID are mostly based on supervised learning that requires quantities of labeled samples in all pairs of camera views…

Cited by 397PDFcodeScholar
2017

RGB-Infrared Cross-Modality Person Re-Identification

ICCV 2017poster

Person re-identification (Re-ID) is an important problem in video surveillance, aiming to match pedestrian images across camera views. Currently, most works focus on RGB-based Re-ID. However, in some applications, RGB images are not suitable, e.g. in a dark environment or at night. Infrared (IR) ima…

Cited by 896PDFScholar
2015

Jointly Learning Heterogeneous Features for RGB-D Activity Recognition

CVPR 2015poster

In this paper, we focus on heterogeneous feature learning for RGB-D activity recognition. Considering that features from different channels could share some similar hidden structures, we propose a joint learning model to simultaneously explore the shared and feature-specific components as an instanc…

Cited by 687SourcePDFScholar
2015

Multi-Scale Learning for Low-Resolution Person Re-Identification

ICCV 2015poster

In real world person re-identification (re-id), images of people captured at very different resolutions from different locations need be matched. Existing re-id models typically normalise all person images to the same size. However, a low-resolution (LR) image contains much less information about a…

Cited by 196PDFScholar