← Search

Guanbin Li

123 accepted papers

2026

3DGS-HPC: Distractor-free 3D Gaussian Splatting with Hybrid Patch-wise Classification

ICML 2026poster

3D Gaussian Splatting (3DGS) has demonstrated remarkable performance in novel view synthesis and 3D scene reconstruction, but its quality often degrades in real-world environments due to transient distractors, such as moving objects and varying shadows. Existing methods commonly rely on semantic cue…

Cited by 0SourceScholar
2026

CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation

ICML 2026poster

Recent video generation models have revealed the emergence of Chain-of-Frame (CoF) reasoning, enabling frame-by-frame visual inference. With this capability, video models have been successfully applied to various visual tasks (*e.g.*, maze solving, visual puzzles). However, their potential to enhanc…

Cited by 0SourceScholar
2026

Cross-Modal Attention Calibration for LVLM Hallucination Mitigation

CVPR 2026

Large vision-language models (LVLMs) have shown remarkable capabilities in visual-language understanding. Despite their success, LVLMs still suffer from generating hallucinations in complex generation tasks, leading to inconsistencies between visual inputs and generated content. To address this issu

Cited by 0SourceScholar
2026

DDP-WM: Disentangled Dynamics Prediction for Efficient World Models

ICML 2026poster

World models are essential for autonomous robotic planning. However, the substantial computational overhead of existing dense Transformer-based models significantly hinders real-time deployment. To address this efficiency-performance bottleneck, we introduce DDP-WM, a novel world model centered on t…

Cited by 0SourceScholar
2026

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

CVPR 2026

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing approaches are mostly adapted from text-to-image diffusion models, which struggle

Cited by 0SourceScholar
2026

FLARE: A Failure-Aware Framework for Autonomous Correction and Recovery in Visual-Language Robotic Manipulation

CVPR 2026

Vision-Language-Action Models (VLAs) have demonstrated significant promise in generalizing to complex, long-horizon robotic manipulation tasks. However, their performance remains brittle, as they are typically trained on trajectory-monotonic, failure-free demonstrations. This reliance on "perfect" d

Cited by 0SourceScholar
2026

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

CVPR 2026

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlusions and during cross-task transfer. To address these challenges, we propose Task-aware Virtual View Exploration (TVVE)

Cited by 0SourcecodeScholar
2026

LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

CVPR 2026

Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex urban environments.While recent advances have achieved progress through large-scale memory graphs and lookahead path planning, they remain limited b

Cited by 0SourceScholar
2026

Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter

ICLR 2026poster

Personalized text-to-image generation aims to synthesize images of user-provided concepts in diverse contexts. Despite recent progress in multi-concept personalization, most are limited to object concepts and struggle to customize abstract concepts (e.g., pose, lighting). Some methods have begun ex…

Cited by 0SourcecodeScholar
2026

OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning

CVPR 2026

Streaming video reasoning requires models to operate in a setting where history grows without bound while meaningful evidence remains scarce. In such a landscape, relevant signal is like an oasis -- small, critical, and easily lost in a desert of redundancy. Enlarging memory only widens the desert;

Cited by 0SourcecodeScholar
2026

OptiMVMap: Offline Vectorized Map Construction via Optimal Multi-vehicle Perspectives

CVPR 2026

Offline vectorized maps constitute critical infrastructure for high-precision autonomous driving and mapping services. Existing approaches rely predominantly on single ego-vehicle trajectories, which fundamentally suffer from viewpoint insufficiency: while memory-based methods extend observation tim

Cited by 0SourcecodeScholar
2026

PhyScene3D: Physically Consistent 3D Interactive Tabletop Scene Generation

ICML 2026poster

Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense object hierarchies and irregular affordances. Existing methods, ranging from decoupled symbolic solvers to end-to-end regress…

Cited by 0SourceScholar
2026

SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling

CVPR 2026

Recent advances in image editing allow impressive manipulation of objects, existing methods still struggle to handle spatial movement in complex scenes, such as objects span different depth layers or are partially occluded. Most image editing methods focus solely on prior information from 2D dataset

Cited by 0SourceScholar
2026

StreamRAG: Enhancing Real-Time Video Understanding with Retrieval Augmentation

CVPR 2026

The transition of Retrieval-Augmented Generation (RAG) from offline video analysis to online, streaming scenarios presents a set of critical, unexplored challenges. These include the need for on-the-fly semantic segmentation of continuous video, the inherent tension between low-latency processing an

Cited by 0SourceScholar
2025

AdaDrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous Driving

ICCV 2025poster

Effectively integrating Large Language Models (LLMs) into autonomous driving requires a balance between leveraging high-level reasoning and maintaining real-time efficiency. Existing approaches either activate LLMs too frequently, causing excessive computational overhead, or use fixed schedules, fai…

2025

Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering

ICCV 2025poster

Embodied Question Answering (EQA) is a challenging task in embodied intelligence that requires agents to dynamically explore 3D environments, actively gather visual information, and perform multi-step reasoning to answer questions. However, current EQA approaches suffer from critical limitations in…

2025

Bridging Knowledge Gap Between Image Inpainting and Large-Area Visible Watermark Removal

AAAI 2025technical

Visible watermark removal which involves watermark cleaning and background content restoration is pivotal to evaluate the resilience of watermarks. Existing deep neural network (DNN)-based models still struggle with large-area watermarks and are overly dependent on the quality of watermark mask pred…

Cited by 0SourcePDFScholar
2025

DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh

CVPR 2025poster

Text-driven avatar generation has gained significant attention owing to its convenience. However, existing methods typically model the human body with all garments as a single 3D model, limiting its usability, such as clothing replacement, and reducing user control over the generation process. To ov…

Cited by 2SourcePDFScholar
2025

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

CVPR 2025poster

3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. However, existing methods usually rely on global scene perception from pure 3D point c…

2025

DeepShield: Fortifying Deepfake Video Detection with Local and Global Forgery Analysis

ICCV 2025poster

Recent advances in deep generative models have made it easier to manipulate face videos, raising significant concerns about their potential misuse for fraud and misinformation. Existing detectors often perform well in in-domain scenarios but fail to generalize across diverse manipulation techniques…

Cited by 0SourcePDFScholar
2025

DreamFuse: Adaptive Image Fusion with Diffusion Transformer

ICCV 2025poster

Image fusion seeks to seamlessly integrate foreground objects with background scenes, producing realistic and harmonious fused images. Unlike existing methods that directly insert objects into the background, adaptive and interactive fusion remains a challenging yet appealing task. It requires the f…

Cited by 0SourcePDFScholar
2025

DreamLayer: Simultaneous Multi-Layer Generation via Diffusion Model

ICCV 2025poster

Text-driven image generation using diffusion models has recently gained significant attention. To enable more flexible image manipulation and editing, recent research has expanded from single image generation to transparent layer generation and multi-layer compositions. However, existing approaches…

Cited by 0SourcePDFScholar
2025

Empowering Large Language Models with 3D Situation Awareness

CVPR 2025poster

Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer in 3D scenes can change, resulting in different descriptions…

Cited by 0SourcePDFScholar
2025

FakeRadar: Probing Forgery Outliers to Detect Unknown Deepfake Videos

ICCV 2025poster

In this paper, we propose FakeRadar, a novel deepfake video detection framework designed to address the challenges of cross-domain generalization in real-world scenarios. Existing detection methods typically rely on manipulation-specific cues, performing well on known forgery types but exhibiting se…

Cited by 0SourcePDFScholar
2025

Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference

ICCV 2025poster

Video Multimodal Large Language Models (Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the underlying LLMs, existing Video-MLLMs typically exhibit suboptimal performance on long video scenarios. To understand e…

2025

GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection

NeurIPS 2025poster

Fine-grained open-vocabulary object detection (FG-OVD) aims to detect novel object categories described by attribute-rich texts. While existing open-vocabulary detectors show promise at the base-category level, they underperform in fine-grained settings due to the semantic entanglement of subjects a…

Cited by 0SourceScholar
2025

GeoSplatting: Towards Geometry Guided Gaussian Splatting for Physically-based Inverse Rendering

ICCV 2025poster

Recent 3D Gaussian Splatting (3DGS) representations have demonstrated remarkable performance in novel view synthesis; further, material-lighting disentanglement on 3DGS warrants relighting capabilities and its adaptability to broader applications. While the general approach to the latter operation l…

Cited by 0SourcePDFScholar
2025

GlassWizard: Harvesting Diffusion Priors for Glass Surface Detection

ICCV 2025poster

Glass Surface Detection (GSD) is a critical task in computer vision, enabling precise interactions with transparent surfaces and enhancing both safety and object recognition accuracy. However, current research still faces challenges in both recognition performance and generalization capability. Than…

Cited by 0SourcePDFScholar
2025

Hierarchically Controlled Deformable 3D Gaussians for Talking Head Synthesis

AAAI 2025technical

Audio-driven talking head synthesis is a critical task in digital human modeling. While recent advances using diffusion models and Neural Radiance Fields (NeRF) have improved visual quality, they often require substantial computational resources, limiting practical deployment. We present a novel fra…

Cited by 1SourcePDFScholar
2025

LLM-driven Multimodal and Multi-Identity Listening Head Generation

CVPR 2025poster

Generating natural listener responses in conversational scenarios is crucial for creating engaging digital humans and avatars. Recent work has shown that large language models (LLMs) can be effectively leveraged for this task, demonstrating remarkable capabilities in generating contextually appropri…

Cited by 0SourcePDFScholar
2025

LaneDiffusion: Improving Centerline Graph Learning via Prior Injected BEV Feature Generation

ICCV 2025poster

Centerline graphs, crucial for path planning in autonomous driving, are traditionally learned using deterministic methods. However, these methods often lack spatial reasoning and struggle with occluded or invisible centerlines. Generative approaches, despite their potential, remain underexplored in…

2025

Pseudo-Label Reconstruction for Partial Multi-Label Learning

IJCAI 2025

In Partial Multi-Label Learning (PML), each instance is associated with a candidate label set containing multiple relevant labels along with other false positive labels. Currently, most PML methods directly extract instance correlation from instance features while ignoring the candidate labels, whic

Cited by 0SourcePDFScholar
2025

ReferSplat: Referring Segmentation in 3D Gaussian Splatting

ICML 2025oral

We introduce Referring 3D Gaussian Splatting Segmentation (R3DGS), a new task that aims to segment target objects in a 3D Gaussian scene based on natural language descriptions, which often contain spatial relationships or object attributes. This task requires the model to identify newly described o…

2025

Rethinking Query-based Transformer for Continual Image Segmentation

CVPR 2025poster

Class-incremental/Continual image segmentation (CIS) aims to train an image segmenter in stages, where the set of available categories differs at each stage. To leverage the built-in objectness of query-based transformers, which mitigates catastrophic forgetting of mask proposals, current methods of…

2025

Screening, Rectifying, and Re-Screening: A Unified Framework for Tuning Vision-Language Models with Noisy Labels

IJCAI 2025

Pre-trained vision-language models have shown remarkable potential for downstream tasks. However, their fine-tuning under noisy labels remains an open problem due to challenges like self-confirmation bias and the limitations of conventional small-loss criteria. In this paper, we propose a unified fr

Cited by 0SourcePDFScholar
2025

Sim-DETR: Unlock DETR for Temporal Sentence Grounding

ICCV 2025poster

Temporal sentence grounding aims to identify exact moments in a video that correspond to a given textual query, typically addressed with detection transformer (DETR) solutions. However, we find that typical strategies designed to enhance DETR do not improve, and may even degrade, its performance in…

Cited by 0SourcePDFScholar
2025

Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method

CVPR 2025poster

Existing Vision-Language Navigation (VLN) methods primarily focus on single-stage navigation, limiting their effectiveness in multi-stage and long-horizon tasks within complex and dynamic environments. To address these limitations, we propose a novel VLN task, named Long-Horizon Vision-Language Navi…

Cited by 5SourcePDFScholar
2025

VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving

ICCV 2025poster

Recent advancements in language-grounded autonomous driving have been significantly promoted by the sophisticated cognition and reasoning capabilities of large language models (LLMs). However, current LLM-based approaches encounter critical challenges: (1) Failure analysis reveals that frequent coll…

2025

VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction

CVPR 2025poster

Virtual Try-On (VTON) is a transformative technology in e-commerce and fashion design, enabling realistic digital visualization of clothing on individuals. In this work, we propose VTON 360, a novel 3D VTON method that addresses the open challenge of achieving high-fidelity VTON that supports any-vi…

Cited by 1SourcePDFScholar
2024

AlignSAM: Aligning Segment Anything Model to Open Context via Reinforcement Learning

CVPR 2024poster

Powered by massive curated training data Segment Anything Model (SAM) has demonstrated its impressive generalization capabilities in open-world scenarios with the guidance of prompts. However the vanilla SAM is class-agnostic and heavily relies on user-provided prompts to segment objects of interest…

2024

Customize your NeRF: Adaptive Source Driven 3D Scene Editing via Local-Global Iterative Training

CVPR 2024poster

In this paper we target the adaptive source driven 3D scene editing task by proposing a CustomNeRF model that unifies a text description or a reference image as the editing prompt. However obtaining desired editing results conformed with the editing prompt is nontrivial since there exist two signifi…

Cited by 14SourcePDFScholar
2024

Decoupled Pseudo-labeling for Semi-Supervised Monocular 3D Object Detection

CVPR 2024poster

We delve into pseudo-labeling for semi-supervised monocular 3D object detection (SSM3OD) and discover two primary issues: a misalignment between the prediction quality of 3D and 2D attributes and the tendency of depth supervision derived from pseudo-labels to be noisy leading to significant optimiza…

Cited by 7SourcePDFScholar
2024

FedDiv: Collaborative Noise Filtering for Federated Learning with Noisy Labels

AAAI 2024technical

Federated Learning with Noisy Labels (F-LNL) aims at seeking an optimal server model via collaborative distributed learning by aggregating multiple client models trained with local noisy or clean samples. On the basis of a federated learning framework, recent advances primarily adopt label noise fil…

2024

Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection

CVPR 2024poster

Open vocabulary object detection (OVD) aims at seeking an optimal object detector capable of recognizing objects from both base and novel categories. Recent advances leverage knowledge distillation to transfer insightful knowledge from pre-trained large-scale vision-language models to the task of ob…

Cited by 16SourcePDFScholar
2024

MMAPS: End-to-End Multi-Grained Multi-Modal Attribute-Aware Product Summarization

COLING 2024main

Given the long textual product information and the product image, Multi-modal Product Summarization (MPS) aims to increase customers’ desire to purchase by highlighting product characteristics with a short textual summary. Existing MPS methods can produce promising results. Nevertheless, they still…

2024

NeRF-HuGS: Improved Neural Radiance Fields in Non-static Scenes Using Heuristics-Guided Segmentation

CVPR 2024poster

Neural Radiance Field (NeRF) has been widely recognized for its excellence in novel view synthesis and 3D scene reconstruction. However their effectiveness is inherently tied to the assumption of static scenes rendering them susceptible to undesirable artifacts when confronted with transient distrac…

2024

OVER-NAV: Elevating Iterative Vision-and-Language Navigation with Open-Vocabulary Detection and StructurEd Representation

CVPR 2024poster

Recent advances in Iterative Vision-and-Language Navigation(IVLN) introduce a more meaningful and practical paradigm of VLN by maintaining the agent's memory across tours of scenes. Although the long-term memory aligns better with the persistent nature of the VLN task it poses more challenges on how…

Cited by 8SourcePDFScholar
2024

Open-Vocabulary Segmentation with Semantic-Assisted Calibration

CVPR 2024poster

This paper studies open-vocabulary segmentation (OVS) through calibrating in-vocabulary and domain-biased embedding space with generalized contextual prior of CLIP. As the core of open-vocabulary understanding alignment of visual content with the semantics of unbounded text has become the bottleneck…

Cited by 32SourcePDFScholar
2024

Removing Interference and Recovering Content Imaginatively for Visible Watermark Removal

AAAI 2024technical

Visible watermarks, while instrumental in protecting image copyrights, frequently distort the underlying content, complicating tasks like scene interpretation and image editing. Visible watermark removal aims to eliminate the interference of watermarks and restore the background content. However, ex…

Cited by 4SourcePDFScholar
2024

UniCell: Universal Cell Nucleus Classification via Prompt Learning

AAAI 2024technical

The recognition of multi-class cell nuclei can significantly facilitate the process of histopathological diagnosis. Numerous pathological datasets are currently available, but their annotations are inconsistent. Most existing methods require individual training on each dataset to deduce the relevant…

2024

UniFL: Improve Latent Diffusion Model via Unified Feedback Learning

NeurIPS 2024poster

Latent diffusion models (LDM) have revolutionized text-to-image generation, leading to the proliferation of various advanced models and diverse downstream applications. However, despite these significant advancements, current diffusion models still suffer from several limitations, including inferior…

Cited by 1SourcePDFScholar
2024

Variance-Insensitive and Target-Preserving Mask Refinement for Interactive Image Segmentation

AAAI 2024technical

Point-based interactive image segmentation can ease the burden of mask annotation in applications such as semantic segmentation and image editing. However, fully extracting the target mask with limited user inputs remains challenging. We introduce a novel method, Variance-Insensitive and Target-Pres…

Cited by 3SourcePDFScholar
2024

VersVideo: Leveraging Enhanced Temporal Diffusion Models for Versatile Video Generation

ICLR 2024poster

Creating stable, controllable videos is a complex task due to the need for significant variation in temporal dynamics and cross-frame temporal consistency. To address this, we enhance the spatial-temporal capability and introduce a versatile video generation model, VersVideo, which leverages textual…

2024

WhodunitBench: Evaluating Large Multimodal Agents via Murder Mystery Games

NeurIPS 2024spotlight

Recently, large language models (LLMs) have achieved superior performance, empowering the development of large multimodal agents (LMAs). An LMA is anticipated to execute practical tasks requires various capabilities including multimodal perception, interaction, reasoning, and decision making. Howeve…

Cited by 1SourcePDFScholar
2023

Adapting Object Size Variance and Class Imbalance for Semi-supervised Object Detection

AAAI 2023technical

Semi-supervised object detection (SSOD) attracts extensive research interest due to its great significance in reducing the data annotation effort. Collecting high-quality and category-balanced pseudo labels for unlabeled images is critical to addressing the SSOD problem. However, most of the existin…

Cited by 13SourcePDFScholar
2023

Advancing Visual Grounding With Scene Knowledge: Benchmark and Method

CVPR 2023poster

Visual grounding (VG) aims to establish fine-grained alignment between vision and language. Ideally, it can be a testbed for vision-and-language models to evaluate their understanding of the images and texts and their reasoning abilities over their joint space. However, most existing VG datasets are…

2023

Affine-Consistent Transformer for Multi-Class Cell Nuclei Detection

ICCV 2023poster

Multi-class cell nuclei detection is a fundamental prerequisite in the diagnosis of histopathology. It is critical to efficiently locate and identify cells with diverse morphology and distributions in digital pathological images. Most existing methods take complex intermediate representations as lea…

Cited by 15PDFcodeScholar
2023

Being Comes From Not-Being: Open-Vocabulary Text-to-Motion Generation With Wordless Training

CVPR 2023highlight

Text-to-motion generation is an emerging and challenging problem, which aims to synthesize motion with the same semantics as the input text. However, due to the lack of diverse labeled training data, most approaches either limit to specific types of text annotations or require online optimizations t…

2023

Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image Segmentation

ICCV 2023poster

Parameter efficient tuning (PET) has received considerable attention owing to its applicability to reduce the number of parameters that need to be updated while maintaining competitive performance and providing better hardware resource savings. Although substantial progress has been made, most exist…

Cited by 73PDFcodeScholar
2023

De-biased Teacher: Rethinking IoU Matching for Semi-supervised Object Detection

AAAI 2023technical

Most of the recent research in semi-supervised object detection follows the pseudo-labeling paradigm evolved from the semi-supervised image classification task. However, the training paradigm of the two-stage object detector inevitably makes the pseudo-label learning process for unlabeled images ful…

2023

DenseLight: Efficient Control for Large-scale Traffic Signals with Dense Feedback

IJCAI 2023poster

Traffic Signal Control (TSC) aims to reduce the average travel time of vehicles in a road network, which in turn enhances fuel utilization efficiency, air quality, and road safety, benefiting society as a whole. Due to the complexity of long-horizon control and coordination, most prior TSC methods l…

2023

Divide and Adapt: Active Domain Adaptation via Customized Learning

CVPR 2023highlight

Active domain adaptation (ADA) aims to improve the model adaptation performance by incorporating the active learning (AL) techniques to label a maximally-informative subset of target samples. Conventional AL methods do not consider the existence of domain shift, and hence, fail to identify the truly…

2023

Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection

ICCV 2023poster

Current semi-supervised object detection (SSOD) algorithms typically assume class balanced datasets (PASCAL VOC etc.) or slightly class imbalanced datasets (MSCOCO, etc). This assumption can be easily violated since real world datasets can be extremely class imbalanced in nature, thus making the per…

Cited by 13PDFcodeScholar
2023

Identity-Preserving Talking Face Generation With Landmark and Appearance Priors

CVPR 2023poster

Generating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for training or fine-tuning. Existing person-generic methods have difficulty in generating realistic and lip-synced videos whi…

2023

Long-term Wind Power Forecasting with Hierarchical Spatial-Temporal Transformer

IJCAI 2023poster

Wind power is attracting increasing attention around the world due to its renewable, pollution-free, and other advantages. However, safely and stably integrating the high permeability intermittent power energy into electric power systems remains challenging. Accurate wind power forecasting (WPF) can…

2023

Parametric Implicit Face Representation for Audio-Driven Facial Reenactment

CVPR 2023poster

Audio-driven facial reenactment is a crucial technique that has a range of applications in film-making, virtual avatars and video conferences. Existing works either employ explicit intermediate face representations (e.g., 2D facial landmarks or 3D face models) or implicit ones (e.g., Neural Radiance…

Cited by 20SourcePDFScholar
2023

RankMatch: Fostering Confidence and Consistency in Learning with Noisy Labels

ICCV 2023poster

Learning with noisy labels (LNL) is one of the most important and challenging problems in weakly-supervised learning. Recent advances adopt the sample selection strategy to mitigate the interference of noisy labels and use small-loss criteria to select clean samples. However, the one-dimensional los…

Cited by 14PDFScholar
2023

SCoDA: Domain Adaptive Shape Completion for Real Scans

CVPR 2023poster

3D shape completion from point clouds is a challenging task, especially from scans of real-world objects. Considering the paucity of 3D shape ground truths for real scans, existing works mainly focus on benchmarking this task on synthetic data, e.g. 3D computer-aided design models. However, the doma…

2023

Semi-DETR: Semi-Supervised Object Detection With Detection Transformers

CVPR 2023poster

We analyze the DETR-based framework on semi-supervised object detection (SSOD) and observe that (1) the one-to-one assignment strategy generates incorrect matching when the pseudo ground-truth bounding box is inaccurate, leading to training inefficiency; (2) DETR-based detectors lack deterministic c…

Cited by 61SourcePDFScholar
2023

SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-training

ICCV 2023poster

Skeleton sequence representation learning has shown great advantages for action recognition due to its promising ability to model human joints and topology. However, the current methods usually require sufficient labeled data for training computationally expensive models. Moreover, these methods ign…

Cited by 59PDFcodeScholar
2023

Towards Real-World Burst Image Super-Resolution: Benchmark and Method

ICCV 2023poster

Despite substantial advances, single-image super-resolution (SISR) is always in a dilemma to reconstruct high-quality images with limited information from one input image, especially in realistic scenarios. In this paper, we establish a large-scale real-world burst super-resolution dataset, i.e., Re…

Cited by 16PDFcodeScholar
2023

Towards Unifying Medical Vision-and-Language Pre-Training via Soft Prompts

ICCV 2023poster

Medical vision-and-language pre-training (Med-VLP) has shown promising improvements on many downstream medical tasks owing to its applicability to extracting generic representations from medical images and texts. Practically, there exist two typical types, i.e., the fusion-encoder type and the dual-…

Cited by 39PDFcodeScholar
2022

A Causal Debiasing Framework for Unsupervised Salient Object Detection

AAAI 2022technical

Unsupervised Salient Object Detection (USOD) is a promising yet challenging task that aims to learn a salient object detection model without any ground-truth labels. Self-supervised learning based methods have achieved remarkable success recently and have become the dominant approach in USOD. Howeve…

Cited by 28SourcePDFScholar
2022

A Causal Inference Look at Unsupervised Video Anomaly Detection

AAAI 2022technical

Unsupervised video anomaly detection, a task that requires no labeled normal/abnormal training data in any form, is challenging yet of great importance to both industrial applications and academic research. Existing methods typically follow an iterative pseudo label generation process. However, they…

Cited by 48SourcePDFScholar
2022

Centrality and Consistency: Two-Stage Clean Samples Identification for Learning with Instance-Dependent Noisy Labels

ECCV 2022poster

"Deep models trained with noisy labels are prone to over-fitting and struggle in generalization. Most existing solutions are based on an ideal assumption that the label noise is class-conditional, i.e., instances of the same class share the same noise model, and are independent of features. While in…

2022

Divide and Contrast: Source-free Domain Adaptation via Adaptive Contrastive Learning

NeurIPS 2022accept

We investigate a practical domain adaptation task, called source-free domain adaptation (SFUDA), where the source pretrained model is adapted to the target domain without access to the source data. Existing techniques mainly leverage self-supervised pseudo-labeling to achieve class-wise global align…

2022

Double-Check Soft Teacher for Semi-Supervised Object Detection

IJCAI 2022poster

In the semi-supervised object detection task, due to the scarcity of labeled data and the diversity and complexity of objects to be detected, the quality of pseudo-labels generated by existing methods for unlabeled data is relatively low, which severely restricts the performance of semi-supervised o…

2022

Dual Adversarial Adaptation for Cross-Device Real-World Image Super-Resolution

CVPR 2022oral

Due to the sophisticated imaging process, an identical scene captured by different cameras could exhibit distinct imaging patterns, introducing distinct proficiency among the super-resolution (SR) models trained on images from different devices. In this paper, we investigate a novel and practical ta…

Cited by 21PDFcodeScholar
2022

Multi-level Consistency Learning for Semi-supervised Domain Adaptation

IJCAI 2022poster

Semi-supervised domain adaptation (SSDA) aims to apply knowledge learned from a fully labeled source domain to a scarcely labeled target domain. In this paper, we propose a Multi-level Consistency Learning (MCL) framework for SSDA. Specifically, our MCL regularizes the consistency of different views…

2022

Neighborhood Collective Estimation for Noisy Label Identification and Correction

ECCV 2022poster

"Learning with noisy labels (LNL) aims at designing strategies to improve model performance and generalization by mitigating the effects of model overfitting to noisy labels. The key success of LNL lies in identifying as many clean samples as possible from massive noisy data, while rectifying the wr…

2022

Unsupervised Domain Adaptive Salient Object Detection through Uncertainty-Aware Pseudo-Label Learning

AAAI 2022technical

Recent advances in deep learning significantly boost the performance of salient object detection (SOD) at the expense of labeling larger-scale per-pixel annotations. To relieve the burden of labor-intensive labeling, deep unsupervised SOD methods have been proposed to exploit noisy labels generated…

2022

X-Trans2Cap: Cross-Modal Knowledge Transfer Using Transformer for 3D Dense Captioning

CVPR 2022poster

3D dense captioning aims to describe individual objects by natural language in 3D scenes, where 3D scenes are usually represented as RGB-D scans or point clouds. However, only exploiting single modal information, e.g., point cloud, previous approaches fail to produce faithful descriptions. Though ag…

Cited by 96PDFcodeScholar
2021

Bottom-Up Shift and Reasoning for Referring Image Segmentation

CVPR 2021poster

Referring image segmentation aims to segment the referent that is the corresponding object or stuff referred by a natural language expression in an image. Its main challenge lies in how to effectively and efficiently differentiate between the referent and other objects of the same category as the re…

Cited by 101PDFcodeScholar
2021

Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor Segmentation

CVPR 2021poster

Language-queried video actor segmentation aims to predict the pixel-level mask of the actor which performs the actions described by a natural language query in the target frames. Existing methods adopt 3D CNNs over the video clip as a general encoder to extract a mixed spatio-temporal feature for th…

Cited by 58PDFScholar
2021

Cross-Domain Adaptive Clustering for Semi-Supervised Domain Adaptation

CVPR 2021poster

In semi-supervised domain adaptation, a few labeled samples per class in the target domain guide features of the remaining target samples to aggregate around them. However, the trained model cannot produce a highly discriminative feature representation for the target domain because the training data…

Cited by 157PDFcodeScholar
2021

Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting

CVPR 2021poster

Crowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited information of RGB images and cannot well discover potential pedestrians in unconstrained scenarios. In this work, we f…

Cited by 165PDFcodeScholar
2021

LapsCore: Language-Guided Person Search via Color Reasoning

ICCV 2021poster

The key point of language-guided person search is to construct the cross-modal association between visual and textual input. Existing methods focus on designing multimodal attention mechanisms and novel cross-modal loss functions to learn such association implicitly. We propose a representation lear…

Cited by 89PDFScholar
2021

Towards Interpretable Deep Networks for Monocular Depth Estimation

ICCV 2021poster

Deep networks for Monocular Depth Estimation (MDE) have achieved promising performance recently and it is of great importance to further understand the interpretability of these networks. Existing methods attempt to provide post-hoc explanations by investigating visual cues, which may not explore th…

Cited by 16PDFcodeScholar
2021

Trash To Treasure: Harvesting OOD Data With Cross-Modal Matching for Open-Set Semi-Supervised Learning

ICCV 2021poster

Open-set semi-supervised learning (open-set SSL) investigates a challenging but practical scenario where out-of-distribution (OOD) samples are contained in the unlabeled data. While the mainstream technique seeks to completely filter out the OOD samples for semi-supervised learning (SSL), we propose…

Cited by 76PDFScholar
2021

Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

IJCAI 2021poster

In this paper, we introduce a novel task, referred to as Weakly-Supervised Spatio-Temporal Anomaly Detection (WSSTAD) in surveillance video. Specifically, given an untrimmed video, WSSTAD aims to localize a spatio-temporal tube (i.e., a sequence of bounding boxes at consecutive times) that encloses…

Cited by 75SourcePDFScholar
2020

A Real-Time Cross-Modality Correlation Filtering Method for Referring Expression Comprehension

CVPR 2020poster

Referring expression comprehension aims to localize the object instance described by a natural language expression. Current referring expression methods have achieved good performance. However, none of them is able to achieve real-time inference without accuracy drop. The reason for the relatively s…

Cited by 240PDFScholar
2020

Collaborative Training between Region Proposal Localization and Classification for Domain Adaptive Object Detection

ECCV 2020poster

Object detectors are usually trained with large amount of labeled data, which is expensive and labor-intensive. Pre-trained detectors applied to unlabeled dataset always suffer from the difference of dataset distribution, also called domain shift. Domain adaptation for object detection tries to adap…

2020

Linguistic Structure Guided Context Modeling for Referring Image Segmentation

ECCV 2020poster

Referring image segmentation aims to predict the foreground mask of the object referred by a natural language sentence. Multimodal context of the sentence is crucial to distinguish the referent from the background. Existing methods either insufficiently or redundantly model the multimodal context. To…

2020

Peeking into occluded joints: A novel framework for crowd pose estimation

ECCV 2020poster

Although occlusion widely exists in nature and remains a fundamental challenge for pose estimation, existing heatmap-based approaches suffer serious degradation on occlusions. Their intrinsic problem is that they directly localize the joints based on visual information; however, the invisible joints…

2020

Referring Image Segmentation via Cross-Modal Progressive Comprehension

CVPR 2020poster

Referring image segmentation aims at segmenting the foreground masks of the entities that can well match the description given in the natural language expression. Previous approaches tackle this problem using implicit feature interaction and fusion between visual and linguistic modalities, but usual…

Cited by 225PDFcodeScholar
2019

ClusterNet: Deep Hierarchical Cluster Network With Rigorously Rotation-Invariant Representation for Point Cloud Analysis

CVPR 2019poster

Current neural networks for 3D object recognition are vulnerable to 3D rotation. Existing works mostly rely on massive amounts of rotation-augmented data to alleviate the problem, which lacks solid guarantee of the 3D rotation invariance. In this paper, we address the issue by introducing a novel po…

Cited by 217PDFScholar
2019

Crowd Counting With Deep Structured Scale Integration Network

ICCV 2019poster

Automatic estimation of the number of people in unconstrained crowded scenes is a challenging task and one major difficulty stems from the huge scale variation of people. In this paper, we propose a novel Deep Structured Scale Integration Network (DSSINet) for crowd counting, which addresses the sca…

Cited by 305PDFScholar
2019

Fashion Retrieval via Graph Reasoning Networks on a Similarity Pyramid

ICCV 2019oral

Matching clothing images from customers and online shopping stores has rich applications in E-commerce. Existing algorithms encoded an image as a global feature vector and performed retrieval with the global representation. However, discriminative local information on clothes are submerged in this g…

Cited by 118PDFScholar
2019

Larger Norm More Transferable: An Adaptive Feature Norm Approach for Unsupervised Domain Adaptation

ICCV 2019oral

Domain adaptation enables the learner to safely generalize into novel environments by mitigating domain shifts across distributions. Previous works may not effectively uncover the underlying reasons that would lead to the drastic model degradation on the target task. In this paper, we empirically re…

Cited by 656PDFcodeScholar
2019

Multivariate-Information Adversarial Ensemble for Scalable Joint Distribution Matching

ICML 2019oral

A broad range of cross-$m$-domain generation researches boil down to matching a joint distribution by deep generative models (DGMs). Hitherto algorithms excel in pairwise domains while as $m$ increases, remain struggling to scale themselves to fit a joint distribution. In this paper, we propose a dom…

2019

Semi-Supervised Skin Detection by Network With Mutual Guidance

ICCV 2019poster

We present a new data-driven method for robust skin detection from a single human portrait image. Unlike previous methods, we incorporate human body as a weak semantic guidance into this task, considering acquiring large-scale of human labeled skin data is commonly expensive and time-consuming. To b…

Cited by 32PDFScholar
2019

Semi-Supervised Video Salient Object Detection Using Pseudo-Labels

ICCV 2019poster

Deep learning-based video salient object detection has recently achieved great success with its performance significantly outperforming any other unsupervised methods. However, existing data-driven approaches heavily rely on a large quantity of pixel-wise annotated video frames to deliver such promi…

Cited by 154PDFScholar
2018

Flow Guided Recurrent Neural Encoder for Video Salient Object Detection

CVPR 2018poster

Image saliency detection has recently witnessed significant progress due to deep convolutional neural networks. However, extending state-of-the-art saliency detectors from image to video is challenging. The performance of salient object detection suffers from object or camera motion and the dramatic…

Cited by 205SourcePDFScholar
2018

Interpretable Video Captioning via Trajectory Structured Localization

CVPR 2018poster

Automatically describing open-domain videos with natural language are attracting increasing interest in the field of artificial intelligence. Most existing methods simply borrow ideas from image captioning and obtain a compact video representation from an ensemble of global image feature before feed…

Cited by 70SourcePDFScholar
2018

Visual Question Reasoning on General Dependency Tree

CVPR 2018poster

The collaborative reasoning for understanding each image-question pair is very critical but under-explored for an interpretable Visual Question Answering (VQA) system. Although very recent works also tried the explicit compositional processes to assemble multiple sub-tasks embedded in the questions…

Cited by 42SourcePDFScholar
2017

Attention-Aware Face Hallucination via Deep Reinforcement Learning

CVPR 2017poster

Face hallucination is a domain-specific super-resolution problem with the goal to generate high-resolution (HR) faces from low-resolution (LR) input images. In contrast to existing methods that often learn a single patch-to-patch mapping from LR to HR images and are regardless of the contextual inte…

Cited by 249PDFScholar
2017

Multi-Label Image Recognition by Recurrently Discovering Attentional Regions

ICCV 2017poster

This paper proposes a novel deep architecture to address multi-label image recognition, a fundamental and practical task towards general visual understanding. Current solutions for this task usually rely on an extra step of extracting hypothesis regions (i.e., region proposals), resulting in redunda…

Cited by 394PDFScholar