← Search

Hang Xu

154 accepted papers

2026

2D-CrossScan Mamba: Enhancing State Space Models with Spatially Consistent Multi-Path 2D Information Propagation

AAAI 2026technical

Despite recent progress in adapting State Space Models such as Mamba to vision tasks, their intrinsic 1D scanning mechanism imposes limitations when applied to inherently 2D-structured data like images. Existing adaptations, including VMamba and 2DMamba, either suffer from inconsistency between scan

Cited by 0SourcePDFScholar
2026

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

CVPR 2026

Recent advances in Visual-Language-Action (VLA) models have shown promising potential for robotic manipulation tasks.However, real-world robotic tasks often involve long-horizon, multi-step problem-solving and require generalization for continual skill acquisition, extending beyond single actions or

Cited by 0SourceScholar
2026

CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning

CVPR 2026

Graphical User Interface (GUI) Agents, benefiting from recent advances in multimodal large language models (MLLM), have achieved significant development. However, due to the frequent updates of GUI applications, adapting to new tasks without forgetting old tasks in GUI continual learning remains an

Cited by 0SourceScholar
2026

Condition-Number Adaptive-Weight PINN (A-PINN): A High-Fidelity and Real-Time Forward Kinematics Solver for Stewart Platforms

RA-L 2026

Parallel kinematic mechanisms (PKMs) are widely adopted in precision and heavy-load applications, yet obtaining the requisite forward kinematics (FK) for closed-loop control remains a challenge. FK in PKMs typically lack a unique closed-form solution, necessitating iterative numerical solvers that a

Cited by 0SourceScholar
2026

JUMP-Hand: Learning Joint-wise Uncertainty to Gate Mixture of View Experts for Multi-View 3D Hand Reconstruction

CVPR 2026

We propose JUMP-Hand, a novel multi-view 3D hand reconstruction method that explicitly models probabilistic joint-wise uncertainty as a gating mechanism for multi-view fusion. Existing approaches usually rely on naive pooling or implicit attention, overlooking that each hand joint exhibits varying v

Cited by 0SourcecodeScholar
2026

MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image Generation

CVPR 2026

Reinforcement learning (RL) has demonstrated significant potential for post-training language models and autoregressive visual generative models, but adapting RL to masked generative models (MGMs) remains challenging. The core factor is that policy optimization requires the probability likelihood of

Cited by 0SourcecodeScholar
2026

MoEG-HOI: Mixture of Expert Groups for One-Stage Hand-Object Interaction Motion Generation with Hand-Finger-Joint Semantic Guidance

AAAI 2026technical

In this paper, MoEG-HOI is proposed as a novel method for the challenging 3D hand-object interaction (HOI) motion generation task, by introducing Mixture-of-Experts (MoE) to this field for the first time. Almost all the mainstream approaches in HOI motion generation leverage diffusion model as its s

Cited by 0SourcePDFScholar
2026

Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving

CVPR 2026

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-language models are weak at spatial grounding and understanding, and VLA systems bui

Cited by 0SourceScholar
2026

SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation

ICLR 2026poster

In this paper, we introduce SemHiTok, a unified image Tokenizer via Semantic-Guided Hierarchical codebook (SGHC) that provides consistent discrete representations for multimodal understanding and generation. Recently, unified image tokenizers have sparked exploration within the research community, w…

Cited by 0SourceScholar
2026

Shared Autonomy Assisted by Impedance-Driven Anisotropic Guidance Field

RA-L 2026

Shared autonomy (SA) enables robots to infer human intent and assist in its achievement. While most research focuses on improving intent inference, it overlooks whether humans can understand the robot's intent in return. Without such mutual understanding, collaboration becomes less effective, degrad

Cited by 0SourceScholar
2026

Swimming under Constraints: A Safe Reinforcement Learning Framework for Quadrupedal Bio-Inspired Propulsion

ICRA 2026poster

Bio-inspired aquatic propulsion offers high thrust and maneuverability but is prone to destabilizing forces such as lift fluctuations, which are further amplified by six-degree-of-freedom (6-DoF) fluid coupling. We formulate quadrupedal swimming as a constrained optimization problem that maximizes f…

2026

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning

ICML 2026poster

Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However, most existing integration strategies remain passive: geometry is exposed as a global stream and fused in an indiscriminate manner, which often induces…

Cited by 0SourceScholar
2026

Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots

ICRA 2026poster

In recent years, Multimodal Large Language Models (MLLMs) have demonstrated the ability to serve as high-level planners, enabling robots to follow complex human instructions. However, their effectiveness, especially in long-horizon tasks involving dual-arm humanoid robots, remains limited. This limi…

2026

UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity

ICLR 2026poster

Recently, considerable progress has been made in all-in-one image restoration. Generally, existing methods can be degradation-agnostic or degradation-aware. However, the former are limited in leveraging degradation estimation-based priors, and the latter suffer from the inevitable error in degradati…

Cited by 0SourcecodeScholar
2026

UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding

ICLR 2026poster

Despite the impressive progress on understanding and generating images shown by the recent unified architectures, the integration of 3D tasks remains challenging and largely unexplored. In this paper, we introduce UniUGG, the first unified understanding and generation framework for 3D modalities. Ou…

Cited by 0SourcecodeScholar
2025

4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration

NeurIPS 2025poster

Leveraging diverse robotic data for pretraining remains a critical challenge. Existing methods typically model the dataset’s action distribution using simple observations as inputs. However, these inputs are often incomplete, resulting in a dispersed conditional action distribution—an issue we refer…

Cited by 0SourceScholar
2025

ACE: Anti-Editing Concept Erasure in Text-to-Image Models

CVPR 2025poster

Recent advance in text-to-image diffusion models have significantly facilitated the generation of high-quality images, but also raising concerns about the illegal creation of harmful content, such as copyrighted images. Existing concept erasure methods achieve superior results in preventing the prod…

2025

Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising

ICASSP 2025accepted

Recent advances in diffusion models have greatly improved text-driven video generation. However, training models for long video generation demands significant computational power and extensive data, leading most video diffusion models to be limited to a small number of frames. Existing training-free…

Cited by 0SourceScholar
2025

DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance

ICASSP 2025accepted

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods frequently integrate semantic information from images or simply concatenate images, which often leads to low fidelity and flickering in the generated videos. T…

Cited by 25SourceScholar
2025

EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation

CVPR 2025poster

Handling complex or nonlinear motion patterns has long posed challenges for video frame interpolation. Although recent advances in diffusion-based methods offer improvements over traditional optical flow-based approaches, they still struggle to generate sharp, temporally consistent frames in scenari…

Cited by 3SourcePDFScholar
2025

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

CVPR 2025poster

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language Models to perceive and generate images, texts, and speeches end-to-end with publicly available data remains challenging…

Cited by 23SourcePDFScholar
2025

EasyControl: Adding Control to Video Diffusion for Controllable Video Generation and Interpolation

ICASSP 2025accepted

The diffusion model is widely leveraged for either controllable video generation or video interpolation. As each field has its task-specific problems, it is difficult to merely develop a single model for completing both tasks simultaneously. Moreover, most existing works only support image condition…

Cited by 0SourceScholar
2025

FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors

ICCV 2025poster

Interactive image editing allows users to modify images through visual interaction operations such as drawing, clicking, and dragging. Existing methods construct such supervision signals from videos, as they capture how objects change with various physical interactions. However, these models are usu…

2025

FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise Alignment

ICCV 2025poster

Domain Adaptation (DA) for dense prediction tasks is an important topic, which enhances the dense prediction model's performance when tested on its unseen domain. Recently, with the development of Diffusion-based Dense Prediction (DDP) models, the exploration of DA designs tailored to this framework…

2025

FreqPrior: Improving Video Diffusion Models with Frequency Filtering Gaussian Noise

ICLR 2025poster

Text-driven video generation has advanced significantly due to developments in diffusion models. Beyond the training and sampling phases, recent studies have investigated noise priors of diffusion models, as improved noise priors yield better generation results. One recent approach employs the Fouri…

Cited by 0SourcePDFScholar
2025

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

NeurIPS 2025poster

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate 3D representations into models to improve spatial understanding, we aim to unlo…

Cited by 0SourceScholar
2025

G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

ICLR 2025poster

Large language models (LLMs) have shown remarkable proficiency in human-level reasoning and generation capabilities, which encourages extensive research on their application in mathematical problem solving. However, current work has been largely focused on text-based mathematical problems, with limi…

2025

Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data

NAACL 2025long

Contrastive Language-Image Pre-training (CLIP) has become the standard for cross- modal image-text representation learning. Improving CLIP typically requires additional data and retraining with new loss functions, but these demands raise resource and time costs, limiting practical use. In this work,…

2025

HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models

CVPR 2025poster

High-resolution image inputs allow Large Vision-Language Models (LVLMs) to capture finer visual details, improving comprehension. However, the increased training and computational costs associated with such inputs pose significant challenges. A common approach to mitigate these costs involves slicin…

Cited by 8SourcePDFScholar
2025

ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

ICCV 2025poster

In this paper, we introduce ILLUME, a unified multimodal large language model (MLLM) that seamlessly integrates multimodal understanding and generation capabilities within a single large language model through a unified next-token prediction formulation.To address the large dataset size typically re…

Cited by 0SourcePDFScholar
2025

INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning

NeurIPS 2025poster

Large Multimodal Models (LMMs) have made significant breakthroughs with the advancement of instruction tuning. However, while existing models can understand images and videos at a holistic level, they still struggle with instance-level understanding that requires a more fine-grained comprehension an…

Cited by 0SourceScholar
2025

Online Competitive Information Gathering for Partially Observable Trajectory Games

RSS 2025poster

To cooperate or compete rationally in continuous, partially observable multi-agent spaces, game theoretic agents must make plans that optimally gather information about their opponents. These problems are modeled by partially observable stochastic games (POSGs), but planning in fully continuous POSG…

Cited by 0PDFScholar
2025

PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor Space

NeurIPS 2025poster

3D human pose lifting from a single RGB image is a challenging task in 3D vision. Existing methods typically establish a direct joint-to-joint mapping from 2D to 3D poses based on 2D features. This formulation suffers from two fundamental limitations: inevitable error propagation from input predicte…

Cited by 0SourceScholar
2025

SeePhys: Does Seeing Help Thinking? – Benchmarking Vision-Based Physics Reasoning

NeurIPS 2025poster

We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline, incorporating 21 categories of highly heterogeneous diagrams. In cont…

Cited by 0SourcecodeScholar
2025

Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization

NeurIPS 2025poster

Unified vision-language models have made significant progress in multimodal understanding and generation, yet they largely fall short in producing multimodal interleaved outputs, which is a crucial capability for tasks like visual storytelling and step-by-step visual reasoning. In this work, we prop…

Cited by 0SourceScholar
2025

UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting

ICLR 2025poster

Recent advancements in multi-modal 3D pre-training methods have shown promising efficacy in learning joint representations of text, images, and point clouds. However, adopting point clouds as 3D representation fails to fully capture the intricacies of the 3D world and exhibits a noticeable gap betwe…

Cited by 0SourcePDFScholar
2025

VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning

ICCV 2025poster

In recent years, video question answering based on multimodal large language models (MLLM) has garnered considerable attention, due to the benefits from the substantial advancements in LLMs. However, these models have a notable deficiency in the domains of video temporal grounding and reasoning, pos…

2024

"Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation"

ECCV 2024poster

"Multimodal large language models (MLLMs) have shown impressive reasoning abilities. However, they are also more vulnerable to jailbreak attacks than their LLM predecessors. Although still capable of detecting the unsafe responses, we observe that safety mechanisms of the pre-aligned LLMs in MLLMs c…

Cited by 48SourcePDFScholar
2024

Any-Size-Diffusion: Toward Efficient Text-Driven Synthesis for Any-Size HD Images

AAAI 2024technical

Stable diffusion, a generative model used in text-to-image synthesis, frequently encounters resolution-induced composition problems when generating images of varying sizes. This issue primarily stems from the model being trained on pairs of single-scale images and their corresponding text descriptio…

2024

BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models

CVPR 2024poster

Diffusion models have made tremendous progress in text-driven image and video generation. Now text-to-image foundation models are widely applied to various downstream image synthesis tasks such as controllable image generation and image editing while downstream video synthesis tasks are less explore…

2024

CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation

ACL 2024findings

Understanding and following natural language instructions while navigating through complex, real-world environments poses a significant challenge for general-purpose robots. These environments often include obstacles and pedestrians, making it essential for autonomous agents to possess the capabilit…

2024

DetCLIPv3: Towards Versatile Generative Open-vocabulary Object Detection

CVPR 2024poster

Existing open-vocabulary object detectors typically require a predefined set of categories from users significantly confining their application scenarios. In this paper we introduce DetCLIPv3 a high-performing detector that excels not only at both open-vocabulary object detection but also generating…

Cited by 12SourcePDFScholar
2024

DreamControl: Control-Based Text-to-3D Generation with 3D Self-Prior

CVPR 2024poster

3D generation has raised great attention in recent years. With the success of text-to-image diffusion models the 2D-lifting technique becomes a promising route to controllable 3D generation. However these methods tend to present inconsistent geometry which is also known as the Janus problem. We obse…

2024

Dynamic Discounted Counterfactual Regret Minimization

ICLR 2024spotlight

Counterfactual regret minimization (CFR) is a family of iterative algorithms showing promising results in solving imperfect-information games. Recent novel CFR variants (e.g., CFR+, DCFR) have significantly improved the convergence rate of the vanilla CFR. The key to these CFR variants’ performance…

2024

Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis

ICLR 2024poster

The rapid development of large language models (LLMs) has not only provided numerous opportunities but also presented significant challenges. This becomes particularly evident when LLMs inadvertently generate harmful or toxic content, either unintentionally or because of intentional inducement. Exis…

Cited by 36SourcePDFScholar
2024

Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models

CVPR 2024poster

The rise of multimodal large language models (MLLMs) has spurred interest in language-based driving tasks. However existing research typically focuses on limited tasks and often omits key multi-view and temporal information which is crucial for robust autonomous driving. To bridge these gaps we intr…

2024

HumanRefiner: Benchmarking Abnormal Human Generation and Refining with Coarse-to-fine Pose-Reversible Guidance

ECCV 2024poster

"Text-to-image diffusion models have significantly advanced in conditional image generation. However, these models usually struggle with accurately rendering images featuring humans, resulting in distorted limbs and other anomalies. This issue primarily stems from the insufficient recognition and ev…

2024

Implicit Concept Removal of Diffusion Models

ECCV 2024poster

"Text-to-image (T2I) diffusion models often inadvertently generate unwanted concepts such as watermarks and unsafe images. These concepts, termed “implicit concepts”, can be unintentionally learned during training and then be generated uncontrollably during inference. Existing removal methods still…

2024

Ins-DetCLIP: Aligning Detection Model to Follow Human-Language Instruction

ICLR 2024poster

This paper introduces Instruction-oriented Object Detection (IOD), a new task that enhances human-computer interaction by enabling object detectors to understand user instructions and locate relevant objects. Unlike traditional open-vocabulary object detection tasks that rely on users providing a li…

Cited by 3SourcePDFScholar
2024

JointDreamer: Ensuring Geometry Consistency and Text Congruence in Text-to-3D Generation via Joint Score Distillation

ECCV 2024poster

"Score Distillation Sampling (SDS) by well-trained 2D diffusion models has shown great promise in text-to-3D generation. However, this paradigm distills view-agnostic 2D image distributions into the rendering distribution of 3D representation for each view independently, overlooking the coherence ac…

2024

LaneGraph2Seq: Lane Topology Extraction with Language Model via Vertex-Edge Encoding and Connectivity Enhancement

AAAI 2024technical

Understanding road structures is crucial for autonomous driving. Intricate road structures are often depicted using lane graphs, which include centerline curves and connections forming a Directed Acyclic Graph (DAG). Accurate extraction of lane graphs relies on precisely estimating vertex and edge i…

2024

LayerDiff: Exploring Text-guided Multi-layered Composable Image Synthesis via Layer-Collaborative Diffusion Model

ECCV 2024poster

"Despite the success of generating high-quality images given any text prompts by diffusion-based generative models, prior work directly generates the entire images, but cannot provide object-wise manipulation capability. To support wider real applications like professional graphic design and digital…

2024

MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing

ECCV 2024poster

"The diffusion model is widely leveraged for either video generation or video editing. As each field has its task-specific problems, it is difficult to merely develop a single diffusion for completing both tasks simultaneously. Video diffusion sorely relying on the text prompt can be adapted to unif…

2024

Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent

IJCAI 2024poster

Counterfactual regret minimization (CFR) is a family of algorithms for effectively solving imperfect-information games. It decomposes the total regret into counterfactual regrets, utilizing local regret minimization algorithms, such as Regret Matching (RM) or RM+, to minimize them. Recent research e…

2024

OpenOcc: Open Vocabulary 3D Scene Reconstruction via Occupancy Representation

IROS 2024poster

3D reconstruction has been widely used in autonomous navigation fields of mobile robotics. However, the former research can only provide the basic geometry structure without the capability of open-world scene understanding, limiting advanced tasks like human interaction and visual navigation. Moreov…

Cited by 2SourcecodeScholar
2024

PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation

NeurIPS 2024poster

Language-guided robotic manipulation is a challenging task that requires an embodied agent to follow abstract user instructions to accomplish various complex manipulation tasks. Previous work generally maps instructions and visual perceptions directly to low-level executable actions, neglecting the…

Cited by 1SourcePDFScholar
2024

PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion

ECCV 2024poster

"Current large-scale diffusion models represent a giant leap forward in conditional image synthesis, capable of interpreting diverse cues like text, human poses, and edges. However, their reliance on substantial computational resources and extensive data collection remains a bottleneck. On the other…

2024

Rethinking Boundary Discontinuity Problem for Oriented Object Detection

CVPR 2024poster

Oriented object detection has been developed rapidly in the past few years where rotation equivariance is crucial for detectors to predict rotated boxes. It is expected that the prediction can maintain the corresponding rotation when objects rotate but severe mutation in angular prediction is someti…

2024

Self-Adaptive Reality-Guided Diffusion for Artifact-Free Super-Resolution

CVPR 2024poster

Artifact-free super-resolution (SR) aims to translate low-resolution images into their high-resolution counterparts with a strict integrity of the original content eliminating any distortions or synthetic details. While traditional diffusion-based SR techniques have demonstrated remarkable abilities…

2024

SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM

NeurIPS 2024poster

Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze video data. However, current Vid-LLMs struggle to simultaneously retain high-quality frame-level semantic information (i.e.…

Cited by 2SourcePDFScholar
2024

TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields

ICLR 2024poster

Recent works learn 3D representation explicitly under text-3D guidance. However, limited text-3D data restricts the vocabulary scale and text control of generations. Generators may easily fall into a stereotype concept for certain text prompts, thus losing open-vocabulary generation ability. To tack…

Cited by 12SourcePDFScholar
2024

TopoLogic: An Interpretable Pipeline for Lane Topology Reasoning on Driving Scenes

NeurIPS 2024poster

As an emerging task that integrates perception and reasoning, topology reasoning in autonomous driving scenes has recently garnered widespread attention. However, existing work often emphasizes "perception over reasoning": they typically boost reasoning performance by enhancing the perception of la…

2024

UNIT: Unifying Image and Text Recognition in One Vision Encoder

NeurIPS 2024poster

Currently, vision encoder models like Vision Transformers (ViTs) typically excel at image recognition tasks but cannot simultaneously support text recognition like human visual recognition. To address this limitation, we propose UNIT, a novel training framework aimed at UNifying Image and Text recog…

Cited by 3SourcePDFScholar
2024

VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation

NeurIPS 2024poster

Recent advancements utilizing large-scale video data for learning video generation models demonstrate significant potential in understanding complex physical dynamics. It suggests the feasibility of leveraging diverse robot trajectory data to develop a unified, dynamics-aware model to enhance robot…

Cited by 1SourcePDFScholar
2023

CLIP2: Contrastive Language-Image-Point Pretraining From Real-World Point Cloud Data

CVPR 2023poster

Contrastive Language-Image Pre-training, benefiting from large-scale unlabeled text-image pairs, has demonstrated great performance in open-world vision understanding tasks. However, due to the limited Text-3D data pairs, adapting the success of 2D Vision-Language Models (VLM) to the 3D space remain…

Cited by 107SourcePDFScholar
2023

CO3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving

ICLR 2023poster

Unsupervised contrastive learning for indoor-scene point clouds has achieved great successes. However, unsupervised representation learning on outdoor-scene point clouds remains challenging because previous methods need to reconstruct the whole scene and capture partial views for the contrastive obj…

2023

CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object Detection

NeurIPS 2023poster

Open-vocabulary 3D Object Detection (OV-3DDet) aims to detect objects from an arbitrary list of categories within a 3D scene, which remains seldom explored in the literature. There are primarily two fundamental problems in OV-3DDet, *i.e.*, localizing and classifying novel objects. This paper aims a…

2023

ConQueR: Query Contrast Voxel-DETR for 3D Object Detection

CVPR 2023highlight

Although DETR-based 3D detectors simplify the detection pipeline and achieve direct sparse predictions, their performance still lags behind dense detectors with post-processing for 3D object detection from point clouds. DETRs usually adopt a larger number of queries than GTs (e.g., 300 queries v.s.…

2023

DetCLIPv2: Scalable Open-Vocabulary Object Detection Pre-Training via Word-Region Alignment

CVPR 2023poster

This paper presents DetCLIPv2, an efficient and scalable training framework that incorporates large-scale image-text pairs to achieve open-vocabulary object detection (OVD). Unlike previous OVD frameworks that typically rely on a pre-trained vision-language model (e.g., CLIP) or exploit image-text p…

2023

DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment

ICCV 2023poster

Cross-modal garment synthesis and manipulation will significantly benefit the way fashion designers generate garments and modify their designs via flexible linguistic interfaces. However, despite the significant progress that has been made in generic image synthesis using diffusion models, producing…

Cited by 17PDFScholar
2023

DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination Capability

ICCV 2023poster

Recently, large-scale diffusion models, e.g., Stable diffusion and DallE2, have shown remarkable results on image synthesis. On the other hand, large-scale cross-modal pre-trained models (e.g., CLIP, ALIGN, and FILIP) are competent for various downstream tasks by learning to align vision and languag…

Cited by 3PDFScholar
2023

FULLER: Unified Multi-modality Multi-task 3D Perception via Multi-level Gradient Calibration

ICCV 2023poster

Multi-modality fusion and multi-task learning are becoming trendy in 3D autonomous driving scenario, considering robust prediction and computation budget. However, naively extending the existing framework to the domain of multi-modality multi-task learning remains ineffective and even poisonous due…

Cited by 10PDFScholar
2023

Gaussian Label Distribution Learning for Spherical Image Object Detection

CVPR 2023poster

Spherical image object detection emerges in many applications from virtual reality to robotics and automatic driving, while many existing detectors use ln-norms loss for regression of spherical bounding boxes. There are two intrinsic flaws for ln-norms loss, i.e., independent optimization of paramet…

Cited by 9SourcePDFScholar
2023

GrowCLIP: Data-Aware Automatic Model Growing for Large-scale Contrastive Language-Image Pre-Training

ICCV 2023poster

Cross-modal pre-training has shown impressive performance on a wide range of downstream tasks, benefiting from massive image-text pairs collected from the Internet. In practice, online data are growing constantly, highlighting the importance of the ability of pre-trained model to learn from data tha…

Cited by 5PDFcodeScholar
2023

MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation

ICCV 2023poster

Recently, semantic segmentation models trained with image-level text supervision have shown promising results in challenging open-world scenarios. However, these models still face difficulties in learning fine-grained semantic alignment at the pixel level and predicting accurate object masks. To add…

Cited by 19PDFScholar
2023

Mixed Autoencoder for Self-Supervised Visual Representation Learning

CVPR 2023poster

Masked Autoencoder (MAE) has demonstrated superior performance on various vision tasks via randomly masking image patches and reconstruction. However, effective data augmentation strategies for MAE still remain open questions, different from those in contrastive learning that serve as the most impor…

Cited by 51SourcePDFScholar
2023

NLIP: Noise-Robust Language-Image Pre-training

AAAI 2023technical

Large-scale cross-modal pre-training paradigms have recently shown ubiquitous success on a wide range of downstream tasks, e.g., zero-shot classification, retrieval and image captioning. However, their successes highly rely on the scale and quality of web-crawled data that naturally contain much inc…

Cited by 33SourcePDFScholar
2023

OpenLane-V2: A Topology Reasoning Benchmark for Unified 3D HD Mapping

NeurIPS 2023poster

Accurately depicting the complex traffic scene is a vital component for autonomous vehicles to execute correct judgments. However, existing benchmarks tend to oversimplify the scene by solely focusing on lane perception tasks. Observing that human drivers rely on both lanes and traffic signals to op…

2023

PARTNER: Level up the Polar Representation for LiDAR 3D Object Detection

ICCV 2023poster

Recently, polar-based representation has shown promising properties in perceptual tasks. In addition to Cartesian-based approaches, which separate point clouds unevenly, representing point clouds as polar grids has been recognized as an alternative due to (1) its advantage in robust performance unde…

Cited by 10PDFcodeScholar
2023

PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video Retrieval

ICCV 2023poster

Text-video retrieval is a fundamental task with high practical value in multi-modal research. Inspired by the great success of pre-trained image-text models with large-scale data, such as CLIP, many methods are proposed to transfer the strong representation learning capability of CLIP to text-video…

Cited by 19PDFScholar
2023

Poisoning the Well: Can We Simultaneously Attack a Group of Learning Agents?

IJCAI 2023poster

Reinforcement Learning's (RL) ubiquity has instigated research on potential threats to its training and deployment. Many works study single-learner training-time attacks that "pre-programme" behavioral triggers into a strategy. However, attacks on collections of learning agents remain largely overlo…

Cited by 2SourcePDFScholar
2023

SLAMB: Accelerated Large Batch Training with Sparse Communication

ICML 2023poster

Distributed training of large deep neural networks requires frequent exchange of massive data between machines, thus communication efficiency is a major concern. Existing compressed communication methods are either not compatible with large batch optimization algorithms, or do not provide sufficient…

Cited by 8SourcePDFScholar
2023

SUIT: Learning Significance-Guided Information for 3D Temporal Detection

IROS 2023poster

3D object detection from LiDAR point cloud is of critical importance for autonomous driving and robotics. While sequential point cloud has the potential to enhance 3D perception through temporal information, utilizing these temporal features effectively and efficiently remains a challenging problem.…

Cited by 3SourceScholar
2023

Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning

ICLR 2023top-25%

There is a rising interest in further exploring the zero-shot learning potential of large pre-trained language models (PLMs). A new paradigm called data-generation-based zero-shot learning has achieved impressive success. In this paradigm, the synthesized data from the PLM acts as the carrier of kno…

2023

Sph2Pob: Boosting Object Detection on Spherical Images with Planar Oriented Boxes Methods

IJCAI 2023poster

Object detection on panoramic/spherical images has been developed rapidly in the past few years, where IoU-calculator is a fundamental part of various detector components, i.e. Label Assignment, Loss and NMS. Due to the low efficiency and non-differentiability of spherical Unbiased IoU, spherical ap…

2023

Task-customized Masked Autoencoder via Mixture of Cluster-conditional Experts

ICLR 2023top-25%

Masked Autoencoder (MAE) is a prevailing self-supervised learning method that achieves promising results in model pre-training. However, when the various downstream tasks have data distributions different from the pre-training data, the semantically irrelevant pre-training information might result i…

Cited by 21SourcePDFScholar
2023

Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images

ICCV 2023poster

Generating 3D faces from textual descriptions has a multitude of applications, such as gaming, movie and robotics. Recent progresses have demonstrated the success of unconditional 3D face generation and text-to-3D shape generation. However, due to the limited text-3D face data pairs, text-driven 3D…

Cited by 18PDFcodeScholar
2023

Translating Images to Road Network: A Non-Autoregressive Sequence-to-Sequence Approach

ICCV 2023oral

The extraction of road network is essential for the generation of high-definition maps since it enables the precise localization of road landmarks and their interconnections. However, generating road network poses a significant challenge due to the conflicting underlying combination of Euclidean (e.…

Cited by 8PDFScholar
2023

ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic Consistency

ICLR 2023poster

Recently, great success has been made in learning visual representations from text supervision, facilitating the emergence of text-supervised semantic segmentation. However, existing works focus on pixel grouping and cross-modal semantic alignment, while ignoring the correspondence among multiple au…

2023

Visual Exemplar Driven Task-Prompting for Unified Perception in Autonomous Driving

CVPR 2023poster

Multi-task learning has emerged as a powerful paradigm to solve a range of tasks simultaneously with good efficiency in both computation resources and inference time. However, these algorithms are designed for different tasks mostly not within the scope of autonomous driving, thus making it hard to…

Cited by 21SourcePDFScholar
2022

Arch-Graph: Acyclic Architecture Relation Predictor for Task-Transferable Neural Architecture Search

CVPR 2022poster

Neural Architecture Search (NAS) aims to find efficient models for multiple tasks. Beyond seeking solutions for a single task, there are surging interests in transferring network design knowledge across multiple tasks. In this line of research, effectively modeling task correlations is vital yet hig…

Cited by 25PDFcodeScholar
2022

AutoBERT-Zero: Evolving BERT Backbone from Scratch

AAAI 2022technical

Transformer-based pre-trained language models like BERT and its variants have recently achieved promising performance in various natural language processing (NLP) tasks. However, the conventional paradigm constructs the backbone by purely stacking the manually designed global self-attention layers,…

Cited by 44SourcePDFScholar
2022

AutoCFR: Learning to Design Counterfactual Regret Minimization Algorithms

AAAI 2022technical

Counterfactual regret minimization (CFR) is the most commonly used algorithm to approximately solving two-player zero-sum imperfect-information games (IIGs). In recent years, a series of novel CFR variants such as CFR+, Linear CFR, DCFR have been proposed and have significantly improved the converge…

2022

CODA: A Real-World Road Corner Case Dataset for Object Detection in Autonomous Driving

ECCV 2022poster

"Contemporary deep-learning object detection methods for autonomous driving usually assume prefixed categories of common traffic participants, such as pedestrians and cars. Most existing detectors are unable to detect uncommon objects and corner cases (e.g., a dog crossing a street), which may lead…

2022

Continual Object Detection via Prototypical Task Correlation Guided Gating Mechanism

CVPR 2022poster

Continual learning is a challenging real-world problem for constructing a mature AI system when data are provided in a streaming fashion. Despite recent progress in continual classification, the researches of continual object detection are impeded by the diverse sizes and numbers of objects in each…

Cited by 44PDFcodeScholar
2022

DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world Detection

NeurIPS 2022accept

Open-world object detection, as a more general and challenging goal, aims to recognize and localize objects described by arbitrary category names. The recent work GLIP formulates this problem as a grounding problem by concatenating all category names of detection datasets into sentences, which leads…

Cited by 178SourcePDFScholar
2022

DevNet: Self-Supervised Monocular Depth Learning via Density Volume Construction

ECCV 2022poster

"Self-supervised depth learning from monocular images normally relies on the 2D pixel-wise photometric relation between temporally adjacent image frames. However, they neither fully exploit the 3D point-wise geometric correspondences, nor effectively tackle the ambiguities in the photometric warping…

2022

Effective Adaptation in Multi-Task Co-Training for Unified Autonomous Driving

NeurIPS 2022accept

Aiming towards a holistic understanding of multiple downstream tasks simultaneously, there is a need for extracting features with better transferability. Though many latest self-supervised pre-training methods have achieved impressive performance on various vision tasks under the prevailing pretrain…

Cited by 38SourcePDFScholar
2022

FILIP: Fine-grained Interactive Language-Image Pre-Training

ICLR 2022poster

Unsupervised large-scale vision-language pre-training has shown promising advances on various downstream tasks. Existing methods often model the cross-modal interaction either via the similarity of the global feature of each modality which misses sufficient information, or finer-grained interactions…

Cited by 672SourcePDFScholar
2022

Generative Negative Text Replay for Continual Vision-Language Pretraining

ECCV 2022poster

"Vision-language pre-training (VLP) has attracted increasing attention recently. With a large amount of image-text pairs, VLP models trained with contrastive loss have achieved impressive performance in various tasks, especially the zero-shot generalization on downstream datasets. In practical appli…

Cited by 26SourcePDFScholar
2022

Laneformer: Object-Aware Row-Column Transformers for Lane Detection

AAAI 2022technical

We present Laneformer, a conceptually simple yet powerful transformer-based architecture tailored for lane detection that is a long-standing research topic for visual perception in autonomous driving. The dominant paradigms rely on purely CNN-based architectures which often fail in incorporating rel…

Cited by 60SourcePDFScholar
2022

Learning Ego 3D Representation As Ray Tracing

ECCV 2022poster

"A self-driving perception model aims to extract 3D semantic representations from multiple cameras collectively into the bird’s-eye-view (BEV) coordinate frame of the ego car in order to ground downstream planner. Existing perception methods often rely on error-prone depth estimation of the whole sc…

2022

MPPNet: Multi-Frame Feature Intertwining with Proxy Points for 3D Temporal Object Detection

ECCV 2022poster

"Accurate and reliable 3D detection is vital for many applications including autonomous driving vehicles and service robots. In this paper, we present a flexible and high-performance 3D detection frame-work, named MPPNet, for 3D temporal object detection with point cloud sequences. We propose a nove…

2022

ManiTrans: Entity-Level Text-Guided Image Manipulation via Token-Wise Semantic Alignment and Generation

CVPR 2022oral

Existing text-guided image manipulation methods aim to modify the appearance of the image or to edit a few objects in a virtual or simple scenario, which is far from practical application. In this work, we study a novel task on text-guided image manipulation on the entity level in the real world. Th…

Cited by 19PDFcodeScholar
2022

ONCE-3DLanes: Building Monocular 3D Lane Detection

CVPR 2022poster

We present ONCE-3DLanes, a real-world autonomous driving dataset with lane layout annotation in 3D space. Conventional 2D lane detection from a monocular image yields poor performance of following planning and control tasks in autonomous driving due to the case of uneven road. Predicting the 3D lane…

Cited by 76PDFcodeScholar
2022

Open-World Semantic Segmentation via Contrasting and Clustering Vision-Language Embedding

ECCV 2022poster

"To bridge the gap between supervised semantic segmentation and real-world applications that acquire one model to recognize arbitrary new concepts, recent zero-shot segmentation attracts a lot of attention by exploring the relationships between unseen and seen object categories, yet requiring large…

2022

PANDORA: A Panoramic Detection Dataset for Object with Orientation

ECCV 2022poster

"Panoramic images have become increasingly popular as omnidirectional panoramic technology has advanced. Many datasets and works resort to object detection to better understand the content of the panoramic image. These datasets and detectors use a Bounding Field of View (BFoV) as a bounding box in p…

2022

Point2Seq: Detecting 3D Objects As Sequences

CVPR 2022poster

We present a simple and effective framework, named Point2Seq, for 3D object detection from point clouds. In contrast to previous methods that normally predict attributes of 3D objects all at once, we expressively model the interdependencies between attributes of 3D objects, which in turn enables a b…

Cited by 18PDFcodeScholar
2022

RCLane: Relay Chain Prediction for Lane Detection

ECCV 2022poster

"Lane detection is an important component of many real-world autonomous systems. Despite a wide variety of lane detection approaches have been proposed, reporting steady benchmark improvements over time, lane detection remains a largely unsolved problem. This is because most of the existing lane det…

Cited by 32SourcePDFScholar
2022

Revisiting Over-smoothing in BERT from the Perspective of Graph

ICLR 2022spotlight

Recently over-smoothing phenomenon of Transformer-based models is observed in both vision and language fields. However, no existing work has delved deeper to further investigate the main cause of this phenomenon. In this work, we make the attempt to analyze the over-smoothing problem from the perspe…

Cited by 83SourcePDFScholar
2022

Task-Customized Self-Supervised Pre-training with Scalable Dynamic Routing

AAAI 2022technical

Self-supervised learning (SSL), especially contrastive methods, has raised attraction recently as it learns effective transferable representations without semantic annotations. A common practice for self-supervised pre-training is to use as much data as possible. For a specific downstream task, howe…

Cited by 23SourcePDFScholar
2022

Unbiased IoU for Spherical Image Object Detection

AAAI 2022technical

As one of the fundamental components of object detection, intersection-over-union (IoU) calculations between two bounding boxes play an important role in samples selection, NMS operation and evaluation of object detection algorithms. This procedure is well-defined and solved for planar images, while…

Cited by 12SourcePDFScholar
2022

Visual-Language Navigation Pretraining via Prompt-based Environmental Self-exploration

ACL 2022long

Vision-language navigation (VLN) is a challenging task due to its large searching space in the environment. To address this problem, previous works have proposed some methods of fine-tuning a large model that pretrained on large-scale datasets. However, the conventional fine-tuning methods require e…

2022

Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark

NeurIPS 2022accept

Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale datasets and benchmarks in Chinese hinders the development of Chinese VLP models an…

2022

ZeroGen: Efficient Zero-shot Learning via Dataset Generation

EMNLP 2022main

There is a growing interest in dataset generation recently due to the superior generative capacity of large pre-trained language models (PLMs). In this paper, we study a flexible and efficient zero-short learning method, ZeroGen.Given a zero-shot task, we first generate a dataset from scratch using…

2021

Ada-Segment: Automated Multi-loss Adaptation for Panoptic Segmentation

AAAI 2021technical

Panoptic segmentation that unifies instance segmentation and semantic segmentation has recently attracted increasing attention. While most existing methods focus on designing novel architectures, we steer toward a different perspective: performing automated multi-loss adaptation (named Ada-Segment)…

Cited by 9SourcePDFScholar
2021

Adversarial Robustness for Unsupervised Domain Adaptation

ICCV 2021poster

Extensive Unsupervised Domain Adaptation (UDA) studies have shown great success in practice by learning transferable representations across a labeled source domain and an unlabeled target domain with deep models. However, current work focuses on improving the generalization ability of UDA models on…

Cited by 46PDFScholar
2021

C3-SemiSeg: Contrastive Semi-Supervised Segmentation via Cross-Set Learning and Dynamic Class-Balancing

ICCV 2021poster

The semi-supervised semantic segmentation methods utilize the unlabeled data to increase the feature discriminative ability to alleviate the burden of the annotated data. However, the dominant consistency learning diagram is limited by a) the misalignment between features from labeled and unlabeled…

Cited by 98PDFScholar
2021

DeepReduce: A Sparse-tensor Communication Framework for Federated Deep Learning

NeurIPS 2021poster

Sparse tensors appear frequently in federated deep learning, either as a direct artifact of the deep neural network’s gradients, or as a result of an explicit sparsification process. Existing communication primitives are agnostic to the peculiarities of deep learning; consequently, they impose unne…

2021

DetCo: Unsupervised Contrastive Learning for Object Detection

ICCV 2021poster

We present DetCo, a simple yet effective self-supervised approach for object detection. Unsupervised pre-training methods have been recently designed for object detection, but they are usually deficient in image classification, or the opposite. Unlike them, DetCo transfers well on downstream instanc…

Cited by 408PDFcodeScholar
2021

Effective Sparsification of Neural Networks With Global Sparsity Constraint

CVPR 2021poster

Weight pruning is an effective technique to reduce the model size and inference time for deep neural networks in real world deployments. However, since magnitudes and relative importance of weights are very different for different layers of a neural network, existing methods rely on either manual tu…

Cited by 82PDFcodeScholar
2021

EfficientBERT: Progressively Searching Multilayer Perceptron via Warm-up Knowledge Distillation

EMNLP 2021finding

Pre-trained language models have shown remarkable results on various NLP tasks. Nevertheless, due to their bulky size and slow inference speed, it is hard to deploy them on edge devices. In this paper, we have a critical insight that improving the feed-forward network (FFN) in BERT has a higher gain…

2021

Exploring Geometry-Aware Contrast and Clustering Harmonization for Self-Supervised 3D Object Detection

ICCV 2021poster

Current 3D object detection paradigms highly rely on extensive annotation efforts, which makes them not practical in many real-world industrial applications. Inspired by that a human driver can keep accumulating experiences from self-exploring the roads without any tutor's guidance, we first step fo…

Cited by 86PDFcodeScholar
2021

G-DetKD: Towards General Distillation Framework for Object Detectors via Contrastive and Semantic-Guided Feature Imitation

ICCV 2021poster

In this paper, we investigate the knowledge distillation (KD) strategy for object detection and propose an effective framework applicable to both homogeneous and heterogeneous student-teacher pairs. The conventional feature imitation paradigm introduces imitation masks to focus on informative foregr…

Cited by 31PDFScholar
2021

How to Save your Annotation Cost for Panoptic Segmentation?

AAAI 2021technical

How to properly reduce the annotation cost for panoptic segmentation? How to leverage and optimize the cost-quality trade-off for training data and model? These questions are key challenges towards a label-efficient and scalable panoptic segmentation system due to its expensive instance/semantic pix…

Cited by 5SourcePDFScholar
2021

Joint-DetNAS: Upgrade Your Detector With NAS, Pruning and Dynamic Distillation

CVPR 2021poster

We propose Joint-DetNAS, a unified NAS framework for object detection, which integrates 3 key components: Neural Architecture Search, pruning, and Knowledge Distillation. Instead of naively pipelining these techniques, our Joint-DetNAS optimizes them jointly. The algorithm consists of two core proce…

Cited by 40PDFcodeScholar
2021

Learning Transferable Features for Point Cloud Detection via 3D Contrastive Co-training

NeurIPS 2021poster

Most existing point cloud detection models require large-scale, densely annotated datasets. They typically underperform in domain adaptation settings, due to geometry shifts caused by different physical environments or LiDAR sensor configurations. Therefore, it is challenging but valuable to learn t…

Cited by 34SourcePDFScholar
2021

Loss Function Discovery for Object Detection via Convergence-Simulation Driven Search

ICLR 2021poster

Designing proper loss functions for vision tasks has been a long-standing research direction to advance the capability of existing models. For object detection, the well-established classification and regression loss functions have been carefully designed by considering diverse learning challenges (…

2021

MultiSiam: Self-Supervised Multi-Instance Siamese Representation Learning for Autonomous Driving

ICCV 2021poster

Autonomous driving has attracted much attention over the years but turns out to be harder than expected, probably due to the difficulty of labeled data collection for model training. Self-supervised learning (SSL), which leverages unlabeled data only for representation learning, might be a promising…

Cited by 65PDFcodeScholar
2021

NASOA: Towards Faster Task-Oriented Online Fine-Tuning With a Zoo of Models

ICCV 2021poster

Fine-tuning from pre-trained ImageNet models has been a simple, effective, and popular approach for various computer vision tasks. The common practice of fine-tuning is to adopt a default hyperparameter setting with a fixed pre-trained model, while both of them are not optimized for specific tasks a…

Cited by 10PDFcodeScholar
2021

One Million Scenes for Autonomous Driving: ONCE Dataset

NeurIPS 2021poster

Current perception models in autonomous driving have become notorious for greatly relying on a mass of annotated data to cover unseen cases and address the long-tail problem. On the other hand, learning from unlabeled large-scale collected data and incrementally self-training powerful recognition mo…

Cited by 332SourcecodeScholar
2021

Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-Modal Pretraining

ICCV 2021poster

Nowadays, customer's demands for E-commerce are more diversified, which introduces more complications to the product retrieval industry. Previous methods are either subject to single-modal input or perform supervised image-level product retrieval, thus fail to accommodate real-life scenarios where e…

Cited by 75PDFcodeScholar
2021

Pyramid R-CNN: Towards Better Performance and Adaptability for 3D Object Detection

ICCV 2021poster

We present a flexible and high-performance framework, named Pyramid R-CNN, for two-stage 3D object detection from point clouds. Current approaches generally rely on the points or voxels of interest for RoI feature extraction on the second stage, but cannot effectively handle the sparsity and non-uni…

Cited by 201PDFcodeScholar
2021

SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving

NeurIPS 2021poster

Aiming at facilitating a real-world, ever-evolving and scalable autonomous driving system, we present a large-scale dataset for standardizing the evaluation of different self-supervised and semi-supervised approaches by learning from raw data, which is the first and largest dataset to date. Existing…

Cited by 82SourcecodeScholar
2021

SOFT: Softmax-free Transformer with Linear Complexity

NeurIPS 2021spotlight

Vision transformers (ViTs) have pushed the state-of-the-art for various visual recognition tasks by patch-wise image tokenization followed by self-attention. However, the employment of self-attention modules results in a quadratic complexity in both computation and memory usage. Various attempts on…

Cited by 198SourcePDFScholar
2021

Segmenting Transparent Objects in the Wild with Transformer

IJCAI 2021poster

This work presents a new fine-grained transparent object segmentation dataset, termed Trans10K-v2, extending Trans10K-v1, the first large-scale transparent object segmentation dataset. Unlike Trans10K-v1 that only has two limited categories, our new dataset has several appealing benefits. (1) It h…

2021

SparseBERT: Rethinking the Importance Analysis in Self-attention

ICML 2021spotlight

Transformer-based models are popularly used in natural language processing (NLP). Its core component, self-attention, has aroused widespread interest. To understand the self-attention mechanism, a direct method is to visualize the attention map of a pre-trained model. Based on the patterns observed,…

2021

TransNAS-Bench-101: Improving Transferability and Generalizability of Cross-Task Neural Architecture Search

CVPR 2021poster

Recent breakthroughs of Neural Architecture Search (NAS) extend the field's research scope towards a broader range of vision tasks and more diversified search spaces. While existing NAS methods mostly design architectures on a single task, algorithms that look beyond single-task search are surging t…

Cited by 82PDFScholar
2020

AABO: Adaptive Anchor Box Optimization for Object Detection via Bayesian Sub-sampling

ECCV 2020poster

Most state-of-the-art object detection systems follow an anchor-based diagram. Anchor boxes are densely proposed over the images and the network is trained to predict the boxes position offset as well as the classification confidence. Existing systems pre-define anchor box shapes and sizes and ad-ho…

Cited by 24SourcePDFScholar
2020

Auto-Panoptic: Cooperative Multi-Component Architecture Search for Panoptic Segmentation

NeurIPS 2020poster

Panoptic segmentation is posed as a new popular test-bed for the state-of-the-art holistic scene understanding methods with the requirement of simultaneously segmenting both foreground things and background stuff. The state-of-the-art panoptic segmentation network exhibits high structural complexity…

2020

Bridging the Gap between Sample-based and One-shot Neural Architecture Search with BONAS

NeurIPS 2020poster

Neural Architecture Search (NAS) has shown great potentials in finding better neural network designs. Sample-based NAS is the most reliable approach which aims at exploring the search space and evaluating the most promising architectures. However, it is computationally very costly. As a remedy, the…

2020

CATCH: Context-based Meta Reinforcement Learning for Transferrable Architecture Search

ECCV 2020poster

Neural Architecture Search (NAS) achieved many breakthroughs in recent years. In spite of its remarkable progress, many algorithms are restricted to particular search spaces. They also lack efficient mechanisms to reuse knowledge when confronting multiple tasks. These challenges preclude their appli…

Cited by 24SourcePDFScholar
2020

Combining CGAN and Mil for Hotspot Segmentation in Bone Scintigraphy

ICASSP 2020accepted

Bone scintigraphy is widely used to diagnose bone tumor and metastasis. Accurate hotspot segmentation from bone scintigraphy is of great importance for tumor metastasis diagnosis. In this paper, we propose a new framework to detect and extract hotspots in thoracic region by integrating the technique…

Cited by 0SourceScholar
2020

CurveLane-NAS: Unifying Lane-Sensitive Architecture Search and Adaptive Point Blending

ECCV 2020poster

We address the curve lane detection problem which poses more real-world challenges than conventional lane detection for better facilitating modern assisted/autonomous driving systems. Current hand-designed lane detection methods are not robust enough to capture the curve lanes especially the remote…

Cited by 237SourcePDFScholar
2020

JGR-P2O: Joint Graph Reasoning based Pixel-to-Offset Prediction Network for 3D Hand Pose Estimation from a Single Depth Image

ECCV 2020poster

State-of-the-art single depth image-based 3D hand pose estimation methods are based on dense predictions, including voxel-to-voxel predictions, point-to-point regression, and pixel-wise estimations. Despite the good performance, those methods have a few issues in nature, such as the poor trade-off b…

2020

Robust Visual Tracking with Context-Based Active Occlusion Recognition

ICASSP 2020accepted

Occlusion is a great challenge for target model update in visual tracking. The target template may be corrupted by non-object information in the process of online learning due to occlusion. In this paper, we propose a context-based active occlusion recognition framework that can be integrated with v…

Cited by 0SourceScholar
2020

SP-NAS: Serial-to-Parallel Backbone Search for Object Detection

CVPR 2020poster

Advanced object detectors usually adopt a backbone network designed and pretrained by ImageNet classification. Recently neural architecture search (NAS) has emerged to automatically design a task-specific backbone to bridge the gap between the tasks of classification and detection. In this paper, we…

Cited by 77PDFScholar
2019

Auto-FPN: Automatic Network Architecture Adaptation for Object Detection Beyond Classification

ICCV 2019poster

Abstract Neural architecture search (NAS) has shown great potential in automating the manual process of designing a good CNN architecture for image classification. In this paper, we study NAS for object detection, a core computer vision task that classifies and localizes object instances in an image…

Cited by 259PDFScholar
2019

Reasoning-RCNN: Unifying Adaptive Global Reasoning Into Large-Scale Object Detection

CVPR 2019oral

In this paper, we address the large-scale object detection problem with thousands of categories, which poses severe challenges due to long-tail data distributions, heavy occlusions, and class ambiguities. However, the dominant object detection paradigm is limited by treating each object region separ…

Cited by 117PDFcodeScholar
2018

Hybrid Knowledge Routed Modules for Large-scale Object Detection

NeurIPS 2018poster

Abstract The dominant object detection approaches treat the recognition of each region separately and overlook crucial semantic correlations between objects in one scene. This paradigm leads to substantial performance drop when facing heavy long-tail problems, where very few samples are available fo…