← Search

Xiangyang Xue

103 accepted papers

2026

ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation

CVPR 2026

Recent advances in robot manipulation have leveraged pre-trained vision-language models (VLMs) and explored integrating 3D spatial signals into these models for effective action prediction, giving rise to the promising vision-language-action (VLA) paradigm. However, most existing approaches overlook

Cited by 0SourcecodeScholar
2026

CME-CAD: Heterogeneous Collaborative Multi-Expert Reinforcement Learning for CAD Code Generation

CVPR 2026

Computer-Aided Design (CAD) is essential in industrial design, but the complexity of traditional CAD modeling and workflows presents significant challenges for automating the generation of high-precision, editable CAD models. Existing methods, such as 3D reconstruction from sketches, often produce n

Cited by 0SourceScholar
2026

Decomposition of Concept-Level Rules in Visual Scenes

ICLR 2026poster

Human cognition is compositional, and one can parse a visual scene into independent concepts and the corresponding concept-changing rules. By contrast, many vision-language systems process images holistically, with limited support for explicit decomposition. And previous methods of decomposing conce…

Cited by 0SourceScholar
2026

DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video

CVPR 2026

Reliable 4D object detection, which refers to 3D object detection in streaming video, is crucial for perceiving and understanding the real world. Existing open-set 4D object detection methods typically make predictions on a frame-by-frame basis without modeling temporal consistency, or rely on compl

Cited by 0SourcecodeScholar
2026

DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving

CVPR 2026

Dynamic scene reconstruction in autonomous driving remains a fundamental challenge due to significant temporal variations, moving objects, and complex scene dynamics. Existing feed-forward 3D models have demonstrated strong performance in static reconstruction but still struggle to capture dynamic m

Cited by 0SourceScholar
2026

EgoSound: Benchmarking Sound Understanding in Egocentric Videos

CVPR 2026

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in vision-language understanding. Yet, human perception is inherently multisensory, integrating sight, sound, and motion to reason about the world. Among these modalities, sound provides indispensable cues about spat

Cited by 0SourcecodeScholar
2026

Envision, Attend, Then Respond: Counterfactual Hallucination Mitigation in Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) often hallucinate when visual evidence conflicts with world knowledge, i.e., in counterfactual scenarios. We propose Envision-Attend-Respond (EnAR), a training-free framework that leverages visual priors to steer the model's attention toward counterfactual elemen

Cited by 0SourcecodeScholar
2026

From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code Generation

AAAI 2026technical

Computer-Aided Design (CAD) plays a vital role in engineering and manufacturing, yet current CAD workflows require extensive domain expertise and manual modeling effort. Recent advances in large language models (LLMs) have made it possible to generate code from natural language, opening new opportun

Cited by 0SourcePDFScholar
2026

LOG-Nav: Efficient Layout-Aware Object-Goal Navigation with Hierarchical Planning

AAAI 2026technical

We introduce LOG-Nav, an efficient layout-aware object-goal navigation approach designed for complex multi-room indoor environments. By planning hierarchically leveraging a global topologigal map with layout information and local imperative approach with detailed scene representation memory, LOG-Na

Cited by 0SourcePDFScholar
2026

Learning Global Representation from Queries for Vectorized HD Map Construction

ICML 2026poster

The online construction of vectorized high-definition (HD) maps is a cornerstone of modern autonomous driving systems. State-of-the-art approaches, particularly those based on the DETR framework, formulate this as an instance detection problem. However, their reliance on independent, learnable objec…

Cited by 0SourceScholar
2026

OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-To-Robot Action Transfer

ICRA 2026poster

We present OCRA, an Object-Centric framework for video-based human-to-Robot Action transfer that learns directly from human demonstration videos to enable robust manipulation. Object-centric learning emphasizes task-relevant objects and their interactions while filtering out irrelevant background, p…

2026

SCOOP'D: Learning Mixed-Liquid-Solid Scooping Via Sim2Real Generative Policy

ICRA 2026poster

Scooping items with tools such as spoons and ladles is common in daily life, ranging from assistive feeding to retrieving items from environmental disaster sites. However, developing a general and autonomous robotic scooping policy is challenging since it requires reasoning about complex tool-object…

2026

Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

RSS 2026poster

Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely on VLM backbones optimized for Visual Question Answering (VQA…

Cited by 0SourceScholar
2026

V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence

CVPR 2026

Cross-view object correspondence, exemplified by the representative task of ego-exo object correspondence, aims to establish consistent associations of the same object across different viewpoints (e.g., ego-centric and exo-centric). This task poses significant challenges due to drastic viewpoint and

Cited by 0SourcecodeScholar
2025

Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning

ICML 2025poster

Abstract visual reasoning (AVR) enables humans to quickly discover and generalize abstract rules to new scenarios. Designing intelligent systems with human-like AVR abilities has been a long-standing topic in the artificial intelligence community. Deep AVR solvers have recently achieved remarkable s…

2025

CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image

CVPR 2025highlight

This paper tackles category-level pose estimation of ar- ticulated objects in robotic manipulation tasks and intro- duces a new benchmark dataset. While recent methods es- timate part poses and sizes at the category level, they often rely on geometric cues and complex multi-stage pipelines that firs…

Cited by 0SourcePDFScholar
2025

CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning

NeurIPS 2025poster

Computer-Aided Design (CAD) is pivotal in industrial manufacturing, with orthographic projection reasoning foundational to its entire workflow—encompassing design, manufacturing, and simulation. However, prevailing deep-learning approaches employ standard 3D reconstruction pipelines as an alternativ…

Cited by 0SourcecodeScholar
2025

ChatReID: Open-ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language Models

ICCV 2025poster

Person re-identification (Re-ID) is a crucial task in computer vision, aiming to recognize individuals across non-overlapping camera views. While recent advanced vision-language models (VLMs) excel in logical reasoning and multi-task generalization, their applications in Re-ID tasks remain limited.…

Cited by 0SourcePDFScholar
2025

CocoER: Aligning Multi-Level Feature by Competition and Coordination for Emotion Recognition

CVPR 2025poster

With the explosion of human-machine interaction, emotion recognition has reignited attention. Previous works focus on improving visual feature fusion and reasoning from multiple image levels. Although it is non-trivial to deduce a person's emotion by integrating multi-level feature (head, body a…

2025

Content and Salient Semantics Collaboration for Cloth-Changing Person Re-Identification

ICASSP 2025accepted

Cloth-changing person re-identification aims at recognizing the same person with clothing changes across non-overlapping cameras. Advanced methods either resort to identity-related auxiliary modalities (e.g., sketches, silhouettes, and keypoints) or clothing labels to mitigate the impact of clothes.…

Cited by 0SourceScholar
2025

Foundation Model Driven Appearance Extraction for Robust Multiple Object Tracking

AAAI 2025technical

Multiple Object Tracking (MOT) is a fundamental task in computer vision. Existing methods utilize motion information or appearance information to perform object tracking. However, these algorithms still struggle with special circumstances, such as occlusion and blurring in complex scenes. Inspired b…

2025

MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model

CVPR 2025poster

We introduce MVGenMaster, a multi-view diffusion model enhanced with 3D priors to address versatile Novel View Synthesis (NVS) tasks. MVGenMaster leverages 3D priors that are warped using metric depth and camera poses, significantly enhancing both generalization and 3D consistency in NVS. Our model…

2025

One-Shot Heterogeneous Federated Learning with Local Model-Guided Diffusion Models

ICML 2025poster

In recent years, One-shot Federated Learning (OSFL) methods based on Diffusion Models (DMs) have garnered increasing attention due to their remarkable performance. However, most of these methods require the deployment of foundation models on client devices, which significantly raises the computation…

Cited by 0SourcePDFScholar
2025

PC-BEV: An Efficient Polar-Cartesian BEV Fusion Framework for LiDAR Semantic Segmentation

AAAI 2025technical

Although multiview fusion has demonstrated potential in LiDAR segmentation, its dependence on computationally intensive point-based interactions, arising from the lack of fixed correspondences between views such as range view and Bird's-Eye View (BEV), hinders its practical deployment. This paper ch…

2025

RAG-6DPose: Retrieval-Augmented 6D Pose Estimation via Leveraging CAD as Knowledge Base

IROS 2025

Accurate 6D pose estimation is key for robotic manipulation, enabling precise object localization for tasks like grasping. We present RAG-6DPose, a retrieval-augmented approach that leverages 3D CAD models as a knowledge base by integrating both visual and geometric cues. Our RAG-6DPose roughly cont

Cited by 1SourcecodeScholar
2025

ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning

CVPR 2025poster

Open-vocabulary 3D visual grounding and reasoning aim to localize objects in a scene based on implicit language descriptions, even when they are occluded. This ability is crucial for tasks such as vision-language navigation and autonomous robotics. However, current methods struggle because they rely…

Cited by 0SourcePDFScholar
2025

Spatial-Temporal Aware Visuomotor Diffusion Policy Learning

ICCV 2025poster

Visual imitation learning is effective for robots to learn versatile tasks. However, many existing methods rely on behavior cloning with supervised historical trajectories, limiting their 3D spatial and 4D spatiotemporal awareness. Consequently, these methods struggle to capture the 3D structures an…

Cited by 0SourcePDFScholar
2025

TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making

NeurIPS 2025poster

In daily life, people often move through spaces to find objects that meet their needs, posing a key challenge in embodied AI. Traditional Demand-Driven Navigation (DDN) handles one need at a time but does not reflect the complexity of real-world tasks involving multiple needs and personal choices. T…

Cited by 0SourceScholar
2025

Topo-Field: Topometric Mapping With Brain-Inspired Hierarchical Layout-Object-Position Fields

RA-L 2025

Mobile robots require comprehensive scene understanding to operate effectively in diverse environments, enriched with contextual information such as layouts, objects, and their relationships. Although advances like neural radiance fields (NeRFs) offer high-fidelity 3D reconstructions, they are compu

Cited by 3SourcecodeScholar
2025

Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency

CVPR 2025highlight

Recent advances in image inpainting increasingly use generative models to handle large irregular masks. However, these models can create unrealistic inpainted images due to two main issues: (1) Unwanted object insertion: Even with unmasked areas as context, generative models may still generate arbit…

2025

You Only Estimate Once: Unified, One-stage, Real-Time Category-Level Articulated Object 6D Pose Estimation for Robotic Grasping

ICRA 2025

This paper addresses the problem of category-level pose estimation for articulated objects in robotic manipulation tasks. Recent works have shown promising results in estimating part pose and size at the category level. However, these approaches primarily follow a complex multi-stage pipeline that f

Cited by 4SourceScholar
2024

Automated Label Unification for Multi-Dataset Semantic Segmentation with GNNs

NeurIPS 2024poster

Deep supervised models possess significant capability to assimilate extensive training data, thereby presenting an opportunity to enhance model performance through training on multiple datasets. However, conflicts arising from different label spaces among datasets may adversely affect model performa…

2024

EAFormer: Scene Text Segmentation with Edge-Aware Transformers

ECCV 2024poster

"Scene text segmentation aims at cropping texts from scene images, which is usually used to help generative models edit or remove texts. The existing text segmentation methods tend to involve various text-related supervisions for better performance. However, most of them ignore the importance of tex…

2024

Exploring One-Shot Semi-supervised Federated Learning with Pre-trained Diffusion Models

AAAI 2024technical

Recently, semi-supervised federated learning (semi-FL) has been proposed to handle the commonly seen real-world scenarios with labeled data on the server and unlabeled data on the clients. However, existing methods face several challenges such as communication costs, data heterogeneity, and training…

Cited by 25SourcePDFScholar
2024

FastOcc: Accelerating 3D Occupancy Prediction by Fusing the 2D Bird’s-Eye View and Perspective View

ICRA 2024poster

In autonomous driving, 3D occupancy prediction outputs voxel-wise status and semantic labels for more comprehensive understandings of 3D scenes compared with traditional perception tasks, such as 3D object detection and bird’s-eye view (BEV) semantic segmentation. Recent researchers have extensively…

Cited by 35SourceScholar
2024

FedRA: A Random Allocation Strategy for Federated Tuning to Unleash the Power of Heterogeneous Clients

ECCV 2024poster

"With the increasing availability of Foundation Models, federated tuning has garnered attention in the field of federated learning, utilizing data and computation resources from multiple clients to collaboratively fine-tune foundation models. However, in real-world federated scenarios, there often e…

2024

Federated Adaptive Prompt Tuning for Multi-Domain Collaborative Learning

AAAI 2024technical

Federated learning (FL) enables multiple clients to collaboratively train a global model without disclosing their data. Previous researches often require training the complete model parameters. However, the emergence of powerful pre-trained models makes it possible to achieve higher performance with…

2024

Generating and Reweighting Dense Contrastive Patterns for Unsupervised Anomaly Detection

AAAI 2024technical

Recent unsupervised anomaly detection methods often rely on feature extractors pretrained with auxiliary datasets or on well-crafted anomaly-simulated samples. However, this might limit their adaptability to an increasing set of anomaly detection tasks due to the priors in the selection of auxiliary…

Cited by 18SourcePDFScholar
2024

Improving Neural Surface Reconstruction with Feature Priors from Multi-View Images

ECCV 2024poster

"Recent advancements in Neural Surface Reconstruction (NSR) have significantly improved multi-view reconstruction when coupled with volume rendering. However, relying solely on photometric consistency in image space falls short of addressing complexities posed by real-world data, including occlusion…

2024

Improving Viewpoint-Independent Object-Centric Representations through Active Viewpoint Selection

NeurIPS 2024poster

Given the complexities inherent in visual scenes, such as object occlusion, a comprehensive understanding often requires observation from multiple viewpoints. Existing multi-viewpoint object-centric learning methods typically employ random or sequential viewpoint selection strategies. While applicab…

Cited by 0SourcePDFScholar
2024

LAC-Net: Linear-Fusion Attention-Guided Convolutional Network for Accurate Robotic Grasping Under the Occlusion

IROS 2024poster

This paper addresses the challenge of perceiving complete object shapes through visual perception. While prior studies have demonstrated encouraging outcomes in segmenting the visible parts of objects within a scene, amodal segmentation, in particular, has the potential to allow robots to infer the…

Cited by 1SourcecodeScholar
2024

MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing

NeurIPS 2024poster

Novel View Synthesis (NVS) and 3D generation have recently achieved prominent improvements. However, these works mainly focus on confined categories or synthetic 3D assets, which are discouraged from generalizing to challenging in-the-wild scenes and fail to be employed with 2D synthesis directly. M…

2024

Make a Strong Teacher with Label Assistance: A Novel Knowledge Distillation Approach for Semantic Segmentation

ECCV 2024poster

"In this paper, we introduce a novel knowledge distillation approach for the semantic segmentation task. Unlike previous methods that rely on power-trained teachers or other modalities to provide additional knowledge, our approach does not require complex teacher models or information from extra sen…

2024

OpenAnnotate3D: Open-Vocabulary Auto-Labeling System for Multi-modal 3D Data

ICRA 2024poster

In the era of big data and large models, automatic annotating functions for multi-modal data are of great significance for real-world AI-driven applications, such as autonomous driving and embodied AI. Unlike traditional closed-set annotation, open-vocabulary annotation is essential to achieve human…

Cited by 13SourcecodeScholar
2024

TaMMa: Target-driven Multi-subscene Mobile Manipulation

CoRL 2024poster

For everyday service robotics, the ability to navigate back and forth based on tasks in multi-subscene environments and perform delicate manipulations is crucial and highly practical. While existing robotics primarily focus on complex tasks within a single scene or simple tasks across scalable scene…

Cited by 2SourceScholar
2024

Towards Generative Abstract Reasoning: Completing Raven’s Progressive Matrix via Rule Abstraction and Selection

ICLR 2024poster

Endowing machines with abstract reasoning ability has been a long-term research topic in artificial intelligence. Raven's Progressive Matrix (RPM) is widely used to probe abstract visual reasoning in machine intelligence, where models will analyze the underlying rules and select one image from candi…

2023

Chinese Text Recognition with A Pre-Trained CLIP-Like Model Through Image-IDS Aligning

ICCV 2023oral

Scene text recognition has been studied for decades due to its broad applications. However, despite Chinese characters possessing different characteristics from Latin characters, such as complex inner structures and large categories, few methods have been proposed for Chinese Text Recognition (CTR).…

Cited by 30PDFcodeScholar
2023

Grad-PU: Arbitrary-Scale Point Cloud Upsampling via Gradient Descent With Learned Distance Functions

CVPR 2023poster

Most existing point cloud upsampling methods have roughly three steps: feature extraction, feature expansion and 3D coordinate prediction. However, they usually suffer from two critical issues: (1) fixed upsampling rate after one-time training, since the feature expansion unit is customized for each…

2023

Improving Empathetic Dialogue Generation by Dynamically Infusing Commonsense Knowledge

ACL 2023findings

In empathetic conversations, individuals express their empathy towards others. Previous work has mainly focused on generating empathetic responses by utilizing the speaker’s emotion. Besides, external commonsense knowledge has been applied to enhance the system’s understandings of the speaker’s situ…

2023

Language Guided Robotic Grasping with Fine-Grained Instructions

IROS 2023poster

Given a single RGB image and the attribute-rich language instructions, this paper investigates the novel problem of using Fine-grained instructions for the Language guided robotic Grasping (FLarG). This problem is made challenging by learning fine-grained language descriptions to ground target objec…

Cited by 12SourcecodeScholar
2023

Learning Versatile 3D Shape Generation with Improved Auto-regressive Models

ICCV 2023poster

Auto-Regressive (AR) models have achieved impressive results in 2D image generation by modeling joint distributions in the grid space. While this approach has been extended to the 3D domain for powerful shape generation, it still has two limitations: expensive computations on volumetric grids and am…

Cited by 1PDFScholar
2023

Multi-to-Single Knowledge Distillation for Point Cloud Semantic Segmentation

ICRA 2023poster

3D point cloud semantic segmentation is one of the fundamental tasks for environmental understanding. Although significant progress has been made in recent years, the performance of classes with few examples or few points is still far from satisfactory. In this paper, we propose a novel multi-to-sin…

Cited by 6SourcecodeScholar
2023

Orientation-Independent Chinese Text Recognition in Scene Images

IJCAI 2023poster

Scene text recognition (STR) has attracted much attention due to its broad applications. The previous works pay more attention to dealing with the recognition of Latin text images with complex backgrounds by introducing language models or other auxiliary networks. Different from Latin texts, many ve…

2023

PourIt!: Weakly-Supervised Liquid Perception from a Single Image for Visual Closed-Loop Robotic Pouring

ICCV 2023poster

Liquid perception is critical for robotic pouring tasks. It usually requires the robust visual detection of flowing liquid. However, while recent works have shown promising results in liquid perception, they typically require labeled data for model training, a process that is both time-consuming and…

Cited by 6PDFScholar
2023

Towards Accurate Video Text Spotting with Text-wise Semantic Reasoning

IJCAI 2023poster

Video text spotting (VTS) aims at extracting texts from videos, where text detection, tracking and recognition are conducted simultaneously. There have been some works that can tackle VTS; however, they may ignore the underlying semantic relationships among texts within a frame. We observe that the…

2023

Training-free Diffusion Model Adaptation for Variable-Sized Text-to-Image Synthesis

NeurIPS 2023poster

Diffusion models (DMs) have recently gained attention with state-of-the-art performance in text-to-image synthesis. Abiding by the tradition in deep learning, DMs are trained and evaluated on the images with fixed sizes. However, users are demanding for various images with specific sizes and various…

Cited by 31SourcePDFScholar
2022

DST: Dynamic Substitute Training for Data-Free Black-Box Attack

CVPR 2022poster

With the wide applications of deep neural network models in various computer vision tasks, more and more works study the model vulnerability to adversarial examples. For data-free black box attack scenario, existing methods are inspired by the knowledge distillation, and thus usually train a substit…

Cited by 22PDFcodeScholar
2022

Density-Preserving Deep Point Cloud Compression

CVPR 2022poster

Local density of point clouds is crucial for representing local details, but has been overlooked by existing point cloud compression methods. To address this, we propose a novel deep point cloud compression method that preserves local density information. Our method works in an auto-encoder fashion:…

Cited by 70PDFcodeScholar
2022

H4D: Human 4D Modeling by Learning Neural Compositional Representation

CVPR 2022poster

Despite the impressive results achieved by deep learning based 3D reconstruction, the techniques of directly learning to model 4D human captures with detailed geometry have been less studied. This work presents a novel framework that can effectively learn a compact and compositional representation f…

Cited by 26PDFScholar
2022

High-Fidelity Portrait Editing Via Exploring Differentiable Guided Sketches from the Latent Space

ICASSP 2022accepted

This paper studies the task of sketch-guided high-fidelity portrait editing. Advanced unconditional generators, such as StyleGAN, can generate a high-quality portrait image with great diversity. In previous researches, StyleGAN has successfully been utilized for color-guided image editing through la…

Cited by 0SourceScholar
2022

I Know What You Draw: Learning Grasp Detection Conditioned on a Few Freehand Sketches

ICRA 2022poster

In this paper, we are interested in the problem of generating target grasps by understanding freehand sketches. The sketch is useful for the persons who cannot formulate language and the cases where a textual description is not available on the fly. However, very few works are aware of the usability…

Cited by 7SourceScholar
2022

Learning 6-DoF Object Poses to Grasp Category-Level Objects by Language Instructions

ICRA 2022poster

This paper studies the task of any objects grasping from the known categories by free-form language instructions. This task demands the technique in computer vision, natural language processing, and robotics. We bring these disciplines together on this open challenge, which is essential to human-rob…

Cited by 23SourceScholar
2022

LoRD: Local 4D Implicit Representation for High-Fidelity Dynamic Human Modeling

ECCV 2022poster

"Recent progress in 4D implicit representation focuses on globally controlling the shape and motion with low dimensional latent vectors, which is prone to missing surface details and accumulating tracking error. While many deep local representations have shown promising results for 3D shape modeling…

2022

RCLane: Relay Chain Prediction for Lane Detection

ECCV 2022poster

"Lane detection is an important component of many real-world autonomous systems. Despite a wide variety of lane detection approaches have been proposed, reporting steady benchmark improvements over time, lane detection remains a largely unsolved problem. This is because most of the existing lane det…

Cited by 32SourcePDFScholar
2022

SAR-Net: Shape Alignment and Recovery Network for Category-Level 6D Object Pose and Size Estimation

CVPR 2022poster

Given a single scene image, this paper proposes a method of Category-level 6D Object Pose and Size Estimation (COPSE) from the point cloud of the target object, without external real pose-annotated training data. Specifically, beyond the visual cues in RGB images, we rely on the shape information pr…

Cited by 83PDFScholar
2022

SGM3D: Stereo Guided Monocular 3D Object Detection

RA-L 2022

Monocular 3D object detection aims to predict the object location, dimension and orientation in 3D space alongside the object category given only a monocular image. It poses a great challenge due to its ill-posed property, which is a critical lack of depth information in the 2D image plane. While ex

Cited by 39SourcecodeScholar
2022

Text Gestalt: Stroke-Aware Scene Text Image Super-resolution

AAAI 2022technical

In the last decade, the blossom of deep learning has witnessed the rapid development of scene text recognition. However, the recognition of low-resolution scene text images remains a challenge. Even though some super-resolution methods have been proposed to tackle this problem, they usually treat te…

2022

Unsupervised Learning of Compositional Scene Representations from Multiple Unspecified Viewpoints

AAAI 2022technical

Visual scenes are extremely rich in diversity, not only because there are infinite combinations of objects and background, but also because the observations of the same scene may vary greatly with the change of viewpoints. When observing a visual scene that contains multiple objects from multiple vi…

2021

Delving into Data: Effectively Substitute Training for Black-box Attack

CVPR 2021poster

Deep models have shown their vulnerability when processing adversarial samples. As for the black-box attack, without access to the architecture and weights of the attacked model, training a substitute model for adversarial attacks has attracted wide attention. Previous substitute training approaches…

Cited by 90PDFScholar
2021

Depth-Conditioned Dynamic Message Propagation for Monocular 3D Object Detection

CVPR 2021poster

The objective of this paper is to learn context- and depth-aware feature representation to solve the problem of monocular 3D object detection. We make following contributions: (i) rather than appealing to the complicated pseudo-LiDAR based approach, we propose a depth-conditioned dynamic message pro…

Cited by 156PDFcodeScholar
2021

Learning Compositional Representation for 4D Captures With Neural ODE

CVPR 2021poster

Learning based representation has become the key to the success of many computer vision systems. While many 3D representations have been proposed, it is still an unaddressed problem how to represent a dynamically changing 3D object. In this paper, we introduce a compositional representation for 4D c…

Cited by 33PDFScholar
2021

Learning Dynamic Alignment via Meta-Filter for Few-Shot Learning

CVPR 2021poster

Few-shot learning (FSL), which aims to recognise new classes by adapting the learned knowledge with extremely limited few-shot (support) examples, remains an important open problem in computer vision. Most of the existing methods for feature alignment in few-shot learning only consider image-level o…

Cited by 150PDFScholar
2021

Progressive Coordinate Transforms for Monocular 3D Object Detection

NeurIPS 2021poster

Recognizing and localizing objects in the 3D space is a crucial ability for an AI agent to perceive its surrounding environment. While significant progress has been achieved with expensive LiDAR point clouds, it poses a great challenge for 3D object detection given only a monocular image. While ther…

2021

The Devil Is in the Task: Exploiting Reciprocal Appearance-Localization Features for Monocular 3D Object Detection

ICCV 2021poster

Low-cost monocular 3D object detection plays a fundamental role in autonomous driving, whereas its accuracy is still far from satisfactory. Our objective is to dig into the 3D object detection task and reformulate it as the sub-tasks of object localization and appearance perception, which benefits t…

Cited by 58PDFScholar
2021

The Image Local Autoregressive Transformer

NeurIPS 2021poster

Recently, AutoRegressive (AR) models for the whole image generation empowered by transformers have achieved comparable or even better performance compared to Generative Adversarial Networks (GANs). Unfortunately, directly applying such AR models to edit/change local image regions, may suffer from th…

Cited by 13SourcePDFScholar
2021

Zero-Shot Chinese Character Recognition with Stroke-Level Decomposition

IJCAI 2021poster

Chinese character recognition has attracted much research interest due to its wide applications. Although it has been studied for many years, some issues in this field have not been completely resolved yet, \textit{e.g.} the zero-shot problem. Previous character-based and radical-based methods have…

2020

3DCFS: Fast and Robust Joint 3D Semantic-Instance Segmentation via Coupled Feature Selection

ICRA 2020poster

We propose a novel fast and robust 3D point clouds segmentation framework via coupled feature selection, named 3DCFS, that jointly performs semantic and instance segmentation. Inspired by the human scene perception process, we design a novel coupled feature selection module, named CFSM, that adaptiv…

Cited by 16SourcecodeScholar
2020

DeepSFM: Structure From Motion Via Deep Bundle Adjustment

ECCV 2020poster

Structure from motion (SfM) is an essential computer vision problem which has not been well handled by deep learning. One of the promising trends is to apply explicit structural constraint, e.g. 3D cost volume, into the network. However, existing methods usually assume accurate camera poses either f…

Cited by 127SourcePDFScholar
2020

FM2u-Net: Face Morphological Multi-Branch Network for Makeup-Invariant Face Verification

CVPR 2020poster

It is challenging in learning a makeup-invariant face verification model, due to (1) insufficient makeup/non-makeup face training pairs, (2) the lack of diverse makeup faces, and (3) the significant appearance changes caused by cosmetics. To address these challenges, we propose a unified Face Morpho…

Cited by 23PDFcodeScholar
2020

Is normalization indispensable for training deep neural network?

NeurIPS 2020oral

Normalization operations are widely used to train deep neural networks, and they can improve both convergence and generalization in most tasks. The theories for normalization's effectiveness and new forms of normalization have always been hot topics in research. To better understand normalization, o…

2020

Neural Pose Transfer by Spatially Adaptive Instance Normalization

CVPR 2020poster

Pose transfer has been studied for decades, in which the pose of a source mesh is applied to a target mesh. Particularly in this paper, we are interested in transferring the pose of source human mesh to deform the target human mesh, while the source and target meshes may have different identity info…

Cited by 71PDFcodeScholar
2020

Sketch-BERT: Learning Sketch Bidirectional Encoder Representation From Transformers by Self-Supervised Learning of Sketch Gestalt

CVPR 2020poster

Previous researches of sketches often considered sketches in pixel format and leveraged CNN based models in the sketch understanding. Fundamentally, a sketch is stored as a sequence of data points, a vector format representation, rather than the photo-realistic image of pixels. SketchRNN studied a g…

Cited by 80PDFScholar
2020

Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models

ICLR 2020spotlight

The impressive performance of neural networks on natural language processing tasks attributes to their ability to model complicated word and phrase compositions. To explain how the model handles semantic compositions, we study hierarchical explanation of neural network predictions. We identify non-a…

Cited by 128SourceScholar
2019

Generative Modeling of Infinite Occluded Objects for Compositional Scene Representation

ICML 2019oral

We present a deep generative model which explicitly models object occlusions for compositional scene representation. Latent representations of objects are disentangled into location, size, shape, and appearance, and the visual scene can be generated compositionally by integrating these representatio…

Cited by 23SourcePDFScholar
2019

SSF-DAN: Separated Semantic Feature Based Domain Adaptation Network for Semantic Segmentation

ICCV 2019poster

Despite the great success achieved by supervised fully convolutional models in semantic segmentation, training the models requires a large amount of labor-intensive work to generate pixel-level annotations. Recent works exploit synthetic data to train the model for semantic segmentation, but the dom…

Cited by 208PDFScholar
2019

Towards Instance-Level Image-To-Image Translation

CVPR 2019poster

Unpaired Image-to-image Translation is a new rising and challenging vision problem that aims to learn a mapping between unaligned image pairs in diverse domains. Recent advances in this field like MUNIT and DRIT mainly focus on disentangling content and style/attribute from a given image first, then…

Cited by 130PDFcodeScholar
2018

ExFuse: Enhancing Feature Fusion for Semantic Segmentation

ECCV 2018poster

Modern semantic segmentation frameworks usually combine low-level and high-level features from pre-trained backbone convolutional models to boost performance. In this paper, we first point out that a simple fusion of low-level and high-level features could be less effective because of the gap in sem…

Cited by 665SourcePDFScholar
2018

Pose-Normalized Image Generation for Person Re-identification

ECCV 2018poster

Person Re-identification (re-id) faces two major challenges: the lack of cross-view paired training data and learning discriminative identity-sensitive and view-invariant features in the presence of large pose variations. In this work, we address both problems by proposing a novel deep person image…

2017

DSOD: Learning Deeply Supervised Object Detectors From Scratch

ICCV 2017poster

We present Deeply Supervised Object Detector (DSOD), a framework that can learn object detectors from scratch. State-of-the-art object objectors rely heavily on the off-the-shelf networks pre-trained on large-scale classification datasets like ImageNet, which incurs learning bias due to the differen…

Cited by 820PDFcodeScholar
2017

Multi-Scale Deep Learning Architectures for Person Re-Identification

ICCV 2017poster

Person Re-identification (re-id) aims to match people across non-overlapping camera views in a public space. It is a challenging problem because many people captured in surveillance videos wear similar clothes. Consequently, the differences in their appearance are often subtle and only detectable at…

Cited by 376PDFScholar
2017

Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach

ICCV 2017poster

In this paper, we study the task of 3D human pose estimation in the wild. This task is challenging due to lack of training data, as existing datasets are either in the wild images with 2D pose or in the lab images with 3D pose. We propose a weakly-supervised transfer learning method that uses mixed…

Cited by 749PDFcodeScholar
2015

Multiple Granularity Descriptors for Fine-Grained Categorization

ICCV 2015poster

Fine-grained categorization, which aims to distinguish subordinate-level categories such as bird species or dog breeds, is an extremely challenging task. This is due to two main issues: how to localize discriminative regions for recognition and how to learn sophisticated features for representation.…

Cited by 286PDFScholar