← Search

Nanning Zheng

126 accepted papers

2026

Decoupling Primitive with Experts: Dynamic Feature Alignment for Compositional Zero-Shot Learning

ICLR 2026poster

Compositional Zero-Shot Learning (CZSL) investigates compositional generalization capacity to recognize unknown state-object pairs based on learned primitive concepts. Existing CZSL methods typically derive primitives features through a simple composition-prototype mapping, which is suboptimal for a…

Cited by 0SourceScholar
2026

EVOKE: Efficient and High-Fidelity EEG-to-Video Reconstruction via Decoupling Implicit Neural Representation

AAAI 2026technical

Visual neural decoding is an important research topic at the intersection of cognitive neuroscience and machine learning. While recent progress has been made in EEG-based neural decoding, reconstructing dynamic visual content remains challenging. In the field of EEG decoding, current models either u

Cited by 0SourcePDFScholar
2026

More Natural, More Real: Object-aware Gaussian Splatting for 3D Visual Decoding from Human Brain

CVPR 2026

Exploring human visual perception and understanding of the stereoscopic world represents a significant topic in computational neuroscience. Recent studies have provided rich Brain-3D datasets, conducted preliminary explorations into 3D visual reconstruction. However, existing research struggles to c

Cited by 0SourceScholar
2026

RefDiffMap: Diffusion-Guided Progressive Refinement for Vectorized HD Map Construction

RA-L 2026

High-definition (HD) map learning serves as an essential component of autonomous driving scene understanding, providing structured priors for planning and prediction. Recent transformer-based methods regress vectorized map elements via deformable attention over Bird's-Eye View (BEV) features. They t

Cited by 0SourceScholar
2026

RefDiffMap: Diffusion-Guided Progressive Refinement for Vectorized HD Map Construction

ICRA 2026poster

High-definition (HD) map learning serves as an essential component of autonomous driving scene understanding, providing structured priors for planning and prediction. Recent transformer-based methods regress vectorized map elements via deformable attention over Bird’s-Eye View (BEV) features. They t…

Cited by 0SourceScholar
2026

Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion

CVPR 2026

Latent Diffusion Models (LDMs) inherently follow a coarse-to-fine generation process, where high-level semantic structure is generated slightly earlier than fine-grained texture. This indicates the preceding semantics potentially benefit the texture generation by providing a semantic anchor. Recent

Cited by 0SourcecodeScholar
2026

UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space

AAAI 2026technical

In the field of human-object interaction (HOI), detection and generation are two dual tasks that have traditionally been addressed separately, hindering the development of comprehensive interaction understanding. To address this, we propose UniHOI, which jointly models HOI detection and generation v

Cited by 0SourcePDFScholar
2026

Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models

CVPR 2026

The performance of Latent Diffusion Models (LDMs) is critically dependent on the quality of their visual tokenizers. While recent works have explored incorporating Vision Foundation Models (VFMs) into the tokenizers training via distillation, we empirically find this approach inevitably weakens the

Cited by 0SourcecodeScholar
2025

Beyond Brain Decoding: Visual-Semantic Reconstructions to Mental Creation Extension Based on fMRI

ICCV 2025poster

Decoding visual information from fMRI signals is an important pathway to understand how the brain represents the world, and is a cutting-edge field of artificial general intelligence. Decoding fMRI should not be limited to reconstructing visual stimuli, but also further transforming them into descri…

Cited by 0SourcePDFScholar
2025

Beyond Image Classification: A Video Benchmark and Dual-Branch Hybrid Discrimination Framework for Compositional Zero-Shot Learning

CVPR 2025poster

Human reasoning naturally combines concepts to identify unseen compositions, a capability that Compositional Zero-Shot Learning (CZSL) aims to replicate in machine learning models. However, we observe that focusing solely on typical image classification tasks in CZSL may limit models' compositional…

Cited by 0SourcePDFScholar
2025

Beyond Single-Modal Boundary: Cross-Modal Anomaly Detection through Visual Prototype and Harmonization

CVPR 2025poster

Anomaly detection is a significant task for its application and research value. While existing methods have made impressive progress within the same modality, cross-modal anomaly detection remains an open and challenging problem. In this paper, we propose a cross-modal anomaly detection model that i…

2025

DAMap: Distance-aware MapNet for High Quality HD Map Construction

ICCV 2025poster

High-definition (HD) map is an important component to support navigation and planning for autonomous driving vehicles. Predicting map elements with high quality (high classification and localization scores) is crucial to the safety of autonomous driving vehicles. However, current methods perform poo…

Cited by 0SourcePDFScholar
2025

ForestLPR: LiDAR Place Recognition in Forests Attentioning Multiple BEV Density Images

CVPR 2025highlight

Place recognition is essential to maintain global consistency in large-scale localization systems. While research in urban environments has progressed significantly using LiDARs or cameras, applications in natural forest-like environments remain largely underexplored. Furthermore, forests present pa…

2025

FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection Transformers

ICCV 2025poster

Detecting 3D objects accurately from multi-view 2D images is a challenging yet essential task in the field of autonomous driving. Current methods resort to integrating depth prediction to recover the spatial information for object query decoding, which necessitates explicit supervision from LiDAR po…

Cited by 0SourcePDFScholar
2025

Modeling Human-like Driving Behavior Based on Maximum Entropy Deep Inverse Reinforcement Learning

IROS 2025

Modeling expert driving behavior is crucial for the successful implementation of human-like autonomous driving. In this paper, we propose a new sampling-based Maximum Entropy Deep Inverse Reinforcement Learning (MEDIRL) framework. It leverages naturalistic human driving data to train the reward mode

Cited by 1SourceScholar
2025

PlaneHEC: Efficient Hand-Eye Calibration for Multi-View Robotic Arm via Any Point Cloud Plane Detection

ICRA 2025

Hand-eye calibration is an important task in vision-guided robotic systems and is crucial for determining the transformation matrix between the camera coordinate system and the robot end-effector. Existing methods, for multi-view robotic systems, usually rely on accurate geometric models or manual a

Cited by 0SourceScholar
2025

Refiner: Fine-grained Cross-modal Concepts Refinement for Compositional Zero-Shot Learning

ICASSP 2025accepted

Recent Compositional Zero-Shot Learning (CZSL) methods increasingly adopt the pre-trained vision-language models to capture the contextual relations between image and text spaces. However, the single-class-token design from Transformer-based encoder inevitably captures contextual information from un…

Cited by 0SourceScholar
2025

SAMap: Semantic Alignment for HD Map Detection Domain Generalization Under Varying Weather and Lighting

IROS 2025

High-definition (HD) maps are crucial for autonomous driving systems. Despite recent advances in learning-based HD map prediction methods, these approaches experience significant performance degradation when encountering unseen weather or lighting conditions due to feature distribution discrepancies

Cited by 0SourceScholar
2025

See Through Their Minds: Learning Transferable Brain Decoding Models from Cross-Subject fMRI

AAAI 2025technical

Deciphering visual content from fMRI sheds light on the human vision system, but data scarcity and noise limit brain decoding model performance. Traditional approaches rely on subject-specific models, which are sensitive to training sample size. In this paper, we address data scarcity by proposing s…

2025

Towards Extrinsic Dexterity Grasping in Unrestricted Environments

IROS 2025

Grasping large and flat objects (e.g., a book or a pan) is often regarded as an ungraspable task, which poses significant challenges due to the unreachable grasping poses. Prior research has exploited environmental interactions through Extrinsic Dexterity, utilizing external structures such as walls

Cited by 0SourcecodeScholar
2025

Unveiling Multi-View Anomaly Detection: Intra-view Decoupling and Inter-view Fusion

AAAI 2025technical

Anomaly detection has garnered significant attention for its extensive industrial application value. Most existing methods focus on single-view scenarios and fail to detect anomalies hidden in blind spots, leaving a gap in addressing the demands of multi-view detection in practical applications. Ens…

2025

VideoVLA: Video Generators Can Be Generalizable Robot Manipulators

NeurIPS 2025poster

Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained understanding models for perception and instruction following, their abi…

Cited by 0SourceScholar
2024

Breaking through the learning plateaus of in-context learning in Transformer

ICML 2024poster

In-context learning, i.e., learning from context examples, is an impressive ability of Transformer. Training Transformers to possess this in-context learning skill is computationally intensive due to the occurrence of *learning plateaus*, which are periods within the training process where there is…

Cited by 1SourcePDFScholar
2024

Can LLMs Learn From Mistakes? An Empirical Study on Reasoning Tasks

EMNLP 2024finding

Towards enhancing the chain-of-thought (CoT) reasoning of large language models (LLMs), much existing work has revealed the effectiveness of straightforward learning on annotated/generated CoT paths. However, there is less evidence yet that reasoning capabilities can be enhanced through a reverse le…

2024

Complementing Onboard Sensors with Satellite Maps: A New Perspective for HD Map Construction

ICRA 2024poster

High-definition (HD) maps play a crucial role in autonomous driving systems. Recent methods have attempted to construct HD maps in real-time using vehicle onboard sensors. Due to the inherent limitations of onboard sensors, which include sensitivity to detection range and susceptibility to occlusion…

Cited by 18SourcecodeScholar
2024

Diffusion Model with Cross Attention as an Inductive Bias for Disentanglement

NeurIPS 2024spotlight

Disentangled representation learning strives to extract the intrinsic factors within the observed data. Factoring these representations in an unsupervised manner is notably challenging and usually requires tailored loss functions or specific structural designs. In this paper, we introduce a new pers…

Cited by 6SourcePDFScholar
2024

GSO-Net: Grid Surface Optimization via Learning Geometric Constraints

AAAI 2024technical

In the context of surface representations, we find a natural structural similarity between grid surface and image data. Motivated by this inspiration, we propose a novel approach: encoding grid surfaces as geometric images and using image processing methods to address surface optimization-related pr…

2024

IS-DARTS: Stabilizing DARTS through Precise Measurement on Candidate Importance

AAAI 2024technical

Among existing Neural Architecture Search methods, DARTS is known for its efficiency and simplicity. This approach applies continuous relaxation of network representation to construct a weight-sharing supernet and enables the identification of excellent subnets in just a few GPU days. However, perfo…

2024

Make Your LLM Fully Utilize the Context

NeurIPS 2024poster

While many contemporary large language models (LLMs) can process lengthy input, they still struggle to fully utilize information within the long context, known as the *lost-in-the-middle* challenge. We hypothesize that it stems from insufficient explicit supervision during the long-context training,…

2024

Molecule Design by Latent Prompt Transformer

NeurIPS 2024spotlight

This work explores the challenging problem of molecule design by framing it as a conditional generative modeling task, where target biological properties or desired chemical constraints serve as conditioning variables. We propose the Latent Prompt Transformer (LPT), a novel generative model comprisi…

Cited by 2SourcePDFScholar
2024

Neural P$^3$M: A Long-Range Interaction Modeling Enhancer for Geometric GNNs

NeurIPS 2024poster

Geometric graph neural networks (GNNs) have emerged as powerful tools for modeling molecular geometry. However, they encounter limitations in effectively capturing long-range interactions in large molecular systems. To address this challenge, we introduce **Neural P$^3$M**, a versatile enhancer of g…

2024

PMT: Progressive Mean Teacher via Exploring Temporal Consistency for Semi-Supervised Medical Image Segmentation

ECCV 2024poster

"Semi-supervised learning has emerged as a widely adopted technique in the field of medical image segmentation. The existing works either focuses on the construction of consistency constraints or the generation of pseudo labels to provide high-quality supervisory signals, whose main challenge mainly…

2024

POAQL: A Partially Observable Altruistic Q-Learning Method for Cooperative Multi-Agent Reinforcement Learning

ICRA 2024poster

Multi-Agent Path Finding (MAPF) is an important issue in multi-agent cooperation. Many studies apply MultiAgent Reinforcement Learning (MARL) to solve MAPF in partially observable settings. The objective of cooperative MARL is to maximize the cumulative team reward. Nevertheless, in partially observ…

Cited by 2SourceScholar
2024

Self-Consistency Training for Density-Functional-Theory Hamiltonian Prediction

ICML 2024poster

Predicting the mean-field Hamiltonian matrix in density functional theory is a fundamental formulation to leverage machine learning for solving molecular science problems. Yet, its applicability is limited by insufficient labeled data for training. In this work, we highlight that Hamiltonian predict…

Cited by 5SourcePDFScholar
2024

Symmetric Consistency with Cross-Domain Mixup for Cross-Modality Cardiac Segmentation

ICASSP 2024accepted

Accurate cardiac segmentation in cross-modality images plays an important role in the quantitative analysis of the heart to diagnose cardiovascular diseases. However, achieving high performance in cross-modality segmentation is hindered by the time-consuming annotation and modality gap. While some a…

Cited by 0SourceScholar
2024

TPR: Topology-Preserving Reservoirs for Generalized Zero-Shot Learning

NeurIPS 2024poster

Pre-trained vision-language models (VLMs) such as CLIP have shown excellent performance for zero-shot classification. Based on CLIP, recent methods design various learnable prompts to evaluate the zero-shot generalization capability on a base-to-novel setting. This setting assumes test samples are a…

Cited by 0SourcePDFScholar
2024

Task-Driven Autonomous Driving: Balanced Strategies Integrating Curriculum Reinforcement Learning and Residual Policy

RA-L 2024

Achieving fully autonomous driving in urban traffic scenarios is a significant challenge that necessitates balancing safety, efficiency, and compliance with traffic regulations. In this letter, we introduce a novel Curriculum Residual Hierarchical Reinforcement Learning (CR-HRL) framework. It integr

Cited by 6SourceScholar
2024

Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis

CVPR 2024poster

Significant progress has been made in scene text detection models since the rise of deep learning but scene text layout analysis which aims to group detected text instances as paragraphs has not kept pace. Previous works either treated text detection and grouping using separate models or train a mod…

Cited by 1SourcePDFScholar
2024

V-DETR: DETR with Vertex Relative Position Encoding for 3D Object Detection

ICLR 2024poster

We introduce a highly performant 3D object detector for point clouds using the DETR framework. The prior attempts all end up with suboptimal results because they fail to learn accurate inductive biases from the limited scale of training data. In particular, the queries often attend to points that ar…

2024

Voxel or Pillar: Exploring Efficient Point Cloud Representation for 3D Object Detection

AAAI 2024technical

Efficient representation of point clouds is fundamental for LiDAR-based 3D object detection. While recent grid-based detectors often encode point clouds into either voxels or pillars, the distinctions between these approaches remain underexplored. In this paper, we quantify the differences between t…

Cited by 8SourcePDFScholar
2023

Closing the gap between the upper bound and lower bound of Adam's iteration complexity

NeurIPS 2023poster

Recently, Arjevani et al. [1] establish a lower bound of iteration complexity for the first-order optimization under an $L$-smooth condition and a bounded noise variance assumption. However, a thorough review of existing literature on Adam's convergence reveals a noticeable gap: none of them meet…

Cited by 24SourcePDFScholar
2023

Controllable Clothoid Path Generation for Autonomous Vehicles

RA-L 2023

This letter proposes a novel and simple smooth path generation algorithm for autonomous vehicles. The proposed method can rapidly generate feasible and curvature continuous paths connecting any two given states with null curvature. The generated path comprises straight lines, circular arcs and cloth

Cited by 7SourceScholar
2023

DBQ-SSD: Dynamic Ball Query for Efficient 3D Object Detection

ICLR 2023poster

Many point-based 3D detectors adopt point-feature sampling strategies to drop some points for efficient inference. These strategies are typically based on fixed and handcrafted rules, making it difficult to handle complicated scenes. Different from them, we propose a Dynamic Ball Query (DBQ) network…

2023

DETR Does Not Need Multi-Scale or Locality Design

ICCV 2023poster

This paper presents an improved DETR detector that maintains a "plain" nature: using a single-scale feature map and global cross-attention calculations without specific locality constraints, in contrast to previous leading DETR-based detectors that reintroduce architectural inductive biases of multi…

Cited by 30PDFcodeScholar
2023

DisDiff: Unsupervised Disentanglement of Diffusion Probabilistic Models

NeurIPS 2023poster

Targeting to understand the underlying explainable factors behind observations and modeling the conditional generation process on these factors, we connect disentangled representation learning to diffusion probabilistic models (DPMs) to take advantage of the remarkable modeling ability of DPMs. We p…

2023

Does Deep Learning Learn to Abstract? A Systematic Probing Framework

ICLR 2023poster

Abstraction is a desirable capability for deep learning models, which means to induce abstract concepts from concrete instances and flexibly apply them beyond the learning context. At the same time, there is a lack of clear understanding about both the presence and further characteristics of this ca…

2023

Efficient Safety-Enhanced Velocity Planning for Autonomous Driving With Chance Constraints

RA-L 2023

Velocity planning is an important module of autonomous driving, which aims to generate the velocity profile given a reference path. However, most existing algorithms fail to adequately address the uncertainty inherent in driving contexts, leading to potentially risky situations. To this end, we prop

Cited by 15SourceScholar
2023

Geometric Transformer with Interatomic Positional Encoding

NeurIPS 2023poster

The widespread adoption of Transformer architectures in various data modalities has opened new avenues for the applications in molecular modeling. Nevertheless, it remains elusive that whether the Transformer-based architecture can do molecular modeling as good as equivariant GNNs. In this paper,…

2023

How Do In-Context Examples Affect Compositional Generalization?

ACL 2023long

Compositional generalization–understanding unseen combinations of seen primitives–is an essential reasoning capability in human intelligence. The AI community mainly studies this capability by fine-tuning neural networks on lots of training samples, while it is still unclear whether and how in-conte…

2023

InteractionNet: Joint Planning and Prediction for Autonomous Driving with Transformers

IROS 2023poster

Planning and prediction are two important modules of autonomous driving and have experienced tremendous advancement recently. Nevertheless, most existing methods regard planning and prediction as independent and ignore the correlation between them, leading to the lack of consideration for interactio…

Cited by 6SourcecodeScholar
2023

Learning Trajectories are Generalization Indicators

NeurIPS 2023poster

This paper explores the connection between learning trajectories of Deep Neural Networks (DNNs) and their generalization capabilities when optimized using (stochastic) gradient descent algorithms. Instead of concentrating solely on the generalization error of the DNN post-training, we present a nov…

Cited by 4SourcePDFScholar
2023

Long-Term Dynamic Window Approach for Kinodynamic Local Planning in Static and Crowd Environments

RA-L 2023

Local planning for a differential wheeled robot is designed to generate kinodynamic feasible actions that guide the robot to a goal position along the navigation path while avoiding obstacles. Reactive, predictive, and learning-based methods are widely used in local planning. However, few of them ca

Cited by 15SourcecodeScholar
2023

MMRDN: Consistent Representation for Multi-View Manipulation Relationship Detection in Object-Stacked Scenes

ICRA 2023poster

Manipulation relationship detection (MRD) aims to guide the robot to grasp objects in the right order, which is important to ensure the safety and reliability of grasping in object stacked scenes. Previous works infer manipulation relationship by deep neural network trained with data collected from…

Cited by 2SourceScholar
2023

MixPHM: Redundancy-Aware Parameter-Efficient Tuning for Low-Resource Visual Question Answering

CVPR 2023poster

Recently, finetuning pretrained vision-language models (VLMs) has been a prevailing paradigm for achieving state-of-the-art performance in VQA. However, as VLMs scale, it becomes computationally expensive, storage inefficient, and prone to overfitting when tuning full model parameters for a specific…

2023

Prioritized Planning for Target-Oriented Manipulation via Hierarchical Stacking Relationship Prediction

IROS 2023poster

In scenarios involving grasping multiple targets, the learning of stacking relationships between objects is fundamental for robots to execute safely and efficiently. However, current methods lack subdivision for the hierarchy of stacking relationship types. In scenes where objects are mostly stacked…

Cited by 5SourceScholar
2023

Quadric Representations for LiDAR Odometry, Mapping and Localization

RA-L 2023

Current LiDAR odometry, mapping and localization methods leverage point-wise representations of 3D scenes and achieve high accuracy in autonomous driving tasks. However, the space-inefficiency of methods that use point-wise representations limits their development and usage in practical applications

Cited by 13SourceScholar
2023

Skill-Based Few-Shot Selection for In-Context Learning

EMNLP 2023long main

*In-context learning* is the paradigm that adapts large language models to downstream tasks by providing a few examples. *Few-shot selection*---selecting appropriate examples for each test instance separately---is important for in-context learning. In this paper, we propose **Skill-KNN**, a skill-ba…

Cited by 0SourceScholar
2023

StructVPR: Distill Structural Knowledge With Weighting Samples for Visual Place Recognition

CVPR 2023poster

Visual place recognition (VPR) is usually considered as a specific image retrieval problem. Limited by existing training frameworks, most deep learning-based works cannot extract sufficiently stable global features from RGB images and rely on a time-consuming re-ranking step to exploit spatial struc…

Cited by 24SourcePDFScholar
2022

A Continuous Learning Approach for Probabilistic Human Motion Prediction

ICRA 2022poster

Human Motion Prediction (HMP) plays a crucial role in safe Human-Robot-Interaction (HRI). Currently, the majority of HMP algorithms are trained by massive pre-collected data. As the training data only contains a few pre-defined motion patterns, these methods cannot handle the unfamiliar motion patte…

Cited by 3SourceScholar
2022

Asymmetric Relation Consistency Reasoning for Video Relation Grounding

ECCV 2022poster

"Video relation grounding has attracted growing attention in the fields of video understanding and multimodal learning. While the past years have witnessed remarkable progress in this issue, the difficulties of multi-instance and complex temporal reasoning make it still a challenging task. In this p…

Cited by 5SourcePDFScholar
2022

Construct Effective Geometry Aware Feature Pyramid Network for Multi-Scale Object Detection

AAAI 2022technical

Feature Pyramid Network (FPN) has been widely adopted to exploit multi-scale features for scale variation in object detection. However, intrinsic defects in most of the current methods with FPN make it difficult to adapt to the feature of different geometric objects. To address this issue, we introd…

Cited by 7SourcePDFScholar
2022

Could Giant Pre-trained Image Models Extract Universal Representations?

NeurIPS 2022accept

Frozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few parameters available for adapting to downstream tasks, which is problematic in computer vision where tasks vary significan…

Cited by 11SourcePDFScholar
2022

Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning

ICML 2022spotlight

Due to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the corresp…

Cited by 15SourcePDFScholar
2022

LGD: Label-Guided Self-Distillation for Object Detection

AAAI 2022technical

In this paper, we propose the first self-distillation framework for general object detection, termed LGD (Label-Guided self-Distillation). Previous studies rely on a strong pretrained teacher to provide instructive knowledge that could be unavailable in real-world scenarios. Instead, we generate an…

2022

Learning Disentangled Classification and Localization Representations for Temporal Action Localization

AAAI 2022technical

A common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that thi…

Cited by 20SourcePDFScholar
2022

Learning To Refactor Action and Co-Occurrence Features for Temporal Action Localization

CVPR 2022poster

The main challenge of Temporal Action Localization is to retrieve subtle human actions from various co-occurring ingredients, e.g., context and background, in an untrimmed video. While prior approaches have achieved substantial progress through devising advanced action detectors, they still suffer f…

Cited by 57PDFScholar
2022

Parametric Path Optimization for Wheeled Robots Navigation

ICRA 2022poster

Collision risk and smoothness are the most important factors in global path planning. Currently, planning methods that reduce global path collision risk and improve its smoothness through numerical optimization have achieved good results. However, these methods cannot always optimize the path. The r…

Cited by 3SourceScholar
2022

Pedestrian Intention Prediction Based on Traffic-Aware Scene Graph Model

IROS 2022poster

Anticipating the future behavior of pedestrians is a crucial part of deploying Automated Driving Systems (ADS) in urban traffic scenarios. Most recent works utilize a convolutional neural network (CNN) to extract visual information, which is then input to a recurrent neural network (RNN) along with…

Cited by 10SourceScholar
2022

REGRAD: A Large-Scale Relational Grasp Dataset for Safe and Object-Specific Robotic Grasping in Clutter

RA-L 2022

Despite the impressive progress achieved in robotic grasping, robots are not skilled in sophisticated tasks (e.g. search and grasp a specified target in clutter). Such tasks involve not only grasping but the comprehensive perception of the world (e.g. the object relationships). Recently, encouraging

Cited by 51SourcecodeScholar
2022

SE(3) Equivariant Graph Neural Networks with Complete Local Frames

ICML 2022spotlight

Group equivariance (e.g. SE(3) equivariance) is a critical physical symmetry in science, from classical and quantum physics to computational biology. It enables robust and accurate prediction under arbitrary reference transformations. In light of this, great efforts have been put on encoding this sy…

2022

Social Interpretable Tree for Pedestrian Trajectory Prediction

AAAI 2022technical

Understanding the multiple socially-acceptable future behaviors is an essential task for many vision applications. In this paper, we propose a tree-based method, termed as Social Interpretable Tree (SIT), to address this multi-modal prediction task, where a hand-crafted tree is built depending on th…

2022

TCL: Tightly Coupled Learning Strategy for Weakly Supervised Hierarchical Place Recognition

RA-L 2022

Visual place recognition (VPR) is a key issue for robotics and autonomous systems. For the trade-off between time and performance, most of methods use the coarse-to-fine hierarchical architecture, which consists of retrieving top-N candidates using global features, and re-ranking top-N with local fe

Cited by 10SourceScholar
2022

Towards Building A Group-based Unsupervised Representation Disentanglement Framework

ICLR 2022poster

Disentangled representation learning is one of the major goals of deep learning, and is a key step for achieving explainable and generalizable models. The key idea of the state-of-the-art VAE-based unsupervised representation disentanglement methods is to minimize the total correlation of the joint…

2022

TransVPR: Transformer-Based Place Recognition With Multi-Level Attention Aggregation

CVPR 2022oral

Visual place recognition is a challenging task for applications such as autonomous driving navigation and mobile robot localization. Distracting elements presenting in complex scenes often lead to deviations in the perception of visual place. To address this problem, it is crucial to integrate infor…

Cited by 168PDFScholar
2021

A Global-Local Coupling Two-Stage Path Planning Method for Mobile Robots

RA-L 2021

The path planning of mobile robots is an optimization problem that is difficult to solve directly owing to its nonlinear characteristics. This letter proposes the “global-local” Coupling Two-Stage Path Planning (CTSP) method. First, the globally optimal solution in the configuration space is given b

Cited by 53SourceScholar
2021

ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action Localization

AAAI 2021technical

The object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foregroun…

Cited by 86SourcePDFScholar
2021

Co-evolution Transformer for Protein Contact Prediction

NeurIPS 2021poster

Proteins are the main machinery of life and protein functions are largely determined by their 3D structures. The measurement of the pairwise proximity between amino acids of a protein, known as inter-residue contact map, well characterizes the structural information of a protein. Protein contact pre…

2021

Correlation-Based Robust Linear Regression with Iterative Outlier Removal

ICASSP 2021accepted

Here we consider linear regression from the view of correlation and propose a robust regression algorithm. The main idea of this work is from the fact that the inliers lying in a low dimensional subspace are mostly correlated, and the presence of outliers leads to the decrease of correlation. We des…

Cited by 0SourceScholar
2021

Dynamic Grained Encoder for Vision Transformers

NeurIPS 2021poster

Transformers, the de-facto standard for language modeling, have been recently applied for vision tasks. This paper introduces sparse queries for vision transformers to exploit the intrinsic spatial redundancy of natural images and save computational costs. Specifically, we propose a Dynamic Grained…

2021

End-to-End Object Detection With Fully Convolutional Network

CVPR 2021poster

Mainstream object detectors based on the fully convolutional network has achieved impressive performance. While most of them still need a hand-designed non-maximum suppression (NMS) post-processing, which impedes fully end-to-end training. In this paper, we give the analysis of discarding NMS, where…

Cited by 269PDFcodeScholar
2021

Enriching Local and Global Contexts for Temporal Action Localization

ICCV 2021poster

Effectively tackling the problem of temporal action localization (TAL) necessitates a visual representation that jointly pursues two confounding goals, i.e., fine-grained discrimination for temporal localization and sufficient visual invariance for action classification. We address this challenge by…

Cited by 148PDFcodeScholar
2021

Instance-Conditional Knowledge Distillation for Object Detection

NeurIPS 2021poster

Knowledge distillation has shown great success in classification, however, it is still challenging for detection. In a typical image for detection, representations from different locations may have different contributions to detection targets, making the distillation hard to balance. In this paper,…

2021

Meta Pairwise Relationship Distillation for Unsupervised Person Re-Identification

ICCV 2021poster

Unsupervised person re-identification (Re-ID) remains challenging due to the lack of ground-truth labels. Existing methods often rely on estimated pseudo labels via iterative clustering and classification, and they are unfortunately highly susceptible to performance penalties incurred by the inaccur…

Cited by 51PDFcodeScholar
2021

Practical Relative Order Attack in Deep Ranking

ICCV 2021poster

Recent studies unveil the vulnerabilities of deep ranking models, where an imperceptible perturbation can trigger dramatic changes in the ranking result. While previous attempts focus on manipulating absolute ranks of certain candidates, the possibility of adjusting their relative order remains unde…

Cited by 21PDFcodeScholar
2021

REGNet: REgion-based Grasp Network for End-to-end Grasp Detection in Point Clouds

ICRA 2021poster

Reliable robotic grasping in unstructured environments is a crucial but challenging task. The main problem is to generate the optimal grasp of novel objects from partial noisy observations. This paper presents an end-to-end grasp detection network taking one single-view point cloud as input to tackl…

Cited by 104SourcecodeScholar
2021

Unlimited Neighborhood Interaction for Heterogeneous Trajectory Prediction

ICCV 2021poster

Understanding complex social interactions among agents is a key challenge for trajectory prediction. Most existing methods consider the interactions between pairwise traffic agents or in a local area, while the nature of interactions is unlimited, involving an uncertain number of agents and non-loca…

Cited by 35PDFcodeScholar
2021

Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and Context

AAAI 2021technical

Weakly-supervised Temporal Action Localization (WS-TAL) methods learn to localize temporal starts and ends of action instances in a video under only video-level supervision. Existing WS-TAL methods rely on deep features learned for action recognition. However, due to the mismatch between classificat…

Cited by 32SourcePDFScholar
2020

A Boundary Based Out-of-Distribution Classifier for Generalized Zero-Shot Learning

ECCV 2020poster

Generalized Zero-Shot Learning (GZSL) is a challenging topic that has promising prospects in many realistic scenarios. Using a gating mechanism that discriminates the unseen samples from the seen samples can decompose the GZSL problem to a conventional Zero-Shot Learning (ZSL) problem and a supervis…

Cited by 107SourcePDFScholar
2020

CoBigICP: Robust and Precise Point Set Registration using Correntropy Metrics and Bidirectional Correspondence

IROS 2020poster

In this paper, we propose a novel probabilistic variant of iterative closest point (ICP) dubbed as CoBigICP. The method leverages both local geometrical information and global noise characteristics. Locally, the 3D structure of both target and source clouds are incorporated into the objective functi…

Cited by 15SourcecodeScholar
2020

Compositional Generalization by Learning Analytical Expressions

NeurIPS 2020spotlight

Compositional generalization is a basic and essential intellective capability of human beings, which allows us to recombine known parts readily. However, existing neural network based models have been proven to be extremely deficient in such a capability. Inspired by work in cognition which argues c…

2020

Fine-Grained Dynamic Head for Object Detection

NeurIPS 2020poster

The Feature Pyramid Network (FPN) presents a remarkable approach to alleviate the scale variance in object representation by performing instance-level assignments. Nevertheless, this strategy ignores the distinct characteristics of different sub-regions in an instance. To this end, we propose a fine…

2020

Rethinking Learnable Tree Filter for Generic Feature Transform

NeurIPS 2020poster

The Learnable Tree Filter presents a remarkable approach to model structure-preserving relations for semantic segmentation. Nevertheless, the intrinsic geometric constraint forces it to focus on the regions with close spatial distance, hindering the effective long-range interactions. To relax the ge…

2020

Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition

CVPR 2020poster

Skeleton-based human action recognition has attracted great interest thanks to the easy accessibility of the human skeleton data. Recently, there is a trend of using very deep feedforward neural networks to model the 3D coordinates of joints without considering the computational efficiency. In this…

Cited by 635PDFcodeScholar
2019

A Multi-task Convolutional Neural Network for Autonomous Robotic Grasping in Object Stacking Scenes

IROS 2019poster

Autonomous robotic grasping plays an important role in intelligent robotics. However, how to help the robot grasp specific objects in object stacking scenes is still an open problem, because there are two main challenges for autonomous robots: (1) it is a comprehensive task to know what and how to g…

Cited by 87SourceScholar
2019

Compressing Unknown Images With Product Quantizer for Efficient Zero-Shot Classification

CVPR 2019poster

For Zero-Shot Learning (ZSL), the Nearest Neighbor (NN) search is generally conducted for classification, which may cause unacceptable computational complexity for large-scale datasets. To compress zero-shot classes by the trained quantizer for efficient search, it tends to induce large quantization…

Cited by 48PDFScholar
2019

Learnable Tree Filter for Structure-preserving Feature Transform

NeurIPS 2019poster

Learning discriminative global features plays a vital role in semantic segmentation. And most of the existing methods adopt stacks of local convolutions or non-local blocks to capture long-range context. However, due to the absence of spatial structure preservation, these operators ignore the object…

2019

Precise Correntropy-based 3D Object Modelling With Geometrical Traffic Prior

IROS 2019poster

Robust 3D perception using LiDAR is of prime importance for robotics, and its fundamental core lies in precise object modelling resisting to noise and outliers. In this paper, a precise 3D object modelling algorithm is designed especially for the intelligent vehicles. The proposed algorithm is advan…

Cited by 1SourceScholar
2019

ROI-based Robotic Grasp Detection for Object Overlapping Scenes

IROS 2019poster

Grasp detection considering the affiliations between grasps and their owner in object overlapping scenes is a necessary and challenging task for the practical use of the robotic grasping approach. In this paper, a robotic grasp detection algorithm named ROI-GD is proposed to provide a feasible solut…

Cited by 213SourceScholar
2019

SR-LSTM: State Refinement for LSTM Towards Pedestrian Trajectory Prediction

CVPR 2019poster

In crowd scenarios, reliable trajectory prediction of pedestrians requires insightful understanding of their social behaviors. These behaviors have been well investigated by plenty of studies, while it is hard to be fully expressed by hand-craft rules. Recent studies based on LSTM networks have show…

Cited by 636PDFScholar
2019

Task-oriented Grasping in Object Stacking Scenes with CRF-based Semantic Model

IROS 2019poster

In task-oriented grasping, the robot is supposed to manipulate the objects in a task-compatible manner, which is more important but more challenging than just stably grasping. However, most of existing works perform task-oriented grasping only in single object scenes. This greatly limits their pract…

Cited by 24SourceScholar
2019

Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation Networks

ICCV 2019poster

Weakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video…

Cited by 135PDFcodeScholar
2018

Adding Attentiveness to the Neurons in Recurrent Neural Networks

ECCV 2018poster

Recurrent neural networks (RNNs) are capable of modeling the temporal dynamics of complex sequential information. However, the structures of existing RNN neurons mainly focus on controlling the contributions of current and historical information but do not explore the different importance levels of…

Cited by 105SourcePDFScholar
2018

Grassmann Pooling as Compact Homogeneous Bilinear Pooling for Fine-Grained Visual Classification

ECCV 2018poster

Designing discriminative and invariant features is the key to visual recognition. Recently, the bilinear pooled feature matrix of Convolutional Neural Network (CNN) has shown to achieve state-of-the-art performance on a range of fine-grained visual recognition tasks. The bilinear feature matrix coll…

Cited by 121SourcePDFScholar
2018

Transductive Semi-Supervised Deep Learning using Min-Max Features

ECCV 2018poster

In this paper, we propose Transductive Semi-Supervised Deep Learning (TSSDL) method that is effective for training Deep Convolutional Neural Network (DCNN) models. The method applies transductive learning principle to DCNN training, introduces confidence levels on unlabeled image samples to overcome…

Cited by 301SourcePDFScholar
2018

Where and Why Are They Looking? Jointly Inferring Human Attention and Intentions in Complex Tasks

CVPR 2018poster

This paper addresses a new problem - jointly inferring human attention, intentions, and tasks from videos. Given an RGB-D video where a human performs a task, we answer three questions simultaneously: 1) where the human is looking - attention prediction; 2) why the human is looking there - intention…

Cited by 82SourcePDFScholar
2017

ER3: A Unified Framework for Event Retrieval, Recognition and Recounting

CVPR 2017poster

We develop a unified framework for complex event retrieval, recognition and recounting. The framework is based on a compact video representation that exploits the temporal correlations in image features. Our feature alignment procedure identifies and removes the feature redundancies across frames an…

Cited by 28PDFScholar
2017

Point to Set Similarity Based Deep Feature Learning for Person Re-Identification

CVPR 2017poster

Person re-identification (Re-ID) remains a challenging problem due to significant appearance changes caused by variations in view angle, background clutter, illumination condition and mutual occlusion. To address these issues, conventional methods usually focus on proposing robust feature represent…

Cited by 191PDFScholar
2017

View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition From Skeleton Data

ICCV 2017poster

Skeleton-based human action recognition has recently attracted increasing attention due to the popularity of 3D skeleton data. One main challenge lies in the large view variations in captured human actions. We propose a novel view adaptation scheme to automatically regulate observation viewpoints du…

Cited by 683PDFcodeScholar
2016

Person Re-Identification by Multi-Channel Parts-Based CNN With Improved Triplet Loss Function

CVPR 2016poster

Person re-identification across cameras remains a very challenging problem, especially when there are no overlapping fields of view between cameras. In this paper, we present a novel multi-channel parts-based convolutional neural network (CNN) model under the triplet framework for person re-identifi…

Cited by 1630PDFScholar
2016

Similarity Learning With Spatial Constraints for Person Re-Identification

CVPR 2016poster

Pose variation remains one of the major factors that adversely affect the accuracy of person re-identification. Such variation is not arbitrary as body parts (e.g. head, torso, legs) have relative stable spatial distribution. Breaking down the variability of global appearance regarding the spatial d…

Cited by 398PDFScholar
2015

Similarity Learning on an Explicit Polynomial Kernel Feature Map for Person Re-Identification

CVPR 2015poster

In this paper, we address the person re-identification problem, discovering the correct matches for a probe person image from a set of gallery person images. We follow the learning-to-rank methodology and learn a similarity function to maximize the difference between the similarity scores of matched…

Cited by 258SourcePDFScholar