← Search

He Wang

151 accepted papers

2026

AVGGT: Rethinking Global Attention for Accelerating VGGT

CVPR 2026

Models such as VGGT and \pi^3 have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-attention variants offer partial speedups, yet lack a systematic analysis of how global attention contributes to multi-view

Cited by 0SourceScholar
2026

Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

ICML 2026poster

Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. How can we incorporate 3D information to enhance VLA capabilities? We conduct a pilot study across different observation spaces and visual representation…

Cited by 0SourceScholar
2026

CLAR: Learning 3D Representations for Robotic Manipulation by Fusing Masked Reconstruction with Multi-Level Contrastive Alignment

ICRA 2026poster

The spatial information inherent in 3D point clouds is crucial for robotic manipulation. However, existing 3D pre-training methods face a fundamental trade-off: Masked Autoencoding (MAE) excels at capturing spatial-geometric features but lacks semantics, whereas contrastive learning, while able to d…

2026

DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics Model

ICLR 2026poster

Achieving generalized in-hand object rotation remains a significant challenge in robotics, largely due to the difficulty of transferring policies from simulation to the real world. The complex, contact-rich dynamics of dexterous manipulation create a "reality gap" that has limited prior work to cons…

Cited by 0SourcecodeScholar
2026

Embodied Navigation Foundation Model

ICLR 2026poster

Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments. To achieve such intelligence, recent advanced works leverage Vision-Language Models (VLMs), which demonstrate strong generalizability and possess a wel…

Cited by 0SourcecodeScholar
2026

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning

RSS 2026poster

Extrinsic dexterity leverages environmental contact to overcome the limitations of prehensile manipulation. However, achieving such dexterity in cluttered scenes remains challenging and underexplored, as it requires selectively exploiting contact among multiple interacting objects with inherently co…

Cited by 0SourceScholar
2026

HDFlow: Hierarchical Diffusion-Flow Planning for Long-horizon Tasks

ICML 2026spotlight

Recent advances in generative models have shown promise in generating behavior plans for long-horizon, sparse reward tasks. While these approaches have achieved promising results, they often lack a principled framework for hierarchical decomposition and struggle with the computational demands of rea…

Cited by 0SourcecodeScholar
2026

Humanoid Generative Pre-Training for Zero-Shot Motion Tracking

CVPR 2026

We introduce Humanoid-GPT, a GPT-style Transformer with causal attention trained on a billion-scale motion corpus for whole-body control. Unlike prior shallow MLP trackers constrained by scarce data and an agility-generalization trade-off, Humanoid-GPT is pre-trained on a 2B-frame retargeted corpus

Cited by 0SourcecodeScholar
2026

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

RSS 2026poster

Recent robot foundation models largely rely on large-scale behavior cloning, which imitates expert actions but discards transferable dynamics knowledge embedded in heterogeneous embodied data. While the Unified World Model (UWM) formulation has the potential to leverage such diverse data, existing i…

Cited by 0SourceScholar
2026

Layered 4D-Rotor Gaussian Splatting: A Compressed Representation for Long Dynamic Scenes

CVPR 2026

We address the challenge of reconstructing long dynamic scenes from multi-view videos in a storage-efficient manner. Recent advances in Gaussian Splatting and its extensions to dynamic scenes have demonstrated impressive visual quality, but remain limited to short duration (<10 s), large storage siz

Cited by 0SourceScholar
2026

NavGSim: High-Fidelity Gaussian Splatting Simulator for Large-Scale Navigation

ICRA 2026poster

Simulating realistic environments for robots is widely recognized as a critical challenge in robot learning, particularly in terms of rendering and physical simulation. This challenge becomes even more pronounced in navigation tasks, where trajectories often extend across multiple rooms or even enti…

2026

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

ICLR 2026poster

Spatial reasoning is a key aspect of cognitive psychology and remains a bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic spatial relations, such as distinguishing left from right, near from far, and object co…

Cited by 0SourcecodeScholar
2026

RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI

ICRA 2026poster

The emerging field of Vision-Language-Action (VLA) for humanoid robots faces several fundamental challenges, including the high cost of data acquisition, the lack of a standardized benchmark, and the significant gap between simulation and the real world. To overcome these obstacles, we propose RealM…

2026

Robust Differentiable Collision Detection for General Objects

ICRA 2026poster

Collision detection is a core component of robotics applications such as simulation, control, and planning. Traditional algorithms like GJK+EPA compute textit{witness points}—the closest or deepest-penetration pairs between two objects—but are inherently non-differentiable, preventing gradient flow …

2026

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

RSS 2026poster

While Vision-Language-Action (VLA) models excel in generalist manipulation, they often lack fine-grained spatial awareness and struggle with viewpoint generalization. This limitation largely stems from the reliance on pretrained RGB encoders, which lack explicit geometric cues and prioritize semanti…

Cited by 0SourceScholar
2026

TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking

ICRA 2026poster

Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously following moving targets is essential. Recent advances have enabled language-guided tracking in complex and unstructured…

2026

Unifying Precise Keyframes and Semantic Control via Multi-level Diffusion

CVPR 2026

Text-conditioned human motion in-betweening leverages keyframes for spatio-temporal control, with text providing high-level semantic guidance for the transitions. However, existing methods are unable to establish a coherent alignment between textual semantics and the spatio-temporal constraints prov

Cited by 0SourceScholar
2026

Unleashing Humanoid Reaching Potential Via Real-World-Ready Skill Space

ICRA 2026poster

Humans possess a large reachable space in the 3D world, enabling interactions with objects at varying heights and distances. However, realizing such large-space reaching on humanoids is a complex whole-body control (WBC) problem. Learning from scratch often leads to optimization difficulty and poor …

2026

Unleashing Humanoid Reaching Potential via Real-World-Ready Skill Space

RA-L 2026

Humans possess a large reachable space in the 3D world, enabling interactions with objects at varying heights and distances. However, realizing such large-space reaching on humanoids is a complex whole-body control (WBC) problem. Learning from scratch often leads to optimization difficulty and poor

Cited by 23SourcecodeScholar
2026

UrbanVLA: A Vision-Language-Action Model for Urban Micromobility

ICRA 2026poster

Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This task is particularly challenging due to the dynamic and unstructured nature of real-world city areas, yet most existing na…

2026

Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and Generation

CVPR 2026

Joint Energy-based Models (JEMs) are well known for their ability to unify classification and generation within a single framework. Despite their promising generative and discriminative performance, their robustness remains far inferior to adversarial training (AT), which, conversely, achieves stron

Cited by 0SourcecodeScholar
2025

AccCtr: Accelerating Training-Free Conditional Control For Diffusion Models

IJCAI 2025

In current training-free Conditional Diffusion Models (CDM), the sampling process is steered by the gradient, which measures the discrepancy between the guidance and the condition extracted by a pre-trained condition extraction network. These methods necessitate small guidance steps, resulting in lo

Cited by 0SourcePDFScholar
2025

Adaptive Layered-Trust Robust Defense Mechanism for Personalized Federated Learning

ICASSP 2025accepted

Personalized Federated Learning (PFL) is confronted with escalating security threats, yet existing defense strategies primarily concentrate on traditional federated learning, lacking robust defense mechanisms tailored for PFL. To fortify the robustness of PFL against stealthy malicious attacks, we p…

Cited by 0SourceScholar
2025

Aligning Text-to-Image Diffusion Models to Human Preference by Classification

NeurIPS 2025spotlight

Text-to-image diffusion models are typically trained on large-scale web data, often resulting in outputs that misalign with human preferences. Inspired by preference learning in large language models, we propose ABC (Alignment by Classification), a simple yet effective framework for aligning diffus…

Cited by 0SourceScholar
2025

An Investigation Into the Application of an RL-GA-Based Multi-Modal Motion Somatosensory Optimization Control Strategy for a Novel Rehabilitation Robot

RA-L 2025

The increasing prevalence of balance disorders presents significant challenges to rehabilitation therapy, prompting the development of rehabilitation robots as effective tools for motion training. To enhance the patient experience and rehabilitation outcomes in human-robot collaborative training, a

Cited by 3SourceScholar
2025

BODex: Scalable and Efficient Robotic Dexterous Grasp Synthesis Using Bilevel Optimization

ICRA 2025

Robotic dexterous grasping is important for interacting with the environment. To unleash the potential of data-driven models for dexterous grasping, a large-scale, highquality dataset is essential. While gradient-based optimization offers a promising way for constructing such datasets, previous work

Cited by 23SourcecodeScholar
2025

CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition

ICASSP 2025accepted

Code-switching automatic speech recognition (ASR) aims to transcribe speech that contains two or more languages accurately. To better capture language-specific speech representations and address language confusion in code-switching ASR, the mixture-of-experts (MoE) architecture and an additional lan…

Cited by 0SourceScholar
2025

Chumor 2.0: Towards Better Benchmarking Chinese Humor Understanding from (Ruo Zhi Ba)

ACL 2025finding

Existing humor datasets and evaluations predominantly focus on English, leaving limited resources for culturally nuanced humor in non-English languages like Chinese. To address this gap, we construct **Chumor**, the first and the largest Chinese humor explanation dataset. **Chumor** is sourced from…

2025

Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

CVPR 2025poster

Automatic detection and prevention of open-set failures are crucial in closed-loop robotic systems. Recent studies often struggle to simultaneously identify unexpected failures reactively after they occur and prevent foreseeable ones proactively. To this end, we propose Code-as-Monitor (CaM), a nove…

Cited by 7SourcePDFScholar
2025

DISCO: DISCrete nOise for Conditional Control in Text-to-Image Diffusion Models

NeurIPS 2025poster

A major challenge in using diffusion models is aligning outputs with user-defined conditions. Existing conditional generation methods fall into two major categories: classifier-based guidance, which requires differentiable target models and gradient-based correction; and classifier-free guidance, wh…

Cited by 0SourceScholar
2025

DexVLG: Dexterous Vision-Language-Grasp Model at Scale

ICCV 2025poster

As large models gain traction, vision-language models are enabling robots to tackle increasingly complex tasks. However, limited by the difficulty of data collection, progress has mainly focused on controlling simple gripper end-effectors. There is little research on functional grasping with large m…

Cited by 0SourcePDFScholar
2025

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

NeurIPS 2025poster

Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation. However, existing methods are limited to challenging image-based forecasting, which suffers from redundant i…

Cited by 0SourcecodeScholar
2025

DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation

ICCV 2025poster

Non-prehensile manipulation is crucial for handling objects that are too thin, large, or otherwise ungraspable in unstructured environments. While conventional planning-based approaches struggle with complex contact modeling, learning-based methods have recently emerged as a promising alternative. H…

Cited by 0SourcePDFScholar
2025

EMControl: Adding Conditional Control to Text-to-Image Diffusion Models via Expectation-Maximization

AAAI 2025technical

Recent advances in diffusion models focus on efficiently handling conditional generative tasks without extra training. The process involves decomposing the result into two components: 1. unconditional sample, generated in the absence of conditions; 2. condition correction, adjusting unconditional sa…

Cited by 0SourcePDFScholar
2025

Effects of Momentum in Implicit Bias of Gradient Flow for Diagonal Linear Networks

AAAI 2025technical

This paper targets on the regularization effect of momentum-based methods in regression settings and analyzes the popular diagonal linear networks to precisely characterize the implicit bias of continuous versions of heavy-ball (HB) and Nesterov's method of accelerated gradients (NAG). We show that,…

Cited by 0SourcePDFScholar
2025

FetchBot: Learning Generalizable Object Fetching in Cluttered Scenes via Zero-Shot Sim2Real

CoRL 2025oral

Generalizable object fetching in cluttered scenes remains a fundamental and application-critical challenge in embodied AI. Closely packed objects cause inevitable occlusions, making safe action generation particularly difficult. Under such partial observability, effective policies must not only gene…

Cited by 0SourceScholar
2025

GAPartManip: A Large-Scale Part-Centric Dataset for Material-Agnostic Articulated Object Manipulation

ICRA 2025

Effectively manipulating articulated objects in household scenarios is a crucial step toward achieving general embodied artificial intelligence. Mainstream research in 3D vision has primarily focused on manipulation through depth perception and pose detection. However, in real-world environments, th

Cited by 8SourcecodeScholar
2025

GenesisTex2: Stable, Consistent and High-Quality Text-to-Texture Generation

AAAI 2025technical

Large-scale text-guided image diffusion models have demonstrated remarkable results in text-to-image (T2I) generation. However, applying these models to synthesize textures for 3D geometries remains challenging due to the domain gap between 2D images and textures on a 3D surface. Early works that us…

2025

GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data

CoRL 2025poster

Embodied foundation models are gaining increasing attention for their zero-shot generalization, scalability, and adaptability to new tasks through few-shot post-training. However, existing models rely heavily on real-world data, which is costly and labor-intensive to collect. Synthetic data offers a…

Cited by 0SourceScholar
2025

Heavy-Ball Momentum Method in Continuous Time and Discretization Error Analysis

NeurIPS 2025poster

This paper establishes a continuous time approximation, a piece-wise continuous differential equation, for the discrete Heavy-Ball (HB) momentum method with explicit discretization error. Investigating continuous differential equations has been a promising approach for studying the discrete optimiza…

Cited by 0SourceScholar
2025

High-fidelity 3D Object Generation from Single Image with RGBN-Volume Gaussian Reconstruction Model

CVPR 2025highlight

Recently single-view 3D generation via Gaussian splatting has emerged and developed quickly. They learn 3D Gaussians from 2D RGB images generated from pre-trained multi-view diffusion (MVD) models, and have shown a promising avenue for 3D generation through a single image. Despite the current progre…

Cited by 0SourcePDFScholar
2025

Learning Extremely High Density Crowds as Active Matters

CVPR 2025poster

Video-based high-density crowd analysis and prediction has been a long-standing topic in computer vision. It is notoriously difficult due to, but not limited to, the lack of high-quality data and complex crowd dynamics. Consequently, it has been relatively under studied. In this paper, we propose a…

2025

Make a Donut: Hierarchical EMD-Space Planning for Zero-Shot Deformable Manipulation With Tools

RA-L 2025

Deformable object manipulation stands as one of the most captivating yet formidable challenges in robotics. While previous techniques have predominantly relied on learning latent dynamics through demonstrations, typically represented as either particles or images, there exists a pertinent limitation

Cited by 4SourceScholar
2025

MobileH2R: Learning Generalizable Human to Mobile Robot Handover Exclusively from Scalable and Diverse Synthetic Data

CVPR 2025poster

This paper introduces MobileH2R, a framework for learning generalizable vision-based human-to-mobile-robot (H2MR) handover skills. Unlike traditional fixed-base handovers, this task requires a mobile robot to reliably receive objects in a large workspace enabled by its mobility. Our key insight is t…

Cited by 0SourcePDFScholar
2025

Moderating the Generalization of Score-based Generative Model

ICCV 2025poster

Score-based Generative Models (SGMs) have demonstrated remarkable generalization capabilities, e.g. generating unseen, but natural data. However, the greater the generalization power, the more likely the unintended generalization, and the more dangerous the abuse. Despite these concerns, research on…

2025

Na Vid-4D: Unleashing Spatial Intelligence in Egocentric RGB-D Videos for Vision-and-Language Navigation

ICRA 2025

Understanding and reasoning about the 4D space-time is crucial for Vision-and-Language Navigation (VLN). However, previous works lack in-depth exploration in this aspect, resulting in bottlenecked spatial perception and action precision of VLN agents. In this work, we introduce NaVid-4D, a Vision La

Cited by 4SourceScholar
2025

NoiseCtrl: A Sampling-Algorithm-Agnostic Conditional Generation Method for Diffusion Models

CVPR 2025poster

In training-free conditional generative tasks, diffusion models utilize differentiable loss functions to steer the generative reverse process, necessitating modifications to sampling algorithms like DDPM and DDIM. However, such adjustments likely reduce flexibility and reliability. In this paper, we…

Cited by 0SourcePDFScholar
2025

QuadWBG: Generalizable Quadrupedal Whole-Body Grasping

ICRA 2025

Legged robots with advanced manipulation capabilities have the potential to significantly improve household duties and urban maintenance. Despite considerable progress in developing robust locomotion and precise manipulation methods, seamlessly integrating these into cohesive whole-body control for

Cited by 6SourcecodeScholar
2025

RoboHanger: Learning Generalizable Robotic Hanger Insertion for Diverse Garments

RA-L 2025

For the task of hanging clothes, learning how to insert a hanger into a garment is a crucial step, but has rarely been explored in robotics. In this work, we address the problem of inserting a hanger into various unseen garments that are initially laid flat on a table. This task is challenging due t

Cited by 5SourceScholar
2025

SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation

NeurIPS 2025spotlight

While spatial reasoning has made progress in object localization relationships, it often overlooks object orientation—a key factor in 6-DoF fine-grained manipulation. Traditional pose representations rely on pre-defined frames or templates, limiting generalization and semantic grounding. In this pap…

Cited by 0SourceScholar
2025

TASAR: Transfer-based Attack on Skeletal Action Recognition

ICLR 2025poster

Skeletal sequence data, as a widely employed representation of human actions, are crucial in Human Activity Recognition (HAR). Recently, adversarial attacks have been proposed in this area, which exposes potential security concerns, and more importantly provides a good tool for model robustness test…

2025

THE ROBUSTNESS OF DIFFERENTIABLE CAUSAL DISCOVERY IN MISSPECIFIED SCENARIOS

ICLR 2025poster

Causal discovery aims to learn causal relationships between variables from targeted data, making it a fundamental task in machine learning. However, causal discovery algorithms often rely on unverifiable causal assumptions, which are usually difficult to satisfy in real-world data, thereby limiting…

Cited by 0SourcePDFScholar
2025

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

RSS 2025poster

Embodied Navigation is a fundamental capability for intelligent robots, requiring robots to follow human commands and move autonomously within physical environments. Despite significant advancements, most existing navigation approaches are tailored to specific navigation tasks, such as instruction f…

Cited by 12PDFScholar
2025

Unsupervised Sentence Representation Learning with Syntactically Aligned Negative Samples

NAACL 2025findings

Sentence representation learning benefits from data augmentation strategies to improve model performance and generalization, yet existing approaches often encounter issues such as semantic inconsistencies and feature suppression. To address these limitations, we propose a method for generating Synta…

2025

W-ControlUDA: Weather-Controllable Diffusion-assisted Unsupervised Domain Adaptation for Semantic Segmentation

RA-L 2025

Image generation has emerged as a potent strategy to enrich training data for unsupervised domain adaptation (UDA) of semantic segmentation in adverse weathers due to the scarcity of labelled target domain data. Previous UDA works commonly utilize generative adversarial networks (GANs) to translate

Cited by 6SourceScholar
2025

Watch Less, Feel More: Sim-to-Real RL for Generalizable Articulated Object Manipulation via Motion Adaptation and Impedance Control

ICRA 2025

Articulated object manipulation poses a unique challenge compared to rigid object manipulation as the object itself represents a dynamic environment. In this work, we present a novel RL-based pipeline equipped with variable impedance control and motion adaptation leveraging observation history for g

Cited by 4SourcecodeScholar
2024

ASGrasp: Generalizable Transparent Object Reconstruction and 6-DoF Grasp Detection from RGB-D Active Stereo Camera

ICRA 2024poster

In this paper, we tackle the problem of grasping transparent and specular objects. This issue holds importance, yet it remains unsolved within the field of robotics due to failure of recover their accurate geometry by depth cameras. For the first time, we propose ASGrasp, a 6-DoF grasp detection net…

Cited by 4SourcecodeScholar
2024

Category-Level Multi-Part Multi-Joint 3D Shape Assembly

CVPR 2024poster

Shape assembly composes complex shapes geometries by arranging simple part geometries and has wide applications in autonomous robotic assembly and CAD modeling. Existing works focus on geometry reasoning and neglect the actual physical assembly process of matching and fitting joints which are the co…

Cited by 15SourcePDFScholar
2024

D$^3$RoMa: Disparity Diffusion-based Depth Sensing for Material-Agnostic Robotic Manipulation

CoRL 2024poster

Depth sensing is an important problem for 3D vision-based robotics. Yet, a real-world active stereo or ToF depth camera often produces noisy and incomplete depth which bottlenecks robot performances. In this work, we propose D3RoMa, a learning-based depth estimation framework on stereo image pairs t…

Cited by 4SourceScholar
2024

DexGraspNet 2.0: Learning Generative Dexterous Grasping in Large-scale Synthetic Cluttered Scenes

CoRL 2024poster

Grasping in cluttered scenes remains highly challenging for dexterous hands due to the scarcity of data. To address this problem, we present a large-scale synthetic dataset, encompassing 1319 objects, 8270 scenes, and 426 million grasps. Beyond benchmarking, we also explore data-efficient learning s…

Cited by 10SourceScholar
2024

Enhancing Generalizable 6D Pose Tracking of an In-Hand Object With Tactile Sensing

RA-L 2024

When manipulating an object to accomplish complex tasks, humans rely on both vision and touch to keep track of the object's 6D pose. However, most existing object pose tracking systems in robotics rely exclusively on visual signals, which hinder a robot's ability to manipulate objects effectively. T

Cited by 25SourcecodeScholar
2024

GAMMA: Graspability-Aware Mobile MAnipulation Policy Learning based on Online Grasping Pose Fusion

ICRA 2024poster

Mobile manipulation constitutes a fundamental task for robotic assistants and garners significant attention within the robotics community. A critical challenge inherent in mobile manipulation is the effective observation of the target while approaching it for grasping. In this work, we propose a gra…

Cited by 24SourcecodeScholar
2024

Human Motion Prediction Under Unexpected Perturbation

CVPR 2024highlight

We investigate a new task in human motion prediction which is predicting motions under unexpected physical perturbation potentially involving multiple people. Compared with existing research this task involves predicting less controlled unpremeditated and pure reactive motions in response to externa…

Cited by 2SourcePDFScholar
2024

MLCA-AVSR: Multi-Layer Cross Attention Fusion Based Audio-Visual Speech Recognition

ICASSP 2024accepted

While automatic speech recognition (ASR) systems degrade significantly in noisy environments, audio-visual speech recognition (AVSR) systems aim to complement the audio stream with noise-invariant visual cues and improve the system’s robustness. However, current studies mainly focus on fusing the we…

Cited by 0SourceScholar
2024

MaskClustering: View Consensus based Mask Graph Clustering for Open-Vocabulary 3D Instance Segmentation

CVPR 2024poster

Open-vocabulary 3D instance segmentation is cutting-edge for its ability to segment 3D instances without predefined categories. However progress in 3D lags behind its 2D counterpart due to limited annotated 3D data. To address this recent works first generate 2D open-vocabulary masks through 2D mode…

2024

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation

RSS 2024poster

Vision-and-language navigation (VLN) stands as a key research problem of Embodied AI, aiming at enabling agents to navigate in unseen environments following linguistic instructions. In this field, generalization is a long-standing challenge, either to out-of-distribution scenes or from Sim to Real.…

Cited by 80SourcePDFScholar
2024

Noisy Ostracods: A Fine-Grained, Imbalanced Real-World Dataset for Benchmarking Robust Machine Learning and Label Correction Methods

NeurIPS 2024poster

We present the Noisy Ostracods, a noisy dataset for genus and species classification of crustacean ostracods with specialists’ annotations. Over the 71466 specimens collected, 5.58% of them are estimated to be noisy (possibly problematic) at genus level. The dataset is created to addressing a real-w…

2024

Open6DOR: Benchmarking Open-instruction 6-DoF Object Rearrangement and A VLM-based Approach

IROS 2024poster

The integration of large-scale Vision-Language Models (VLMs) with embodied AI can greatly enhance the generalizability and the capacity to follow open instructions for robots. However, existing studies on object manipulation are not up to full consideration of the 6-DoF requirements, let alone estab…

Cited by 8SourceScholar
2024

RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation

CoRL 2024poster

This work proposes a retrieve-and-transfer framework for zero-shot robotic manipulation, dubbed RAM, featuring generalizability across various objects, environments, and embodiments. Unlike existing approaches that learn manipulation from expensive in-domain demonstrations, RAM capitalizes on a retr…

Cited by 26SourcecodeScholar
2024

Rethinking Word-level Adversarial Attack: The Trade-off between Efficiency, Effectiveness, and Imperceptibility

COLING 2024main

Neural language models have demonstrated impressive performance in various tasks but remain vulnerable to word-level adversarial attacks. Word-level adversarial attacks can be formulated as a combinatorial optimization problem, and thus, an attack method can be decomposed into search space and searc…

Cited by 3SourcePDFScholar
2024

SAGE: Bridging Semantic and Actionable Parts for GEneralizable Articulated-Object Manipulation under Language Instructions

RSS 2024poster

To interact with daily-life articulated objects of diverse structures and functionalities, understanding the object parts plays a central role in both user instruction comprehension and task execution. However, the possible discordance between the semantic meaning and physics functionalities of the…

Cited by 0SourcePDFScholar
2024

STOPNet: Multiview-based 6-DoF Suction Detection for Transparent Objects on Production Lines

ICRA 2024poster

In this work, we present STOPNet, a framework for 6-DoF object suction detection on production lines, with a focus on but not limited to transparent objects, which is an important and challenging problem in robotic systems and modern industry. Current methods requiring depth input fail on transparen…

Cited by 2SourceScholar
2024

Sample Design Engineering: An Empirical Study on Designing Better Fine-Tuning Samples for Information Extraction with LLMs

EMNLP 2024industry

Large language models (LLMs) have achieved significant leadership in many NLP tasks, but aligning structured output with generative models in information extraction (IE) tasks remains a challenge. Prompt Engineering (PE) is renowned for improving IE performance through prompt modifications. However,…

2024

ScissorBot: Learning Generalizable Scissor Skill for Paper Cutting via Simulation, Imitation, and Sim2Real

CoRL 2024poster

This paper tackles the challenging robotic task of generalizable paper cutting using scissors. In this task, scissors attached to a robot arm are driven to accurately cut curves drawn on the paper, which is hung with the top edge fixed. Due to the frequent paper-scissor contact and consequent frac…

Cited by 5SourceScholar
2024

SocialCVAE: Predicting Pedestrian Trajectory via Interaction Conditioned Latents

AAAI 2024technical

Pedestrian trajectory prediction is the key technology in many applications for providing insights into human behavior and anticipating human future motions. Most existing empirical models are explicitly formulated by observed human behaviors using explicable mathematical terms with deterministic na…

2024

Successive POI Recommendation via Brain-Inspired Spatiotemporal Aware Representation

AAAI 2024technical

Existing approaches usually perform spatiotemporal representation in the spatial and temporal dimensions, respectively, which isolates the spatial and temporal natures of the target and leads to sub-optimal embeddings. Neuroscience research has shown that the mammalian brain entorhinal-hippocampal s…

Cited by 2SourcePDFScholar
2024

Task-Oriented Dexterous Hand Pose Synthesis Using Differentiable Grasp Wrench Boundary Estimator

IROS 2024poster

This work tackles the problem of task-oriented dexterous hand pose synthesis, which involves generating a static hand pose capable of applying a task-specific set of wrenches to manipulate objects. Unlike previous approaches that focus solely on force-closure grasps, which are unsuitable for non-pre…

Cited by 1SourcecodeScholar
2023

3D-Aware Object Goal Navigation via Simultaneous Exploration and Identification

CVPR 2023poster

Object goal navigation (ObjectNav) in unseen environments is a fundamental task for Embodied AI. Agents in existing works learn ObjectNav policies based on 2D maps, scene graphs, or image sequences. Considering this task happens in 3D space, a 3D-aware agent can advance its ObjectNav capability via…

Cited by 47SourcePDFScholar
2023

A Laplace-inspired Distribution on SO(3) for Probabilistic Rotation Estimation

ICLR 2023top-25%

Estimating the 3DoF rotation from a single RGB image is an important yet challenging problem. Probabilistic rotation regression has raised more and more attention with the benefit of expressing uncertainty information along with the prediction. Though modeling noise using Gaussian-resembling Bingham…

2023

Adaptive Zone-Aware Hierarchical Planner for Vision-Language Navigation

CVPR 2023poster

The task of Vision-Language Navigation (VLN) is for an embodied agent to reach the global goal according to the instruction. Essentially, during navigation, a series of sub-goals need to be adaptively set and achieved, which is naturally a hierarchical navigation process. However, previous methods l…

2023

Affective and Dynamic Beam Search for Story Generation

EMNLP 2023long findings

Storytelling's captivating potential makes it a fascinating research area, with implications for entertainment, education, therapy, and cognitive studies. In this paper, we propose Affective Story Generator (AffGen) for generating interesting narratives. AffGen introduces `intriguing twists' in narr…

Cited by 0SourcecodeScholar
2023

Defending Black-Box Skeleton-Based Human Activity Classifiers

AAAI 2023technical

Skeletal motions have been heavily relied upon for human activity recognition (HAR). Recently, a universal vulnerability of skeleton-based HAR has been identified across a variety of classifiers and data, calling for mitigation. To this end, we propose the first black-box defense method for skeleton…

2023

Delving Into Discrete Normalizing Flows on SO(3) Manifold for Probabilistic Rotation Modeling

CVPR 2023poster

Normalizing flows (NFs) provide a powerful tool to construct an expressive distribution by a sequence of trackable transformations of a base distribution and form a probabilistic model of underlying data.Rotation, as an important quantity in computer vision, graphics, and robotics, can exhibit many…

2023

DexGraspNet: A Large-Scale Robotic Dexterous Grasp Dataset for General Objects Based on Simulation

ICRA 2023poster

Robotic dexterous grasping is the first step to enable human-like dexterous object manipulation and thus a crucial robotic technology. However, dexterous grasping is much more under-explored than object grasping with parallel grippers, partially due to the lack of a large-scale dataset. In this work…

Cited by 122SourcecodeScholar
2023

DiGA: Distil To Generalize and Then Adapt for Domain Adaptive Semantic Segmentation

CVPR 2023poster

Domain adaptive semantic segmentation methods commonly utilize stage-wise training, consisting of a warm-up and a self-training stage. However, this popular approach still faces several challenges in each stage: for warm-up, the widely adopted adversarial training often results in limited performanc…

2023

Few-Shot Physically-Aware Articulated Mesh Generation via Hierarchical Deformation

ICCV 2023poster

We study the problem of few-shot physically-aware articulated mesh generation. By observing an articulated object dataset containing only a few examples, we wish to learn a model that can generate diverse meshes with high visual fidelity and physical validity. Previous mesh generative models either…

Cited by 8PDFcodeScholar
2023

Full-Body Articulated Human-Object Interaction

ICCV 2023poster

Fine-grained capture of 3D Human-Object Interactions (HOIs) boosts human activity understanding and facilitates various downstream visual tasks. Prior models mostly assume that humans interact with rigid objects using only a few body parts, limiting their scope. In this paper, we address the challen…

Cited by 59PDFcodeScholar
2023

GAPartNet: Cross-Category Domain-Generalizable Object Perception and Manipulation via Generalizable and Actionable Parts

CVPR 2023highlight

For years, researchers have been devoted to generalizable object perception and manipulation, where cross-category generalizability is highly desired yet underexplored. In this work, we propose to learn such cross-category skills via Generalizable and Actionable Parts (GAParts). By identifying and d…

2023

GraspNeRF: Multiview-based 6-DoF Grasp Detection for Transparent and Specular Objects Using Generalizable NeRF

ICRA 2023poster

In this work, we tackle 6-DoF grasp detection for transparent and specular objects, which is an important yet challenging problem in vision-based robotic systems, due to the failure of depth cameras in sensing their geometry. We, for the first time, propose a multiview RGB-based 6-DoF grasp detectio…

Cited by 111SourcecodeScholar
2023

Hard No-Box Adversarial Attack on Skeleton-Based Human Action Recognition with Skeleton-Motion-Informed Gradient

ICCV 2023poster

Recently, methods for skeleton-based human activity recognition have been shown to be vulnerable to adversarial attacks. However, these attack methods require either the full knowledge of the victim (i.e. white-box attacks), access to training data (i.e. transfer-based attacks) or frequent model que…

Cited by 17PDFcodeScholar
2023

LeNo: Adversarial Robust Salient Object Detection Networks with Learnable Noise

AAAI 2023technical

Pixel-wise prediction with deep neural network has become an effective paradigm for salient object detection (SOD) and achieved remarkable performance. However, very few SOD models are robust against adversarial attacks which are visually imperceptible for human visual attention. The previous work r…

2023

PartManip: Learning Cross-Category Generalizable Part Manipulation Policy From Point Cloud Observations

CVPR 2023poster

Learning a generalizable object manipulation policy is vital for an embodied agent to work in complex real-world scenes. Parts, as the shared components in different object categories, have the potential to increase the generalization ability of the manipulation policy and achieve cross-category obj…

Cited by 40SourcePDFScholar
2023

Privacy-Preserving Adversarial Facial Features

CVPR 2023poster

Face recognition service providers protect face privacy by extracting compact and discriminative facial features (representations) from images, and storing the facial features for real-time recognition. However, such features can still be exploited to recover the appearance of the original face by b…

Cited by 22SourcePDFScholar
2023

Self-Supervised Category-Level Articulated Object Pose Estimation with Part-Level SE(3) Equivariance

ICLR 2023poster

Category-level articulated object pose estimation aims to estimate a hierarchy of articulation-aware object poses of an unseen articulated object from a known category. To reduce the heavy annotations needed for supervised learning methods, we present a novel self-supervised strategy that solves thi…

2023

Similarizing the Influence of Words with Contrastive Learning to Defend Word-level Adversarial Text Attack

ACL 2023findings

Neural language models are vulnerable to word-level adversarial text attacks, which generate adversarial examples by directly substituting discrete input words. Previous search methods for word-level attacks assume that the information in the important words is more influential on prediction than un…

Cited by 7SourcePDFScholar
2023

The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge

ICASSP 2023accepted

This paper describes our NPU-ASLP system for the Audio-Visual Diarization and Recognition (AVDR) task in the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. Specifically, the weighted prediction error (WPE) and guided source separation (GSS) techniques are used to reduce rever…

Cited by 0SourceScholar
2023

Tracking and Reconstructing Hand Object Interactions from Point Cloud Sequences in the Wild

AAAI 2023technical

In this work, we tackle the challenging task of jointly tracking hand object poses and reconstructing their shapes from depth point cloud sequences in the wild, given the initial poses at frame 0. We for the first time propose a point cloud-based hand joint tracking network, HandTrackNet, to estimat…

2023

UniDexGrasp++: Improving Dexterous Grasping Policy Learning via Geometry-Aware Curriculum and Iterative Generalist-Specialist Learning

ICCV 2023oral

We propose a novel, object-agnostic method for learning a universal policy for dexterous object grasping from realistic point cloud observations and proprioceptive information under a table-top setting, namely UniDexGrasp++. To address the challenge of learning the vision-based policy across thousan…

Cited by 88PDFScholar
2023

UniDexGrasp: Universal Robotic Dexterous Grasping via Learning Diverse Proposal Generation and Goal-Conditioned Policy

CVPR 2023poster

In this work, we tackle the problem of learning universal robotic dexterous grasping from a point cloud observation under a table-top setting. The goal is to grasp and lift up objects in high-quality and diverse ways and generalize across hundreds of categories and even the unseen. Inspired by succe…

Cited by 119SourcePDFScholar
2023

VE-KWS: Visual Modality Enhanced End-to-End Keyword Spotting

ICASSP 2023accepted

The performance of the keyword spotting (KWS) system based on audio modality, commonly measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. Therefore, audio-visual keyword spotting, which leverages complementary relationships over multiple moda…

Cited by 0SourceScholar
2022

ADeLA: Automatic Dense Labeling With Attention for Viewpoint Shift in Semantic Segmentation

CVPR 2022oral

We describe a method to deal with performance drop in semantic segmentation caused by viewpoint changes within multi-camera systems, where temporally paired images are readily available, but the annotations may only be abundant for a few typical views. Existing methods alleviate performance drop via…

Cited by 6PDFScholar
2022

CodedVTR: Codebook-Based Sparse Voxel Transformer With Geometric Guidance

CVPR 2022poster

Transformers have gained much attention by outperforming convolutional neural networks in many 2D vision tasks. However, they are known to have generalization problems and rely on massive-scale pre-training and sophisticated training techniques. When applying to 3D tasks, the irregular data structur…

Cited by 10PDFcodeScholar
2022

Domain Adaptation on Point Clouds via Geometry-Aware Implicits

CVPR 2022poster

As a popular geometric representation, point clouds have attracted much attention in 3D vision, leading to many applications in autonomous driving and robotics. One important yet unsolved issue for learning on point cloud is that point clouds of the same object can have significant geometric variati…

Cited by 64PDFcodeScholar
2022

Domain Randomization-Enhanced Depth Simulation and Restoration for Perceiving and Grasping Specular and Transparent Objects

ECCV 2022poster

"Commercial depth sensors usually generate noisy and missing depths, especially on specular and transparent objects, which poses critical issues to downstream depth or point cloud-based tasks. To mitigate this problem, we propose a powerful RGBD fusion network, SwinDRNet, for depth restoration. We f…

2022

Fine-grained Differentiable Physics: A Yarn-level Model for Fabrics

ICLR 2022poster

Differentiable physics modeling combines physics models with gradient-based learning to provide model explicability and data efficiency. It has been used to learn dynamics, solve inverse problems and facilitate design, and is at its inception of impact. Current successes have concentrated on general…

2022

FisherMatch: Semi-Supervised Rotation Regression via Entropy-Based Filtering

CVPR 2022oral

Estimating the 3DoF rotation from a single RGB image is an important yet challenging problem. Recent works achieve good performance relying on a large amount of expensive-to-obtain labeled data. To reduce the amount of supervision, we for the first time propose a general framework, FisherMatch, for…

Cited by 20PDFcodeScholar
2022

HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction

CVPR 2022poster

We present HOI4D, a large-scale 4D egocentric dataset with rich annotations, to catalyze the research of category-level human-object interaction. HOI4D consists of 2.4M RGB-D egocentric video frames over 4000 sequences collected by 9 participants interacting with 800 different object instances from…

Cited by 183PDFcodeScholar
2022

Harvesting Partially-Disjoint Time-Frequency Information for Improving Degenerate Unmixing Estimation Technique

ICASSP 2022accepted

The degenerate unmixing estimation technique (DUET) is one of the most efficient blind source separation algorithms tackling the challenging situation when the number of sources exceeds the number of microphones. However, as a time-frequency mask-based method, DUET erroneously results in interferenc…

Cited by 0SourceScholar
2022

Learning Category-Level Generalizable Object Manipulation Policy Via Generative Adversarial Self-Imitation Learning From Demonstrations

RA-L 2022

Generalizable object manipulation skills are critical for intelligent and multi-functional robots to work in real-world complex scenes. Despite the recent progress in reinforcement learning, it is still very challenging to learn a generalizable manipulation policy that can handle a category of geome

Cited by 33SourcecodeScholar
2022

Multi-Robot Active Mapping via Neural Bipartite Graph Matching

CVPR 2022poster

We study the problem of multi-robot active mapping, which aims for complete scene map construction in minimum time steps. The key to this problem lies in the goal position estimation to enable more efficient robot movements. Previous approaches either choose the frontier as the goal position via a m…

Cited by 34PDFScholar
2022

Pose Guided Image Generation from Misaligned Sources via Residual Flow Based Correction

AAAI 2022technical

Generating new images with desired properties (e.g. new view/poses) from source images has been enthusiastically pursued recently, due to its wide range of potential applications. One way to ensure high-quality generation is to use multiple sources with complementary information such as different vi…

Cited by 3SourcePDFScholar
2022

Projective Manifold Gradient Layer for Deep Rotation Regression

CVPR 2022poster

Regressing rotations on SO(3) manifold using deep neural networks is an important yet unsolved problem. The gap between the Euclidean network output space and the non-Euclidean SO(3) manifold imposes a severe challenge for neural network learning in both forward and backward passes. While several wo…

Cited by 32PDFcodeScholar
2022

Stable and Efficient Shapley Value-Based Reward Reallocation for Multi-Agent Reinforcement Learning of Autonomous Vehicles

ICRA 2022poster

With the development of sensing and communication technologies in networked cyber-physical systems (CPSs), multi-agent reinforcement learning (MARL)-based methodologies are integrated into the control process of physical systems and demonstrate prominent performance in a wide array of CPS domains, s…

Cited by 33SourceScholar
2021

3DIoUMatch: Leveraging IoU Prediction for Semi-Supervised 3D Object Detection

CVPR 2021poster

3D object detection is an important yet demanding task that heavily relies on difficult to obtain 3D annotations. To reduce the required amount of supervision, we propose 3DIoUMatch, a novel semi-supervised method for 3D object detection applicable to both indoor and outdoor scenes. We leverage a te…

Cited by 155PDFcodeScholar
2021

BASAR:Black-Box Attack on Skeletal Action Recognition

CVPR 2021poster

Skeletal motion plays a vital role in human activity recognition as either an independent data source or a complement. The robustness of skeleton-based activity recognizers has been questioned recently, which shows that they are vulnerable to adversarial attacks when the full-knowledge of the recogn…

Cited by 45PDFcodeScholar
2021

CAPTRA: CAtegory-Level Pose Tracking for Rigid and Articulated Objects From Point Clouds

ICCV 2021poster

In this work, we tackle the problem of category-level online pose tracking for objects from point cloud sequences. For the first time, we propose a unified framework that can handle 9DoF object pose tracking for novel rigid object instances as well as per-part pose tracking for articulated objects f…

Cited by 114PDFcodeScholar
2021

In-game Residential Home Planning via Visual Context-aware Global Relation Learning

AAAI 2021technical

In this paper, we propose an effective global relation learning algorithm to recommend an appropriate location of a building unit for in-game customization of residential home complex. Given a construction layout, we propose a visual context-aware graph generation network that learns the implicit gl…

Cited by 5SourcePDFScholar
2021

Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point Clouds

NeurIPS 2021poster

Category-level object pose estimation aims to find 6D object poses of previously unseen object instances from known categories without access to object CAD models. To reduce the huge amount of pose annotations needed for category-level learning, we propose for the first time a self-supervised learni…

2021

MultiBodySync: Multi-Body Segmentation and Motion Estimation via 3D Scan Synchronization

CVPR 2021poster

We present MultiBodySync, a novel, end-to-end trainable multi-body motion segmentation and rigid registration framework for multiple input 3D point clouds. The two non-trivial challenges posed by this multi-scan multibody setting that we investigate are: (i) guaranteeing correspondence and segmentat…

Cited by 57PDFcodeScholar
2021

Robust Neural Routing Through Space Partitions for Camera Relocalization in Dynamic Indoor Environments

CVPR 2021poster

Localizing the camera in a known indoor environment is a key building block for scene mapping, robot navigation, AR, etc. Recent advances estimate the camera pose via optimization over the 2D/3D-3D correspondences established between the coordinates in 2D/3D camera space and 3D world space. Such a m…

Cited by 32PDFcodeScholar
2021

Single Image 3D Shape Retrieval via Cross-Modal Instance and Category Contrastive Learning

ICCV 2021poster

In this work, we tackle the problem of single image-based 3D shape retrieval (IBSR), where we seek to find the most matched shape of a given single 2D image from a shape repository. Most of the existing works learn to embed 2D images and 3D shapes into a common feature space and perform metric learn…

Cited by 40PDFcodeScholar
2021

Understanding the Robustness of Skeleton-Based Action Recognition Under Adversarial Attack

CVPR 2021poster

Action recognition has been heavily employed in many applications such as autonomous vehicles, surveillance, etc, where its robustness is a primary concern. In this paper, we examine the robustness of state-of-the-art action recognizers against adversarial attack, which has been rarely investigated…

Cited by 55PDFcodeScholar
2021

Unsupervised Image Generation With Infinite Generative Adversarial Networks

ICCV 2021poster

Image generation has been heavily investigated in computer vision, where one core research challenge is to generate images from arbitrarily complex distributions with little supervision. Generative Adversarial Networks (GANs) as an implicit approach have achieved great successes in this direction an…

Cited by 6PDFcodeScholar
2020

A 1 mm-Thick Miniatured Mobile Soft Robot With Mechanosensation and Multimodal Locomotion

RA-L 2020

The miniature soft robots have many promising applications, including micro-manipulations, endoscopy, and microsurgery, etc. Nevertheless, it remains challenging to fabricate a miniatured robot device that is thin, flexible, and can perform multimodal locomotor mobility with sensory capacity. In thi

Cited by 20SourceScholar
2020

Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction

CVPR 2020oral

We propose a novel attention-based framework for 3D human pose estimation from a monocular video. Despite the general success of end-to-end deep learning paradigms, our approach is based on two key observations: (1) temporal incoherence and jitter are often yielded from a single frame prediction; (2…

Cited by 279PDFcodeScholar
2020

Category-Level Articulated Object Pose Estimation

CVPR 2020oral

This paper addresses the task of category-level pose estimation for articulated objects from a single depth image. We present a novel category-level approach that correctly accommodates object instances previously unseen during training. We introduce Articulation-aware Normalized Coordinate Space Hi…

Cited by 245PDFcodeScholar
2020

Human-like Planning for Reaching in Cluttered Environments

ICRA 2020poster

Humans, in comparison to robots, are remarkably adept at reaching for objects in cluttered environments. The best existing robot planners are based on random sampling of configuration space- which becomes excessively high-dimensional with large number of objects. Consequently, most planners often fa…

Cited by 25SourcecodeScholar
2020

PT2PC: Learning to Generate 3D Point Cloud Shapes from Part Tree Conditions

ECCV 2020poster

Generative 3D shape modeling is a fundamental research area in computer vision and interactive computer graphics, with many real-world applications. This paper investigates the novel problem of generating a 3D point cloud geometry for a shape from a symbolic part tree representation. In order to lea…

Cited by 49SourcePDFScholar
2020

SAPIEN: A SimulAted Part-Based Interactive ENvironment

CVPR 2020oral

Building home assistant robots has long been a goal for vision and robotics researchers. To achieve this task, a simulated environment with physically realistic simulation, sufficient articulated objects, and transferability to the real robot is indispensable. Existing environments achieve these req…

Cited by 560PDFcodeScholar
2020

Truth-to-Estimate Ratio Mask: A Post-Processing Method for Speech Enhancement Direct at Low Signal-to-Noise Ratios

ICASSP 2020accepted

This study proposes a bi-directional recurrent neural network (Bi-RNN) post-processing method for speech enhancement (SE) at low signal-to noise ratios (SNR). Current speech enhancement solutions performed badly under low SNR situations. Loizou and Kim proposed a solution to reduce speech distortion…

Cited by 0SourceScholar
2019

Competing Against Nash Equilibria in Adversarially Changing Zero-Sum Games

ICML 2019oral

We study the problem of repeated play in a zero-sum game in which the payoff matrix may change, in a possibly adversarial fashion, on each round; we call these Online Matrix Games. Finding the Nash Equilibrium (NE) of a two player zero-sum game is core to many problems in statistics, optimization, a…

Cited by 45SourcePDFScholar
2019

GSPN: Generative Shape Proposal Network for 3D Instance Segmentation in Point Cloud

CVPR 2019poster

We introduce a novel 3D object proposal approach named Generative Shape Proposal Network (GSPN) for instance segmentation in point cloud data. Instead of treating object proposal as a direct bounding box regression problem, we take an analysis-by-synthesis strategy and generate proposals by reconstr…

Cited by 381PDFcodeScholar
2019

Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation

CVPR 2019oral

The goal of this paper is to estimate the 6D pose and dimensions of unseen object instances in an RGB-D image. Contrary to "instance-level" 6D pose estimation tasks, our problem assumes that no exact object CAD models are available during either training or testing time. To handle different and unse…

Cited by 870PDFcodeScholar
2016

“Congruent” and “Opposite” Neurons: Sisters for Multisensory Integration and Segregation

NeurIPS 2016poster

Experiments reveal that in the dorsal medial superior temporal (MSTd) and the ventral intraparietal (VIP) areas, where visual and vestibular cues are integrated to infer heading direction, there are two types of neurons with roughly the same number. One is “congruent” cells, whose preferred heading…

Cited by 10SourcePDFScholar