← Search

Chen Wang

174 accepted papers

2026

Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning

CVPR 2026

Computational pathology has advanced rapidly in recent years, driven by domain-specific image encoders and growing interest in using vision-language models to answer natural-language questions about diseases. Yet, the core problem behind pathology question-answering remains unsolved, considering tha

Cited by 0SourcecodeScholar
2026

Adapt Data to Model: Adaptive Transformation Optimization for Domain-shared Time Series Foundation Models

ICLR 2026poster

Large time series models (LTMs) have recently demonstrated powerful capabilities for universal forecasting. However, these models still struggle to address the variety and nonstationarity of time series, resulting in an unsatisfying balance between forecasting performance and generalizability. Inste…

Cited by 0SourcecodeScholar
2026

Benchmarking LLMs for Political Science: A United Nations Perspective

AAAI 2026technical

Large Language Models (LLMs) have achieved significant advances in natural language processing, yet their potential for high-stake political decision-making remains largely unexplored. This paper addresses the gap by focusing on the application of LLMs to the United Nations (UN) decision-making proc

Cited by 0SourcePDFScholar
2026

Catalog-Native LLM: Speaking Item-ID dialect with Less Entanglement for Recommendation

ICLR 2026poster

While collaborative filtering delivers predictive accuracy and efficiency, and Large Language Models (LLMs) enable expressive and generalizable reasoning, modern recommendation systems must bring these strengths together. Growing user expectations, such as natural-language queries and transparent ex…

Cited by 0SourceScholar
2026

DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoML

ICML 2026poster

Data cleaning and automated machine learning (AutoML) are both crucial for reliable learning systems, yet are commonly treated as independent or sequential stages. This separation ignores their strong interaction and leads to inefficient use of limited computational budgets. We propose DMCO, a unifi…

Cited by 0SourceScholar
2026

DispViT: Direct Stereo Disparity Regression with a Single-Stream Vision Transformer

ICLR 2026poster

Deep stereo disparity estimation has long been dominated by a \textbf{matching-centric paradigm}, built on constructing cost volumes and iteratively refining local correspondences. Despite its success, this paradigm exhibits an intrinsic vulnerability: visual ambiguities from occlusion or non-Lamber…

Cited by 0SourcecodeScholar
2026

Efficient Testing for Correlation Clustering: Improved Algorithms and Optimal Bounds

ICLR 2026poster

Correlation clustering is an important unsupervised learning problem with broad applications. In this problem, we are given a labeled complete graph $G=(V,E^+ \cup E^-)$, and the optimal clustering is defined as a partition of the vertices that minimizes the $+$ edges between clusters and $-$ edges…

Cited by 0SourceScholar
2026

Evolution of Benchmark: Black-Box Optimization Benchmark Design through Large Language Model

ICML 2026poster

Benchmark Design in Black-Box Optimization (BBO) is a fundamental yet open-ended topic. Early BBO benchmarks are predominantly human-crafted, introducing expert bias and constraining diversity. Automating this design process can relieve the human-in-the-loop burden while enhancing diversity and obje…

Cited by 0SourceScholar
2026

Fingerprinting Pre-trained Encoders under Arbitrary Downstream Fine-Tuning via Adversarial Shifting

ICML 2026poster

In the pre-training-fine-tuning paradigm, pre-trained encoders have become high-value intellectual property (IP) due to their immense training costs, necessitating robust protection. Existing fingerprinting or watermarking methods typically rely on pre-defined samples and labels, or require intrusiv…

Cited by 0SourceScholar
2026

GeoEvo: Identity-Aware Potential Game with Geometric Evolution for Personalized Multimodal Federated Learning

ICML 2026poster

We reconceptualize Personalized Multimodal Federated Learning (PMFL) by treating missing modalities as intrinsic structural identities that constrain each client to a distinct Riemannian submanifold, rather than deficiencies to be compensated. To resolve the tension between identity preservation and…

Cited by 0SourceScholar
2026

Instance Generation for Meta-Black-Box Optimization Through Latent Space Reverse Engineering

AAAI 2026technical

To relieve intensive human-expertise required to design optimization algorithms, recent Meta-Black-Box Optimization (MetaBBO) researches leverage generalization strength of meta-learning to train neural network-based algorithm design policies over a predefined training problem set, which automates t

Cited by 0SourcePDFScholar
2026

Learning When to Jump for Off-road Navigation

RSS 2026poster

Low speed does not always guarantee safety in off-road driving. For instance, crossing a ditch may be risky at a low speed due to the risk of getting stuck, yet safe at a higher speed with a controlled, accelerated jump. Achieving such behavior requires path planning that explicitly models complex m…

Cited by 0SourceScholar
2026

Learning-Augmented Moment Estimation on Time-Decay Models

ICLR 2026poster

Motivated by the prevalence and success of machine learning, a line of recent work has studied learning-augmented algorithms in the streaming model. These results have shown that for natural and practical oracles implemented with machine learning models, we can obtain streaming algorithms with impro…

Cited by 0SourcecodeScholar
2026

Online Learning with Recency: Algorithms for Sliding-window Streaming Multi-armed Bandits

ICML 2026poster

Motivated by the recency effect in online learning, we study algorithms for single-pass \emph{sliding-window streaming multi-armed bandits (MABs)} in this paper. In this setting, we are given $n$ arms with unknown sub-Gaussian reward distributions and a parameter $W$. The arms arrive in a single-pas…

Cited by 0SourceScholar
2026

Shared Autonomy Assisted by Impedance-Driven Anisotropic Guidance Field

RA-L 2026

Shared autonomy (SA) enables robots to infer human intent and assist in its achievement. While most research focuses on improving intent inference, it overlooks whether humans can understand the robot's intent in return. Without such mutual understanding, collaboration becomes less effective, degrad

Cited by 0SourceScholar
2026

Thinking in Structures: Evaluating Spatial Intelligence through Reasoning on Constrained Manifolds

ICML 2026poster

Spatial intelligence is crucial for vision--language models (VLMs) in the physical world, yet many benchmarks evaluate largely unconstrained scenes where models can exploit 2D shortcuts. We introduce SSI-Bench, a VQA benchmark for spatial reasoning on constrained manifolds, built from complex real-w…

Cited by 0SourceScholar
2026

TuneAhead: Predicting Fine-tuning Performance Before Training Begins

ICML 2026poster

Fine-tuning large language models (LLMs) is compute-intensive and error-prone: model performance depends sensitively on data quality and hyperparameter choices, and naïve runs can even degrade model performance. This raises a fundamental question: Can we predict fine-tuning performance before traini…

Cited by 0SourceScholar
2026

UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching

CVPR 2026

Existing feed-forward networks excel at predicting a single set of physical properties from visual appearance, but this point-estimate paradigm fundamentally fails to capture the real world's inherent physical ambiguity. We address this by reframing physics prediction as a task of learning a control

Cited by 0SourceScholar
2026

VSTYLE: A BENCHMARK FOR VOICE STYLE ADAPTATION WITH SPOKEN INSTRUCTIONS

ICASSP 2026poster

Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress has focused on semantic accuracy and instruction following, the ability of SLMs to adapt their speaking style based on sp…

Cited by 0SourcePDFScholar
2026

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

AAAI 2026technical

Recent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains a critical unsolved challenge. To address this gap, we introduce Video SimpleQA, the first comprehensive benchmark tailo

Cited by 0SourcePDFScholar
2026

tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction

CVPR 2026

We propose tttLRM, a novel large 3D reconstruction model that leverages a Test-Time Training (TTT) layer to enable long-context, autoregressive 3D reconstruction with linear computational complexity, further scaling the model's capability. Our framework efficiently compresses multiple image observat

Cited by 0SourcecodeScholar
2025

Adaptive Dynamic Programming-Based Fixed-Time Optimal Control for Wheeled Mobile Robot

RA-L 2025

In this study, the adaptive dynamic programming (ADP)-based fixed-time optimal trajectory tracking control is investigated for wheeled mobile robots. An ADP-based fixed-time optimal tracking controller is developed based on the critic-only neural network ADP technique, which guarantees the robot tra

Cited by 5SourceScholar
2025

BEHAVIOR Robot Suite: Streamlining Real-World Whole-Body Manipulation for Everyday Household Activities

CoRL 2025poster

Real-world household tasks present significant challenges for mobile manipulation robots. An analysis of existing robotics benchmarks reveals that successful task performance hinges on three key whole-body control capabilities: bimanual coordination, stable and precise navigation, and extensive end-…

Cited by 0SourcecodeScholar
2025

BagIt! An Adaptive Dual-Arm Manipulation of Fabric Bags for Object Bagging

RA-L 2025

Bagging tasks, commonly found in industrial scenarios, are challenging considering deformable bags' complicated and unpredictable nature. This paper presents an automated bagging system from the proposed adaptive Structure-of-Interest (SOI) manipulation strategy for dual robot arms. The system dynam

Cited by 0SourceScholar
2025

BridgeDepth: Bridging Monocular and Stereo Reasoning with Latent Alignment

ICCV 2025poster

Monocular and stereo depth estimation offer complementary strengths: monocular methods capture rich contextual priors but lack geometric precision, while stereo approaches leverage epipolar geometry yet struggle with ambiguities such as reflective or textureless surfaces. Despite post-hoc synergies,…

2025

Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition

ICASSP 2025accepted

Recent advancements in event-based zero-shot object recognition have demonstrated promising results. However, these methods heavily depend on extensive training and are inherently constrained by the characteristics of CLIP. To the best of our knowledge, this research is the first study to explore th…

Cited by 0SourceScholar
2025

Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models

ICRA 2025

Learning to perform manipulation tasks from human videos is a promising approach for teaching robots. However, many manipulation tasks require changing control parameters during task execution, such as force, which visual data alone cannot capture. In this work, we leverage sensing devices such as a

Cited by 7SourcecodeScholar
2025

DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single Image

ICLR 2025poster

Reconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, co…

2025

DIMO: Diverse 3D Motion Generation for Arbitrary Objects

ICCV 2025poster

We present DIMO, a generative approach capable of generating diverse 3D motions for arbitrary objects from a single image. The core idea of our work is to leverage the rich priors in well-trained video models to extract the common motion patterns and then embed them into a shared low-dimensional lat…

Cited by 0SourcePDFScholar
2025

Decoupling Memories, Muting Neurons: Towards Practical Machine Unlearning for Large Language Models

ACL 2025finding

Machine Unlearning (MU) has emerged as a promising solution for removing the influence of data that an owner wishes to unlearn from Large Language Models (LLMs). However, existing MU methods, which require tuning the entire model parameters on the unlearned data with random labels or perturbed gradi…

2025

EgoMimic: Scaling Imitation Learning via Egocentric Video

ICRA 2025

The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through

Cited by 136SourcecodeScholar
2025

Enhancing Scene Coordinate Regression With Efficient Keypoint Detection and Sequential Information

RA-L 2025

Scene Coordinate Regression (SCR) is a visual localization technique that utilizes deep neural networks (DNN) to directly regress 2D-3D correspondences for camera pose estimation. However, current SCR methods often face challenges in handling repetitive textures and meaningless areas due to their re

Cited by 3SourcecodeScholar
2025

Express What You See: Can Multimodal LLMs Decode Visual Ciphers with Intuitive Semiosis Comprehension?

ACL 2025finding

Bridging the gap between visual and language remains a pivotal challenge for the multimodal community. Traditional VQA benchmarks encounter a modality gap and over-reliance on language priors, whereas human cognition excels at intuitive semiosis, associating abstract visual symbols to linguistic sem…

Cited by 0SourcePDFScholar
2025

First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training

NeurIPS 2025poster

Improving Multi-modal Large Language Models (MLLMs) in the post-training stage typically relies on supervised fine-tuning (SFT) or reinforcement learning (RL), which require expensive and manually annotated multi-modal data--an ultimately unsustainable resource. This limitation has motivated a growi…

Cited by 0SourcecodeScholar
2025

Fully Dynamic Adversarially Robust Correlation Clustering in Polylogarithmic Update Time

AISTATS 2025poster

We study the dynamic correlation clustering problem with *adaptive* edge label flips. In correlation clustering, we are given a $n$-vertex complete graph whose edges are labeled either $(+)$ or $(-)$, and the goal is to minimize the total number of $(+)$ edges between clusters and the number of $(-)…

Cited by 0SourceScholar
2025

Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation

ICASSP 2025accepted

Retrieval-Augmented Generation (RAG) is a state-of-the-art technique that mitigates issues such as hallucinations and knowledge staleness in Large Language Models (LLMs) by retrieving relevant knowledge from an external database to assist in content generation. Existing research has demonstrated pot…

Cited by 0SourceScholar
2025

Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment

ACL 2025finding

Direct Preference Optimization (DPO) has become a prominent method for aligning Large Language Models (LLMs) with human preferences. While DPO has enabled significant progress in aligning English LLMs, multilingual preference alignment is hampered by data scarcity. To address this, we propose a nove…

2025

Language Imbalance Driven Rewarding for Multilingual Self-improving

ICLR 2025poster

Large Language Models (LLMs) have achieved state-of-the-art performance across numerous tasks. However, these advancements have predominantly benefited "first-class" languages such as English and Chinese, leaving many other languages underrepresented. This imbalance, while limiting broader applicati…

2025

Latent Swap Joint Diffusion for 2D Long-Form Latent Generation

ICCV 2025poster

This paper introduces Swap Forward (SaFa), a modality-agnostic and efficient method to generate seamless and coherent long spectrum and panorama using a latent swap joint diffusion process across multi-views. We first investigate spectrum aliasing problem in spectrum-based audio generation caused by…

2025

Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models

EMNLP 2025

Recent advances in text-only “slow-thinking” reasoning have prompted efforts to transfer this capability to vision-language models (VLMs), for training visual reasoning models (VRMs). However, such transfer faces critical challenges: Effective “slow thinking” in VRMs requires visual reflection, the

2025

MapEx: Indoor Structure Exploration with Probabilistic Information Gain from Global Map Predictions

ICRA 2025

Exploration is a critical challenge in robotics, centered on understanding unknown environments. In this work, we focus on structured indoor environments, which often exhibit predictable, repeating patterns. Conventional frontier-based exploration approaches have difficulty leveraging this predictab

Cited by 27SourcecodeScholar
2025

MetaBox-v2: A Unified Benchmark Platform for Meta-Black-Box Optimization

NeurIPS 2025poster

Meta-Black-Box Optimization (MetaBBO) streamlines the automation of optimization algorithm design through meta-learning. It typically employs a bi-level structure: the meta-level policy undergoes meta-training to reduce the manual effort required in developing algorithms for low-level optimization t…

Cited by 0SourcecodeScholar
2025

Nearly Tight Bounds for Exploration in Streaming Multi-Armed Bandits with Known Optimality Gap

AAAI 2025technical

We investigate the sample-memory-pass trade-offs for pure exploration in multi-pass streaming multi-armed bandits (MABs) with the *a priori* knowledge of the optimality gap ?_[2]. Here, and throughout, the optimality gap ?_[i] is defined as the mean reward gap between the best and the i-th best arms…

Cited by 2SourcePDFScholar
2025

On the Price of Differential Privacy for Hierarchical Clustering

ICLR 2025poster

Hierarchical clustering is a fundamental unsupervised machine learning task with the aim of organizing data into a hierarchy of clusters. Many applications of hierarchical clustering involve sensitive user information, therefore motivating recent studies on differentially private hierarchical cluste…

2025

PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors

CVPR 2025poster

This paper presents PCDreamer, a novel method for point cloud completion. Traditional methods typically extract features from partial point clouds to predict missing regions, but the large solution space often leads to unsatisfactory results. More recent approaches have started to use images as extr…

Cited by 1SourcePDFScholar
2025

PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation

NeurIPS 2025poster

Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limitations, we introduce PhysCtrl, a novel framework for physics-grounded image-to-video generation with physical parameters…

Cited by 0SourceScholar
2025

Physics-Informed LSTM for Shape and Contact Force Prediction of a Flexible Surgical Robot*

IROS 2025

Real-time morphological perception and precise end force feedback prediction of surgical robots constitute critical technical elements for ensuring safety and efficacy in complex interventional procedures such as Endoscopic Retrograde Cholangiopancreatography (ERCP). In this paper, we design a minia

Cited by 0SourceScholar
2025

RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

ICML 2025poster

Reasoning stands as a cornerstone of intelligence, enabling the synthesis of existing knowledge to solve complex problems. Despite remarkable progress, existing reasoning benchmarks often fail to rigorously evaluate the nuanced reasoning capabilities required for complex, real-world problemsolving,…

2025

Real-Time Neural Denoising with Render-Aware Knowledge Distillation

AAAI 2025technical

Real-time Monte Carlo (MC) ray tracing with low sampling rates demands a denoising algorithm that adeptly balances the trade-off between quality and efficiency. Previous works have paid much attention on designing delicate denoising architecture while ignoring model compression. In this work, we pre…

Cited by 0SourcePDFScholar
2025

Relative Error Fair Clustering in the Weak-Strong Oracle Model

ICML 2025poster

We study fair clustering problems in a setting where distance information is obtained from two sources: a strong oracle providing exact distances, but at a high cost, and a weak oracle providing potentially inaccurate distance estimates at a low cost. The goal is to produce a near-optimal fair clust…

Cited by 0SourcePDFScholar
2025

SLU-DQN: A Model for Anticipatory Steam Detection for Steamer-Filling in Baijiu Intelligent Distillation Systems

IROS 2025

The true implementation of the Anticipatory Steam Detection for Steamer-Filling(ASDSF) process in baijiu intelligent distillation systems, which involves predicting and precisely spreading distillers’ grains before steam emerges, remains a critical unresolved challenge. In this study, we introduce t

Cited by 0SourceScholar
2025

See the World, Discover Knowledge: A Chinese Factuality Evaluation for Large Vision Language Models

ACL 2025finding

The evaluation of factual accuracy in large vision language models (LVLMs) has lagged behind their rapid development, making it challenging to fully reflect these models’ knowledge capacity and reliability. In this paper, we introduce the first factuality-based visual question-answering benchmark in…

Cited by 0SourcePDFScholar
2025

SuperPC: A Single Diffusion Model for Point Cloud Completion, Upsampling, Denoising, and Colorization

CVPR 2025poster

Point cloud (PC) processing tasks--such as completion, upsampling, denoising, and colorization--are crucial in applications like autonomous driving and 3D reconstruction. Despite substantial advancements, prior approaches often address each of these tasks independently, with separate models focused…

Cited by 1SourcePDFScholar
2025

Taxonomy-Guided Zero-Shot Recommendations with LLMs

COLING 2025main

With the emergence of large language models (LLMs) and their ability to perform a variety of tasks, their application in recommender systems (RecSys) has shown promise. However, we are facing significant challenges when deploying LLMs into RecSys, such as limited prompt length, unstructured item inf…

2025

Time Travel is Cheating: Going Live with DeepFund for Real-Time Fund Investment Benchmarking

NeurIPS 2025poster

Large Language Models (LLMs) have demonstrated notable capabilities across financial tasks, including financial report summarization, earnings call transcript analysis, and asset classification. However, their real-world effectiveness in managing complex fund investment remains inadequately assessed…

Cited by 0SourcecodeScholar
2025

V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative Perception

NeurIPS 2025spotlight

Modern autonomous vehicle perception systems often struggle with occlusions and limited perception range. Previous studies have demonstrated the effectiveness of cooperative perception in extending the perception range and overcoming occlusions, thereby enhancing the safety of autonomous driving. In…

Cited by 0SourceScholar
2025

Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation

CVPR 2025poster

Faithfully reconstructing textured shapes and physical properties from videos presents an intriguing yet challenging problem. Significant efforts have been dedicated to advancing such a system identification problem in this area. Previous methods often rely on heavy optimization pipelines with a dif…

Cited by 0SourcePDFScholar
2025

iKap: Kinematics-Aware Planning with Imperative Learning

ICRA 2025

Trajectory planning in robotics aims to generate collision-free pose sequences that can be reliably executed. Recently, vision-to-planning systems have gained increasing attention for their efficiency and ability to interpret and adapt to surrounding environments. However, traditional modular system

Cited by 5SourceScholar
2024

A Multi-Scale Convolutional Hybrid Attention Residual Network for Enhancing Underwater Image and Identifying Underwater Multi-Scene Sea Cucumber

RA-L 2024

At present, the use of underwater robots to replace underwater manual work is a future development direction. The complex and changeable underwater environment brings great difficulties to the operation of robots. In order to improve the problem of color distortion and degradation of sea cucumber im

Cited by 2SourceScholar
2024

AI-Olympics: Exploring the Generalization of Agents through Open Competitions

IJCAI 2024poster

Between 2021 and 2023, AI-Olympics---a series of online AI competitions, was hosted by the online evaluation platform Jidi in collaboration with the IJCAI committee. In these competitions, an agent is required to accomplish diverse sports tasks in a two-dimensional continuous world, while competing…

Cited by 2SourcePDFScholar
2024

AirShot: Efficient Few-Shot Detection for Autonomous Exploration

IROS 2024poster

Few-shot object detection has drawn increasing attention in the field of robotic exploration, where robots are required to find unseen objects with a few online provided examples. Despite recent efforts have been made to yield online processing capabilities, slow inference speeds of low-powered robo…

Cited by 8SourcecodeScholar
2024

Automated Creation of Digital Cousins for Robust Policy Learning

CoRL 2024poster

Training robot policies in the real world can be unsafe, costly, and difficult to scale. Simulation serves as an inexpensive and potentially limitless source of training data, but suffers from the semantics and physics disparity between simulated and real-world environments. These discrepancies can…

Cited by 11SourcecodeScholar
2024

BLSP-Emo: Towards Empathetic Large Speech-Language Models

EMNLP 2024main

The recent release of GPT-4o showcased the potential of end-to-end multimodal models, not just in terms of low latency but also in their ability to understand and generate expressive speech with rich emotions. While the details are unknown to the open research community, it likely involves significa…

2024

Collaborative Tooth Motion Diffusion Model in Digital Orthodontics

AAAI 2024technical

Tooth motion generation is an essential task in digital orthodontic treatment for precise and quick dental healthcare, which aims to generate the whole intermediate tooth motion process given the initial pathological and target ideal tooth alignments. Most prior works for multi-agent motion plannin…

Cited by 2SourcePDFScholar
2024

Data-free Distillation of Diffusion Models with Bootstrapping

ICML 2024poster

Diffusion models have demonstrated great potential for generating diverse images. However, their performance often suffers from slow generation due to iterative denoising. Knowledge distillation has been recently proposed as a remedy which can reduce the number of inference steps to one or a few, wi…

Cited by 2SourcePDFScholar
2024

Design-Modeling and Control of a Novel Wearable Exoskeleton for Lower-Limb Enhancement

RA-L 2024

In this paper, a novel powered lower limb exoskeleton prototype called PTEXO for reducing user burden and enhancing following comfort is presented. The PTEXO is designed with a new control strategy, Enhanced Sensitivity Amplification Control (ESAC), and improves comfort of lower-limb locomotion thro

Cited by 5SourceScholar
2024

DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation

RSS 2024poster

Imitation learning from human hand motion data presents a promising avenue for imbuing robots with human-like dexterity in real-world manipulation tasks. Despite this potential, substantial challenges persist, particularly with the portability of existing hand motion capture (mocap) systems and the…

Cited by 120SourcePDFScholar
2024

Facilitating Message Passing with Potential Links for Knowledge Graph Completion

ICASSP 2024accepted

Knowledge graph completion (KGC) aims at inferring missing links between two entities. Most previous models focus on learning representations for entities and relations via graph neural networks. In this formalism, representations heavily rely on structural information. However, it is common for Kno…

Cited by 0SourceScholar
2024

GeoReF: Geometric Alignment Across Shape Variation for Category-level Object Pose Refinement

CVPR 2024poster

Object pose refinement is essential for robust object pose estimation. Previous work has made significant progress towards instance-level object pose refinement. Yet category-level pose refinement is a more challenging problem due to large shape variations within a category and the discrepancies bet…

Cited by 4SourcePDFScholar
2024

Learning-on-the-Drive: Self-supervised Adaptive Long-range Perception for High-speed Offroad Driving

IROS 2024poster

Autonomous offroad driving is essential for applications like emergency rescue, military operations, and agriculture. Despite progress, systems struggle with high-speed vehicles exceeding 10m/s due to the need for accurate long-range (> 50m) perception for safe navigation. Current approaches are lim…

Cited by 2SourceScholar
2024

LogiCity: Advancing Neuro-Symbolic AI with Abstract Urban Simulation

NeurIPS 2024poster

Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing benchmarks for NeSy AI fail to provide long-horizon reasoning tasks with complex multi-agent interactions. Furthermore, t…

2024

Map It Anywhere: Empowering BEV Map Prediction using Large-scale Public Datasets

NeurIPS 2024poster

Top-down Bird's Eye View (BEV) maps are a popular perception representation for ground robot navigation due to their richness and flexibility for downstream tasks. While recent methods have shown promise for predicting BEV maps from First-Person View (FPV) images, their generalizability is limited t…

2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

PhysORD: A Neuro-Symbolic Approach for Physics-infused Motion Prediction in Off-road Driving

IROS 2024poster

Motion prediction is critical for autonomous off-road driving, however, it presents significantly more challenges than on-road driving because of the complex interaction between the vehicle and the terrain. Traditional physics-based approaches encounter difficulties in accurately modeling dynamic sy…

Cited by 8SourcecodeScholar
2024

Practical Measurements of Translucent Materials with Inter-Pixel Translucency Prior

CVPR 2024poster

Material appearance is a key component of photorealism with a pronounced impact on human perception. Although there are many prior works targeting at measuring opaque materials using light-weight setups (e.g. consumer-level cameras) little attention is paid on acquiring the optical properties of tra…

Cited by 1SourcePDFScholar
2024

ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation

CoRL 2024poster

Representing robotic manipulation tasks as constraints that associate the robot and the environment is a promising way to encode desired robot behaviors. However, it remains unclear how to formulate the constraints such that they are 1) versatile to diverse tasks, 2) free of manual labeling, and 3)…

Cited by 97SourceScholar
2024

Real-Time Estimation for the Swimming Direction of Robotic Fish Based on IMU Sensors*

ICRA 2024poster

An increasing number of underwater robots inspired by Carangidae are developed, which is characterized by high efficiency and flexibility. However, estimating the swimming direction of these robotic fish is challenging due to the constant swinging of the head during movement, which complicates preci…

Cited by 0SourceScholar
2024

Salient Sparse Visual Odometry With Pose-Only Supervision

RA-L 2024

Visual Odometry (VO) is vital for the navigation of autonomous systems, providing accurate position and orientation estimates at reasonable costs. While traditional VO methods excel in some conditions, they struggle with challenges like variable lighting and motion blur. Deep learning-based VO, thou

Cited by 14SourceScholar
2024

TRANSIC: Sim-to-Real Policy Transfer by Learning from Online Correction

CoRL 2024poster

Learning in simulation and transferring the learned policy to the real world has the potential to enable generalist robots. The key challenge of this approach is to address simulation-to-reality (sim-to-real) gaps. Previous methods often require domain-specific knowledge *a priori*. We argue that a…

Cited by 28SourcecodeScholar
2024

United We Stand, Divided We Fall: Fingerprinting Deep Neural Networks via Adversarial Trajectories

NeurIPS 2024poster

In recent years, deep neural networks (DNNs) have witnessed extensive applications, and protecting their intellectual property (IP) is thus crucial. As a non-invasive way for model IP protection, model fingerprinting has become popular. However, existing single-point based fingerprinting methods are…

Cited by 0SourcePDFScholar
2024

iMTSP: Solving Min-Max Multiple Traveling Salesman Problem with Imperative Learning

IROS 2024poster

This paper considers a Min-Max Multiple Traveling Salesman Problem (MTSP), where the goal is to find a set of tours, one for each agent, to collectively visit all the cities while minimizing the length of the longest tour. Though MTSP has been widely studied, obtaining near-optimal solutions for lar…

Cited by 4SourcecodeScholar
2023

Animal3D: A Comprehensive Dataset of 3D Animal Pose and Shape

ICCV 2023poster

Accurately estimating the 3D pose and shape is an essential step towards understanding animal behavior, and can potentially benefit many downstream applications, such as wildlife conservation. However, research in this area is held back by the lack of a comprehensive and diverse dataset with high-qu…

Cited by 24PDFScholar
2023

Boundary Unlearning: Rapid Forgetting of Deep Networks via Shifting the Decision Boundary

CVPR 2023poster

The practical needs of the "right to be forgotten" and poisoned data removal call for efficient machine unlearning techniques, which enable machine learning models to unlearn, or to forget a fraction of training data and its lineage. Recent studies on machine unlearning for deep neural networks (DNN…

Cited by 88SourcePDFScholar
2023

Commdre: Document-Level Relation Extraction with Self-Supervised Commonsense Learning

ICASSP 2023accepted

Document-level relation extraction (DocRE) is a more challenging task for which multi-label and multi-entity problems need to be resolved effectively than its sentence-level counterpart. It aims at extracting relationships between two entities at once while taking into account significant cross-sent…

Cited by 0SourceScholar
2023

HS-Pose: Hybrid Scope Feature Extraction for Category-Level Object Pose Estimation

CVPR 2023poster

In this paper, we focus on the problem of category-level object pose estimation, which is challenging due to the large intra-category shape variation. 3D graph convolution (3D-GC) based methods have been widely used to extract local geometric features, but they have limitations for complex shaped ob…

2023

MimicPlay: Long-Horizon Imitation Learning by Watching Human Play

CoRL 2023oral

Imitation learning from human demonstrations is a promising paradigm for teaching robots manipulation skills in the real world. However, learning complex long-horizon tasks often requires an unattainable amount of demonstrations. To reduce the high data requirement, we resort to human play data - vi…

Cited by 187SourcecodeScholar
2023

Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task Similarity

NeurIPS 2023poster

Multi-agent reinforcement learning (MARL) has primarily focused on solving a single task in isolation, while in practice the environment is often evolving, leaving many related tasks to be solved. In this paper, we investigate the benefits of meta-learning in solving multiple MARL tasks collectively…

Cited by 8SourcePDFScholar
2023

NIRVANA: Neural Implicit Representations of Videos With Adaptive Networks and Autoregressive Patch-Wise Modeling

CVPR 2023poster

Implicit Neural Representations (INR) have recently shown to be powerful tool for high-quality video compression. However, existing works are limiting as they do not explicitly exploit the temporal redundancy in videos, leading to a long encoding time. Additionally, these methods have fixed architec…

2023

NOIR: Neural Signal Operated Intelligent Robots for Everyday Activities

CoRL 2023poster

We present Neural Signal Operated Intelligent Robots (NOIR), a general-purpose, intelligent brain-robot interface system that enables humans to command robots to perform everyday activities through brain signals. Through this interface, humans communicate their intended objects of interest and actio…

Cited by 18SourceScholar
2023

NeRF-Loc: Transformer-Based Object Localization Within Neural Radiance Fields

RA-L 2023

Neural Radiance Fields (NeRFs) have become a widely-applied scene representation technique in recent years, showing advantages for robot navigation and manipulation tasks. To further advance the utility of NeRFs for robotics, we propose a transformer-based framework, <monospace xmlns:mml="http://www

Cited by 14SourceScholar
2023

Off-Policy Evaluation With Online Adaptation for Robot Exploration in Challenging Environments

RA-L 2023

Autonomous exploration has many important applications. However, classic information gain-based or frontier-based exploration only relies on the robot current state to determine the immediate exploration goal, which lacks the capability of predicting the value of future states and thus leads to inef

Cited by 18SourceScholar
2023

POMDP-Guided Active Force-Based Search for Robotic Insertion

IROS 2023poster

In robotic insertion tasks where the uncertainty exceeds the allowable tolerance, a good search strategy is essential for successful insertion and significantly influences efficiency. The commonly used blind search method is time-consuming and does not exploit the rich contact information. In this p…

Cited by 3SourceScholar
2023

PRRD: Pixel-Region Relation Distillation For Efficient Semantic Segmentation

ICASSP 2023accepted

Current state-of-the-art semantic segmentation methods usually require high computational resources for accurate segmentation. Knowledge distillation has been one promising way to achieve a good trade-off between accuracy and efficiency. However, current distillation methods focus on transferring th…

Cited by 0SourceScholar
2023

PanoGRF: Generalizable Spherical Radiance Fields for Wide-baseline Panoramas

NeurIPS 2023poster

Achieving an immersive experience enabling users to explore virtual environments with six degrees of freedom (6DoF) is essential for various applications such as virtual reality (VR). Wide-baseline panoramas are commonly used in these applications to reduce network bandwidth and storage requirements…

Cited by 9SourcePDFScholar
2023

Primitive Skill-Based Robot Learning from Human Evaluative Feedback

IROS 2023poster

Reinforcement learning (RL) algorithms face significant challenges when dealing with long-horizon robot manipulation tasks in real-world environments due to sample inefficiency and safety issues. To overcome these challenges, we propose a novel framework, SEED, which leverages two approaches: reinfo…

Cited by 12SourcecodeScholar
2023

PyPose: A Library for Robot Learning With Physics-Based Optimization

CVPR 2023poster

Deep learning has had remarkable success in robotic perception, but its data-centric nature suffers when it comes to generalizing to ever-changing environments. By contrast, physics-based optimization generalizes better, but it does not perform as well in complicated tasks due to the lack of high-le…

2023

Sequential Dexterity: Chaining Dexterous Policies for Long-Horizon Manipulation

CoRL 2023poster

Many real-world manipulation tasks consist of a series of subtasks that are significantly different from one another. Such long-horizon, complex tasks highlight the potential of dexterous hands, which possess adaptability and versatility, capable of seamlessly transitioning between different modes o…

Cited by 48SourcecodeScholar
2023

Streaming Algorithms and Lower Bounds for Estimating Correlation Clustering Cost

NeurIPS 2023poster

Correlation clustering is a fundamental optimization problem at the intersection of machine learning and theoretical computer science. Motivated by applications to big data processing, recent years have witnessed a flurry of results on this problem in the streaming model. In this model, the algori…

Cited by 4SourcePDFScholar
2023

Vision-based Six-Dimensional Peg-in-Hole for Practical Connector Insertion

ICRA 2023poster

We study six-dimensional (6D) perceptive peg-in-hole problem for practical connector insertion task in this paper. To enable the manipulator system to handle different types of pegs in complex environment, we develop a perceptive robotic assembly system that utilizes an in-hand RGB-D camera for peg-…

Cited by 11SourceScholar
2023

VoxDet: Voxel Learning for Novel Instance Detection

NeurIPS 2023spotlight

Detecting unseen instances based on multi-view templates is a challenging problem due to its open-world nature. Traditional methodologies, which primarily rely on $2 \mathrm{D}$ representations and matching techniques, are often inadequate in handling pose variations and occlusions. To solve this, w…

2023

VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

CoRL 2023oral

Large language models (LLMs) are shown to possess a wealth of actionable knowledge that can be extracted for robot manipulation in the form of reasoning and planning. Despite the progress, most still rely on pre-defined motion primitives to carry out the physical interactions with the environment, w…

Cited by 564SourcecodeScholar
2022

A Dual Representation Framework for Robot Learning with Human Guidance

CoRL 2022poster

The ability to interactively learn skills from human guidance and adjust behavior according to human preference is crucial to accelerating robot learning. But human guidance is an expensive resource, calling for methods that can learn efficiently. In this work, we argue that learning is more efficie…

Cited by 15SourceScholar
2022

A Mean-Field Game Approach to Cloud Resource Management with Function Approximation

NeurIPS 2022accept

Reinforcement learning (RL) has gained increasing popularity for resource management in cloud services such as serverless computing. As self-interested users compete for shared resources in a cluster, the multi-tenancy nature of serverless platforms necessitates multi-agent reinforcement learning (M…

Cited by 26SourcePDFScholar
2022

A Snake-Inspired Multi-Segmented Magnetic Soft Robot Towards Medical Applications

RA-L 2022

Magnetically-actuated soft robots have potential for medical application but require further innovation on functionality and biocompatibility. In this letter, a multi-segmented snake-inspired soft robot with dissolvable and biocompatible segments is designed. The actuation response under external ma

Cited by 49SourceScholar
2022

AirDOS: Dynamic SLAM benefits from Articulated Objects

ICRA 2022poster

Dynamic Object-aware SLAM (DOS) exploits object-level information to enable robust motion estimation in dynamic environments. Existing methods mainly focus on identifying and excluding dynamic objects from the optimization. In this paper, we show that feature-based visual SLAM systems can also benef…

Cited by 64SourcecodeScholar
2022

AirDet: Few-Shot Detection without Fine-Tuning for Autonomous Exploration

ECCV 2022poster

"Few-shot object detection has attracted increasing attention and rapidly progressed in recent years. However, the requirement of an exhaustive offline fine-tuning stage in existing methods is time-consuming and significantly hinders their usage in online applications such as autonomous exploration…

2022

AirObject: A Temporally Evolving Graph Embedding for Object Identification

CVPR 2022poster

Object encoding and identification are vital for robotic tasks such as autonomous exploration, semantic scene understanding, and re-localization. Previous approaches have attempted to either track objects or generate descriptors for object identification. However, such systems are limited to a "fixe…

Cited by 6PDFcodeScholar
2022

BEHAVIOR-1K: A Benchmark for Embodied AI with 1,000 Everyday Activities and Realistic Simulation

CoRL 2022oral

We present BEHAVIOR-1K, a comprehensive simulation benchmark for human-centered robotics. BEHAVIOR-1K includes two components, guided and motivated by the results of an extensive survey on "what do you want robots to do for you?". The first is the definition of 1,000 everyday activities, grounded in…

Cited by 205SourceScholar
2022

Discrete Cross-Modal Alignment Enables Zero-Shot Speech Translation

EMNLP 2022main

End-to-end Speech Translation (ST) aims at translating the source language speech into target language text without generating the intermediate transcriptions. However, the training of end-to-end methods relies on parallel ST data, which are difficult and expensive to obtain. Fortunately, the superv…

2022

Electric Sense Based Pose Estimation and Localization for Small Underwater Robots

RA-L 2022

Accurate pose estimation and localization technology is always a challenge for small underwater robots, since the underwater lighting conditions could limit the use of cameras while the cramped environments restrict the use of sonars. In nature, some fishes perceive other creatures by sensing the we

Cited by 19SourceScholar
2022

GOCA: Guided Online Cluster Assignment for Self-Supervised Video Representation Learning

ECCV 2022poster

"Clustering is a ubiquitous tool in unsupervised learning. Most of the existing self-supervised representation learning methods typically cluster samples based on visually dominant features. While this works well for image-based selfsupervision, it often fails for videos, which require understanding…

2022

Revisiting Domain Generalized Stereo Matching Networks From a Feature Consistency Perspective

CVPR 2022poster

Despite recent stereo matching networks achieving impressive performance given sufficient training data, they suffer from domain shifts and generalize poorly to unseen domains. We argue that maintaining feature consistency between matching pixels is a vital factor for promoting the generalization ca…

Cited by 79PDFcodeScholar
2022

Robotic Interestingness via Human-Informed Few-Shot Object Detection

IROS 2022poster

Interestingness recognition is crucial for decision making in autonomous exploration for mobile robots. Previous methods proposed an unsupervised online learning approach that can adapt to environments and detect interesting scenes quickly, but lack the ability to adapt to human-informed interesting…

Cited by 3SourceScholar
2022

SADN: Learned Light Field Image Compression with Spatial-Angular Decorrelation

ICASSP 2022accepted

Light field image becomes one of the most promising media types for immersive video applications. In this paper, we propose a novel end-to-end spatial-angular-decorrelated network (SADN) for high-efficiency light field image compression. Different from the existing methods that exploit either spatia…

Cited by 0SourceScholar
2022

Single-pass Streaming Lower Bounds for Multi-armed Bandits Exploration with Instance-sensitive Sample Complexity

NeurIPS 2022accept

Motivated by applications to process massive datasets, we study streaming algorithms for pure exploration in Stochastic Multi-Armed Bandits (MABs). This problem was first formulated by Assadi and Wang [STOC 2020] as follows: A collection of $n$ arms with unknown rewards are arriving one by one in a…

Cited by 11SourcePDFScholar
2021

Co-GAIL: Learning Diverse Strategies for Human-Robot Collaboration

CoRL 2021poster

We present a method for learning human-robot collaboration policy from human-human collaboration demonstrations. An effective robot assistant must learn to handle diverse human behaviors shown in the demonstrations and be robust when the humans adjust their strategies during online task execution. O…

Cited by 49SourceScholar
2021

Decentralized Circle Formation Control for Fish-like Robots in the Real-world via Reinforcement Learning

ICRA 2021poster

In this paper, the circle formation control problem is addressed for a group of cooperative underactuated fish-like robots involving unknown nonlinear dynamics and disturbances. Based on the reinforcement learning and cognitive consistency theory, we propose a decentralized controller without the kn…

Cited by 27SourceScholar
2021

Encirclement Guaranteed Cooperative Pursuit with Robust Model Predictive Control

IROS 2021poster

This paper studies a novel encirclement guaranteed cooperative pursuit problem involving N pursuers and a single evader in an unbounded two-dimensional game domain. Throughout the game, the pursuers are required to maintain encirclement of the evader, i.e., the evader should always stay inside the c…

Cited by 12SourceScholar
2021

FOP: Factorizing Optimal Joint Policy of Maximum-Entropy Multi-Agent Reinforcement Learning

ICML 2021spotlight

Value decomposition recently injects vigorous vitality into multi-agent actor-critic methods. However, existing decomposed actor-critic methods cannot guarantee the convergence of global optimum. In this paper, we present a novel multi-agent actor-critic method, FOP, which can factorize the optimal…

Cited by 102SourcePDFScholar
2021

Generalization Through Hand-Eye Coordination: An Action Space for Learning Spatially-Invariant Visuomotor Control

IROS 2021poster

Imitation Learning (IL) is an effective framework to learn visuomotor skills from offline demonstration data. However, IL methods often fail to generalize to new scene configurations not covered by training data. On the other hand, humans can manipulate objects in varying conditions. Key to such cap…

Cited by 35SourceScholar
2021

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

CoRL 2021oral

Imitating human demonstrations is a promising approach to endow robots with various manipulation capabilities. While recent advances have been made in imitation learning and batch (offline) reinforcement learning, a lack of open-source human datasets and reproducible learning methods make assessing…

Cited by 523SourcecodeScholar
2020

6-PACK: Category-level 6D Pose Tracker with Anchor-Based Keypoints

ICRA 2020poster

We present 6-PACK, a deep learning approach to category-level 6D object pose tracking on RGB-D data. Our method tracks in real time novel object instances of known object categories such as bowls, laptops, and mugs. 6-PACK learns to compactly represent an object by a handful of 3D keypoints, based o…

Cited by 190SourcecodeScholar
2020

Autonomous Obstacle Avoidance for UAV based on Fusion of Radar and Monocular Camera

IROS 2020poster

UAVs face many challenges in autonomous obstacle avoidance in large outdoor scenarios, specifically the long communication distance from ground stations. The computing power of onboard computers is limited, and the unknown obstacles cannot be accurately detected. In this paper, an autonomous obstacl…

Cited by 45SourceScholar
2020

Intensity Scan Context: Coding Intensity and Geometry Relations for Loop Closure Detection

ICRA 2020poster

Loop closure detection is an essential and challenging problem in simultaneous localization and mapping (SLAM). It is often tackled with light detection and ranging (LiDAR) sensor due to its view-point and illumination invariant properties. Existing works on 3D loop closure detection often leverage…

Cited by 328SourcecodeScholar
2020

Motion Planning for Heterogeneous Unmanned Systems under Partial Observation from UAV

IROS 2020poster

For heterogeneous unmanned systems composed of unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs), using UAVs serve as eyes to assist UGVs in motion planning is a promising research direction due to the UAVs’ vast view scope. However, its limitations on flight altitude prevent the U…

Cited by 7SourceScholar
2020

Stacking Networks Dynamically for Image Restoration Based on the Plug-and-Play Framework

ECCV 2020poster

Recently, stacked networks show powerful performance in Image Restoration, such as challenging motion deblurring problems. However, the number of stacking levels is a hyper-parameter fine-tuned manually, making the stacking levels static during training without theoretical explanations for optimal s…

Cited by 12SourcePDFScholar
2020

Supervised Contrastive Learning

NeurIPS 2020poster

Contrastive learning applied to self-supervised representation learning has seen a resurgence in recent years, leading to state of the art performance in the unsupervised training of deep image models. Modern batch contrastive approaches subsume or significantly outperform traditional contrastive lo…

2020

SwingBot: Learning Physical Features from In-hand Tactile Exploration for Dynamic Swing-up Manipulation

IROS 2020poster

Several robot manipulation tasks are extremely sensitive to variations of the physical properties of the manipulated objects. One such task is manipulating objects by using gravity or arm accelerations, increasing the importance of mass, center of mass, and friction information. We present SwingBot,…

Cited by 120SourceScholar
2020

TartanAir: A Dataset to Push the Limits of Visual SLAM

IROS 2020poster

We present a challenging dataset, the TartanAir, for robot navigation tasks and more. The data is collected in photo-realistic simulation environments with the presence of moving objects, changing light and various weather conditions. By collecting data in simulations, we are able to obtain multi-mo…

Cited by 406SourcecodeScholar
2020

Visual Memorability for Robotic Interestingness via Unsupervised Online Learning

ECCV 2020poster

In this paper, we explore the problem of interesting scene prediction for mobile robots. This area is currently underexplored but is crucial for many practical applications such as autonomous exploration and decision making. Inspired by industrial demands, we first propose a novel translation-invari…

2019

DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion

CVPR 2019poster

A key technical challenge in performing 6D object pose estimation from RGB-D image is to fully leverage the two complementary data sources. Prior works either extract information from the RGB image and depth separately or use costly post-processing steps, limiting their performances in highly clutte…

Cited by 1292PDFScholar
2019

TendencyRL: Multi-stage Discriminative Hints for Efficient Goal-Oriented Reverse Curriculum Learning

IROS 2019poster

Deep reinforcement learning algorithms have been proven successful in a variety of simulation tasks with dense reward feedback. However, real-world RL applications, e.g. robotic manipulation, remain challenging as most of them are multi-stage and a positive reward can only be received when the final…

Cited by 4SourceScholar
2019

Vision-based Automatic Control of a 5-Fingered Assistive Robotic Manipulator for Activities of Daily Living

IROS 2019poster

Assistive Robotic Manipulators (ARMs) play an important role for people with upper-limb disabilities and the elderly by helping them complete Activities of Daily Living (ADLs). However, as the objects to handle in ADLs differ in size, shape and manipulation constraints, many two or three fingered en…

Cited by 10SourceScholar
2018

Correlation Flow: Robust Optical Flow Using Kernel Cross-Correlators

ICRA 2018poster

Robust velocity and position estimation is crucial for autonomous robot navigation. The optical flow based methods for autonomous navigation have been receiving increasing attentions in tandem with the development of micro unmanned aerial vehicles. This paper proposes a kernel cross-correlator (KCC)…

Cited by 26SourcecodeScholar
2018

Efficient Global Point Cloud Registration by Matching Rotation Invariant Features Through Translation Search

ECCV 2018poster

Three-dimensional rigid point cloud registration has many applications in computer vision and robotics. Local methods tend to fail, causing global methods to be needed, when the relative transformation is large or the overlap ratio is small. Most existing global methods utilize BnB optimization over…

Cited by 91SourcePDFScholar
2018

Robust Target-Relative Localization with Ultra-Wideband Ranging and Communication

ICRA 2018poster

In this paper we propose a method to achieve relative positioning and tracking of a target by a quadcopter using Ultra-wideband (UWB) ranging sensors, which are strategically installed to help retrieve both relative position and bearing between the quadcopter and target. To achieve robust localizati…

Cited by 77SourceScholar
2018

Robust image stitching with multiple registrations

ECCV 2018poster

Panorama creation is one of the most widely deployed techniques in computer vision. In addition to industry applications such as Google Street View, it is also used by millions of consumers in smartphones and other cameras. Traditionally, the problem is decomposed into three phases: registration, wh…

Cited by 77SourcePDFScholar
2017

CSMA/CA-based electrocommunication system design for underwater robot groups

IROS 2017poster

Underwater communication is particularly challenging for small submarine robots that have limited power and size constraints. Inspired by weakly electric fish, a novel electric current communication (termed electrocommunication) system has been developed for small underwater robots in our previous s…

Cited by 17SourceScholar
2017

Locality Sensitive Hashing based deepmatching for optical flow estimation

ICASSP 2017accepted

DeepMatching (DM) is one of the state-of-art matching algorithms to compute quasi-dense correspondences between images. Recent optical flow methods use DeepMatching to find initial image correspondences and achieves outstanding performance. However, the key building block of DeepMatching, the correl…

Cited by 0SourceScholar
2017

Ultra-wideband aided fast localization and mapping system

IROS 2017poster

This paper proposes an ultra-wideband (UWB) aided localization and mapping system that leverages on inertial sensor and depth camera. Inspired by the fact that visual odometry (VO) system, regardless of its accuracy in the short term, still faces challenges with accumulated errors in the long run or…

Cited by 103SourcecodeScholar
2016

Adaptive control for robot navigation in human environments based on social force model

ICRA 2016

In this paper, we introduce a novel control scheme based on the social force model for robots navigating in human environments. Social proxemics potential field is constructed based on the theory of proxemics and used to generate social interaction force for design of robot motion control. A combine

Cited by 15SourceScholar
2016

Speed evaluation of a freely swimming robotic fish with an artificial lateral line

ICRA 2016

Artificial lateral line has been drawing an increasing attention recently for its potential applications in robotics. Experiments are usually conducted with a bioinspired robot in a controlled environment, where the sensing platform is held stationary or slowly driven with a simple linear motion. In

Cited by 32SourceScholar
2015

Semantic Object Segmentation via Detection in Weakly Labeled Video

CVPR 2015poster

Semantic object segmentation in video is an important step for large-scale multimedia analysis. In many cases, however, semantic objects are only tagged at video-level, making them difficult to be located and segmented. To address this problem, this paper proposes an approach to segment semantic obj…

Cited by 94SourcePDFScholar