← Search

Wei Yang

128 accepted papers

2026

Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual Learning

ICLR 2026poster

While scaling individual Large Language Models (LLMs) has delivered remarkable progress, the next frontier lies in scaling collaboration through multi-agent systems (MAS). However, purely autonomous MAS remain ``closed-world'' systems, constrained by the static knowledge horizon of pre-trained model…

Cited by 0SourcecodeScholar
2026

Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Model

RSS 2026poster

Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for evaluating Vision-Language-Action (VLA) systems. However, their reliability as evaluation benchmarks is often limited by the representation gap between low-dimensional actions and high-dimensional video synthesis. Thi…

Cited by 0SourceScholar
2026

From Assistant to Independent Developer — Are GPTs Ready for Software Development?

ICLR 2026poster

Large language models (LLMs) have demonstrated remarkable capability in function-level code generation tasks. Unlike isolated functions, real-world applications demand reasoning over the entire software system: developers must orchestrate how different components interact, maintain consistency acro…

Cited by 0SourceScholar
2026

Hypergraph-State Collaborative Reasoning for Multi-Object Tracking

CVPR 2026

Motion reasoning serves as the cornerstone of multi-object tracking (MOT), as it enables consistent association of targets across frames. However, existing motion estimation approaches face two major limitations: (1) instability caused by noisy or probabilistic predictions, and (2) vulnerability und

Cited by 0SourcecodeScholar
2026

KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits

ICML 2026poster

High-performance GPU kernels are critical for efficient LLM serving, yet their optimization remains a bottleneck requiring deep system expertise. While code LLMs show promise in generating functionally correct code, kernel optimization is intrinsically a search problem over a vast optimization space…

Cited by 0SourceScholar
2026

LATO: 3D Mesh Flow Matching with Structured TOpology Preserving LAtents

ICML 2026poster

In this paper, we introduce LATO, a novel topology-preserving latent representation that enables scalable, flow matching-based synthesis of explicit 3D meshes. LATO represents a mesh as a Vertex Displacement Field (VDF) anchored on surface, incorporating a sparse voxel Variational Autoencoder (VAE) …

Cited by 0SourceScholar
2026

LangField4D: Learning Identity-Adaptive and Spatio-Temporal Continuous 4D Language Fields for Dynamic Scenes

CVPR 2026

Constructing a 4D language field that supports open-vocabulary queries is essential for semantic perception and interaction in dynamic environments. Existing 4D Gaussian-based approaches face two major challenges. First, the assumption of a static identity per Gaussian leads to semantic inconsistenc

Cited by 0SourceScholar
2026

Learning to Deliberate: Meta-policy Collaboration for Agentic LLMs with Multi-agent Reinforcement Learning

AAAI 2026technical

Multi-agent systems of large language models (LLMs) show promise for complex reasoning, but their effectiveness is often limited by fixed collaboration protocols. These frameworks typically focus on macro-level orchestration while overlooking agents’ internal deliberative capabilities. This critical

Cited by 0SourcePDFScholar
2026

LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing

ICLR 2026poster

Recent attempts to combine low-rank adaptation (LoRA) with mixture-of-experts (MoE) for multi-task adaptation of Large Language Models (LLMs) often replace whole attention/FFN layers with switch experts or append parallel expert branches, undermining parameter efficiency and limiting task specializa…

Cited by 0SourceScholar
2026

MeshRipple: Structured Autoregressive Generation of Artist-Meshes

CVPR 2026

Meshes serve as a primary representation for 3D assets. Autoregressive mesh generators serialize faces into sequences and train on truncated segments with sliding-window inference to cope with memory limits. However, this mismatch breaks long-range geometric dependencies, producing holes and fragmen

Cited by 0SourceScholar
2026

Multi-Sensor Fusion and Multi-Agent Control for Autonomous Environmental Exploration With Rat Robots

RA-L 2026

This letter presents a bio-robotic multi-agent framework that integrates multi-sensor fusion with multi-agent control strategies for rat robots. The sensors equipped with rat robots, including laser-ranging sensor, camera and infrared sensor. We develop two multi-agent control strategies: a leader-f

Cited by 0SourceScholar
2026

PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization

ICML 2026poster

Unifying the complementary strengths of diverse Vision Foundation Models (VFMs) into a single efficient model is highly desirable but challenged by the negative transfer inherent in monolithic distillation. To address these feature conflicts, we introduce \textbf{PRISM}, a novel dual-stream Mixture-…

Cited by 0SourceScholar
2026

ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation

CVPR 2026

The ability to extrapolate dynamic 3D scenes beyond the observed timeframe is fundamental to advancing physical world understanding and predictive modeling. Existing dynamic 3D reconstruction methods have achieved high-fidelity rendering of temporal interpolation, but typically lack physical consist

Cited by 0SourceScholar
2026

``Someone Hid It!'': Query-Agnostic Black-Box Attacks on LLM-Based Retrieval

ICML 2026poster

Large language models (LLMs) have been serving as effective backbones for retrieval systems, including Retrieval-Augmentation-Generation (RAG), Dense Information Retriever (IR), and Agent Memory Retrieval. Recent studies have demonstrated that such LLM-based Retrieval (LLMR) is vulnerable to adversa…

Cited by 0SourceScholar
2025

AAKR: Adversarial Attack-based Knowledge Retention for Continual Semantic Segmentation

AAAI 2025technical

In the context of Continual Semantic Segmentation (CSS), replay-based methods tend to achieve better performance than knowledge distillation-based ones, as the former utilizes additional data to transfer old knowledge. However, this advantage is at the cost of necessitating additional space for sto…

Cited by 0SourcePDFScholar
2025

Automatic Mathematic In-Context Example Generation for LLM Using Multi-Modal Consistency

COLING 2025main

Large Language Models (LLMs) have advanced Natural Language Processing (NLP) tasks but are limited in mathematical reasoning. To address this, few-shot examples are used in prompts for in-context learning. However, existing methods require annotated datasets, resulting in higher computational costs…

2025

Coarse-to-Fine Grounded Memory for LLM Agent Planning

EMNLP 2025

Recent advancements in Large Language Models (LLMs) have driven growing interest in LLM-based agents for complex planning tasks. To avoid costly agent training, many studies adopted memory mechanism that enhances LLM with offline experiences or online trajectory analysis. However, existing works foc

Cited by 0SourcePDFScholar
2025

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference Scoped Exploration

CoRL 2025poster

Hand–object motion-capture (MoCap) repositories provide abundant, contact-rich human demonstrations for scaling dexterous manipulation on robots. Yet demonstration inaccuracy and embodiment gaps between human and robot hands challenge direct policy learning. Existing pipelines adapt a three-stage wo…

Cited by 0SourceScholar
2025

Energy-Efficient Omnidirectional Locomotion for Wheeled Quadrupeds via Predictive Energy-Aware Nominal Gait Selection

IROS 2025

Wheeled-legged robots combine the efficiency of wheels with the versatility of legs, but face significant energy optimization challenges when navigating diverse environments. In this work, we present a hierarchical control framework that integrates predictive power modeling with residual reinforceme

Cited by 0SourceScholar
2025

Exoskeleton Gait Adaptation Framework via Hm-DMP and PI2 Optimization for Dynamic Patient Mobility Matching

IROS 2025

Repetitive gait training with lower-limb exoskeletons enhances neuroplasticity and reduces muscle atrophy by promoting patient engagement in active rehabilitation training. Importantly, the therapeutic efficacy of such engagement critically depends on providing patients with task difficulty levels m

Cited by 0SourceScholar
2025

GA-S3: Comprehensive Social Network Simulation with Group Agents

ACL 2025finding

Social network simulation is developed to provide a comprehensive understanding of social networks in the real world, which can be leveraged for a wide range of applications such as group behavior emergence, policy optimization, and business strategy development. However, billions of individuals and…

2025

Leaving No OOD Instance Behind: Instance-Level OOD Fine-Tuning for Anomaly Segmentation

NeurIPS 2025poster

Out-of-distribution (OOD) fine-tuning has emerged as a promising approach for anomaly segmentation. Current OOD fine-tuning strategies typically employ global-level objectives, aiming to guide segmentation models to accurately predict a large number of anomaly pixels. However, these strategies often…

Cited by 0SourceScholar
2025

MaGS: Reconstructing and Simulating Dynamic 3D Objects with Mesh-adsorbed Gaussian Splatting

ICCV 2025poster

3D reconstruction and simulation, although interrelated, have distinct objectives: reconstruction requires a flexible 3D representation that can adapt to diverse scenes, while simulation needs a structured representation to model motion principles effectively. This paper introduces the Mesh-adsorbed…

Cited by 0SourcePDFScholar
2025

OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval

ACL 2025long

Vision-language retrieval-augmented generation (RAG) has become an effective approach for tackling Knowledge-Based Visual Question Answering (KB-VQA), which requires external knowledge beyond the visual content presented in images. The effectiveness of Vision-language RAG systems hinges on multimoda…

2025

Optimized View and Geometry Distillation from Multi-view Diffuser

IJCAI 2025

Generating multi-view images from a single input view using image-conditioned diffusion models is a recent advancement and has shown considerable potential. However, issues such as the lack of consistency in synthesized views and over-smoothing in extracted geometry persist. Previous methods integra

2025

RP-PGD: Boosting Segmentation Robustness with a Region-and-Prototype Based Adversarial Attack

AAAI 2025technical

Adversarial attack and defense have been extensively explored in classification tasks, but their study in semantic segmentation remains limited. Moreover, current attacks fail to act as strong underlying attacks for adversarial training (AT), making it difficult to achieve segmentation robustness ag…

Cited by 0SourcePDFScholar
2025

Ref-GS: Directional Factorization for 2D Gaussian Splatting

CVPR 2025poster

In this paper, we introduce Ref-GS, a novel approach for directional light factorization in 2D Gaussian splatting, which enables photorealistic view-dependent appearance rendering and precise geometry recovery. Ref-GS builds upon the deferred rendering of Gaussian splatting and applies directional e…

2025

SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding

CVPR 2025poster

Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding,…

Cited by 1SourcePDFScholar
2025

StdGEN: Semantic-Decomposed 3D Character Generation from Single Images

CVPR 2025poster

We present StdGEN, an innovative pipeline for generating semantically decomposed high-quality 3D characters from single images, enabling broad applications in virtual reality, gaming, and filmmaking, etc. Unlike previous methods which struggle with limited decomposability, unsatisfactory quality, an…

2025

Stop Diverse OOD Attacks: Knowledge Ensemble for Reliable Defense

AAAI 2025technical

Enhancing defense through model ensemble is an emerging trend, where the challenge lies in how to use ensemble knowledge to counter Out-of-Distribution (OOD) attacks. In this paper, we propose the Reliable Defense Ensemble (REE) to address this issue. REE optimizes the ensemble knowledge of models t…

Cited by 0SourcePDFScholar
2025

Structured Spectral Reasoning for Frequency-Adaptive Multimodal Recommendation

NeurIPS 2025poster

Multimodal recommendation aims to integrate collaborative signals with heterogeneous content such as visual and textual information, but remains challenged by modality-specific noise, semantic inconsistency, and unstable propagation over user–item graphs. These issues are often exacerbated by naive…

Cited by 0SourceScholar
2025

Temporal Coherent Object Flow for Multi-Object Tracking

AAAI 2025technical

Multi-object tracking is a challenging vision task that requires simultaneous reasoning about object detection and object association. Conventional solutions use frame as the basic unit and typically rely on a motion predictor that exploits the appearance features to associate detected candidates, l…

Cited by 0SourcePDFScholar
2025

Tip the Scales: Achieving Balance in Adversarial Examples Across Modalities

ICASSP 2025accepted

In the field of multimodal learning, controlling the training of unimodal encoders from different perspectives is a primary approach to addressing Training Imbalance. However, the inherent capacity limitations of the modality affect the model’s capability. Therefore, generating adversarial examples…

Cited by 0SourceScholar
2025

VT-Refine: Learning Bimanual Assembly with Visuo-Tactile Feedback via Simulation Fine-Tuning

CoRL 2025poster

Humans excel at bimanual assembly tasks by adapting to rich tactile feedback—a capability that remains difficult to replicate in robots through behavioral cloning alone, due to the suboptimality and limited diversity of human demonstrations. In this work, we present VT-Refine, a visuo-tactile policy…

Cited by 0SourcecodeScholar
2025

Video Anomaly Detection with Motion and Appearance Guided Patch Diffusion Model

AAAI 2025technical

A recent endeavor in one class of video anomaly detection is to leverage diffusion models and posit the task as a generation problem, where the diffusion model is trained to recover normal patterns exclusively, thus reporting abnormal patterns as outliers. Yet, existing attempts neglect the various…

2024

AMD: Anatomical Motion Diffusion with Interpretable Motion Decomposition and Fusion

AAAI 2024technical

Generating realistic human motion sequences from text descriptions is a challenging task that requires capturing the rich expressiveness of both natural language and human motion. Recent advances in diffusion models have enabled significant progress in human motion synthesis. However, existing metho…

Cited by 3SourcePDFScholar
2024

Attacking Transformers with Feature Diversity Adversarial Perturbation

AAAI 2024technical

Understanding the mechanisms behind Vision Transformer (ViT), particularly its vulnerability to adversarial perturbations, is crucial for addressing challenges in its real-world applications. Existing ViT adversarial attackers rely on labels to calculate the gradient for perturbation, and exhibit lo…

Cited by 5SourcePDFScholar
2024

Attacks on Continual Semantic Segmentation by Perturbing Incremental Samples

AAAI 2024technical

As an essential computer vision task, Continual Semantic Segmentation (CSS) has received a lot of attention. However, security issues regarding this task have not been fully studied. To bridge this gap, we study the problem of attacks in CSS in this paper. We first propose a new task, namely, attack…

Cited by 2SourcePDFScholar
2024

Coupled Mamba: Enhanced Multimodal Fusion with Coupled State Space Model

NeurIPS 2024poster

The essence of multi-modal fusion lies in exploiting the complementary information inherent in diverse modalities.However, most prevalent fusion methods rely on traditional neural architectures and are inadequately equipped to capture the dynamics of interactions across modalities, particularly in p…

Cited by 7SourcePDFScholar
2024

DiffusionTrack: Diffusion Model for Multi-Object Tracking

AAAI 2024technical

Multi-object tracking (MOT) is a challenging vision task that aims to detect individual objects within a single frame and associate them across multiple frames. Recent MOT approaches can be categorized into two-stage tracking-by-detection (TBD) methods and one-stage joint detection and tracking (JDT…

2024

Dynamic Feature Pruning and Consolidation for Occluded Person Re-identification

AAAI 2024technical

Occluded person re-identification (ReID) is a challenging problem due to contamination from occluders. Existing approaches address the issue with prior knowledge cues, such as human body key points and semantic segmentations, which easily fail in the presence of heavy occlusion and other humans as o…

2024

Entangled View-Epipolar Information Aggregation for Generalizable Neural Radiance Fields

CVPR 2024poster

Generalizable NeRF can directly synthesize novel views across new scenes eliminating the need for scene-specific retraining in vanilla NeRF. A critical enabling factor in these approaches is the extraction of a generalizable 3D representation by aggregating source-view features. In this paper we pro…

2024

FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects

CVPR 2024highlight

We present FoundationPose a unified foundation model for 6D object pose estimation and tracking supporting both model-based and model-free setups. Our approach can be instantly applied at test-time to a novel object without finetuning as long as its CAD model is given or a small number of reference…

2024

GenSeg: On Generating Unified Adversary for Segmentation

IJCAI 2024poster

Great advancements in semantic, instance, and panoptic segmentation have been made in recent years, yet the top-performing models remain vulnerable to imperceptible adversarial perturbation. Current attacks on segmentation primarily focus on a single task, and these methods typically rely on iterati…

2024

HateModerate: Testing Hate Speech Detectors against Content Moderation Policies

NAACL 2024findings

To protect users from massive hateful content, existing works studied automated hate speech detection. Despite the existing efforts, one question remains: Do automated hate speech detectors conform to social media content policies? A platform’s content policies are a checklist of content moderated b…

2024

Learning Pseudo 3D Guidance for View-consistent Texturing with 2D Diffusion

ECCV 2024poster

"Text-driven 3D texturing requires the generation of high-fidelity texture that conforms to given geometry and description. Recently, the high-quality text-to-image generation ability of 2D diffusion model has significantly promoted this task, by converting it into a texture optimization process gui…

2024

PTDE: Personalized Training with Distilled Execution for Multi-Agent Reinforcement Learning

IJCAI 2024poster

Centralized Training with Decentralized Execution (CTDE) has emerged as a widely adopted paradigm in multi-agent reinforcement learning, emphasizing the utilization of global information for learning an enhanced joint Q-function or centralized critic. In contrast, our investigation delves into harne…

Cited by 14SourcePDFScholar
2024

Progressive Text-to-Image Diffusion with Soft Latent Direction

AAAI 2024technical

In spite of the rapidly evolving landscape of text-to-image generation, the synthesis and manipulation of multiple entities while adhering to specific relational constraints pose enduring challenges. This paper introduces an innovative progressive synthesis and editing operation that systematically…

2024

SynH2R: Synthesizing Hand-Object Motions for Learning Human-to-Robot Handovers

ICRA 2024poster

Vision-based human-to-robot handover is an important and challenging task in human-robot interaction. Recent work has attempted to train robot policies by interacting with dynamic virtual humans in simulated environments, where the policies can later be transferred to the real world. However, a majo…

Cited by 20SourceScholar
2024

TIKP: Text-to-Image Knowledge Preservation for Continual Semantic Segmentation

AAAI 2024technical

Continual Semantic Segmentation (CSS) is an emerging trend, where catastrophic forgetting has been a perplexing problem. In this paper, we propose a Text-to-Image Knowledge Preservation (TIKP) framework to address this issue. TIKP applies Text-to-Image techniques to CSS by automatically generating p…

Cited by 5SourcePDFScholar
2024

Towards Detailed Text-to-Motion Synthesis via Basic-to-Advanced Hierarchical Diffusion Model

AAAI 2024technical

Text-guided motion synthesis aims to generate 3D human motion that not only precisely reflects the textual description but reveals the motion details as much as possible. Pioneering methods explore the diffusion model for text-to-motion synthesis and obtain significant superiority. However, these me…

Cited by 15SourcePDFScholar
2024

Trimodal Navigable Region Segmentation Model: Grounding Navigation Instructions in Urban Areas

RA-L 2024

In this study, we develop a model that enables mobilities to have more friendly interactions with users. Specifically, we focus on the referring navigable regions task in which a model grounds navigable regions of the road using the mobility's camera image and natural language navigation instruction

Cited by 6SourceScholar
2023

ASM: Adaptive Skinning Model for High-Quality 3D Face Modeling

ICCV 2023poster

The research fields of parametric face model and 3D face reconstruction have been extensively studied. However, a critical question remains unanswered: how to tailor the face model for specific reconstruction settings. We argue that reconstruction with multi-view uncalibrated images demands a new mo…

Cited by 6PDFScholar
2023

Adaptive Patch Deformation for Textureless-Resilient Multi-View Stereo

CVPR 2023poster

In recent years, deep learning-based approaches have shown great strength in multi-view stereo because of their outstanding ability to extract robust visual features. However, most learning-based methods need to build the cost volume and increase the receptive field enormously to get a satisfactory…

2023

AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System

RSS 2023poster

Vision-based teleoperation offers the possibility to endow robots with human-level intelligence to physically interact with the environment, while only requiring low-cost camera sensors. However, current vision-based teleoperation systems are designed and engineered towards a particular robot model…

Cited by 114SourcePDFScholar
2023

C2F2NeUS: Cascade Cost Frustum Fusion for High Fidelity and Generalizable Neural Surface Reconstruction

ICCV 2023poster

There is an emerging effort to combine the two popular 3D frameworks using Multi-View Stereo (MVS) and Neural Implicit Surfaces (NIS) with a specific focus on the few-shot / sparse view setting. In this paper, we introduce a novel integration scheme that combines the multi-view stereo with neural si…

Cited by 26PDFScholar
2023

Compact Transformer Tracker with Correlative Masked Modeling

AAAI 2023technical

Transformer framework has been showing superior performances in visual object tracking for its great strength in information aggregation across the template and search image with the well-known attention mechanism. Most recent advances focus on exploring attention mechanism variants for better infor…

2023

Dual Memory Units with Uncertainty Regulation for Weakly Supervised Video Anomaly Detection

AAAI 2023technical

Learning discriminative features for effectively separating abnormal events from normality is crucial for weakly supervised video anomaly detection (WS-VAD) tasks. Existing approaches, both video and segment level label oriented, mainly focus on extracting representations for anomaly data while negl…

2023

Dynamic Transformers Provide a False Sense of Efficiency

ACL 2023long

Despite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference. Multi-exit is a mainstream approach to address this issue by making a trade-off between efficiency and accuracy, where the saving of computation comes…

2023

FGNet: Towards Filling the Intra-class and Inter-class Gaps for Few-shot Segmentation

IJCAI 2023poster

Current few-shot segmentation (FSS) approaches have made tremendous achievements based on prototypical learning techniques. However, due to the scarcity of the support data provided, FSS methods still suffer from the intra-class and inter-class gaps. In this paper, we propose a uniform network to fi…

2023

Fine-Grained Private Knowledge Distillation

ICASSP 2023accepted

Knowledge distillation has emerged as a scalable and effective way for privacy-preserving machine learning. One remaining drawback is that it consumes privacy in a client-level manner. In order to attain fine-grained privacy accountant and improve utility, this work proposes a model-free reverse k-N…

Cited by 0SourceScholar
2023

Learning Human-to-Robot Handovers From Point Clouds

CVPR 2023highlight

We propose the first framework to learn control policies for vision-based human-to-robot handovers, a critical task for human-robot interaction. While research in Embodied AI has made significant progress in training robot agents in simulated environments, interacting with humans remains challenging…

Cited by 50SourcePDFScholar
2023

Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination

AAAI 2023technical

We study the problem of training a Reinforcement Learning (RL) agent that is collaborative with humans without using human data. Although such agents can be obtained through self-play training, they can suffer significantly from the distributional shift when paired with unencountered partners, such…

2023

NeMF: Inverse Volume Rendering with Neural Microflake Field

ICCV 2023poster

Recovering the physical attributes of an object's appearance from its images captured under an unknown illumination is challenging yet essential for photo-realistic rendering.Recent approaches adopt the emerging implicit scene representations and have shown impressive results.However, they unanimous…

Cited by 26PDFcodeScholar
2023

RLogist: Fast Observation Strategy on Whole-Slide Images with Deep Reinforcement Learning

AAAI 2023technical

Whole-slide images (WSI) in computational pathology have high resolution with gigapixel size, but are generally with sparse regions of interest, which leads to weak diagnostic relevance and data inefficiency for each area in the slide. Most of the existing methods rely on a multiple instance learnin…

2023

SYNC: SAFETY-AWARE NEURAL CONTROL FOR STABILIZING STOCHASTIC DELAY-DIFFERENTIAL EQUATIONS

ICLR 2023poster

Stabilization of the systems described by \textit{stochastic delay}-differential equations (SDDEs) under preset conditions is a challenging task in the control community. Here, to achieve this task, we leverage neural networks to learn control policies using the information of the controlled systems…

Cited by 10SourcePDFScholar
2023

The Dark Side of Dynamic Routing Neural Networks: Towards Efficiency Backdoor Injection

CVPR 2023poster

Recent advancements in deploying deep neural networks (DNNs) on resource-constrained devices have generated interest in input-adaptive dynamic neural networks (DyNNs). DyNNs offer more efficient inferences and enable the deployment of DNNs on devices with limited resources, such as mobile devices. H…

2022

Against Backdoor Attacks In Federated Learning With Differential Privacy

ICASSP 2022accepted

The training process of federated learning is known to be vulnerable to adversarial attacks (e.g., backdoor attack). Previous works showed that differential privacy (DP) can be used to defend against backdoor attacks, yet at the cost of vastly losing model utility. To address this issue, we in this…

Cited by 0SourceScholar
2022

Detaching and Boosting: Dual Engine for Scale-Invariant Self-Supervised Monocular Depth Estimation

RA-L 2022

Monocular depth estimation (MDE) in the self-supervised scenario has emerged as a promising method as it refrains from the requirement of ground truth depth. Despite continuous efforts, MDE is still sensitive to scale changes especially when all the training samples are from one single camera. Meanw

Cited by 1SourcecodeScholar
2022

Greedy when Sure and Conservative when Uncertain about the Opponents

ICML 2022spotlight

We develop a new approach, named Greedy when Sure and Conservative when Uncertain (GSCU), to competing online against unknown and nonstationary opponents. GSCU improves in four aspects: 1) introduces a novel way of learning opponent policy embeddings offline; 2) trains offline a single best response…

2022

HandoverSim: A Simulation Framework and Benchmark for Human-to-Robot Object Handovers

ICRA 2022poster

We introduce a new simulation benchmark “Han-doverSim” for human-to-robot object handovers. To simulate the giver's motion, we leverage a recent motion capture dataset of hand grasping of objects. We create training and evaluation environments for the receiver with standardized protocols and metrics…

Cited by 29SourcecodeScholar
2022

HumanNeRF: Efficiently Generated Human Radiance Field From Sparse Inputs

CVPR 2022poster

Recent neural human representations can produce high-quality multi-view rendering but require using dense multi-view inputs and costly training. They are hence largely limited to static models as training each frame is infeasible. We present HumanNeRF - a neural representation with efficient general…

Cited by 227PDFScholar
2022

JueWu-MC: Playing Minecraft with Sample-efficient Hierarchical Reinforcement Learning

IJCAI 2022poster

Learning rational behaviors in open-world games like Minecraft remains to be challenging for Reinforcement Learning (RL) research due to the compound challenge of partial observability, high-dimensional visual perception and delayed reward. To address this, we propose JueWu-MC, a sample-efficient hi…

Cited by 44SourcePDFScholar
2022

Learning Free Gait Transition for Quadruped Robots Via Phase-Guided Controller

RA-L 2022

Gaits and transitions are key components in legged locomotion. For legged robots, describing and reproducing gaits as well as transitions remain longstanding challenges. Reinforcement learning has become a powerful tool to formulate controllers for legged robots. Learning multiple gaits and transiti

Cited by 86SourcecodeScholar
2022

Learning Perceptual Concepts by Bootstrapping From Human Queries

RA-L 2022

When robots operate in human environments, it's critical that humans can quickly teach them new <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">concepts:</i> object-centric properties of the environment that they care about (e.g., objects <italic xml

Cited by 17SourceScholar
2022

Learning Robust Real-World Dexterous Grasping Policies via Implicit Shape Augmentation

CoRL 2022poster

Dexterous robotic hands have the capability to interact with a wide variety of household objects. However, learning robust real world grasping policies for arbitrary objects has proven challenging due to the difficulty of generating high quality training data. In this work, we propose a learning sys…

Cited by 32SourceScholar
2022

Model Predictive Control for Fluid Human-to-Robot Handovers

ICRA 2022poster

Human-robot handover is a fundamental yet challenging task in human-robot interaction and collaboration. Recently, remarkable progressions have been made in human-to-robot handovers of unknown objects by using learning-based grasp generators. However, how to responsively generate smooth motions to t…

Cited by 31SourceScholar
2022

NICGSlowDown: Evaluating the Efficiency Robustness of Neural Image Caption Generation Models

CVPR 2022poster

Neural image caption generation (NICG) models have received massive attention from the research community due to their excellent performance in visual understanding. Existing work focuses on improving NICG model accuracy while efficiency is less explored. However, many real-world applications requir…

Cited by 44PDFcodeScholar
2022

Node-Aligned Graph Convolutional Network for Whole-Slide Image Representation and Classification

CVPR 2022oral

The large-scale whole-slide images (WSIs) facilitate the learning-based computational pathology methods. However, the gigapixel size of WSIs makes it hard to train a conventional model directly. Current approaches typically adopt multiple-instance learning (MIL) to tackle this problem. Among them, M…

Cited by 71PDFcodeScholar
2022

S2worm: A Fast-Moving Untethered Insect-Scale Robot With 2-DoF Transmission Mechanism

RA-L 2022

Designing terrestrial insect-scale robot with high maneuverability and autonomy is becoming an essential challenge in the field of robotics research. Previous work has indicated that compact transmission and integrated control devices can improve the application potential. In this work, an untethere

Cited by 20SourceScholar
2022

Shape Prior Guided Attack: Sparser Perturbations on 3D Point Clouds

AAAI 2022technical

Deep neural networks are extremely vulnerable to malicious input data. As 3D data is increasingly used in vision tasks such as robots, autonomous driving and drones, the internal robustness of the classification models for 3D point cloud has received widespread attention. In this paper, we propose a…

Cited by 22SourcePDFScholar
2022

TestAug: A Framework for Augmenting Capability-based NLP Tests

COLING 2022main

The recently proposed capability-based NLP testing allows model developers to test the functional capabilities of NLP models, revealing functional failures for models with good held-out evaluation scores. However, existing work on capability-based testing requires the developer to compose each indiv…

2021

Adversarial Attacks on Object Detectors with Limited Perturbations

ICASSP 2021accepted

Deep convolutional neural networks are widely witnessed vulnerable to adversarial attacks. Recently, great progress has been achieved in attacking object detectors. However, current attacks neglect the practical utility and rely on global perturbations on the target image with a large number of patc…

Cited by 0SourceScholar
2021

Boosting Offline Reinforcement Learning with Residual Generative Modeling

IJCAI 2021poster

Offline reinforcement learning (RL) tries to learn the near-optimal policy with recorded offline experience without online exploration.Current offline RL research includes: 1) generative modeling, i.e., approximating a policy using fixed data; and 2) learning the state-action value function. While m…

Cited by 15SourcePDFScholar
2021

Continuous Copy-Paste for One-Stage Multi-Object Tracking and Segmentation

ICCV 2021poster

Current one-step multi-object tracking and segmentation (MOTS) methods lag behind recent two-step methods. By separating the instance segmentation stage from the tracking stage, two-step methods can exploit non-video datasets as extra data for training instance segmentation. Moreover, instances belo…

Cited by 29PDFcodeScholar
2021

DexYCB: A Benchmark for Capturing Hand Grasping of Objects

CVPR 2021poster

We introduce DexYCB, a new dataset for capturing hand grasping of objects. We first compare DexYCB with a related one through cross-dataset evaluation. We then present a thorough benchmark of state-of-the-art approaches on three relevant tasks: 2D object and keypoint detection, 6D object pose estima…

Cited by 314PDFcodeScholar
2021

Environment-Independent Wi-Fi Human Activity Recognition with Adversarial Network

ICASSP 2021accepted

Human activity recognition is an essential part of human-computer interaction systems. Environment-robust Wi-Fi-based systems for this task is still a challenging problem, due to the fact that most existing systems may drop in performance when the environment is changed. To address this issue, we in…

Cited by 0SourceScholar
2021

Goal-Auxiliary Actor-Critic for 6D Robotic Grasping with Point Clouds

CoRL 2021poster

6D robotic grasping beyond top-down bin-picking scenarios is a challenging task. Previous solutions based on 6D grasp synthesis with robot motion planning usually operate in an open-loop setting, which are sensitive to grasp synthesis errors. In this work, we propose a new method for learning closed…

Cited by 55SourcecodeScholar
2021

Hiding Numerical Vectors in Local Private and Shuffled Messages

IJCAI 2021poster

Numerical vector aggregation has numerous applications in privacy-sensitive scenarios, such as distributed gradient estimation in federated learning, and statistical analysis on key-value data. Within the framework of local differential privacy, this work gives tight minimax error bounds of O(d s/(n…

Cited by 9SourcePDFScholar
2021

MDANet: Multi-Modal Deep Aggregation Network for Depth Completion

ICRA 2021poster

Depth completion aims to recover the dense depth map from sparse depth data and RGB image respectively. However, due to the huge difference between the multi-modal signal input, vanilla convolutional neural network and simple fusion strategy cannot extract features from sparse data and aggregate mul…

Cited by 17SourcecodeScholar
2021

MapGo: Model-Assisted Policy Optimization for Goal-Oriented Tasks

IJCAI 2021poster

In Goal-oriented Reinforcement learning, relabeling the raw goals in past experience to provide agents with hindsight ability is a major solution to the reward sparsity problem. In this paper, to enhance the diversity of relabeled goals, we develop FGI (Foresight Goal Inference), a new relabeling st…

2021

Mask4D: 4D Convolution Network for Light Field Occlusion Removal

ICASSP 2021accepted

Current light field (LF) occlusion removal approaches usually select only a part of sub-aperture images (SAIs) or simply stack all SAIs to reconstruct the center view, which destroys the spatial layout of SAIs. In this paper, we present a simple yet effective LF occlusion removal method name Mask4D,…

Cited by 0SourceScholar
2021

Optimizing Deeper Transformers on Small Datasets

ACL 2021long

It is a common belief that training deep transformers from scratch requires large datasets. Consequently, for small datasets, people usually use shallow and simple additional layers on top of pre-trained models during fine-tuning. This work shows that this does not always need to be the case: with p…

2021

Pointer Networks for Arbitrary-Shaped Text Spotting

ICASSP 2021accepted

Current text spotting methods perform text detection and text recognition separately. However, in complex scenes where bounding boxes of texts with various shapes are often overlapped, text detection becomes error-prone. By contrast, character detection is more non-ambiguous and easier to learn. In…

Cited by 0SourceScholar
2021

RaP-Net: A Region-wise and Point-wise Weighting Network to Extract Robust Features for Indoor Localization

IROS 2021poster

Feature extraction plays an important role in visual localization. Unreliable features on dynamic objects or repetitive regions will interfere with feature matching and challenge indoor localization greatly. To address the problem, we propose a novel network, RaP-Net, to simultaneously predict regio…

Cited by 7SourcecodeScholar
2021

Reactive Human-to-Robot Handovers of Arbitrary Objects

ICRA 2021poster

Human-robot object handovers have been an actively studied area of robotics over the past decade; however, very few techniques and systems have addressed the challenge of handing over diverse objects with arbitrary appearance, size, shape, and deformability. In this paper, we present a vision-based…

Cited by 94SourceScholar
2021

Revealing the Reciprocal Relations Between Self-Supervised Stereo and Monocular Depth Estimation

ICCV 2021poster

Current self-supervised depth estimation algorithms mainly focus on either stereo or monocular only, neglecting the reciprocal relations between them. In this paper, we propose a simple yet effective framework to improve both stereo and monocular depth estimation by leveraging the underlying complem…

Cited by 34PDFScholar
2021

VK-Net: Category-Level Point Cloud Registration with Unsupervised Rotation Invariant Keypoints

ICASSP 2021accepted

In this paper, we propose VK-Net, a neural network that learns to discover a set of category-specific keypoints from a single point cloud in an unsupervised manner. VK-Net is able to generate semantically consistent and rotation invariant keypoints across objects of the same category and different v…

Cited by 0SourceScholar
2020

Are We Ready for Service Robots? The OpenLORIS-Scene Datasets for Lifelong SLAM

ICRA 2020poster

Service robots should be able to operate autonomously in dynamic and daily changing environments over an extended period of time. While Simultaneous Localization And Mapping (SLAM) is one of the most fundamental problems for robotic autonomy, most existing SLAM works are evaluated with data sequence…

Cited by 174SourcecodeScholar
2020

BioARS: Designing Adaptive and Reconfigurable Bionic Assembly Robotic System with Inchworm Modules

IROS 2020poster

Designing a swarm of robots to address different tasks and adapt to various environments through self-assembly is one of the most challenging topics in the field of robotics research. Here, we present an assembly robotic system with inchworm robots as modules. The system is called BioARS (Bionic Ass…

Cited by 7SourceScholar
2020

Collaborative Interaction Models for Optimized Human-Robot Teamwork

IROS 2020poster

Effective human-robot collaboration requires informed anticipation. The robot must anticipate the human’s actions, but also react quickly and intuitively when its predictions are wrong. The robot must plan its actions to account for the human’s own plan, with the knowledge that the human’s behavior…

Cited by 22SourceScholar
2020

DXSLAM: A Robust and Efficient Visual SLAM System with Deep Features

IROS 2020poster

A robust and efficient Simultaneous Localization and Mapping (SLAM) system is essential for robot autonomy. For visual SLAM algorithms, though the theoretical framework has been well established for most aspects, feature extraction and association is still empirically designed in most cases, and can…

Cited by 155SourcecodeScholar
2020

DexPilot: Vision-Based Teleoperation of Dexterous Robotic Hand-Arm System

ICRA 2020

Teleoperation offers the possibility of imparting robotic systems with sophisticated reasoning skills, intuition, and creativity to perform tasks. However, teleoperation solutions for high degree-of-actuation (DoA), multi-fingered robots are generally cost-prohibitive, while low-cost offerings usual

Cited by 279SourceScholar
2020

Integrating Discrete and Neural Features Via Mixed-Feature Trans-Dimensional Random Field Language Models

ICASSP 2020accepted

There has been a long recognition that discrete features (n-gram features) and neural network based features have complementary strengths for language models (LMs). Improved performance can be obtained by model interpolation, which is, however, a sub-optimal two-step integration of discrete and neur…

Cited by 0SourceScholar
2020

Private FL-GAN: Differential Privacy Synthetic Data Generation Based on Federated Learning

ICASSP 2020accepted

Generative Adversarial Network (GAN) has already made a big splash in the field of generating realistic "fake" data. However, when data is distributed and data-holders are reluctant to share data for privacy reasons, GAN’s training is difficult. To address this issue, we propose private FL-GAN, a di…

Cited by 0SourceScholar
2020

Segment as Points for Efficient Online Multi-Object Tracking and Segmentation

ECCV 2020poster

Current multi-object tracking and segmentation (MOTS) methods follow the tracking-by-detection paradigm and adopt convolutions for feature extraction. However, as affected by the inherent receptive field, convolution based feature extraction inevitably mixes up the foreground features and the backgr…

2020

Towards Playing Full MOBA Games with Deep Reinforcement Learning

NeurIPS 2020poster

MOBA games, e.g., Honor of Kings, League of Legends, and Dota 2, pose grand challenges to AI systems such as multi-agent, enormous state-action space, complex action control, etc. Developing AI for playing MOBA games has raised much attention accordingly. However, existing work falls short in handli…

2019

Visual Semantic Navigation using Scene Priors

ICLR 2019poster

How do humans navigate to target objects in novel scenes? Do we use the semantic/functional priors we have built over years to efficiently search and navigate? For example, to search for mugs, we search cabinets near the coffee machine and for fruits we try the fridge. In this work, we focus on inco…

Cited by 391SourcePDFScholar
2018

3D Human Pose Estimation in the Wild by Adversarial Learning

CVPR 2018poster

Recently, remarkable advances have been achieved in 3D human pose estimation from monocular images because of the powerful Deep Convolutional Neural Networks (DCNNs). Despite their success on large-scale datasets collected in the constrained lab environment, it is difficult to obtain the 3D pose ann…

Cited by 493SourcePDFScholar
2018

Towards End-to-End License Plate Detection and Recognition: A Large Dataset and Baseline

ECCV 2018poster

Most current license plate (LP) detection and recognition approaches are evaluated on a small and usually unrepresentative dataset since there are no publicly available large diverse datasets. In this paper, we introduce CCPD, a large and comprehensive LP dataset. All images are taken manually by wo…

2017

Identity-Aware Textual-Visual Matching With Latent Co-Attention

ICCV 2017poster

Textual-visual matching aims at measuring similarities between sentence descriptions and images. Most existing methods tackle this problem without effectively utilizing identity-level annotations. In this paper, we propose an identity-aware two-stage framework for the textual-visual matching problem…

Cited by 315PDFScholar
2017

Learning Feature Pyramids for Human Pose Estimation

ICCV 2017poster

Articulated human pose estimation is a fundamental yet challenging task in computer vision. The difficulty is particularly pronounced in scale variations of human body parts when camera view changes or severe foreshortening happens. Although pyramid methods are widely used to handle scale changes at…

Cited by 645PDFcodeScholar
2017

Multi-Context Attention for Human Pose Estimation

CVPR 2017poster

In this paper, we propose to incorporate convolutional neural networks with a multi-context attention mechanism into an end-to-end framework for human pose estimation. We adopt stacked hourglass networks to generate attention maps from features at multiple resolutions with various semantics. The Con…

Cited by 909PDFScholar
2017

Training data reduction in deep neural networks with partial mutual information based feature selection and correlation matching based active learning

ICASSP 2017accepted

In this paper, we develop a novel scheme to reduce the amount of training data required for training deep neural networks (DNNs). We first apply a partial mutual information (PMI) technique to seek for the optimal DNN feature set. Then we use a correlation matching based active learning (CMAL) techn…

Cited by 0SourceScholar
2016

End-To-End Learning of Deformable Mixture of Parts and Deep Convolutional Neural Networks for Human Pose Estimation

CVPR 2016oral

Recently, Deep Convolutional Neural Networks (DCNNs) have been applied to the task of human pose estimation, and have shown its potential of learning better feature representations and capturing contextual relationships. However, it is difficult to incorporate domain prior knowledge such as geometri…

Cited by 346PDFScholar
2015

Ambient Occlusion via Compressive Visibility Estimation

CVPR 2015poster

There has been emerging interest on recovering traditionally challenging intrinsic scene properties. In this paper, we present a novel computational imaging solution for recovering the ambient occlusion (AO) map of an object. AO measures how much light from all different directions can reach a surfa…

Cited by 10SourcePDFScholar