← Search

Shuang Li

124 accepted papers

2026

Beyond Accuracy: Latent Perturbations for Cognitive-Aware Diagnosis

ICML 2026poster

Diagnosing rare diseases remains a persistent challenge, often hindered by *cognitive anchoring*: once clinicians settle on a common diagnosis, they often discount alternative explanations, including rare conditions. To address this, we propose a human-centered counterfactual reasoning framework usi…

Cited by 0SourceScholar
2026

Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generation

CVPR 2026

Subject-Driven Text-to-Image (T2I) Generation aims to preserve a subject's identity while editing its context based on a text prompt. A core challenge in this task is the "similarity-controllability paradox", where enhancing textual control often degrades the subject's fidelity, and vice-versa. We a

Cited by 0SourceScholar
2026

Dynamic-Static Collaboration for Unsupervised Domain Adaptive Video-Based Visible-Infrared Person Re-Identification

AAAI 2026technical

Video-based visible-infrared person re-identification (VVI-ReID) aims to match pedestrian sequences across modalities for all-day surveillance. While supervised methods have shown progress, their dependence on large-scale cross-modal annotations limits scalability. We investigate the task of unsuper

Cited by 0SourcePDFScholar
2026

Geometry-aware 4D Video Generation for Robot Manipulation

ICLR 2026poster

Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes, generating videos that are both temporally coherent and geo…

Cited by 0SourcecodeScholar
2026

Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification

AAAI 2026technical

Person re-identification (ReID) aims to retrieve target pedestrian images given either visual queries (image-to-image, I2I) or textual descriptions (text-to-image, T2I). Although both tasks share a common retrieval objective, they pose distinct challenges: I2I emphasizes discriminative identity lea

Cited by 0SourcePDFScholar
2026

Hyperbolic Hierarchical Alignment for Video-Based Visible-Infrared Person Re-Identification

ICML 2026poster

Video-based visible-infrared person re-identification (VVI-ReID) aims to learn robust video-level representations under modality discrepancy. However, existing methods typically rely on Euclidean geometry, which is suboptimal for modeling the complex temporal dynamics within visible and infrared tra…

Cited by 0SourceScholar
2026

Inferring the Invisible: Neuro-Symbolic Rule Discovery for Missing Value Imputation

ICLR 2026poster

One of the central challenges in artificial intelligence is reasoning under partial observability, where key values are missing but essential for understanding and modeling the system. This paper presents a neuro-symbolic framework for latent rule discovery and missing value imputation. In contrast…

Cited by 0SourceScholar
2026

Learning Adaptive Distribution Alignment with Neural Characteristic Function for Graph Domain Adaptation

ICLR 2026poster

Graph Domain Adaptation (GDA) transfers knowledge from labeled source graphs to unlabeled target graphs but is challenged by complex, multi-faceted distributional shifts. Existing methods attempt to reduce distributional shifts by aligning manually selected graph elements (e.g., node attributes or s…

Cited by 0SourcecodeScholar
2026

Learning Structure-Semantic Evolution Trajectories for Graph Domain Adaptation

ICLR 2026poster

Graph Domain Adaptation (GDA) aims to bridge distribution shifts between domains by transferring knowledge from well-labeled source graphs to given unlabeled target graphs. One promising recent approach addresses graph transfer by discretizing the adaptation process, typically through the construct…

Cited by 0SourceScholar
2026

PortraitSR: Artist-Inspired Prior Learning for Progressive Face Super-Resolution

AAAI 2026technical

Face super-resolution (FSR) aims to reconstruct high-resolution (HR) face images from low-resolution (LR) inputs. While recent methods have advanced this task through architectural innovations and generative modeling, but they often leads to semantically inconsistent structures and unrealistic textu

Cited by 0SourcePDFScholar
2026

SPSC: Sparse and Scalable Multi-Modal 3D Occupancy Prediction for Autonomous Driving

AAAI 2026technical

3D semantic occupancy prediction offers a nuanced representation of the surrounding environment, which is crucial for ensuring the safety of autonomous driving. However, fine-grained scene representations inevitably result in cubic growth in data scale, which imposes substantial demands on model arc

Cited by 0SourcePDFScholar
2026

SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation

ICML 2026poster

Generative Recommendation (GR) has emerged as a promising paradigm by formulating item recommendation as a sequence-to-sequence generation task over item identifiers. Recent studies have incorporated multimodal signals to provide richer token-level evidence for generation. However, existing approach…

Cited by 0SourceScholar
2026

Thermal Diffusion Matters: Infrared Spatial-Temporal Video Super-Resolution through Heat Conduction Priors

CVPR 2026

Infrared video acquisition inherently suffers from low spatial resolution and limited frame rates due to the physical constraints of thermal imaging sensors. These limitations make infrared video enhancement uniquely challenging, as it requires restoring spatial details and temporal continuity from

Cited by 0SourcecodeScholar
2026

Thermal-Physics Guided Infrared Image Super-Resolution with Dynamic High-Frequency Amplification

AAAI 2026technical

The practical deployment of infrared imaging is hindered by its inherent output of low-resolution (LR) images. While the super-resolution (SR) technique is a promising remedy, we discover two major challenges concerning infrared image SR: preserving accurate thermal distributions, which are fundamen

Cited by 0SourcePDFScholar
2026

VirtualEnv: A Platform for Embodied AI Research

AAAI 2026technical

As large language models (LLMs) continue to improve in reasoning and decision-making, there is a growing need for realistic and interactive environments where their abilities can be rigorously evaluated. We present VirtualEnv, a next-generation simulation platform built on Unreal Engine 5 that enabl

Cited by 0SourcePDFScholar
2025

Convergence of Mean-Field Langevin Stochastic Descent-Ascent for Distributional Minimax Optimization

ICML 2025spotlight

We study convergence properties of the discrete-time Mean-Field Langevin Stochastic Descent-Ascent (MFL-SDA) algorithm for solving distributional minimax optimization. These problems arise in various applications, such as zero-sum games, generative adversarial networks and distributionally robust le…

Cited by 0SourcePDFScholar
2025

Evolving Minds: Logic-Informed Inference from Temporal Action Patterns

ICML 2025poster

Understanding human mental states—such as intentions and desires—is crucial for natural AI-human collaboration. However, this is challenging because human actions occur irregularly over time, and the underlying mental states that drive these actions are unobserved. To tackle this, we propose a novel…

Cited by 0SourcePDFScholar
2025

From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection

ICML 2025poster

Pretrained vision-language models (VLMs), e.g., CLIP, demonstrate impressive zero-shot capabilities on downstream tasks. Prior research highlights the crucial role of visual augmentation techniques, like random cropping, in alignment with fine-grained class descriptions generated by large language m…

2025

Learning Time-Aware Causal Representation for Model Generalization in Evolving Domains

ICML 2025poster

Endowing deep models with the ability to generalize in dynamic scenarios is of vital significance for real-world deployment, given the continuous and complex changes in data distribution. Recently, evolving domain generalization (EDG) has emerged to address distribution shifts over time, aiming to c…

Cited by 0SourcePDFScholar
2025

Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains

ICLR 2025poster

Large language models (LLMs) have achieved remarkable performance in recent years but are fundamentally limited by the underlying training data. To improve models beyond the training data, recent works have explored how LLMs can be used to generate synthetic data for autonomous self-improvement. How…

Cited by 12SourcePDFScholar
2024

An Unforgeable Publicly Verifiable Watermark for Large Language Models

ICLR 2024poster

Recently, text watermarking algorithms for large language models (LLMs) have been proposed to mitigate the potential harms of text generated by LLMs, including fake news and copyright issues. However, current watermark detection algorithms require the secret key used in the watermark generation proc…

2024

Beyond Entities: A Large-Scale Multi-Modal Knowledge Graph with Triplet Fact Grounding

AAAI 2024technical

Much effort has been devoted to building multi-modal knowledge graphs by visualizing entities on images, but ignoring the multi-modal information of the relation between entities. Hence, in this paper, we aim to construct a new large-scale multi-modal knowledge graph with triplet facts grounded on i…

2024

Enhancing Cross-Modal Fine-Tuning with Gradually Intermediate Modality Generation

ICML 2024poster

Large-scale pretrained models have proven immensely valuable in handling data-intensive modalities like text and image. However, fine-tuning these models for certain specialized modalities, such as protein sequence and cosmic ray, poses challenges due to the significant modality discrepancy and scar…

Cited by 4SourcePDFScholar
2024

Enhancing Human-AI Collaboration Through Logic-Guided Reasoning

ICLR 2024poster

We present a systematic framework designed to enhance human-robot perception and collaboration through the integration of logical rules and Theory of Mind (ToM). Logical rules provide interpretable predictions and generalize well across diverse tasks, making them valuable for learning and decision-m…

Cited by 5SourcePDFScholar
2024

Enhancing Reinforcement Learning with Label-Sensitive Reward for Natural Language Understanding

ACL 2024long

Recent strides in large language models (LLMs) have yielded remarkable performance, leveraging reinforcement learning from human feedback (RLHF) to significantly enhance generation and alignment capabilities. However, RLHF encounters numerous challenges, including the objective mismatch issue, leadi…

2024

Exploring Structured Semantic Priors Underlying Diffusion Score for Test-time Adaptation

NeurIPS 2024poster

Capitalizing on the complementary advantages of generative and discriminative models has always been a compelling vision in machine learning, backed by a growing body of research. This work discloses the hidden semantic structure within score-based generative models, unveiling their potential as eff…

2024

Improving Factuality and Reasoning in Language Models through Multiagent Debate

ICML 2024poster

Large language models (LLMs) have demonstrated remarkable capabilities in language generation, understanding, and few-shot learning in recent years. An extensive body of work has explored how their performance may be further improved through the tools of prompting, ranging from verification, self-co…

2024

Latent Logic Tree Extraction for Event Sequence Explanation from LLMs

ICML 2024poster

Modern high-stakes systems, such as healthcare or robotics, often generate vast streaming event sequences. Our goal is to design an efficient, plug-and-play tool to elicit logic tree-based explanations from Large Language Models (LLMs) to provide customized insights into each observed event sequence…

Cited by 5SourcePDFScholar
2024

Learning Modality Knowledge Alignment for Cross-Modality Transfer

ICML 2024poster

Cross-modality transfer aims to leverage large pretrained models to complete tasks that may not belong to the modality of pretraining data. Existing works achieve certain success in extending classical finetuning to cross-modal scenarios, yet we still lack understanding about the influence of modali…

Cited by 3SourcePDFScholar
2024

MGRL: Mutual-Guidance Representation Learning for Text-to-Image Person Retrieval

ICASSP 2024accepted

Text-to-image person retrieval aims to recognize target pedestrians based on specified text. Existing methods mainly obtain image and text features separately through distinct feature extractors, subsequently embedding them into a unified feature space and calculating their similarity. Despite great…

Cited by 0SourceScholar
2024

On the Robustness of Document-Level Relation Extraction Models to Entity Name Variations

ACL 2024findings

Driven by the demand for cross-sentence and large-scale relation extraction, document-level relation extraction (DocRE) has attracted increasing research interest. Despite the continuous improvement in performance, we find that existing DocRE models which initially perform well may make more mistake…

2024

SEGMENT+: Long Text Processing with Short-Context Language Models

EMNLP 2024main

There is a growing interest in expanding the input capacity of language models (LMs) across various domains. However, simply increasing the context window does not guarantee robust performance across diverse long-input processing tasks, such as understanding extensive documents and extracting detail…

Cited by 1SourcePDFScholar
2024

Strengthened Symbol Binding Makes Large Language Models Reliable Multiple-Choice Selectors

ACL 2024long

Multiple-Choice Questions (MCQs) constitute a critical area of research in the study of Large Language Models (LLMs). Previous works have investigated the selection bias problem in MCQs within few-shot scenarios, in which the LLM’s performance may be influenced by the presentation of answer choices,…

2024

Translate Meanings, Not Just Words: IdiomKB’s Role in Optimizing Idiomatic Translation with Language Models

AAAI 2024technical

To translate well, machine translation (MT) systems and general-purposed language models (LMs) need a deep understanding of both source and target languages and cultures. Therefore, idioms, with their non-compositional nature, pose particular challenges for Transformer-based systems, as literal tran…

2024

Unveiling Latent Causal Rules: A Temporal Point Process Approach for Abnormal Event Explanation

AISTATS 2024poster

In high-stakes systems such as healthcare, it is critical to understand the causal reasons behind unusual events, such as sudden changes in patient’s health. Unveiling the causal reasons helps with quick diagnoses and precise treatment planning. In this paper, we propose an automated method for unco…

2024

Weight Diffusion for Future: Learn to Generalize in Non-Stationary Environments

NeurIPS 2024poster

Enabling deep models to generalize in non-stationary environments is vital for real-world machine learning, as data distributions are often found to continually change. Recently, evolving domain generalization (EDG) has emerged to tackle the domain generalization in a time-varying system, where the…

Cited by 0SourcePDFScholar
2023

AMR-based Network for Aspect-based Sentiment Analysis

ACL 2023long

Aspect-based sentiment analysis (ABSA) is a fine-grained sentiment classification task. Many recent works have used dependency trees to extract the relation between aspects and contexts and have achieved significant improvements. However, further improvement is limited due to the potential mismatch…

Cited by 0SourcePDFScholar
2023

Annotator: A Generic Active Learning Baseline for LiDAR Semantic Segmentation

NeurIPS 2023poster

Active learning, a label-efficient paradigm, empowers models to interactively query an oracle for labeling new data. In the realm of LiDAR semantic segmentation, the challenges stem from the sheer volume of point clouds, rendering annotation labor-intensive and cost-prohibitive. This paper presents…

Cited by 12SourcePDFScholar
2023

Borrowing Knowledge From Pre-trained Language Model: A New Data-efficient Visual Learning Paradigm

ICCV 2023poster

The development of vision models for real-world applications is hindered by the challenge of annotated data scarcity, which has necessitated the adoption of data-efficient visual learning techniques such as semi-supervised learning. Unfortunately, the prevalent cross-entropy supervision is limited b…

Cited by 8PDFcodeScholar
2023

Composing Ensembles of Pre-trained Models via Iterative Consensus

ICLR 2023poster

Large pre-trained models exhibit distinct and complementary capabilities dependent on the data they are trained on. Language models such as GPT-3 are capable of textual reasoning but cannot understand visual information, while vision models such as DALL-E can generate photorealistic photos but fail…

Cited by 33SourcePDFScholar
2023

Compositional Foundation Models for Hierarchical Planning

NeurIPS 2023poster

To make effective decisions in novel environments with long-horizon goals, it is crucial to engage in hierarchical reasoning across spatial and temporal scales. This entails planning abstract subgoal sequences, visually reasoning about the underlying plans, and executing actions in accordance with t…

Cited by 45SourcePDFScholar
2023

ConceptFusion: Open-set multimodal 3D mapping

RSS 2023poster

Building 3D maps of the environment is central to robot navigation, planning, and interaction with objects in a scene. Most existing approaches that integrate semantic concepts with 3D maps largely remain confined to the closed-set setting: they can only reason about a finite set of concepts, pre-de…

2023

Dirichlet-based Uncertainty Calibration for Active Domain Adaptation

ICLR 2023top-25%

Active domain adaptation (DA) aims to maximally boost the model adaptation on a new target domain by actively selecting limited target data to annotate, whereas traditional active learning methods may be less effective since they do not consider the domain shift issue. Despite active DA methods addr…

2023

Discovering Intrinsic Spatial-Temporal Logic Rules to Explain Human Actions

NeurIPS 2023poster

We propose an interpretable model to uncover the behavioral patterns of human movements by analyzing their trajectories. Our approach is based on the belief that human actions are driven by intentions and are influenced by environmental factors such as spatial relationships with surrounding objects.…

Cited by 7SourcePDFScholar
2023

Enhancing Cross-lingual Natural Language Inference by Soft Prompting with Multilingual Verbalizer

ACL 2023findings

Cross-lingual natural language inference is a fundamental problem in cross-lingual language understanding. Many recent works have used prompt learning to address the lack of annotated parallel corpora in XNLI.However, these methods adopt discrete prompting by simply translating the templates to the…

2023

Evolving Standardization for Continual Domain Generalization over Temporal Drift

NeurIPS 2023poster

The capability of generalizing to out-of-distribution data is crucial for the deployment of machine learning models in the real world. Existing domain generalization (DG) mainly embarks on offline and discrete scenarios, where multiple source domains are simultaneously accessible and the distributio…

2023

Exploring the Compositional Generalization in Context Dependent Text-to-SQL Parsing

ACL 2023findings

In the context-dependent Text-to-SQL task, the generated SQL statements are refined iteratively based on the user input utterance from each interaction. The input text from each interaction can be viewed as component modifications to the previous SQL statements, which could be further extracted as t…

2023

FIND: A Function Description Benchmark for Evaluating Interpretability Methods

NeurIPS 2023poster

Labeling neural network submodules with human-legible descriptions is useful for many downstream tasks: such descriptions can surface failures, guide interventions, and perhaps even explain important model behaviors. To date, most mechanistic descriptions of trained networks have involved small mode…

2023

Improving Generalization With Domain Convex Game

CVPR 2023poster

Domain generalization (DG) tends to alleviate the poor generalization capability of deep neural networks by learning model with multiple source domains. A classical solution to DG is domain augmentation, the common belief of which is that diversifying source domains will be conducive to the out-of-d…

2023

Language Semantic Graph Guided Data-Efficient Learning

NeurIPS 2023poster

Developing generalizable models that can effectively learn from limited data and with minimal reliance on human supervision is a significant objective within the machine learning community, particularly in the era of deep neural networks. Therefore, to achieve data-efficient learning, researchers ty…

2023

On the Difficulty of Unpaired Infrared-to-Visible Video Translation: Fine-Grained Content-Rich Patches Transfer

CVPR 2023poster

Explicit visible videos can provide sufficient visual information and facilitate vision applications. Unfortunately, the image sensors of visible cameras are sensitive to light conditions like darkness or overexposure. To make up for this, recently, infrared sensors capable of stable imaging have re…

2023

Open-vocabulary Panoptic Segmentation with Embedding Modulation

ICCV 2023poster

Open-vocabulary segmentation is attracting increasing attention due to its critical applications in the real world. Traditional closed-vocabulary segmentation methods are not able to characterize novel objects, whereas several recent open-vocabulary attempts obtain unsatisfactory results, i.e., nota…

Cited by 34PDFScholar
2023

PoseFusion: Robust Object-in-Hand Pose Estimation with SelectLSTM

IROS 2023poster

Accurate estimation of the relative pose between an object and a robot hand is critical for many manipulation tasks. However, most of the existing object-in-hand pose datasets use two-finger grippers and also assume that the object remains fixed in the hand without any relative movements, which is n…

Cited by 9SourcecodeScholar
2023

RAPL: A Relation-Aware Prototype Learning Approach for Few-Shot Document-Level Relation Extraction

EMNLP 2023long main

How to identify semantic relations among entities in a document when only a few labeled documents are available? Few-shot document-level relation extraction (FSDLRE) is crucial for addressing the pervasive data scarcity problem in real-world scenarios. Metric-based meta-learning is an effective fram…

Cited by 0SourcecodeScholar
2023

SP2 : A Second Order Stochastic Polyak Method

ICLR 2023poster

Recently the SP (Stochastic Polyak step size) method has emerged as a competitive adaptive method for setting the step sizes of SGD. SP can be interpreted as a method specialized to interpolated models, since it solves the interpolation equations. SP solves these equation by using local linearizati…

Cited by 13SourcePDFScholar
2023

Unsupervised Compositional Concepts Discovery with Text-to-Image Generative Models

ICCV 2023poster

Text-to-image generative models have enabled high-resolution image synthesis across different domains, but require users to specify the content they wish to generate. In this paper, we consider the inverse problem - given a collection of different images, can we discover the generative concepts that…

Cited by 14PDFScholar
2023

VBLC: Visibility Boosting and Logit-Constraint Learning for Domain Adaptive Semantic Segmentation under Adverse Conditions

AAAI 2023technical

Generalizing models trained on normal visual conditions to target domains under adverse conditions is demanding in the practical systems. One prevalent solution is to bridge the domain gap between clear- and adverse-condition images to make satisfactory prediction on the target. However, previous me…

2023

“Why Not Looking backward?” A Robust Two-Step Method to Automatically Terminate Bayesian Optimization

NeurIPS 2023poster

Bayesian Optimization (BO) is a powerful method for tackling expensive black-box optimization problems. As a sequential model-based optimization strategy, BO iteratively explores promising solutions until a predetermined budget, either iterations or time, is exhausted. The decision on when to termin…

Cited by 1SourcePDFScholar
2022

Active Learning for Domain Adaptation: An Energy-Based Approach

AAAI 2022technical

Unsupervised domain adaptation has recently emerged as an effective paradigm for generalizing deep neural networks to new target domains. However, there is still enormous potential to be tapped to reach the fully supervised performance. In this paper, we present a novel active learning strategy to a…

2022

Causality Inspired Representation Learning for Domain Generalization

CVPR 2022oral

Domain generalization (DG) is essentially an out-of-distribution problem, aiming to generalize the knowledge learned from multiple source domains to an unseen target domain. The mainstream is to leverage statistical models to model the dependence between data and labels, intending to learn represent…

Cited by 213PDFcodeScholar
2022

Compositional Visual Generation with Composable Diffusion Models

ECCV 2022poster

"Large text-guided diffusion models, such as DALLE-2, are able to generate stunning photorealistic images given natural language descriptions. While such models are highly flexible, they struggle to understand the composition of certain concepts, such as confusing the attributes of different objects…

2022

Explaining Point Processes by Learning Interpretable Temporal Logic Rules

ICLR 2022poster

We propose a principled method to learn a set of human-readable logic rules to explain temporal point processes. We assume that the generative mechanisms underlying the temporal point processes are governed by a set of first-order temporal logic rules, as a compact representation of domain knowledg…

Cited by 26SourcePDFScholar
2022

Generalization of Robot Force-Relevant Skills Through Adapting Compliant Profiles

RA-L 2022

Skill generalization in force fields is quite challenging and has not been fully investigated yet in the domain of robot learning. In this letter, we present a novel adaptation strategy that allows a robot to generalize the learned skill to deal with new task conditions with different force fields.

Cited by 11SourceScholar
2022

Learning Iterative Reasoning through Energy Minimization

ICML 2022spotlight

Deep learning has excelled on complex pattern recognition tasks such as image classification and object recognition. However, it struggles with tasks requiring nontrivial reasoning, such as algorithmic computation. Humans are able to solve such tasks through iterative reasoning – spending more time…

2022

Learning Modal-Invariant and Temporal-Memory for Video-Based Visible-Infrared Person Re-Identification

CVPR 2022poster

Thanks for the cross-modal retrieval techniques, visible-infrared (RGB-IR) person re-identification (Re-ID) is achieved by projecting them into a common space, allowing person Re-ID in 24-hour surveillance systems. However, with respect to the "probe-to-gallery", almost all existing RGB-IR based cro…

Cited by 64PDFcodeScholar
2022

Multifingered Grasping Based on Multimodal Reinforcement Learning

RA-L 2022

In this work, we tackle the challenging problem of grasping novel objects using a high-DoF anthropomorphic hand-arm system. Combining fingertip tactile sensing, joint torques and proprioception, a multimodal agent is trained in simulation to learn the finger motions and to determine when to lift an

Cited by 34SourceScholar
2022

Pre-Trained Language Models for Interactive Decision-Making

NeurIPS 2022accept

Language model (LM) pre-training is useful in many language processing tasks. But can pre-trained LMs be further leveraged for more general machine learning problems? We propose an approach for using LMs to scaffold learning and generalization in general sequential decision-making problems. In this…

Cited by 229SourcePDFScholar
2022

Towards Fewer Annotations: Active Learning via Region Impurity and Prediction Uncertainty for Domain Adaptive Semantic Segmentation

CVPR 2022oral

Self-training has greatly facilitated domain adaptive semantic segmentation, which iteratively generates pseudo labels on unlabeled target data and retrains the network. However, realistic segmentation datasets are highly imbalanced, pseudo labels are typically biased to the majority classes and bas…

Cited by 111PDFcodeScholar
2021

3D Neural Scene Representations for Visuomotor Control

CoRL 2021oral

Humans have a strong intuitive understanding of the 3D environment around us. The mental model of the physics in our brain applies to objects of different materials and enables us to perform a wide range of manipulation tasks that are far beyond the reach of current robots. In this work, we desire t…

Cited by 155SourceScholar
2021

Bi-Classifier Determinacy Maximization for Unsupervised Domain Adaptation

AAAI 2021technical

Unsupervised domain adaptation challenges the problem of transferring knowledge from a well-labelled source domain to an unlabelled target domain. Recently, adversarial learning with bi-classifier has been proven effective in pushing cross-domain distributions close. Prior approaches typically lever…

2021

Improved Contrastive Divergence Training of Energy-Based Models

ICML 2021spotlight

Contrastive divergence is a popular method of training energy-based models, but is known to have difficulties with training stability. We propose an adaptation to improve contrastive divergence training by scrutinizing a gradient term that is difficult to calculate and is often left out for convenie…

Cited by 169SourcePDFScholar
2021

Learning compliant grasping and manipulation by teleoperation with adaptive force control

IROS 2021poster

In this work, we focus on improving the robot’s dexterous capability by exploiting visual sensing and adaptive force control. TeachNet, a vision-based teleoperation learning framework, is exploited to map human hand postures to a multi-fingered robot hand. We augment TeachNet, which is originally ba…

Cited by 12SourceScholar
2021

MetaSAug: Meta Semantic Augmentation for Long-Tailed Visual Recognition

CVPR 2021poster

Real-world training data usually exhibits long-tailed distribution, where several majority classes have a significantly larger number of samples than the remaining minority classes. This imbalance degrades the performance of typical supervised learning algorithms designed for balanced training sets.…

Cited by 201PDFcodeScholar
2021

One-shot Face Reenactment Using Appearance Adaptive Normalization

AAAI 2021technical

The paper proposes a novel generative adversarial network for one-shot face reenactment, which can animate a single face image to a different pose-and-expression (provided by a driving image) while keeping its original appearance. The core of our network is a novel mechanism called appearance adapti…

Cited by 31SourcePDFScholar
2021

SOE-Net: A Self-Attention and Orientation Encoding Network for Point Cloud Based Place Recognition

CVPR 2021poster

We tackle the problem of place recognition from point cloud data and introduce a self-attention and orientation encoding network (SOE-Net) that fully explores the relationship between points and incorporates long-range context into point-wise local descriptors. Local information of each point from e…

Cited by 189PDFcodeScholar
2021

Semantic Concentration for Domain Adaptation

ICCV 2021poster

Domain adaptation (DA) paves the way for label annotation and dataset bias issues by the knowledge transfer from a label-rich source domain to a related but unlabeled target domain. A mainstream of DA methods is to align the feature distributions of the two domains. However, the majority of them foc…

Cited by 117PDFcodeScholar
2021

Transferable Semantic Augmentation for Domain Adaptation

CVPR 2021poster

Domain adaptation has been widely explored by transferring the knowledge from a label-rich source domain to a related but unlabeled target domain. Most existing domain adaptation algorithms attend to adapting feature representations across two domains with the guidance of a shared source-supervised…

Cited by 165PDFcodeScholar
2021

Unsupervised Learning of Compositional Energy Concepts

NeurIPS 2021poster

Humans are able to rapidly understand scenes by utilizing concepts extracted from prior experience. Such concepts are diverse, and include global scene descriptors, such as the weather or lighting, as well as local scene descriptors, such as the color or size of a particular object. So far, unsuperv…

2021

Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration

ICLR 2021spotlight

In this paper, we introduce Watch-And-Help (WAH), a challenge for testing social intelligence in agents. In WAH, an AI agent needs to help a human-like agent perform a complex household task efficiently. To succeed, the AI agent needs to i) understand the underlying goal of the task by watching a si…

2021

Weakly Supervised Human-Object Interaction Detection in Video via Contrastive Spatiotemporal Regions

ICCV 2021poster

We introduce the task of weakly supervised learning for detecting human and object interactions in videos. Our task poses unique challenges as a system does not know what types of human-object interactions are present in a video or the actual spatiotemporal location of the human and object. To addre…

Cited by 13PDFcodeScholar
2020

A Mobile Robot Hand-Arm Teleoperation System by Vision and IMU

IROS 2020poster

In this paper, we present a multimodal mobile teleoperation system that consists of a novel vision-based hand pose regression network (Transteleop) and an IMU (inertial measurement units)-based arm tracking method. Transteleop observes the human hand through a low-cost depth camera and generates not…

Cited by 75SourceScholar
2020

Robust Robotic Pouring using Audition and Haptics

IROS 2020poster

Robust and accurate estimation of liquid height lies as an essential part of pouring tasks for service robots. However, vision-based methods often fail in occluded conditions while audio-based methods cannot work well in a noisy environment. We instead propose a multimodal pouring network (MP-Net) t…

Cited by 24SourcecodeScholar
2019

Generative Adversarial User Model for Reinforcement Learning Based Recommendation System

ICML 2019oral

There are great interests as well as many challenges in applying reinforcement learning (RL) to recommendation systems. In this setting, an online user is the environment; neither the reward function nor the environment dynamics are clearly defined, making the application of RL challenging. In this…

2019

Making Sense of Audio Vibration for Liquid Height Estimation in Robotic Pouring

IROS 2019poster

In this paper, we focus on the challenging perception problem in robotic pouring. Most of the existing approaches either leverage visual or haptic information. However, these techniques may suffer from poor generalization performances on opaque containers or concerning measuring precision. To tackle…

Cited by 43SourceScholar
2019

PointNetGPD: Detecting Grasp Configurations from Point Sets

ICRA 2019poster

In this paper, we propose an end-to-end grasp evaluation model to address the challenging problem of localizing robot grasp configurations directly from the point cloud. Compared to recent grasp evaluation metrics that are based on handcrafted depth features and a convolutional neural network (CNN),…

Cited by 444SourcecodeScholar
2019

Simultaneous Blind Deconvolution and Phase Retrieval with Tensor Iterative Hard Thresholding

ICASSP 2019accepted

Blind deconvolution and phase retrieval are both fundamental problems with a growing interest in signal processing and communications. In this work, we consider the task of simultaneous blind deconvolution and phase retrieval. We show that this non-linear problem can be reformulated as a low-rank te…

Cited by 0SourceScholar
2019

The Landscape of Non-convex Empirical Risk with Degenerate Population Risk

NeurIPS 2019poster

The landscape of empirical risk has been widely studied in a series of machine learning problems, including low-rank matrix factorization, matrix sensing, matrix completion, and phase retrieval. In this work, we focus on the situation where the corresponding population risk is a degenerate non-conve…

Cited by 9SourcePDFScholar
2019

Vision-based Teleoperation of Shadow Dexterous Hand using End-to-End Deep Neural Network

ICRA 2019poster

In this paper, we present TeachNet, a novel neural network architecture for intuitive and markerless vision-based teleoperation of dexterous robotic hands. Robot joint angles are directly generated from depth images of the human hand that produce visually similar robot hand poses in an end-to-end fa…

Cited by 120SourceScholar
2018

Diversity Regularized Spatiotemporal Attention for Video-Based Person Re-Identification

CVPR 2018poster

Video-based person re-identification matches video clips of people across non-overlapping cameras. Most existing methods tackle this problem by encoding each video frame in its entirety and computing an aggregate representation across all frames. In practice, people are often partially occluded, whi…

Cited by 443SourcePDFScholar
2018

Learning Temporal Point Processes via Reinforcement Learning

NeurIPS 2018spotlight

Social goods, such as healthcare, smart city, and information networks, often produce ordered event data in continuous time. The generative processes of these event data can be very complex, requiring flexible models to capture their dynamics. Temporal point processes offer an elegant framework for…

2018

Question-Guided Hybrid Convolution for Visual Question Answering

ECCV 2018poster

In this paper, we propose a novel Question-Guided Hybrid Convolution (QGHC) network for Visual Question Answering (VQA). Most state-of-the-art VQA methods fuse the high-level textual and visual features from the neural network and abandon the visual spatial information when learning multi-modal feat…

Cited by 93SourcePDFScholar
2017

Fake News Mitigation via Point Process Based Intervention

ICML 2017poster

We propose the first multistage intervention framework that tackles fake news in social networks by combining reinforcement learning with a point process network activity model. The spread of fake news and mitigation events within the network is modeled by a multivariate Hawkes process with addition…

Cited by 222SourcePDFScholar
2017

Identity-Aware Textual-Visual Matching With Latent Co-Attention

ICCV 2017poster

Textual-visual matching aims at measuring similarities between sentence descriptions and images. Most existing methods tackle this problem without effectively utilizing identity-level annotations. In this paper, we propose an identity-aware two-stage framework for the textual-visual matching problem…

Cited by 315PDFScholar
2017

Jazz: A companion to music for frequency estimation with missing data

ICASSP 2017accepted

Frequency estimation is a classical problem in signal processing, with applications ranging from sensor array processing to wireless communications and structural health monitoring. Modern algorithms based on atomic norm minimization can cope with missing data but incur a high computational cost. To…

Cited by 0SourceScholar
2017

Joint Detection and Identification Feature Learning for Person Search

CVPR 2017spotlight

Existing person re-identification benchmarks and methods mainly focus on matching cropped pedestrian images between queries and candidates. However, it is different from real-world scenarios where the annotations of pedestrian bounding boxes are unavailable and the target person needs to be searched…

Cited by 1086PDFcodeScholar
2017

Learning Feature Pyramids for Human Pose Estimation

ICCV 2017poster

Articulated human pose estimation is a fundamental yet challenging task in computer vision. The difficulty is particularly pronounced in scale variations of human body parts when camera view changes or severe foreshortening happens. Although pyramid methods are widely used to handle scale changes at…

Cited by 645PDFcodeScholar
2015

COEVOLVE: A Joint Point Process Model for Information Diffusion and Network Co-evolution

NeurIPS 2015oral

Information diffusion in online social networks is affected by the underlying network topology, but it also has the power to change it. Online users are constantly creating new links when exposed to new information sources, and in turn these links are alternating the way information spreads. However…

2015

Efficient Learning of Continuous-Time Hidden Markov Models for Disease Progression

NeurIPS 2015poster

The Continuous-Time Hidden Markov Model (CT-HMM) is an attractive approach to modeling disease progression due to its ability to describe noisy observations arriving irregularly in time. However, the lack of an efficient parameter learning algorithm for CT-HMM restricts its use to very small models…

Cited by 146SourcePDFScholar