← Search

han wang

97 accepted papers

2026

COMPRESSED BC-LISTA VIA LOW-RANK CONVOLUTIONAL DECOMPOSITION

ICASSP 2026poster

We study Sparse Signal Recovery (SSR) methods for multichannel imaging with compressed {forward and backward} operators that preserve reconstruction accuracy. We propose a Compressed Block-Convolutional (C-BC) measurement model based on a low-rank Convolutional Neural Network (CNN) decomposition tha…

Cited by 0SourcePDFScholar
2026

CSC-FMT*: Efficient 3D Path Planning Via Cylindrical Space-Cutting Fast Marching Tree

RA-L 2026

Sampling-based path planning algorithms like the Fast Marching Tree (FMT<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">*</sup>) often suffer from redundant exploration and inefficient sampling in complex 3D environments. To tackle these issues, this pa

Cited by 0SourceScholar
2026

D²G-TO: Task-aware and OOD-guided Discrete Graph Diffusion for Robust CNS Drug Discovery

IJCAI 2026

Central nervous system (CNS) drug discovery is constrained by an immense and sparse chemical search space. Meanwhile, molecules that simultaneously achieve brain penetration, target efficacy, and synthesizability are extremely scarce. However, existing generative models rarely couple rigorous multi-

Cited by 0Scholar
2026

FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI

AAAI 2026technical

As embodied intelligence emerges as a core frontier in artificial intelligence research, simulation platforms must evolve beyond low-level physical interactions to capture complex, human-centered social behaviors. We introduce FreeAskWorld, an interactive simulation framework that integrates large l

Cited by 0SourcePDFScholar
2026

Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology

ICLR 2026poster

Recent years have witnessed remarkable progress in multimodal learning within computational pathology. Existing models primarily rely on vision and language modalities; however, language alone lacks molecular specificity and offers limited pathological supervision, leading to representational bottle…

Cited by 0SourcecodeScholar
2026

Learning Koopman Representations with Controllability Guarantees

ICLR 2026poster

Learning nonlinear dynamical models from data is central to control. Two fundamental challenges exist: (1) how to learn accurate models from limited data, and (2) how to ensure the learned models are suitable for control design of the nominal system. We address both by enforcing a critical \emph{a p…

Cited by 0SourceScholar
2026

MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement

AAAI 2026technical

Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers pro

Cited by 0SourcePDFScholar
2026

Multi-Agent VLMs Guided Self-Training with PNU Loss for Low-Resource Offensive Content Detection

AAAI 2026technical

Accurate detection of offensive content on social media demands high-quality labeled data; however, such data is often scarce due to the low prevalence of offensive instances and the high cost of manual annotation. To address this low-resource challenge, we propose a self-training framework that lev

Cited by 0SourcePDFScholar
2026

Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports from Scratch with Agentic Framework

AAAI 2026technical

Visualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep research and generate comprehensive reports. Despite its progress, existing deep resear

Cited by 0SourcePDFScholar
2026

R-Tuning: Wavelet-Decomposed Replay and Semantic Alignment for Continual Adaptation of Pretrained Time-Series Models

AAAI 2026technical

Pre-trained models have demonstrated exceptional generalization capabilities in time-series forecasting; however, adapting them to evolving data distributions remains a significant challenge. A key hurdle lies in accessing the original training data, as fine-tuning solely on new data often leads to

Cited by 0SourcePDFScholar
2026

RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding

CVPR 2026

Understanding mid-level road semantics, which capture the structural and contextual cues that link low-level perception to high-level planning, is essential for reliable autonomous driving and digital map construction. However, existing benchmarks primarily target perception tasks such as detection

Cited by 0SourcecodeScholar
2026

S${3}$aDPWo: Spatial-, Semantic-, and Shape-Aware Diffusion Policy Toward Autonomous Wound Repair

RA-L 2026

Imitation learning (IL) offers a promising pathway for enabling surgical robots to perform autonomous wound repair. However, existing methods often neglect spatial semantics and wound-shape information, leading to poor generalization and low success rates. This paper presents the <bold>S</bold>patia

Cited by 0SourceScholar
2026

Verifiable Multimodal Reasoning: Fact-level Attribution with Multimodal Sources

ICML 2026poster

Multimodal large language models (MLLMs) are increasingly used for real-world tasks involving multi-step reasoning and long-form generation, where reliability requires grounding model outputs in heterogeneous input sources and verifying individual factual claims. However, existing multimodal groundi…

Cited by 0SourceScholar
2025

$q$-exponential family for policy optimization

ICLR 2025poster

Policy optimization methods benefit from a simple and tractable policy parametrization, usually the Gaussian for continuous action spaces. In this paper, we consider a broader policy family that remains tractable: the $q$-exponential family. This family of policies is flexible, allowing the specif…

2025

A Bounding Box is Worth One Token - Interleaving Layout and Text in a Large Language Model for Document Understanding

ACL 2025finding

Recently, many studies have demonstrated that exclusively incorporating OCR-derived text and spatial layouts with large language models (LLMs) can be highly effective for document understanding tasks. However, existing methods that integrate spatial layouts with text have limitations, such as produc…

2025

AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric Knowledge

NAACL 2025long

Knowledge conflict arises from discrepancies between information in the context of a large language model (LLM) and the knowledge stored in its parameters. This can hurt performance when using standard decoding techniques, which tend to ignore the context. Existing test-time contrastive methods seek…

2025

AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time

EMNLP 2025

This paper presents AlphaOne ( 𝛼1 ), a universal framework for modulating reasoning progress in large reasoning models (LRMs) at test time. 𝛼1 first introduces 𝛼 moment, which represents the scaled thinking phase with a universal parameter 𝛼 .Within this scaled pre- 𝛼 moment phase, it dynamically sc

2025

An Empirical Study of LLM Reasoning Ability Under Strict Output Length Constraint

EMNLP 2025

Recent work has demonstrated the remarkable potential of Large Language Models (LLMs) in test-time scaling. By making models think before answering, they are able to achieve much higher accuracy with extra inference computation.However, in many real-world scenarios, models are used under time constr

Cited by 0SourcePDFScholar
2025

Apollo-Forecast: Overcoming Aliasing and Inference Speed Challenges in Language Models for Time Series Forecasting

AAAI 2025technical

Encoding time series into tokens and using language models for processing has been shown to substantially augment the models' ability to generalize to unseen tasks. However, existing language models for time series forecasting encounter several obstacles, including aliasing distortion and prolonged…

2025

Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling

ICASSP 2025accepted

As small unmanned aerial vehicles (UAVs) become increasingly prevalent, there is growing concern regarding their impact on public safety and privacy, highlighting the need for advanced tracking and trajectory estimation solutions. In response, this paper introduces a novel framework that utilizes au…

Cited by 0SourceScholar
2025

Cross-modal Ship Re-Identification via Optical and SAR Imagery: A Novel Dataset and Method

ICCV 2025poster

Detecting and tracking ground objects using earth observation imagery remains a significant challenge in the field of remote sensing. Continuous maritime ship tracking is crucial for applications such as maritime search and rescue, law enforcement, and shipping analysis. However, most current ship t…

2025

Dual Multi-Scale GCN with Deformable Temporal Kernel for Skeleton-based Action Recognition

ICASSP 2025accepted

Skeleton sequences for action recognition are with complex temporal dynamics due to various factors such as speed variation and different activities. It is crucial and essential to model variation changes in the temporal dimension. In recent years, skeleton sequence is always modeled as a graph stru…

Cited by 0SourceScholar
2025

Dynamic Open-Vocabulary 3D Scene Graphs for Long-Term Language-Guided Mobile Manipulation

RA-L 2025

Enabling mobile robots to perform long-term tasks in dynamic real-world environments is a formidable challenge, especially when the environment changes frequently due to human-robot interactions or the robot's own actions. Traditional methods typically assume static scenes, which limits their applic

Cited by 36SourceScholar
2025

Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM

ICCV 2025poster

The application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in high-quality image-text datasets for fine-tuning image understanding, but there is still a lack of comparable datasets…

2025

Embodied Cognition Augmented End2End Autonomous Driving

NeurIPS 2025poster

In recent years, vision-based end-to-end autonomous driving has emerged as a new paradigm. However, popular end-to-end approaches typically rely on visual feature extraction networks trained under label supervision. This limited supervision framework restricts the generality and applicability of dri…

Cited by 0SourcecodeScholar
2025

Fat-to-Thin Policy Optimization: Offline Reinforcement Learning with Sparse Policies

ICLR 2025poster

Sparse continuous policies are distributions that can choose some actions at random yet keep strictly zero probability for the other actions, which are radically different from the Gaussian. They have important real-world implications, e.g. in modeling safety-critical tasks like medicine. The combin…

2025

Innovative Thinking, Infinite Humor: Humor Research of Large Language Models through Structured Thought Leaps

ICLR 2025poster

Humor is previously regarded as a gift exclusive to humans for the following reasons. Humor is a culturally nuanced aspect of human language, presenting challenges for its understanding and generation. Humor generation necessitates a multi-hop reasoning process, with each hop founded on proper ratio…

Cited by 1SourcePDFScholar
2025

Learning Structured Compressed Sensing with Automatic Resource Allocation

ICASSP 2025accepted

Multidimensional data acquisition often requires extensive time and poses significant challenges for hardware and software regarding data storage and processing. Rather than designing a single compression matrix as in conventional compressed sensing, structured compressed sensing yields dimension-sp…

Cited by 1SourceScholar
2025

Learning to Communicate Through Implicit Communication Channels

ICLR 2025poster

Effective communication is an essential component in collaborative multi-agent systems. Situations where explicit messaging is not feasible have been common in human society throughout history, which motivate the study of implicit communication. Previous works on learning implicit communication most…

Cited by 1SourcePDFScholar
2025

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations

NeurIPS 2025poster

Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR, most existing multimodal datasets provide only monaural audi…

Cited by 0SourcecodeScholar
2025

MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

ICML 2025poster

In recent years, while generative AI has advanced significantly in image generation, video generation continues to face challenges in controllability, length, and detail quality, which hinder its application. We present MimicMotion, a framework for generating high-quality human videos of arbitrary l…

2025

Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks

CVPR 2025poster

Vision Language Models (VLMs) can produce unintended and harmful content when exposed to adversarial attacks, particularly because their vision capabilities create new vulnerabilities. Existing defenses, such as input preprocessing, adversarial training, and response evaluation-based methods, are of…

2025

The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination

ICML 2025poster

Benchmark Data Contamination (BDC)—the inclusion of benchmark testing samples in the training set—has raised increasing concerns in Large Language Model (LLM) evaluation, leading to falsely inflated performance estimates and undermining evaluation reliability. To address this, researchers have propo…

2025

WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?

EMNLP 2025

The rapid advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced capabilities in Document Understanding. However, prevailing benchmarks like DocVQA and ChartQA predominantly comprise scanned or digital documents, inadequately reflecting the intricate challenges posed by

2024

A Novel Medical Image Fusion Framework Integrating Multi-scale Encoder-Decoder with Discrete Wavelet Decomposition

ICASSP 2024accepted

In recent years, many fusion algorithms based on multi-scale transform or neural networks have been proposed to improve medical image fusion (MIF) performance. However, there is still enormous potential to explore the combination of different fusion theories. In this paper, we propose a novel MIF fr…

Cited by 0SourceScholar
2024

ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based Evaluation

NeurIPS 2024poster

Large Language Models (LLMs) can elicit unintended and even harmful content when misaligned with human values, posing severe risks to users and society. To mitigate these risks, current evaluation benchmarks predominantly employ expert-designed contextual scenarios to assess how well LLMs align with…

2024

Carbon Market Simulation with Adaptive Mechanism Design

IJCAI 2024poster

A carbon market is a market-based tool that incentivizes economic agents to align individual profits with the global utility, i.e., reducing carbon emissions to tackle climate change. Cap and trade stands as a critical principle based on allocating and trading carbon allowances (carbon emission cred…

2024

Chinese MentalBERT: Domain-Adaptive Pre-training on Social Media for Chinese Mental Health Text Analysis

ACL 2024findings

In the current environment, psychological issues are prevalent and widespread, with social media serving as a key outlet for individuals to share their feelings. This results in the generation of vast quantities of data daily, where negative emotions have the potential to precipitate crisis situatio…

2024

Disentangling Masked Autoencoders for Unsupervised Domain Generalization

ECCV 2024poster

"Domain Generalization (DG), designed to enhance out-of-distribution (OOD) generalization, is all about learning invariance against domain shifts utilizing sufficient supervision signals. Yet, the scarcity of such labeled data has led to the rise of unsupervised domain generalization (UDG) — a more…

2024

Exploiting the Replay Memory Before Exploring the Environment: Enhancing Reinforcement Learning Through Empirical MDP Iteration

NeurIPS 2024poster

Reinforcement learning (RL) algorithms are typically based on optimizing a Markov Decision Process (MDP) using the optimal Bellman equation. Recent studies have revealed that focusing the optimization of Bellman equations solely on in-sample actions tends to result in more stable optimization, espec…

Cited by 0SourcePDFScholar
2024

Few-Shot Character Understanding in Movies as an Assessment to Meta-Learning of Theory-of-Mind

ICML 2024poster

When reading a story, humans can quickly understand new fictional characters with a few observations, mainly by drawing analogies to fictional and real people they already know. This reflects the few-shot and meta-learning essence of humans' inference of characters' mental states, *i.e.*, theory-of-…

2024

Finite-Time Analysis of On-Policy Heterogeneous Federated Reinforcement Learning

ICLR 2024poster

Federated reinforcement learning (FRL) has emerged as a promising paradigm for reducing the sample complexity of reinforcement learning tasks by exploiting information from different agents. However, when each agent interacts with a potentially different environment, little to nothing is known theor…

Cited by 20SourcePDFScholar
2024

Grasp Manipulation Relationship Detection based on Graph Sample and Aggregation

ICRA 2024poster

In multi-object stacking scenarios, exploring the relationships among objects and determining the correct sequence of operations are crucial for robotic manipulation. However, previous algorithms inefficiently combine global and local information, often focusing solely on the local features of objec…

Cited by 4SourceScholar
2024

Jointly Learning Selection Matrices for Transmitters, Receivers and Fourier Coefficients in Multichannel Imaging

ICASSP 2024accepted

Strategic subsampling has become a focal point due to its effectiveness in compressing data, particularly in the Full Matrix Capture (FMC) approach in ultrasonic imaging. This paper introduces the Joint Deep Probabilistic Subsampling (J-DPS) method, which aims to learn optimal selection matrices sim…

Cited by 0SourceScholar
2024

MMAUD: A Comprehensive Multi-Modal Anti-UAV Dataset for Modern Miniature Drone Threats

ICRA 2024poster

In response to the evolving challenges posed by small unmanned aerial vehicles (UAVs), which possess the potential to transport harmful payloads or independently cause damage, we introduce MMAUD: a comprehensive Multi-Modal Anti-UAV Dataset. MMAUD addresses a critical gap in contemporary threat dete…

Cited by 22SourcecodeScholar
2024

Momentum for the Win: Collaborative Federated Reinforcement Learning across Heterogeneous Environments

ICML 2024poster

We explore a Federated Reinforcement Learning (FRL) problem where $N$ agents collaboratively learn a common policy without sharing their trajectory data. To date, existing FRL work has primarily focused on agents operating in the same or ``similar" environments. In contrast, our problem setup allows…

Cited by 7SourcePDFScholar
2024

Online-Learning-Based Distributionally Robust Motion Control with Collision Avoidance for Mobile Robots

ICRA 2024poster

Collision-free navigation is a critical issue in robotic systems as the environment is often dynamic and uncertain. This paper investigates a data-stream-driven motion control problem for mobile robots to avoid randomly moving obstacles when the probability distribution of the obstacle’s movement is…

Cited by 0SourceScholar
2024

Real-Time Semantic Segmentation in Natural Environments with SAM-assisted Sim-to-Real Domain Transfer

IROS 2024poster

Semantic segmentation plays a pivotal role in many robotic applications requiring high-level scene understanding, such as smart farming, where the precise identification of trees or plants can aid navigation and crop monitoring tasks. While deep-learning-based semantic segmentation approaches have r…

Cited by 0SourcecodeScholar
2024

Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning

ACL 2024long

The surge in Large Language Models (LLMs) has revolutionized natural language processing, but fine-tuning them for specific tasks often encounters challenges in balancing performance and preserving general instruction-following abilities. In this paper, we posit that the distribution gap between tas…

2024

Self-Supervised Reinforcement Learning for Out-of-Distribution Recovery via Auxiliary Reward

ICASSP 2024accepted

Recently, the real-world applications of reinforcement learning (RL) have seen the problem of taking actions in an out-of-distribution (OOD) state. However, most existing research is limited to take actions to narrow the visited training distribution and OOD, and does not consider the efficiency to…

Cited by 0SourceScholar
2024

Soft Self-Consistency Improves Language Models Agents

ACL 2024short

Generations from large language models (LLMs) can be improved by sampling and scoring multiple solutions to select a final answer. Current “sample and select” methods such as self-consistency (SC) rely on majority voting to score answers. However, when tasks have many distinct and valid answers, sel…

Cited by 13SourcePDFScholar
2023

Evaluating GPT-3 Generated Explanations for Hateful Content Moderation

IJCAI 2023poster

Recent research has focused on using large language models (LLMs) to generate explanations for hate speech through fine-tuning or prompting. Despite the growing interest in this area, these generated explanations' effectiveness and potential limitations remain poorly understood. A key concern is tha…

2023

Improved Communication Efficiency in Federated Natural Policy Gradient via ADMM-based Gradient Updates

NeurIPS 2023poster

Federated reinforcement learning (FedRL) enables agents to collaboratively train a global policy without sharing their individual data. However, high communication overhead remains a critical bottleneck, particularly for natural policy gradient (NPG) methods, which are second-order. To address this…

Cited by 33SourcePDFScholar
2023

Look Beneath the Surface: Exploiting Fundamental Symmetry for Sample-Efficient Offline RL

NeurIPS 2023poster

Offline reinforcement learning (RL) offers an appealing approach to real-world tasks by learning policies from pre-collected datasets without interacting with the environment. However, the performance of existing offline RL algorithms heavily depends on the scale and state-action space coverage of d…

2023

MMRDN: Consistent Representation for Multi-View Manipulation Relationship Detection in Object-Stacked Scenes

ICRA 2023poster

Manipulation relationship detection (MRD) aims to guide the robot to grasp objects in the right order, which is important to ensure the safety and reliability of grasping in object stacked scenes. Previous works infer manipulation relationship by deep neural network trained with data collected from…

Cited by 2SourceScholar
2023

PyPose: A Library for Robot Learning With Physics-Based Optimization

CVPR 2023poster

Deep learning has had remarkable success in robotic perception, but its data-centric nature suffers when it comes to generalizing to ever-changing environments. By contrast, physics-based optimization generalizes better, but it does not perform as well in complicated tasks due to the lack of high-le…

2023

Replay Memory as An Empirical MDP: Combining Conservative Estimation with Experience Replay

ICLR 2023poster

Experience replay, which stores transitions in a replay memory for repeated use, plays an important role of improving sample efficiency in reinforcement learning. Existing techniques such as reweighted sampling, episodic learning and reverse sweep update further process the information in the replay…

Cited by 11SourcePDFScholar
2023

TLM: Token-Level Masking for Transformers

EMNLP 2023long main

Structured dropout approaches, such as attention dropout and DropHead, have been investigated to regularize the multi-head attention mechanism in Transformers. In this paper, we propose a new regularization scheme based on token-level rather than structure-level to reduce overfitting. Specifically,…

Cited by 0SourcecodeScholar
2023

The In-Sample Softmax for Offline Reinforcement Learning

ICLR 2023top-25%

Reinforcement learning (RL) agents can leverage batches of previously collected data to extract a reasonable control policy. An emerging issue in this offline RL setting, however, is that the bootstrapping update underlying many of our methods suffers from insufficient action-coverage: standard max…

2023

Versatile LiDAR-Inertial Odometry With SE(2) Constraints for Ground Vehicles

RA-L 2023

LiDAR SLAM has become one of the major localization systems for ground vehicles since LiDAR Odometry And Mapping (LOAM). Many extension works on LOAM mainly leverage one specific constraint to improve the performance, e.g., information from on-board sensors such as loop closure and inertial state; p

Cited by 11SourceScholar
2023

WINNER: Weakly-Supervised hIerarchical decompositioN and aligNment for Spatio-tEmporal Video gRounding

CVPR 2023poster

Spatio-temporal video grounding aims to localize the aligned visual tube corresponding to a language query. Existing techniques achieve such alignment by exploiting dense boundary and bounding box annotations, which can be prohibitively expensive. To bridge the gap, we investigate the weakly-supervi…

Cited by 40SourcePDFScholar
2022

Ask Question First for Enhancing Lifelong Language Learning

COLING 2022main

Lifelong language learning aims to stream learning NLP tasks while retaining knowledge of previous tasks. Previous works based on the language model and following data-free constraint approaches have explored formatting all data as “begin token (B) + context (C) + question (Q) + answer (A)” for diff…

2022

Automatic Multi-Label Prompting: Simple and Interpretable Few-Shot Classification

NAACL 2022long

Prompt-based learning (i.e., prompting) is an emerging paradigm for exploiting knowledge learned by a pretrained language model. In this paper, we propose Automatic Multi-Label Prompting (AMuLaP), a simple yet effective method to automatically select label mappings for few-shot text classification w…

2022

CRPN: Distinguish Novel Categories Via Class-Relevant Region Proposal Network for Few-Shot Object Detection

ICASSP 2022accepted

Few-shot object detection (FSOD) has attracted more attention in computer vision, where only very few training examples are presented during model learning process. A commonly-overlooked issue in FSOD is that novel classes are usually classified as background clutters in the pre-training process. An…

Cited by 0SourceScholar
2022

Incorporating Instructional Prompts into a Unified Generative Framework for Joint Multiple Intent Detection and Slot Filling

COLING 2022main

The joint multiple Intent Detection (ID) and Slot Filling (SF) is a significant challenge in spoken language understanding. Because the slots in an utterance may relate to multi-intents, most existing approaches focus on utilizing task-specific components to capture the relations between intents and…

2022

Language Model Pre-Training with Sparse Latent Typing

EMNLP 2022main

Modern large-scale Pre-trained Language Models (PLMs) have achieved tremendous success on a wide range of downstream tasks. However, most of the LM pre-training objectives only focus on text reconstruction, but have not sought to learn latent-level interpretable representations of sentences. In this…

2022

Multitask Prompted Training Enables Zero-Shot Task Generalization

ICLR 2022spotlight

Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a consequence of implicit multitask learning in language models’ pretraining (Radford et al., 2019). Can zero-shot genera…

2022

PTSEFormer: Progressive Temporal-Spatial Enhanced TransFormer towards Video Object Detection

ECCV 2022poster

"Recent years have witnessed a trend of applying context frames to boost the performance of object detection as video object detection. Existing methods usually aggregate features at one stroke to enhance the feature. These methods, however, usually lack spatial information from neighboring frames a…

2022

REGRAD: A Large-Scale Relational Grasp Dataset for Safe and Object-Specific Robotic Grasping in Clutter

RA-L 2022

Despite the impressive progress achieved in robotic grasping, robots are not skilled in sophisticated tasks (e.g. search and grasp a specified target in clutter). Such tasks involve not only grasping but the comprehensive perception of the world (e.g. the object relationships). Recently, encouraging

Cited by 51SourcecodeScholar
2021

Decomposing Complex Questions Makes Multi-Hop QA Easier and More Interpretable

EMNLP 2021finding

Multi-hop QA requires the machine to answer complex questions through finding multiple clues and reasoning, and provide explanatory evidence to demonstrate the machine’s reasoning process. We propose Relation Extractor-Reader and Comparator (RERC), a three-stage framework based on complex question d…

2021

Entity Resolution in Open-domain Conversations

NAACL 2021industry

In recent years, incorporating external knowledge for response generation in open-domain conversation systems has attracted great interest. To improve the relevancy of retrieved knowledge, we propose a neural entity linking (NEL) approach. Different from formal documents, such as news, conversationa…

Cited by 11SourcePDFScholar
2021

Optimizing NLU Reranking Using Entity Resolution Signals in Multi-domain Dialog Systems

NAACL 2021industry

In dialog systems, the Natural Language Understanding (NLU) component typically makes the interpretation decision (including domain, intent and slots) for an utterance before the mentioned entities are resolved. This may result in intent classification and slot tagging errors. In this work, we propo…

Cited by 2SourcePDFScholar
2021

Self-critical Learning of Influencing Factors for Trajectory Prediction using Gated Graph Convolutional Network

IROS 2021poster

Forecasting future trajectories of multiple pedestrians in a crowded environment is a challenging problem due to the complex interactions among the pedestrians. The interactions can be asymmetric and their influences may vary over time. Moreover, each pedestrian can exhibit different behavior at any…

Cited by 5SourceScholar
2020

Intensity Scan Context: Coding Intensity and Geometry Relations for Loop Closure Detection

ICRA 2020poster

Loop closure detection is an essential and challenging problem in simultaneous localization and mapping (SLAM). It is often tackled with light detection and ranging (LiDAR) sensor due to its view-point and illumination invariant properties. Existing works on 3D loop closure detection often leverage…

Cited by 328SourcecodeScholar
2020

OriNet: Robust 3-D Orientation Estimation With a Single Particular IMU

RA-L 2020

Estimating the robot's heading is a crucial requirement in odometry systems which are attempting to estimate the movement trajectory of a robot. Small errors in the orientation estimation result in a significant difference between the estimated and real trajectory, and failure of the odometry system

Cited by 94SourceScholar
2020

Towards Understanding and Inferring the Crowd: Guided Second Order Attention Networks and Re-identification for Multi-object Tracking

IROS 2020poster

Multi-human tracking in the crowded environment is a challenging problem due to occlusions, pose change, viewpoint variation and cluttered background. In this work, we propose a robust feature learning for tracking-by-detection methods based on second-order attention network that can capture higher-…

Cited by 1SourceScholar
2019

An Attention-aware Bidirectional Multi-residual Recurrent Neural Network (Abmrnn): A Study about Better Short-term Text Classification

ICASSP 2019accepted

Long Short-Term Memory (LSTM) has been proven an efficient way to model sequential data, because of its ability to overcome the gradient diminishing problem during training. However, due to the limited memory capacity in LSTM cells, LSTM is weak in capturing long-time dependency in sequential data.…

Cited by 0SourceScholar
2018

End-to-end Symmetry Preserving Inter-atomic Potential Energy Model for Finite and Extended Systems

NeurIPS 2018poster

Machine learning models are changing the paradigm of molecular modeling, which is a fundamental tool for material science, chemistry, and computational biology. Of particular interest is the inter-atomic potential energy surface (PES). Here we develop Deep Potential - Smooth Edition (DeepPot-SE), an…

2017

Geometric Map-Assisted Localization for Mobile Robots Based on Uniform-Gaussian Distribution

RA-L 2017

Drift and scale ambiguity are two main issues which reduce localization accuracy in monocular visual odometry (MVO). It is necessary to propose a unified model to represent these measurement uncertainties. In this paper, we present a geometric map-assisted localization approach for mobile robots equ

Cited by 18SourceScholar
2017

Image classification: A hierarchical dictionary learning approach

ICASSP 2017accepted

Hierarchical dictionary learning seeks multiple dictionaries at different image scales to capture complementary coherent characteristics. We propose a method to learn a hierarchy of two overcomplete synthesis dictionaries with an image classification goal. The classification objective in some sense…

Cited by 0SourceScholar
2017

Information diffusion in interconnected heterogeneous networks

ICASSP 2017accepted

In this paper, we are interested in modeling the diffusion of information in a multilayer network of agents using a thermodynamic diffusion approach. The state of each agent is viewed as a topic mixture, to describe his/her resources, and represented by a distribution over multiple topics. We observ…

Cited by 0SourceScholar
2015

Single carrier with multi-channel time-frequency domain equalization for underwater acoustic communications

ICASSP 2015accepted

Single-carrier with frequency domain equalization (SC-FDE) has been considered for bandwidth efficiency underwater acoustic (UWA) communication recently due to its reduced computational complexity and low peak-to-average power ratio. A multi-channel time-frequency domain equalization method for pseu…

Cited by 0SourceScholar