← Search

Liang He

105 accepted papers

2026

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation

CVPR 2026

Vision-Language Navigation (VLN) enables embodied agents to reach target locations in unseen environments by following language instructions. Despite recent progress with vision-language models (VLMs), a critical semantic-geometric gap remains: while VLMs excel at language and 2D visual understandin

Cited by 0SourcecodeScholar
2026

Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models

ICLR 2026poster

Large Vision-Language Models (LVLMs) exhibit outstanding performance on vision-language tasks but struggle with hallucination problems. Through in-depth analysis of LVLM activation patterns, we reveal two key findings: 1) truthfulness and visual perception capabilities predominantly engage different…

Cited by 4SourceScholar
2026

FlexProtein: Joint Sequence and Structure Pretraining for Protein Modeling

ICLR 2026poster

Protein foundation models have advanced rapidly, with most approaches falling into two dominant paradigms. Sequence-only language models (e.g., ESM-2) capture sequence semantics at scale but lack structural grounding. MSA-based predictors (e.g., AlphaFold 2/3) achieve accurate folding by exploiting…

Cited by 0SourceScholar
2026

Hyperbolic Additive Margin Softmax with Hierarchical Information for Speaker Verification

ICASSP 2026poster

Speaker embedding learning based on Euclidean space has achieved significant progress, but it is still insufficient in modeling hierarchical information within speaker features. Hyperbolic space, with its negative curvature geometric properties, can efficiently represent hierarchical information wit…

Cited by 0SourcePDFScholar
2026

LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization

AAAI 2026technical

Alignment plays a crucial role in Large Language Models (LLMs) in aligning with human preferences on a specific task/domain. Traditional alignment methods suffer from catastrophic forgetting, where models lose previously learned values when adapting to new preferences or domains. We introduce LifeAl

Cited by 0SourcePDFScholar
2026

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

CVPR 2026

While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image reasoning scenarios. Multi-image reasoning presents fundamental challenges including complex inter-relationships between images and scattered critical in

Cited by 0SourceScholar
2026

MovieGraph-ToM: Evaluating Long-Range Theory of Mind in Large Language Models via Implicit Social-Causal Graphs

AAAI 2026technical

The capacity for social reasoning, particularly Theory of Mind (ToM), is a foundational prerequisite for aligning Large Language Models (LLMs) with human values. However, current evaluations are predominantly confined to simplistic, short-text scenarios, obscuring their true capabilities and potenti

Cited by 0SourcePDFScholar
2026

VecDesigner: Exploring Visual Guidance and Structural Consistency for Semantic Typography

ICML 2026poster

Semantic Typography aims to visualize the meaning of an input word through the form of a character, while preserving its legibility. Existing vector-based methods, which primarily rely on text-driven optimization like Score Distillation Sampling (SDS), often produce glyphs that lack rich semantic de…

Cited by 0SourceScholar
2025

ACE-M3: Automatic Capability Evaluator for Multimodal Medical Models

COLING 2025main

As multimodal large language models (MLLMs) gain prominence in the medical field, the need for precise evaluation methods to assess their effectiveness has become critical. While benchmarks provide a reliable means to evaluate the capabilities of MLLMs, traditional metrics like ROUGE and BLEU employ…

Cited by 0SourcePDFScholar
2025

Aspect Enhancement and Text Simplification in Multimodal Aspect-Based Sentiment Analysis for Multi-Aspect and Multi-Sentiment Scenarios

AAAI 2025technical

Multimodal Aspect-Based Sentiment Analysis (MABSA) plays a pivotal role in the advancement of sentiment analysis technology. Although current methods strive to integrate multimodal information to enhance the performance of sentiment analysis, they still face two critical challenges when dealing with…

Cited by 0SourcePDFScholar
2025

AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation

ACL 2025long

With the proliferation of large language models (LLMs) in the medical domain, there is increasing demand for improved evaluation techniques to assess their capabilities. However, traditional metrics like F1 and ROUGE, which rely on token overlaps to measure quality, significantly overlook the import…

Cited by 0SourcePDFScholar
2025

CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering

CVPR 2025highlight

Multimodal large language models (MLLMs) have garnered widespread attention from researchers due to their remarkable understanding and generation capabilities in visual language tasks (e.g., visual question answering). However, the rapid pace of knowledge updates in the real world makes offline trai…

Cited by 2SourcePDFScholar
2025

Disentangled Modeling of Preferences and Social Influence for Group Recommendation

AAAI 2025technical

The group recommendation (GR) aims to suggest items for a group of users in social networks. Existing work typically considers individual preferences as the sole factor in aggregating group preferences. Actually, social influence is also an important factor in modeling users' contributions to the fi…

2025

DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving

ICCV 2025poster

This paper introduces DriveArena, the first high-fidelity closed-loop simulation system designed for driving agents navigating real-world scenarios. DriveArena comprises two core components: Traffic Manager, a traffic simulator capable of generating realistic traffic flow on any global street map, a…

Cited by 0SourcePDFScholar
2025

Dynamically Causal-Enhanced Exercise Representations for Adaptive Knowledge Tracing

ICASSP 2025accepted

Knowledge tracing assesses students’ mastery and predicts future performance based on historical learning data. Traditional methods primarily rely on predefined static associations between concepts and exercises, which struggle to capture potential causal relationships and dynamic learning patterns,…

Cited by 0SourceScholar
2025

Enhancing Medical Dialogue Generation through Knowledge Refinement and Dynamic Prompt Adjustment

ACL 2025finding

Medical dialogue systems (MDS) have emerged as crucial online platforms for enabling multi-turn, context-aware conversations with patients. However, existing MDS often struggle to (1) identify relevant medical knowledge and (2) generate personalized, medically accurate responses. To address these ch…

2025

Human Simulacra: Benchmarking the Personification of Large Language Models

ICLR 2025poster

Large Language Models (LLMs) are recognized as systems that closely mimic aspects of human intelligence. This capability has attracted the attention of the social science community, who see the potential in leveraging LLMs to replace human participants in experiments, thereby reducing research costs…

2025

Integrating Spectro-Temporal Cross Aggregation and Multi-Scale Dynamic Learning for Audio Deepfake Detection

ICASSP 2025accepted

Audio deepfake refers to the technology of synthesizing speech using deep learning or large model algorithms. Compared to human voice, synthetic deepfake speech exhibits artifacts at global and local levels, which can be leveraged by audio deepfake detection (ADD) to distinguish real and fake speech…

Cited by 0SourceScholar
2025

Lark: Low-Rank Updates After Knowledge Localization for Few-shot Class-Incremental Learning

ICCV 2025poster

For Few-Shot Class-Incremental Learning (FSCIL), direct fine-tuning causes significant parameter shifts, resulting in catastrophic forgetting and increased resource consumption. While, freezing the pre-trained backbone exacerbates the inconsistency between the backbone and the evolving classifier. T…

Cited by 0SourcePDFScholar
2025

Multi-Type Preference Learning: Empowering Preference-Based Reinforcement Learning with Equal Preferences

ICRA 2025

Preference-Based reinforcement learning (PBRL) learns directly from the preferences of human teachers regarding agent behaviors without needing meticulously designed reward functions. However, existing PBRL methods often learn primarily from explicit preferences, neglecting the possibility that teac

Cited by 1SourcecodeScholar
2025

NLSR: Neuron-Level Safety Realignment of Large Language Models Against Harmful Fine-Tuning

AAAI 2025technical

The emergence of fine-tuning-as-a-service has revealed a new vulnerability in large language models (LLMs). A mere handful of malicious data uploaded by users can subtly manipulate the fine-tuning process, leading to a compromised alignment state. Existing methods to counteract fine-tuning attacks t…

2025

Noise Supervised Contrastive Learning and Feature-Perturbed for Anomalous Sound Detection

ICASSP 2025accepted

Unsupervised anomalous sound detection aims to detect unknown anomalous sounds by training a model using only normal audio data. Despite advancements in self-supervised methods, the issue of frequent false alarms when handling samples of the same type from different machines remains unresolved. This…

Cited by 0SourceScholar
2025

Optimizing Question Semantic Space for Dynamic Retrieval-Augmented Multi-hop Question Answering

ACL 2025long

Retrieval-augmented generation (RAG) is usually integrated into large language models (LLMs) to mitigate hallucinations and knowledge obsolescence. Whereas, conventional one-step retrieve-and-read methods are insufficient for multi-hop question answering, facing challenges of retrieval semantic mism…

Cited by 0SourcePDFScholar
2025

P-React: Synthesizing Topic-Adaptive Reactions of Personality Traits via Mixture of Specialized LoRA Experts

ACL 2025finding

Personalized large language models (LLMs) have attracted great attention in many applications, such as emotional support and role-playing. However, existing works primarily focus on modeling explicit character profiles, while ignoring the underlying personality traits that truly shape behaviors and…

Cited by 0SourcePDFScholar
2025

RGR-KBQA: Generating Logical Forms for Question Answering Using Knowledge-Graph-Enhanced Large Language Model

COLING 2025main

In the field of natural language processing, Knowledge Base Question Answering (KBQA) is a challenging task that involves accurately retrieving answers from structured knowledge. Existing methods often face issues when generating query statements using LLMs, as the knowledge introduced may be imprec…

Cited by 0SourcePDFScholar
2025

Sentiment-enhanced Multi-hop Connected Graph Attention Network for Multimodal Aspect-Based Sentiment Analysis

IJCAI 2025

Multimodal aspect-based sentiment analysis aims to extract aspects from different data sources and recognize the corresponding sentiments. While current research has broadly focused on syntax relation-driven semantic comprehension, the impact of the importance of different syntactic relations on sem

Cited by 0SourcePDFScholar
2025

Supervisor Alignment Framework: Enhancing LLM Alignment with Query-Ignoring Strategy and Multi-Agent Interaction

ICASSP 2025accepted

The increasing focus on value alignment in Large Language Models (LLMs) underscores the need to ensure alignment with human morals and avoid biased or harmful outputs. However, LLMs aligned using existing methods are still easily affected by adversarial prompt attacks. Inspired by psychology, this p…

Cited by 0SourceScholar
2024

A Hierarchical Network for Multimodal Document-Level Relation Extraction

AAAI 2024technical

Document-level relation extraction aims to extract entity relations that span across multiple sentences. This task faces two critical issues: long dependency and mention selection. Prior works address the above problems from the textual perspective, however, it is hard to handle these problems solel…

2024

A Regularization-based Transfer Learning Method for Information Extraction via Instructed Graph Decoder

COLING 2024main

Information extraction (IE) aims to extract complex structured information from the text. Numerous datasets have been constructed for various IE tasks, leading to time-consuming and labor-intensive data annotations. Nevertheless, most prevailing methods focus on training task-specific models, while…

2024

BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind

AAAI 2024technical

As a foundational component of cognitive intelligence, theory of mind (ToM) can make AI more closely resemble human thought processes, thereby enhancing their interaction and collaboration with human. In particular, it can significantly improve a model's comprehension of videos in complex scenes. Ho…

2024

Boosting Large Language Models with Continual Learning for Aspect-based Sentiment Analysis

EMNLP 2024finding

Aspect-based sentiment analysis (ABSA) is an important subtask of sentiment analysis, which aims to extract the aspects and predict their sentiments. Most existing studies focus on improving the performance of the target domain by fine-tuning domain-specific models (trained on source domains) based…

Cited by 6SourcePDFScholar
2024

C-LLM: Learn to Check Chinese Spelling Errors Character by Character

EMNLP 2024main

Chinese Spell Checking (CSC) aims to detect and correct spelling errors in sentences. Despite Large Language Models (LLMs) exhibit robust capabilities and are widely applied in various tasks, their performance on CSC is often unsatisfactory. We find that LLMs fail to meet the Chinese character-level…

2024

CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios

EMNLP 2024main

With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in clinical medical scenarios, where models need to be examined very thoroughly. We present CliMedBench, a comprehensive benchmark with 14 expert-guided core clinica…

2024

Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving

NeurIPS 2024poster

Autonomous driving has advanced significantly due to sensors, machine learning, and artificial intelligence improvements. However, prevailing methods struggle with intricate scenarios and causal relationships, hindering adaptability and interpretability in varied environments. To address the above p…

2024

DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models

ICLR 2024poster

Recent advancements in autonomous driving have relied on data-driven approaches, which are widely adopted but face challenges including dataset bias, overfitting, and uninterpretability. Drawing inspiration from the knowledge-driven nature of human driving, we explore the question of how to instill…

2024

DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models

EMNLP 2024finding

Though large language models (LLMs) achieve significant success in recent years, the hallucination issue remains a challenge, and numerous benchmarks are proposed for hallucination detection. Nevertheless, some of these benchmarks are not naturally generated by LLMs but are intentionally induced. Al…

2024

EmoRED: A Dataset for Relation Extraction in Texts with Emoticons

ICASSP 2024accepted

Relation extraction (RE) is a vital task within natural language processing. Previous works predominantly focus on extracting relations from plain text. However, with the evolution of communication habits, many individuals employ symbolic representations, e.g. emoticons, to convey nuanced informatio…

Cited by 0SourceScholar
2024

Generating Persona-Aware Empathetic Responses with Retrieval-Augmented Prompt Learning

ICASSP 2024accepted

Empathetic response generation requires perceiving and understanding the user’s emotion to deliver suitable responses. However, existing models generally lack an ability to respond in a persona-specific way, which has been shown to play a vital role in expressing appropriate empathy. To address this…

Cited by 0SourceScholar
2024

Generative Calibration of Inaccurate Annotation for Label Distribution Learning

AAAI 2024technical

Label distribution learning (LDL) is an effective learning paradigm for handling label ambiguity. When applying LDL, it typically requires datasets annotated with label distributions. However, obtaining supervised data for LDL is a challenging task. Due to the randomness of label annotation, the ann…

Cited by 5SourcePDFScholar
2024

How Do Humans Write Code? Large Models Do It the Same Way Too

EMNLP 2024main

Program-of-Thought (PoT) replaces natural language-based Chain-of-Thought (CoT) as the most popular method in Large Language Models (LLMs) mathematical reasoning tasks by utilizing external tool calls to circumvent computational errors. However, our evaluation of the GPT-4 and Llama series reveals t…

2024

Hypernetwork-Assisted Parameter-Efficient Fine-Tuning with Meta-Knowledge Distillation for Domain Knowledge Disentanglement

NAACL 2024findings

Domain adaptation from labeled source domains to the target domain is important in practical summarization scenarios. However, the key challenge is domain knowledge disentanglement. In this work, we explore how to disentangle domain-invariant knowledge from source domains while learning specific kno…

Cited by 1SourcePDFScholar
2024

Introducing Multilingual Phonetic Information to Speaker Embedding for Speaker Verification

ICASSP 2024accepted

Incorporating frame-level phonetic information during the extraction of speaker embeddings has been shown to enhance the performance of speaker verification systems. However, previous studies have primarily relied on phonetic information obtained from pre-trained models of monolingual automatic spee…

Cited by 0SourceScholar
2024

Joint Multimodal Aspect Sentiment Analysis with Aspect Enhancement and Syntactic Adaptive Learning

IJCAI 2024poster

As an important task in sentiment analysis, joint multimodal aspect sentiment analysis (JMASA) has received increasing attention in recent years. However, previous approaches either i) directly fuse multimodal data without fully exploiting the correlation between multimodal input data, or ii) equall…

Cited by 4SourcePDFScholar
2024

Learning Intrinsic Dimension via Information Bottleneck for Explainable Aspect-based Sentiment Analysis

COLING 2024main

Gradient-based explanation methods are increasingly used to interpret neural models in natural language processing (NLP) due to their high fidelity. Such methods determine word-level importance using dimension-level gradient values through a norm function, often presuming equal significance for all…

Cited by 1SourcePDFScholar
2024

Let’s Rectify Step by Step: Improving Aspect-based Sentiment Analysis with Diffusion Models

COLING 2024main

Aspect-Based Sentiment Analysis (ABSA) stands as a crucial task in predicting the sentiment polarity associated with identified aspects within text. However, a notable challenge in ABSA lies in precisely determining the aspects’ boundaries (start and end indices), especially for long ones, due to us…

2024

MGCL: Multi-Granularity Clue Learning for Emotion-Cause Pair Extraction via Cross-Grained Knowledge Distillation

EMNLP 2024finding

Emotion-cause pair extraction (ECPE) aims to identify emotion clauses and their corresponding cause clauses within a document. Traditional methods often rely on coarse-grained clause-level annotations, which can overlook valuable fine-grained clues. To address this issue, we propose Multi-Granularit…

Cited by 0SourcePDFScholar
2024

MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

AAAI 2024technical

The emergence of various medical large language models (LLMs) in the medical domain has highlighted the need for unified evaluation standards, as manual evaluation of LLMs proves to be time-consuming and labor-intensive. To address this issue, we introduce MedBench, a comprehensive benchmark for the…

2024

MixRED: A Mix-lingual Relation Extraction Dataset

COLING 2024main

Relation extraction is a critical task in the field of natural language processing with numerous real-world applications. Existing research primarily focuses on monolingual relation extraction or cross-lingual enhancement for relation extraction. Yet, there remains a significant gap in understanding…

2024

Multi-View Speaker Embedding Learning for Enhanced Stability and Discriminability

ICASSP 2024accepted

Deep neural network models based on x-vector have become the most popular framework for speaker recognition, and the quality of speaker features (embeddings) is important for open-set tasks such as speaker verification and speaker diarization. Currently, the most popular loss function is based on ma…

Cited by 0SourceScholar
2024

Phase Continuity-Aware Self-Attentive Recurrent Network with Adaptive Feature Selection for Robust VAD

ICASSP 2024accepted

Deep neural network (DNN) applications have significantly progressed in voice activity detection (VAD). Most current DNN-based VAD methods ignore the rich audio information in the phase domain. Therefore, applying this auxiliary information rationally and coping with low signal-to-noise ratio (SNR)…

Cited by 0SourceScholar
2024

SMMA-Net: An Audio Clue-Based Target Speaker Extraction Network with Spectrogram Matching and Mutual Attention

ICASSP 2024accepted

We propose a deep neural network with spectrogram matching and mutual attention (SMMA-Net) for audio clue-based target speaker extraction (TSE). To effectively use the auxiliary speech, we proposed spectrogram matching (SM) strategy and mutual attention (MA) block. We conducted all experiments on th…

Cited by 0SourceScholar
2024

The design of a sensorized laryngoscope training system for pediatric intubation

IROS 2024poster

Intubation is essential for ventilating critically ill patients and involves precise maneuvering of a laryngoscope to place an endotracheal tube (ETT). However, training for this procedure is fraught with challenges. Traditional methods, relying on manikins or training with a single sensing modality…

Cited by 0SourceScholar
2024

Vanessa: Visual Connotation and Aesthetic Attributes Understanding Network for Multimodal Aspect-based Sentiment Analysis

EMNLP 2024finding

Prevailing research concentrates on superficial features or descriptions of images, revealing a significant gap in the systematic exploration of their connotative and aesthetic attributes. Furthermore, the use of cross-modal relation detection modules to eliminate noise from comprehensive image repr…

Cited by 6SourcePDFScholar
2023

A Disentangled-Attention Based Framework with Persona-Aware Prompt Learning for Dialogue Generation

AAAI 2023technical

Endowing dialogue agents with personas is the key to delivering more human-like conversations. However, existing persona-grounded dialogue systems still lack informative details of human conversations and tend to reply with inconsistent and generic responses. One of the main underlying causes is tha…

Cited by 5SourcePDFScholar
2023

DetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds

ICCV 2023poster

Existing offboard 3D detectors always follow a modular pipeline design to take advantage of unlimited sequential point clouds. We have found that the full potential of offboard 3D detectors is not explored mainly due to two reasons: (1) the onboard multi-object tracker cannot generate sufficient com…

Cited by 35PDFcodeScholar
2023

Disentangled CVAEs with Contrastive Learning for Explainable Recommendation

AAAI 2023technical

Modern recommender systems are increasingly expected to provide informative explanations that enable users to understand the reason for particular recommendations. However, previous methods struggle to interpret the input IDs of user--item pairs in real-world datasets, failing to extract adequate ch…

Cited by 7SourcePDFScholar
2023

Generative Label Enhancement with Gaussian Mixture and Partial Ranking

AAAI 2023technical

Label distribution learning (LDL) is an effective learning paradigm for dealing with label ambiguity. When applying LDL, the datasets annotated with label distributions (i.e., the real-valued vectors like the probability distribution) are typically required. Unfortunately, most existing datasets onl…

Cited by 5SourcePDFScholar
2023

HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval Generalization

ICLR 2023poster

Recently, large-scale text retrieval has made impressive progress, facilitating both information retrieval and downstream knowledge-intensive tasks (e.g., open-domain QA and dialogue). With a moderate amount of data, a neural text retriever can outperform traditional methods such as BM25 by a large…

Cited by 10SourcePDFScholar
2023

LoGoNet: Towards Accurate 3D Object Detection With Local-to-Global Cross-Modal Fusion

CVPR 2023poster

LiDAR-camera fusion methods have shown impressive performance in 3D object detection. Recent advanced multi-modal methods mainly perform global fusion, where image features and point cloud features are fused across the whole scene. Such practice lacks fine-grained region-level information, yielding…

2023

Pseudo-Query Generation For Semi-Supervised Visual Grounding With Knowledge Distillation

ICASSP 2023accepted

Visual grounding is a crucial multi-modal job for locating the objects that the referring queries refer to in images. In recent years, both fully-supervised and weakly-supervised algorithms rely on a large number of query annotations. However, collecting queries in natural language is labor-intensiv…

Cited by 0SourceScholar
2023

Real-Time Decentralized Navigation of Nonholonomic Agents Using Shifted Yielding Areas

ICRA 2023poster

We present a lightweight, decentralized algorithm for navigating multiple nonholonomic agents through challenging environments with narrow passages. Our key idea is to allow agents to yield to each other in large open areas instead of narrow passages, to increase the success rate of conventional dec…

Cited by 2SourceScholar
2023

SKD-NER: Continual Named Entity Recognition via Span-based Knowledge Distillation with Reinforcement Learning

EMNLP 2023long main

Continual learning for named entity recognition (CL-NER) aims to enable models to continuously learn new entity types while retaining the ability to recognize previously learned ones. However, the current strategies fall short of effectively addressing the catastrophic forgetting of previously learn…

Cited by 0SourceScholar
2023

Tell Model Where to Attend: Improving Interpretability of Aspect-Based Sentiment Classification via Small Explanation Annotations

ICASSP 2023accepted

Gradient-based explanation methods play an important role in the field of interpreting complex deep neural networks for NLP models. However, the existing work has shown that the gradients of a model are unstable and easily manipulable, which impacts the model’s reliability largely. According to our…

Cited by 0SourceScholar
2023

UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

EMNLP 2023long findings

Text is ubiquitous in our visual world, conveying crucial information, such as in documents, websites, and everyday photographs. In this work, we propose UReader, a first exploration of universal OCR-free visually-situated language understanding based on the Multimodal Large Language Model (MLLM). B…

Cited by 0SourcecodeScholar
2023

Uncertainty-Aware Few-Shot Class-Incremental Learning

ICASSP 2023accepted

In a real-world setting, machine needs to continuously recognize new categories without forgetting. However, the number of new categories may be small. For some difficult categories, even humans cannot recognize only based on few-shot examples. To address the above issues, an innovative uncertainty-…

Cited by 0SourceScholar
2022

A Modular Approach to Design Multi-Channel Bistable Valves for Integrated Pneumatically-Driven Soft Robots via 3D-Printing

RA-L 2022

A pneumatic system that transmits power via the force of compressed air is an essential component of an air-driven soft robot. Pneumatic valves are one of the key parts of this system. However, the development of soft or electronics-free valves for soft robotic applications is in its infancy, with o

Cited by 10SourceScholar
2022

A Multi-Format Transfer Learning Model for Event Argument Extraction via Variational Information Bottleneck

COLING 2022main

Event argument extraction (EAE) aims to extract arguments with given roles from texts, which have been widely studied in natural language processing. Most previous works have achieved good performance in specific EAE datasets with dedicated neural architectures. Whereas, these architectures are usua…

Cited by 20SourcePDFScholar
2022

An Information Minimization Based Contrastive Learning Model for Unsupervised Sentence Embeddings Learning

COLING 2022main

Unsupervised sentence embeddings learning has been recently dominated by contrastive learning methods (e.g., SimCSE), which keep positive pairs similar and push negative pairs apart. The contrast operation aims to keep as much information as possible by maximizing the mutual information between posi…

2022

CUP: Curriculum Learning based Prompt Tuning for Implicit Event Argument Extraction

IJCAI 2022poster

Implicit event argument extraction (EAE) aims to identify arguments that could scatter over the document. Most previous work focuses on learning the direct relations between arguments and the given trigger, while the implicit relations with long-range dependency are not well studied. Moreover, recen…

2022

Curriculum Prompt Learning with Self-Training for Abstractive Dialogue Summarization

EMNLP 2022main

Succinctly summarizing dialogue is a task of growing interest, but inherent challenges, such as insufficient training data and low information density impede our ability to train abstractive models. In this work, we propose a novel curriculum-based prompt learning method with self-training to addres…

2022

Design and Characterization of a 3D-Printed Pneumatically-Driven Bistable Valve With Tunable Characteristics

RA-L 2022

Although research studies in pneumatic soft robots develop rapidly, most pneumatic actuators are still controlled by rigid valves and conventional electronics. The existence of these rigid, electronic components sacrifices the compliance and adaptability of soft robots. Current electronics-free valv

Cited by 13SourceScholar
2022

Design of a 3D-Printed Soft Robotic Hand With Integrated Distributed Tactile Sensing

RA-L 2022

Humans rely on distributed tactile sensing in their hands to achieve robust and dexterous manipulation of delicate objects. Soft robotic hands have received increased attention in recent years due to their adaptability to unknown objects and safe interactions with the environment. However, the integ

Cited by 65SourceScholar
2022

Enhancing Class Understanding Via Prompt-Tuning For Zero-Shot Text Classification

ICASSP 2022accepted

Zero-shot text classification (ZSTC) poses a big challenge due to the lack of labeled data for unseen classes during training. Most studies focus on transferring knowledge from seen classes to unseen classes, which have achieved good performance in most cases. Whereas, it is difficult to transfer kn…

Cited by 0SourceScholar
2022

Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection

ECCV 2022poster

"Multi-modal 3D object detection has been an active research topic in autonomous driving. Nevertheless, it is non-trivial to explore the cross-modal feature fusion between sparse 3D points and dense 2D pixels. Recent approaches either fuse the image features with the point cloud features that are pr…

2022

Multi-Channel Attentive Graph Convolutional Network with Sentiment Fusion for Multimodal Sentiment Analysis

ICASSP 2022accepted

Nowadays, with the explosive growth of multimodal reviews on social media platforms, multimodal sentiment analysis has recently gained popularity because of its high relevance to these social media posts. Although most previous studies design various fusion frameworks for learning an interactive rep…

Cited by 0SourceScholar
2022

Multi-Robot Path Planning Using Medial-Axis-Based Pebble-Graph Embedding

IROS 2022poster

We present a centralized algorithm for labeled, disk-shaped Multi-Robot Path Planning (MPP) in a continuous planar workspace with polygonal boundaries. Our method automatically transform the continuous problem into a discrete, graph-based variant termed the pebble motion problem, which can be solved…

Cited by 3SourceScholar
2022

Multi-Scale Distribution Deep Variational Autoencoder for Explanation Generation

ACL 2022findings

Generating explanations for recommender systems is essential for improving their transparency, as users often wish to understand the reason for receiving a specified recommendation. Previous methods mainly focus on improving the generation quality, but often produce generic explanations that fail to…

Cited by 5SourcePDFScholar
2022

Shifting More Attention to Visual Backbone: Query-Modulated Refinement Networks for End-to-End Visual Grounding

CVPR 2022poster

Visual grounding focuses on establishing fine-grained alignment between vision and natural language, which has essential applications in multimodal reasoning systems. Existing methods use pre-trained query-agnostic visual backbones to extract visual feature maps independently without considering the…

Cited by 88PDFcodeScholar
2021

A Haptic Mouse Design with Stiffening Muscle Layer for Simulating Guarding in Abdominal Palpation Training

ICRA 2021poster

A patient would contract surface muscles as a reaction called muscle guarding when experiencing discomfort and pain during physical palpation. This reaction carries important information about an affected location. Training physicians to regulate palpation forces to elicit just enough muscle tension…

Cited by 9SourceScholar
2021

Automated Cross-prompt Scoring of Essay Traits

AAAI 2021technical

The majority of current research in Automated Essay Scoring (AES) focuses on prompt-specific scoring of either the overall quality of an essay or the quality with regards to certain traits. In real-world applications obtaining labelled data for a target essay prompt is often expensive or unfeasible,…

2021

Co-evolution Transformer for Protein Contact Prediction

NeurIPS 2021poster

Proteins are the main machinery of life and protein functions are largely determined by their 3D structures. The measurement of the pairwise proximity between amino acids of a protein, known as inter-residue contact map, well characterizes the structural information of a protein. Protein contact pre…

2021

Cross-Modal Knowledge Distillation For Fine-Grained One-Shot Classification

ICASSP 2021accepted

Few-shot learning can recognize a novel category based on only a few samples because it learns to learn from a lot of labeled samples during the training process. When data is insufficient, the performance is affected. And it is expensive to obtain a large-scale finegrained dataset with annotation.…

Cited by 0SourceScholar
2021

KERS: A Knowledge-Enhanced Framework for Recommendation Dialog Systems with Multiple Subgoals

EMNLP 2021finding

Recommendation dialogs require the system to build a social bond with users to gain trust and develop affinity in order to increase the chance of a successful recommendation. It is beneficial to divide up, such conversations with multiple subgoals (such as social chat, question answering, recommenda…

2021

Looking Wider for Better Adaptive Representation in Few-Shot Learning

AAAI 2021technical

Building a good feature space is essential for the metric-based few-shot algorithms to recognize a novel class with only a few samples. The feature space is often built by Convolutional Neural Networks (CNNs). However, CNNs primarily focus on local information with the limited receptive field, and t…

Cited by 58SourcePDFScholar
2021

MorphFace: A Hybrid Morphable Face for a Robopatient

RA-L 2021

Physicians use pain expressions shown in a patient's face to regulate their palpation methods during physical examination. Training to interpret patients’ facial expressions with different genders and ethnicities still remains a challenge, taking novices a long time to learn through experience. This

Cited by 10SourceScholar
2020

A Soft Pressure Sensor Skin for Hand and Wrist Orthoses

RA-L 2020

Side effects caused by excessive contact pressure such as discomfort and pressure sores are commonly complained by patients wearing orthoses. These problems leading to low patient compliance decrease the effectiveness of the device. To mitigate side effects, this study describes the design and fabri

Cited by 26SourceScholar
2020

Inner-Approximation of Manipulable and Reachable Regions using Bilinear Matrix Inequalities

IROS 2020poster

Given an articulated robot arm, we present a method to identify two regions with non-empty interiors. The first region is a subset of the configuration space where every point in the region is manipulable. The second region is a subset of the workspace where every point in the region is reachable by…

Cited by 1SourceScholar
2020

Scene Text Recognition with Temporal Convolutional Encoder

ICASSP 2020accepted

Texts from scene images typically consist of several characters and exhibit a characteristic sequence structure. Existing methods capture the structure with the sequence-to-sequence models by an encoder to have the visual representations and then a decoder to translate the features into the label se…

Cited by 0SourceScholar
2020

SentiX: A Sentiment-Aware Pre-Trained Model for Cross-Domain Sentiment Analysis

COLING 2020main

Pre-trained language models have been widely applied to cross-domain NLP tasks like sentiment analysis, achieving state-of-the-art performance. However, due to the variety of users’ emotional expressions across domains, fine-tuning the pre-trained models on the source domain tends to overfit, leadin…

2020

Soft Fingertips With Tactile Sensing and Active Deformation for Robust Grasping of Delicate Objects

RA-L 2020

Soft fingertips have shown significant adaptability for grasping a wide range of object shapes, thanks to elasticity. This ability can be enhanced to grasp soft, delicate objects by adding touch sensing. However, in these cases, the complete restraint and robustness of the grasps have proved to be c

Cited by 63SourceScholar
2017

Deep neural networks based speaker modeling at different levels of phonetic granularity

ICASSP 2017accepted

Recently, a hybrid deep neural network/i-vector framework has been proved effective for speaker verification, where the DNN trained to predict tied-triphone states (senones) is used to produce frame alignments for sufficient statistics extraction. In this work, in order to better understand the impa…

Cited by 0SourceScholar
2016

Efficient Penetration Depth Computation Between Rigid Models Using Contact Space Propagation Sampling

RA-L 2016

We present a novel method to compute the approximate global penetration depth (PD) between two nonconvex geometric models. Our approach consists of two phases: offline precomputation and run-time queries. In the first phase, our formulation uses a novel sampling algorithm to precompute an approximat

Cited by 8SourceScholar