← Search

Lin Li

107 accepted papers

2026

Atom-level Adaptive Receptive Fields: A Pruning-Based Encoder for 2D Molecular Graphs (Student Abstract)

AAAI 2026technical

The two-dimensional (2D) graph structure of a molecule encodes abundant latent property information. A well-designed molecular graph encoder can capture informative low-dimensional dense representations of molecules, which can subsequently be applied to a widerange of downstream tasks. To achieve fi

Cited by 0SourcePDFScholar
2026

Building Reliable Long-Form Generation via Hallucination Rejection Sampling

ICML 2026poster

Large language models (LLMs) have achieved remarkable progress in open-ended text generation, yet they remain prone to hallucinating incorrect or unsupported content, which undermines their reliability. This issue is exacerbated in long-form generation due to hallucination snowballing, a phenomenon …

Cited by 0SourceScholar
2026

C-GNN-PRUNE: A Unified Graph-Based Framework for Structure-Aware Pruning of Mixture-of-Experts Models

AAAI 2026technical

The Mixture-of-Experts (MoE) architecture has emerged as a promising paradigm for scaling large language models (LLMs) by activating only a sparse subset of experts per input. However, its massive parameter size remains a major obstacle to efficient deployment. Existing pruning methods often ignore

Cited by 0SourcePDFScholar
2026

Efficient Offline Reinforcement Learning via Peer-Influenced Constraint

ICLR 2026poster

Offline reinforcement learning (RL) seeks to learn an optimal policy from a fixed dataset, but distributional shift between the dataset and the learned policy often leads to suboptimal real-world performance. Existing methods typically use behavior policy regularization to constrain the learned poli…

Cited by 0SourceScholar
2026

Evolving Quantitative Reasoning through Self-Play in Digital Twin Markets

ICML 2026poster

Large Language Models (LLMs) exhibit strong capabilities in high-level semantic understanding and strategic planning, yet they suffer from persistent quantitative failure modes, such as imprecise computation and the illusion of quantitative coherence, which limit their reliability in high-stakes dec…

Cited by 0SourceScholar
2026

Exploring Selective Avoidance for Online User Behavior Analysis: A Forest of Thought Explanation

AAAI 2026technical

The response behaviors observed in online user-generated content (UGC) frequently demonstrate non-linear characteristics, such as conditional branching and selective avoidance. These patterns present additional challenges for ensuring the trustworthiness of Large Language Model (LLMs) reasoning, par

Cited by 0SourcePDFScholar
2026

Guided Distillation and Risk Adaptive Evolution for Multi-Robot Navigation

AAAI 2026technical

Recent advancements in multi-robot navigation have explored methods that combine Large Language Models (LLMs) for tasks like scene understanding or high-level decision-making. However, these approaches face challenges with high inference latency and potential hallucinations. To address these challen

Cited by 0SourcePDFScholar
2026

Modeling Item-Level Dynamic Variability with Residual Diffusion for Bundle Recommendation

AAAI 2026technical

Existing solutions for bundle recommendation (BR) have achieved remarkable effectiveness for predicting the user’s preference for prebuilt bundles. However, bundle-item (B-I) affiliation will vary dynamically in real scenarios. For ex ample, a bundle themed as ‘casual outfit’ may add ‘hat’ or re

Cited by 0SourcePDFScholar
2026

Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech Separation

ICML 2026poster

Current audio-visual speech separation (AVSS) models typically rely on implicit multimodal fusion, but the absence of explicit modality alignment and reliability modeling often causes semantic misalignment and contaminates speech representations. The brain addresses this with a hierarchy: top-down a…

Cited by 0SourceScholar
2026

OTora: A Unified Red Teaming Framework for Reasoning-Level Denial-of-Service in LLM Agents

ICML 2026poster

Large Language Models (LLMs) are increasingly deployed as autonomous agents that execute tool-augmented, multi-step tasks, where latency is a critical factor for real-world applications. Yet an overlooked threat is Reasoning-Level Denial-of-Service (R-DoS), in which an attacker preserves task correc…

Cited by 0SourceScholar
2026

PAAL: Pattern-Anchor Alignment for Continual Knowledge Graph Embedding Under Structural Distribution Shift

IJCAI 2026

Continual knowledge graph embedding (CKGE) has gained popularity for managing dynamic knowledge graphs. Unlike general graph continual-learning approaches, CKGE focuses on retaining triple-level knowledge, thereby overcoming the inability of static models to accommodate continuously arriving facts.

Cited by 0Scholar
2026

PV-Ground: Text-Guided Point-Voxel Interaction for 3D Visual Grounding

CVPR 2026

3D visual grounding (VG) aims to localize target objects in 3D scenes based on free-form textual descriptions. Existing 3D VG models predominantly employ point-based backbones for point cloud feature extraction. Such methods require aggressive downsampling of the input point cloud, which sacrifices

Cited by 0SourcecodeScholar
2026

Path-Decoupled Hyperbolic Flow Matching for Few-Shot Adaptation

ICML 2026poster

Recent advances in cross-modal few-shot adaptation treat visual-semantic alignment as a continuous feature transport problem via Flow Matching (FM). However, we argue that Euclidean-based FM overlooks fundamental limitations of flat geometry, where polynomial volume growth fails to accommodate diver…

Cited by 0SourceScholar
2026

Q-SAM: Unlocking Sharpness-Aware Minimization for Generalization in Offline Reinforcement Learning

ICML 2026poster

Generalization remains a central challenge in offline reinforcement learning (RL), where policies are trained solely from static datasets and must perform reliably under distribution shift. While most existing offline RL methods focus on reducing training loss using standard optimizers such as Adam,…

Cited by 0SourceScholar
2026

Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension

AAAI 2026technical

Recent advances in multi-modal large language models (MLLMs) have significantly improved object-level grounding and region captioning. However, they remain limited in visual relation understanding, struggling even with binary relation detection, let alone N-ary relations involving multiple semantic

Cited by 0SourcePDFScholar
2026

Time Series Reasoning via Process-Verifiable Thinking Data Synthesis and Scheduling for Tailored LLM Reasoning

ICML 2026poster

Time series is a pervasive data type across various application domains, rendering the reasonable solving of diverse time series tasks a long-standing goal. Recent advances in large language models (LLMs), especially their reasoning abilities unlocked through reinforcement learning (RL), have opened…

Cited by 0SourceScholar
2025

A Noisy Label Filter based on GMM Binary Classification for Speaker Verification

ICASSP 2025accepted

Noisy labels are inevitable in real-world datasets. These noisy labels cause deep neural networks to gradient descent towards the wrong direction, leading to performance degradation. In this paper, We propose an efficient method for filtering out noisy labels during training. We calculate an embeddi…

Cited by 3SourceScholar
2025

A Recursive Total Least Squares Solution for Bearing-Only Target Motion Analysis and Circumnavigation

IROS 2025

Bearing-only Target Motion Analysis (TMA) is a promising technique for passive tracking in various applications as a bearing angle is easy to measure. Despite its advantages, bearing-only TMA is challenging due to the nonlinearity of the bearing measurement model and the lack of range information, w

Cited by 0SourceScholar
2025

A Survey on Multi-View Knowledge Graph: Generation, Fusion, Applications and Future Directions

IJCAI 2025

Knowledge Graphs (KGs) have revolutionized structured knowledge representation, yet their capacity to model real-world complexity and heterogeneity remains fundamentally constrained. The emerging paradigm of Multi-View Knowledge Graphs (MVKGs) addresses this gap through multi-view learning, but exis

Cited by 0SourcePDFScholar
2025

Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual Knowledge

CVPR 2025poster

Does seeing always mean knowing? Large Vision-Language Models (LVLMs) integrate separately pre-trained vision and language components, often using CLIP-ViT as vision backbone. However, these models frequently encounter a core issue of "cognitive misalignment" between the vision encoder (VE) and the…

2025

CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

CVPR 2025highlight

Interleaved image-text generation has emerged as a crucial multimodal task, aiming at creating sequences of interleaved visual and textual content given a query. Despite notable advancements in recent multimodal large language models (MLLMs), generating integrated image-text sequences that exhibit n…

2025

Cooperative Circumnavigation for Multi-Quadrotor Systems via Onboard Sensing

RA-L 2025

A cooperative circumnavigation framework is proposed for multi-quadrotor systems to enclose and track a moving target without reliance on external localization systems. The distinct relationships between quadrotor-quadrotor and quadrotor-target interactions are evaluated using a heterogeneous percep

Cited by 0SourceScholar
2025

Dynamic Language Group-based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical Routing

ICASSP 2025accepted

The Mixture of Experts (MoE) model is a promising approach for handling code-switching speech recognition (CS-ASR) tasks. However, the existing CS-ASR work on MoE has yet to leverage the advantages of MoE’s parameter scaling ability fully. This work proposes DLG-MoE, a Dynamic Language Group-based M…

Cited by 0SourceScholar
2025

Dynamic Uncertainty Estimation for Offline Reinforcement Learning

AAAI 2025technical

Offline reinforcement learning confronts the distributional shift challenge, a consequence of learning policy from static datasets. Current methods primarily handle this issue by aligning the learned policy with the behavior policy or conservatively estimating Q-values for out-of-distribution (OOD)…

Cited by 0SourcePDFScholar
2025

Hierarchy Coverage Path Planning With Proactive Extremum Prevention in Unknown Environments

RA-L 2025

The local extremum is a crucial factor that affects the efficiency of online coverage path planning (CPP). Most online CPP methods generate coverage motions point by point in unknown environments. However, these solutions ignore efficient global coverage and probably result in local extremum. This l

Cited by 0SourceScholar
2025

Improving Generalization in Offline Reinforcement Learning via Latent Distribution Representation Learning

AAAI 2025technical

Dealing with the distribution shift is a significant challenge when building offline reinforcement learning (RL) models that can generalize from a static dataset to out-of-distribution (OOD) scenarios. Previous approaches have employed pessimism or conservatism strategies. More recently, data-driven…

Cited by 0SourcePDFScholar
2025

Indirect Alignment and Relationship Preservation for Domain Generalization

IJCAI 2025

Domain generalization (DG) aims to train models on multiple source domains to generalize effectively to unseen target domains, addressing performance degradation caused by domain shifts. Many existing methods rely on direct feature alignment, which disrupts natural sequence relationships, causes mis

Cited by 0SourcePDFScholar
2025

Interaction-Centric Knowledge Infusion and Transfer for Open Vocabulary Scene Graph Generation

NeurIPS 2025poster

Open-vocabulary scene graph generation (OVSGG) extends traditional SGG by recognizing novel objects and relationships beyond predefined categories, leveraging the knowledge from pre-trained large-scale models. Existing OVSGG methods always adopt a two-stage pipeline: 1) Infusing knowledge into large…

Cited by 0SourceScholar
2025

L2M2: A Hierarchical Framework Integrating Large Language Model and Multi-agent Reinforcement Learning

IJCAI 2025

Multi-agent reinforcement learning (MARL) has demonstrated remarkable success in collaborative tasks, yet faces significant challenges in scaling to complex scenarios requiring sustained planning and coordination across long horizons. While hierarchical approaches help decompose these tasks, they ty

Cited by 0SourcePDFScholar
2025

LGNet: Linear Graph Representation for Efficient Cold-Start Recommendations

ICASSP 2025accepted

Graph Convolutional Networks (GCNs) demonstrate significant potential in recommendation systems but face difficulties with the cold-start problem, especially in integrating new nodes during inference. The typical solution leverages meta-learning for few-shot learning, though it often fails to fully…

Cited by 0SourceScholar
2025

MATO: A Model-Agnostic Training Optimization for Aspect Sentiment Triplet Extraction

NAACL 2025long

As an important fine-grained sentiment analysis task, aspect sentiment triplet extraction (ASTE) aims to identify three elements, i.e., aspect, opinion and sentiment polarity as a triplet. Advanced ASTE researches have mostly explored triplet-wise ability to achieve superior improvement. However, ex…

2025

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

EMNLP 2025

Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences. Typically, a reward model is trained or supplied to act as a proxy for humans in evaluating generated responses during the reinforcement training phase. Howe

2025

Risk-aware Direct Preference Optimization under Nested Risk Measure

NeurIPS 2025poster

When fine-tuning pre-trained Large Language Models (LLMs) to align with human values and intentions, maximizing the estimated reward can lead to superior performance, but it also introduces potential risks due to deviations from the reference model's intended behavior. Most existing methods typicall…

Cited by 0SourcecodeScholar
2025

SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow

ICASSP 2025accepted

Recently, flow matching based speech synthesis has significantly enhanced the quality of synthesized speech while reducing the number of inference steps. In this paper, we introduce SlimSpeech, a lightweight and efficient speech synthesis system based on rectified flow. We have built upon the existi…

Cited by 4SourceScholar
2025

Subgraph Information Bottleneck with Causal Dependency for Stable Molecular Relational Learning

IJCAI 2025

Molecular Relational Learning (MRL) is widely applied in molecular sciences. Recent studies attempt to retain molecular core information (e.g., substructures) by Graph Information Bottleneck but primarily focus on information compression without considering the causal dependencies of chemical reacti

Cited by 0SourcePDFScholar
2025

Zero-Shot Learning for Materials Science Texts: Leveraging Duck Typing Principles

AAAI 2025technical

Materials science text mining (MSTM), involving tasks like property extraction and synthesis action retrieval, is pivotal for advancing research by deriving critical insights from scientific literature. Descriptors, serving as essential task labels, often vary in meaning depending on researchers' us…

2024

A Coarse-to-Fine Place Recognition Approach using Attention-guided Descriptors and Overlap Estimation

ICRA 2024poster

Place recognition is a challenging but crucial task in robotics. Current description-based methods may be limited by representation capabilities, while pairwise similarity-based methods require exhaustive searches, which is time-consuming. In this paper, we present a novel coarse-to-fine approach to…

Cited by 1SourcecodeScholar
2024

An Integrated Position-velocity-force Method for Safety-enhanced Shared Control in Robot-assisted Surgical Cutting

ICRA 2024poster

Numerous studies have emphasized the application of autonomous intelligence in human-robot shared control to enhance surgical convenience and efficiency. However, the neglect of human dominance may reduce surgical safety. This paper developed a safety-enhanced human-robot shared control method by in…

Cited by 0SourceScholar
2024

Improving Generalization in Offline Reinforcement Learning via Adversarial Data Splitting

ICML 2024poster

Offline Reinforcement Learning (RL) commonly suffers from the out-of-distribution (OOD) overestimation issue due to the distribution shift. Prior work gradually shifts their focus from suppressing OOD overestimation to avoiding overly conservative learning from suboptimal behavior policies to improv…

2024

Improving Multi-Speaker ASR With Overlap-Aware Encoding And Monotonic Attention

ICASSP 2024accepted

End-to-end (E2E) multi-speaker speech recognition with the serialized output training (SOT) strategy demonstrates good performance in modeling diverse speaker scenarios. However, the E2E architecture doesn’t explicitly address the modeling of overlapping speech areas, potentially limiting the model’…

Cited by 0SourceScholar
2024

MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis

AAAI 2024technical

The style transfer task in Text-to-Speech (TTS) refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However, most existing style transfer approaches are either based on fixed emotional labels or reference speech clips, whi…

2024

Majority Rules Guided Aspect-Category Based Sentiment Analysis via Label Prior Knowledge

COLING 2024main

As an important fine-grained task of sentiment analysis, Aspect-Category based Sentiment Analysis (ACSA) aims to identify the sentiment polarities of pre-defined categories in text. However, due to subjectivity, the highly semantically similar text has polysemous sentiments to different people, lead…

2024

MemoryFormer : Minimize Transformer Computation by Removing Fully-Connected Layers

NeurIPS 2024poster

In order to reduce the computational complexity of large language models, great efforts have been made to to improve the efficiency of transformer models such as linear attention and flash-attention. However, the model size and corresponding computational complexity are constantly scaled up in pursu…

Cited by 0SourcePDFScholar
2024

Multivariate Fourier Distribution Perturbation: Domain Shifts with Uncertainty in Frequency Domain

ICASSP 2024accepted

Diversifying training data techniques have achieved tremendous success in Domain Generalization (DG) tasks. The key to diversifying domain data is by increasing the types of domain styles. After investigating this issue from the perspective of the Fourier transform, the domain cue is found to be imp…

Cited by 0SourceScholar
2024

OODRobustBench: a Benchmark and Large-Scale Analysis of Adversarial Robustness under Distribution Shift

ICML 2024poster

Existing works have made great progress in improving adversarial robustness, but typically test their method only on data from the same distribution as the training data, i.e. in-distribution (ID) testing. As a result, it is unclear how such robustness generalizes under input distribution shifts, i.…

2024

One Prompt Word is Enough to Boost Adversarial Robustness for Pre-trained Vision-Language Models

CVPR 2024poster

Large pre-trained Vision-Language Models (VLMs) like CLIP despite having remarkable generalization ability are highly vulnerable to adversarial examples. This work studies the adversarial robustness of VLMs from the novel perspective of the text prompt instead of the extensively studied model weight…

2024

Reflow-TTS: A Rectified Flow Model for High-Fidelity Text-to-Speech

ICASSP 2024accepted

The diffusion models including Denoising Diffusion Probabilistic Models (DDPM) and score-based generative models have demonstrated excellent performance in speech synthesis tasks. However, its effectiveness comes at the cost of numerous sampling steps, resulting in prolonged sampling time required t…

Cited by 0SourceScholar
2024

SR-HuBERT : An Efficient Pre-Trained Model for Speaker Verification

ICASSP 2024accepted

Recently, pre-trained models (PTMs) have been extensively applied in speaker verification (SV) and greatly boosted system performance. However, mainstream PTMs currently concentrate on using frame-level universal representations. In this paper, we propose a novel pre-training framework that jointly…

Cited by 0SourceScholar
2024

Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement Learning

NeurIPS 2024poster

A challenging problem in seeking to bring multi-agent reinforcement learning (MARL) techniques into real-world applications, such as autonomous driving and drone swarms, is how to control multiple agents safely and cooperatively to accomplish tasks. Most existing safe MARL methods learn the centrali…

Cited by 1SourcePDFScholar
2024

Visual Pivoting Unsupervised Multimodal Machine Translation in Low-Resource Distant Language Pairs

EMNLP 2024finding

Unsupervised multimodal machine translation (UMMT) aims to leverage vision information as a pivot between two languages to achieve better performance on low-resource language pairs. However, there is presently a challenge: how to handle alignment between distant language pairs (DLPs) in UMMT. To thi…

2024

VulLibGen: Generating Names of Vulnerability-Affected Packages via a Large Language Model

ACL 2024long

Security practitioners maintain vulnerability reports (e.g., GitHub Advisory) to help developers mitigate security risks. An important task for these databases is automatically extracting structured information mentioned in the report, e.g., the affected software packages, to accelerate the defense…

2024

ZONE: Zero-Shot Instruction-Guided Local Editing

CVPR 2024poster

Recent advances in vision-language models like Stable Diffusion have shown remarkable power in creative image synthesis and editing.However most existing text-to-image editing methods encounter two obstacles: First the text prompt needs to be carefully crafted to achieve good results which is not in…

2023

A Unified BEV Model for Joint Learning of 3D Local Features and Overlap Estimation

ICRA 2023poster

Pairwise point cloud registration is a critical task for many applications, which heavily depends on finding correct correspondences from the two point clouds. However, the low overlap between input point clouds causes the registration to fail easily, leading to mistaken overlapping and mismatched c…

Cited by 5SourcecodeScholar
2023

Background Disturbance Mitigation for Video Captioning Via Entity-Action Relocation

ICASSP 2023accepted

Video captioning aims to generate sentences to accurately describe the video content, in which video background plays the role of prompts. State-of-the-art methods tend to explore richer video representations adequately, fusing with language to improve caption quality, which has shown great success.…

Cited by 0SourceScholar
2023

Code-Enhanced Fine-Grained Semantic Matching For Tag Recommendation In Software Information Sites

ICASSP 2023accepted

Tag recommendation in software information sites is a significant task to help developers make distinctions among software objects. Most existing methods usually ignore the semantic information of code snippets in software information sites. To tackle this issue, we regard the code as a semantic enh…

Cited by 0SourceScholar
2023

Collision-free Coverage Path Planning for the Variable-speed Curvature-constrained Robot

ICRA 2023poster

Dubins coverage has been extensively researched to address the coverage path planning (CPP) problem of a known environment for the curvature-constrained robot. However, its fixed-speed assumption prevents the robot from accelerating to reduce the time and limits its flexibility to avoid obstacles. T…

Cited by 2SourceScholar
2023

Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization

ICASSP 2023accepted

The clustering algorithm plays a crucial role in speaker diarization systems. However, traditional clustering algorithms suffer from the complex distribution of speaker embeddings and lack of digging potential relationships between speakers in a session. We propose a novel graph-based clustering app…

Cited by 0SourceScholar
2023

Compositional Feature Augmentation for Unbiased Scene Graph Generation

ICCV 2023poster

Scene Graph Generation (SGG) aims to detect all the visual relation triplets <sub, pred, obj> in a given image. With the emergence of various advanced techniques for better utilizing both the intrinsic and extrinsic information in each relation triplet, SGG has achieved great progress over the recen…

Cited by 44PDFcodeScholar
2023

DT-Solver: Automated Theorem Proving with Dynamic-Tree Sampling Guided by Proof-level Value Function

ACL 2023long

Recent advances in neural theorem-proving resort to large language models and tree searches. When proving a theorem, a language model advises single-step actions based on the current proving state and the tree search finds a sequence of correct steps using actions given by the language model. Howeve…

Cited by 35SourcePDFScholar
2023

Descriptive Prompt Paraphrasing for Target-Oriented Multimodal Sentiment Classification

EMNLP 2023long findings

Target-Oriented Multimodal Sentiment Classification (TMSC) aims to perform sentiment polarity on a target jointly considering its corresponding multiple modalities including text, image, and others. Current researches mainly work on either of two types of targets in a decentralized manner. One type…

Cited by 0SourceScholar
2023

Fast Task Allocation of Heterogeneous Robots With Temporal Logic and Inter-Task Constraints

RA-L 2023

This work develops a fast task allocation framework for heterogeneous multi-robot systems subject to both temporal logic and inter-task constraints. The considered inter-task constraints include unrelated tasks, compatible tasks, and exclusive tasks. To specify such inter-task relationships, we exte

Cited by 24SourceScholar
2023

Long Legal Article Question Answering via Cascaded Key Segment Learning (Student Abstract)

AAAI 2023technical

Current sentence-level evidence extraction based methods may lose the discourse coherence of legal articles since they tend to make the extracted sentences scattered over the article. To solve the problem, this paper proposes a Cascaded Answer-guided key segment learning framework for long Legal ar…

Cited by 5SourcePDFScholar
2023

MFF-Net: Towards Efficient Monocular Depth Completion With Multi-Modal Feature Fusion

RA-L 2023

Remarkable progress has been achieved by current depth completion approaches, which produce dense depth maps from sparse depth maps and corresponding color images. However, the performances of these approaches are limited due to the insufficient feature extractions and fusions. In this work, we prop

Cited by 38SourceScholar
2023

MGIA: Mutual Gradient Inversion Attack in Multi-Modal Federated Learning (Student Abstract)

AAAI 2023technical

Recent studies have demonstrated that local training data in Federated Learning can be recovered from gradients, which are called gradient inversion attacks. These attacks display powerful effects on either computer vision or natural language processing tasks. As it is known that there are certain c…

Cited by 5SourcePDFScholar
2023

Meta Learning with Adaptive Loss Weight for Low-Resource Speech Recognition

ICASSP 2023accepted

Model Agnostic Meta-Learning (MAML) is an effective meta-learning algorithm for low-resource automatic speech recognition (ASR). It uses gradient descent to learn the initialization parameters of the model through various languages, making the model quickly adapt to unseen low-resource languages. Bu…

Cited by 0SourceScholar
2023

Multimodal Propaganda Detection Via Anti-Persuasion Prompt enhanced contrastive learning

ICASSP 2023accepted

Propaganda, commonly used in memes disinformation, can influence the thinking of the audience and increase the reach of communication. Usually logical fallacy, as a kind of popular expression of memes, aims to create a logical reasonable illusion where the conclusion cannot be drawn with the use of…

Cited by 0SourceScholar
2023

Set-membership Belief State-based Reinforcement Learning for POMDPs

ICML 2023poster

Reinforcement learning (RL) has made significant progress in areas such as Atari games and robotic control, where the agents have perfect sensing capabilities. However, in many real-world sequential decision-making tasks, the observation data could be noisy or incomplete due to the intrinsic low qua…

Cited by 0SourcePDFScholar
2023

TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language Models

EMNLP 2023long main

Automated theorem proving (ATP) has become an appealing domain for exploring the reasoning ability of the recent successful generative language models. However, current ATP benchmarks are mainly focus on symbolic inference, but rarely involve the understanding of complex number combination reasoni…

Cited by 0SourcecodeScholar
2023

The XMU System for Audio-Visual Diarization and Recognition in MISP Challenge 2022

ICASSP 2023accepted

In this paper, we present our work in track 2 of the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. We built a cascaded system and explored different acoustic front-ends and end-to-end speech recognition back-ends based on multimodal. To promote effective fusion between the d…

Cited by 0SourceScholar
2023

Tree-Like Interaction Learning for Bundle Recommendation

ICASSP 2023accepted

Bundle recommendation suggests a set of items to users against their complex needs, where user-bundle interaction learning is key. It is observed that Gromov’s δ-hyperbolicity of the interaction graph in bundle recommendation is smaller (lower is more hyperbolic) than those in traditional item recom…

Cited by 0SourceScholar
2023

Unsupervised Speaker Verification Using Pre-Trained Model and Label Correction

ICASSP 2023accepted

Recently, the fine-tuning pre-trained model framework has emerged as a promising paradigm for speech-processing tasks. In this study, we present a novel strategy for unsupervised speaker verification using the Sub-structure of Pre-Trained Model (Sub-PTM), which consists of a CNN-based feature extrac…

Cited by 0SourceScholar
2023

Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language Models

NeurIPS 2023poster

Pretrained vision-language models, such as CLIP, have demonstrated strong generalization capabilities, making them promising tools in the realm of zero-shot visual recognition. Visual relation detection (VRD) is a typical task that identifies relationship (or interaction) types between object pairs…

2022

Code Generation From Flowcharts with Texts: A Benchmark Dataset and An Approach

EMNLP 2022finding

Currently, researchers focus on generating codes from the requirement documents. However, current approaches still perform poorly on some requirements needing complex problem-solving skills. In reality, to tackle such complex requirements, instead of directly translating requirement documents into c…

2022

Controlling Underestimation Bias in Reinforcement Learning via Quasi-median Operation

AAAI 2022technical

How to get a good value estimation is one of the key problems in reinforcement learning (RL). Current off-policy methods, such as Maxmin Q-learning, TD3 and TADD, suffer from the underestimation problem when solving the overestimation problem. In this paper, we propose the Quasi-Median Operation, a…

Cited by 16SourcePDFScholar
2022

Graph Convolutional Network Based Semi-Supervised Learning on Multi-Speaker Meeting Data

ICASSP 2022accepted

Unsupervised clustering on speakers is becoming increasingly important for its potential uses in semi-supervised learning. In reality, we are often presented with enormous amounts of unlabeled data from multi-party meetings and discussions. An effective unsupervised clustering approach would allow u…

Cited by 0SourceScholar
2022

Improving Cross-Lingual Speech Synthesis with Triplet Training Scheme

ICASSP 2022accepted

Recent advances in cross-lingual text-to-speech (TTS) made it possible to synthesize speech in a language foreign to a monolingual speaker. However, there is still a large gap between the pronunciation of generated cross-lingual speech and that of native speakers in terms of naturalness and intellig…

Cited by 0SourceScholar
2022

RINet: Efficient 3D Lidar-Based Place Recognition Using Rotation Invariant Neural Network

RA-L 2022

LiDAR-based place recognition (LPR) is one of the basic capabilities of robots, which can retrieve scenes from maps and identify previously visited locations based on 3D point clouds. As robots often pass the same place from different views, LPR methods are supposed to be robust to rotation, which i

Cited by 64SourceScholar
2022

The Devil Is in the Labels: Noisy Label Correction for Robust Scene Graph Generation

CVPR 2022oral

Unbiased SGG has achieved significant progress over recent years. However, almost all existing SGG models have overlooked the ground-truth annotation qualities of prevailing SGG datasets, i.e., they always assume: 1) all the manually annotated positive samples are equally correct; 2) all the un-anno…

Cited by 119PDFcodeScholar
2022

Towards the Quantitative Interpretability Analysis of Citizens Happiness Prediction

IJCAI 2022poster

Evaluating the high-effect factors of citizens' happiness is beneficial to a wide range of policy-making for economics and politics in most countries. Benefiting from the high-efficiency of regression models, previous efforts by sociology scholars have analyzed the effect of happiness factors with h…

2021

ASV-SUBTOOLS: Open Source Toolkit for Automatic Speaker Verification

ICASSP 2021accepted

In this paper, we introduce a new open source toolkit for automatic speaker verification (ASV), named ASV-Subtools. Adopting PyTorch as main deep learning engine and Kaldi toolkit for data processing, ASV-Subtools allows users to develop modern speaker recognizers flexibly and efficiently. The toolk…

Cited by 0SourceScholar
2021

COOPNet: Multi-Modal Cooperative Gender Prediction in Social Media User Profiling

ICASSP 2021accepted

The principal way of performing user profiling is to investigate accumulated social media data. However, the problem of information asymmetry generally exists in user generated contents since users post multi-modal contents in social media freely. In this paper, we propose a novel text-image coopera…

Cited by 0SourceScholar
2021

CREAD: Combined Resolution of Ellipses and Anaphora in Dialogues

NAACL 2021long

Anaphora and ellipses are two common phenomena in dialogues. Without resolving referring expressions and information omission, dialogue systems may fail to generate consistent and coherent responses. Traditionally, anaphora is resolved by coreference resolution and ellipses by query rewrite. In this…

2021

End-To-End Multi-Accent Speech Recognition with Unsupervised Accent Modelling

ICASSP 2021accepted

End-to-end speech recognition has achieved good recognition performance on standard English pronunciation datasets. However, one prominent problem with end-to-end speech recognition systems is that non-native English speakers tend to have complex and varied accents, which reduces the accuracy of Eng…

Cited by 0SourceScholar
2021

Generate & Rank: A Multi-task Framework for Math Word Problems

EMNLP 2021finding

Math word problem (MWP) is a challenging and critical task in natural language processing. Many recent studies formalize MWP as a generation task and have adopted sequence-to-sequence models to transform problem descriptions to mathematical expressions. However, mathematical expressions are prone to…

2021

Noise Robust Named Entity Understanding for Voice Assistants

NAACL 2021industry

Named Entity Recognition (NER) and Entity Linking (EL) play an essential role in voice assistant interaction, but are challenging due to the special difficulties associated with spoken user queries. In this paper, we propose a novel architecture that jointly solves the NER and EL tasks by combining…

Cited by 5SourcePDFScholar
2021

SA-LOAM: Semantic-aided LiDAR SLAM with Loop Closure

ICRA 2021poster

LiDAR-based SLAM system is admittedly more accurate and stable than others, while its loop closure detection is still an open issue. With the development of 3D semantic segmentation for point cloud, semantic information can be obtained conveniently and steadily, essential for high-level intelligence…

Cited by 110SourceScholar
2021

SSC: Semantic Scan Context for Large-Scale Place Recognition

IROS 2021poster

Place recognition gives a SLAM system the ability to correct cumulative errors. Unlike images that contain rich texture features, point clouds are almost pure geometric information which makes place recognition based on point clouds challenging. Existing works usually encode low-level features such…

Cited by 110SourcecodeScholar
2020

Manifold Learning-based Word Representation Refinement Incorporating Global and Local Information

COLING 2020main

Recent studies show that word embedding models often underestimate similarities between similar words and overestimate similarities between distant words. This results in word similarity results obtained from embedding models inconsistent with human judgment. Manifold learning-based methods are wide…

Cited by 2SourcePDFScholar
2020

Spatial Attentional Bilinear 3D Convolutional Network for Video-Based Autism Spectrum Disorder Detection

ICASSP 2020accepted

Video-based Autism Spectrum Disorder (ASD) detection is a challenge to most video classification networks due to the high degree of similarity between categories. Bilinear pooling is a second-order method, which is widely used in fine-grained visual recognition. However, the average summation in bil…

Cited by 0SourceScholar
2019

Training Multi-task Adversarial Network for Extracting Noise-robust Speaker Embedding

ICASSP 2019accepted

Under noisy environments, to achieve the robust performance of speaker recognition is still a challenging task. Motivated by the promising performance of multi-task training in a variety of image processing tasks, we explore the potential of multitask adversarial training for learning a noise-robust…

Cited by 0SourceScholar
2018

Particle Filtering and Inference for Limit Order Books in High Frequency Finance

ICASSP 2018accepted

This paper investigates the on-line analysis of high-frequency financial order book data using Bayesian modelling techniques. Order book data involves evolving queues of orders at different prices, and here we propose that the order book shape is proportional to a gamma or inverse-gamma density func…

Cited by 0SourceScholar
2016

A transfer learning method for PLDA-based speaker verification

ICASSP 2016accepted

Currently, the state-of-the-art speaker verification system is based on i-vector and PLDA. However, PLDA requires tens of thousands of development data from many speakers. This makes it difficult to learn the PLDA parameters for a domain with scarce data. In this paper, we propose an effective trans…

Cited by 0SourceScholar