← Search

shuai zhang

100 accepted papers

2026

A Theoretical Analysis of Mamba’s Training Dynamics: Filtering Relevant Features for Generalization in State Space Models

ICLR 2026poster

The recent empirical success of Mamba and other selective state space models (SSMs) has renewed interest in non-attention architectures for sequence modeling, yet their theoretical foundations remain underexplored. We present a first-step analysis of generalization and learning dynamics for a simpli…

Cited by 3SourceScholar
2026

AStar: Boosting Multimodal Reasoning with Automated Structured Thinking

AAAI 2026technical

Multimodal large language models excel across diverse domains but struggle with complex visual reasoning tasks. To enhance their reasoning capabilities, current approaches typically rely on explicit search or post-training techniques. However, search-based methods suffer from computational inefficie

Cited by 0SourcePDFScholar
2026

CPiRi: Channel Permutation-Invariant Relational Interaction for Multivariate Time Series Forecasting

ICLR 2026poster

Current methods for multivariate time series forecasting can be classified into channel-dependent and channel-independent models. Channel-dependent models learn cross-channel features but often overfit the channel ordering, which hampers adaptation when channels are added or reordered. Channel-indep…

Cited by 0SourcecodeScholar
2026

From Imitation to Discrimination: Toward a Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks

AAAI 2026technical

Reinforcement learning has emerged as a paradigm for post-training large language models, boosting their reasoning capabilities. Such approaches compute an advantage value for each sample, reflecting better or worse performance than expected, thereby yielding both positive and negative signals for t

Cited by 0SourcePDFScholar
2026

Message Tuning Outshines Graph Prompt Tuning: A Prismatic Space Perspective

ICML 2026poster

Graph Foundation Models (GFMs), built upon the *Pre-training and Adaptation* paradigm, have emerged as a research hotspot in graph learning. For GNN-based GFMs, graph prompt tuning has become the prevailing adaptation method for downstream tasks. Although recent methods explain why graph prompt tuni…

Cited by 0SourceScholar
2026

Near-Field Driven Origami-Based Bio-Inspired Jellyfish Robot

ICRA 2026poster

The development of bio-inspired jellyfish robots holds significant benefits for autonomous aquatic systems due to jellyfish’s efficient water jet propulsion. However, the current design of jellyfish robots still faces challenges in balancing high biological fidelity with the demands of lightweight, …

Cited by 0Scholar
2026

OPIC: Enhancing Language Model Merging via Optimizing In-Context Capability

ICML 2026poster

Task-vector–based model merging enables low-cost, training-free multi-task learning for large language models, but suffers from severe performance degradation due to task conflict. Prior mitigation strategies largely rely on validation data for costly hyperparameter tuning, limiting both interpretab…

Cited by 0SourceScholar
2026

SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging

ICML 2026poster

Attention is the critical component of a transformer. Yet the quadratic computational complexity of vanilla full attention in the input size and the inability of its linear attention variant to focus have been challenges for computer vision tasks. We provide a mathematical definition of generalized …

Cited by 0SourceScholar
2026

StereoMamba: Real-Time and Robust Intraoperative Stereo Disparity Estimation Via Long-Range Spatial Dependencies

ICRA 2026poster

Stereo disparity estimation is crucial for obtaining depth information in robot-assisted minimally invasive surgery (RAMIS). While current deep learning methods have made significant advancements, challenges remain in achieving an optimal balance between accuracy, robustness, and inference speed. To…

2026

TINY BUT MIGHTY: A SOFTWARE-HARDWARE CO- DESIGN APPROACH FOR EFFICIENT MULTIMODAL IN- FERENCE ON BATTERY-POWERED SMALL DEVICES

ICLR 2026poster

Large Multimodal Models (LMMs) are inherently modular, consisting of vision and audio encoders, projectors, and large language models. Yet, they are almost always executed monolithically, which underutilizes the heterogeneous accelera- tors (NPUs, GPUs, DSPs) in modern SoCs and leads to high end-to-…

Cited by 0SourceScholar
2026

Theoretical Analysis of Contrastive Learning under Imbalanced Data: From Training Dynamics to a Pruning Solution

ICLR 2026poster

Contrastive learning has emerged as a powerful framework for learning generalizable representations, yet its theoretical understanding remains limited, particularly under imbalanced data distributions that are prevalent in real-world applications. Such an imbalance can degrade representation quality…

Cited by 0SourceScholar
2026

Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices

AAAI 2026technical

There is a growing demand for deploying large generative AI models on mobile devices. For recent popular video generative models, however, the Variational AutoEncoder (VAE) represents one of the major computational bottlenecks. Both large parameter sizes and mismatched kernels cause out-of-memory er

Cited by 0SourcePDFScholar
2025

Adapting to Online Distribution Shifts in Deep Learning: A Black-Box Approach

AISTATS 2025poster

We study the well-motivated problem of online distribution shift in which the data arrive in batches and the distribution of each batch can change arbitrarily over time. Since the shifts can be large or small, abrupt or gradual, the length of the relevant historical data to learn from may vary over…

Cited by 0SourceScholar
2025

AoI-MDP: An AoI Optimized Markov Decision Process Dedicated in the Underwater Task (Student Abstract)

AAAI 2025technical

Ocean exploration places high demands on autonomous underwater vehicles, especially when there's observation delay. We propose age of information optimized Markov decision process (AoI-MDP) to enhance underwater tasks by modeling observation delay as signal delay and including it in the state space.…

2025

Code-switching Mediated Sentence-level Semantic Learning

AAAI 2025technical

Code-switching is a linguistic phenomenon in which different languages are used interactively during conversation. It poses significant performance challenges to natural language processing (NLP) tasks due to the often monolingual nature of the underlying system. We focus on sentence-level semantic…

Cited by 0SourcePDFScholar
2025

Conformal Anomaly Detection in Event Sequences

ICML 2025poster

Anomaly detection in continuous-time event sequences is a crucial task in safety-critical applications. While existing methods primarily focus on developing a superior test statistic, they fail to provide guarantees regarding the false positive rate (FPR), which undermines their reliability in pract…

Cited by 0SourcePDFScholar
2025

Contrastive Learning with Data Misalignment: Feature Purity, Training Dynamics and Theoretical Generalization Guarantees

NeurIPS 2025poster

Contrastive learning is a powerful framework for learning discriminative representations from image-text pairs. Despite its success, its theoretical foundations, especially when the image-text pair exhibits misalignment, remain underexplored. This paper provides the first theoretical analysis of c…

Cited by 0SourceScholar
2025

ERFSL: An Efficient Reward Function Searcher via Large Language Models for Custom-Environment Multi-Objective Reinforcement Learning (Student Abstract)

AAAI 2025technical

We propose ERFSL, an efficient reward function searcher using large language models (LLMs) for custom-environment, multi-objective reinforcement learning (RL). ERFSL generates reward components based on explicit user requirements and rectifies them, and iteratively optimizes the weights of these com…

2025

HFUS-NeRF: Hybrid Representation for Fast Ultrasound Reconstruction in Robotic Ultrasound System

ICRA 2025

Telemedicine is promising in digital healthcare management, such as supporting the coronavirus disease 2019 (COVID-19) pandemic. Three-dimensional (3D) ultrasound reconstruction and new view image synthesis, which can assist in diagnosis and reexamine, have significant potential in teleultrasound, e

Cited by 1SourceScholar
2025

Integrating Drug Substructures and Longitudinal Electronic Health Records for Personalized Drug Recommendation

NeurIPS 2025poster

Drug recommendation systems aim to identify optimal drug combinations for patient care, balancing therapeutic efficacy and safety. Advances in large-scale longitudinal EHRs have enabled learning-based approaches that leverage patient histories such as diagnoses, procedures, and previously prescribed…

Cited by 1SourceScholar
2025

Iterative Substructure Extraction for Molecular Relational Learning with Interactive Graph Information Bottleneck

ICLR 2025poster

Molecular relational learning (MRL) seeks to understand the interaction behaviors between molecules, a pivotal task in domains such as drug discovery and materials science. Recently, extracting core substructures and modeling their interactions have emerged as mainstream approaches within machine le…

Cited by 0SourcePDFScholar
2025

LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling

ICASSP 2025accepted

Singing Voice Conversion (SVC) has emerged as a significant subfield of Voice Conversion (VC), enabling the transformation of one singer’s voice into another while preserving musical elements such as melody, rhythm, and timbre. Traditional SVC methods have limitations in terms of audio quality, data…

Cited by 0SourceScholar
2025

Make Your AUV Adaptive: An Environment-Aware Reinforcement Learning Framework For Underwater Tasks

IROS 2025

This study presents a novel environment-aware reinforcement learning (RL) framework designed to augment the operational capabilities of autonomous underwater vehicles (AUVs) in underwater environments. Departing from traditional RL architectures, the proposed framework integrates an environment-awar

Cited by 3SourceScholar
2025

MalDetectFormer: Leveraging Sparse SpatioTemporal Information for Effective Malicious Traffic Detection

AAAI 2025technical

Malicious traffic detection is one of the main challenges in the field of cybersecurity. Although modern deep learning methods have made progress in identifying malicious traffic, they often overlook the persistent nature of attack behaviors, making it difficult to distinguish between malicious and…

2025

Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models

NeurIPS 2025poster

Since the seminal work of TabPFN, research on tabular foundation models (TFMs) based on in-context learning (ICL) has challenged long-standing paradigms in machine learning. Without seeing any real-world data, models pretrained on purely synthetic datasets generalize remarkably well across diverse d…

Cited by 0SourceScholar
2025

Multi-level Relevance Document Identifier Learning for Generative Retrieval

ACL 2025long

Generative Retrieval (GR) introduces a new information retrieval paradigm that directly generates unique document identifiers (DocIDs). The key challenge of GR lies in creating effective yet discrete DocIDs that preserve semantic relevance for similar documents while differentiating dissimilar ones.…

2025

Never too Prim to Swim: An LLM-Enhanced RL-based Adaptive S-Surface Controller for AUVs under Extreme Sea Conditions

IROS 2025

The adaptivity and maneuvering capabilities of Autonomous Underwater Vehicles (AUVs) have drawn significant attention in oceanic research, due to the unpredictable disturbances and strong coupling among the AUV’s degrees of freedom. In this paper, we developed large language model (LLM)-enhanced rei

Cited by 10SourceScholar
2025

PALMBENCH: A COMPREHENSIVE BENCHMARK OF COMPRESSED LARGE LANGUAGE MODELS ON MOBILE PLATFORMS

ICLR 2025poster

Deploying large language models (LLMs) locally on mobile devices is advantageous in scenarios where transmitting data to remote cloud servers is either undesirable due to privacy concerns or impractical due to network connection. Recent advancements have facilitated the local deployment of LLMs. How…

Cited by 2SourcePDFScholar
2025

Pandora’s Box or Aladdin’s Lamp: A Comprehensive Analysis Revealing the Role of RAG Noise in Large Language Models

ACL 2025long

Retrieval-Augmented Generation (RAG) has emerged as a crucial method for addressing hallucinations in large language models (LLMs). While recent research has extended RAG models to complex noisy scenarios, these explorations often confine themselves to limited noise types and presuppose that noise i…

2025

RadialRouter: Structured Representation for Efficient and Robust Large Language Models Routing

EMNLP 2025

The rapid advancements in large language models (LLMs) have led to the emergence of routing techniques, which aim to efficiently select the optimal LLM from diverse candidates to tackle specific tasks, optimizing performance while reducing costs. Current LLM routing methods are limited in effectiven

Cited by 0SourcePDFScholar
2025

RetrieverGuard: Empowering Information Retrieval to Combat LLM-Generated Misinformation

NAACL 2025findings

Large language models (LLMs) have demonstrated impressive capabilities in generating human-like text and have been shown to store factual knowledge within their extensive parameters. However, models like ChatGPT can still actively or passively generate false or misleading information, increasing the…

Cited by 0SourcePDFScholar
2025

Self-Training with Dynamic Weighting for Robust Gradual Domain Adaptation

NeurIPS 2025poster

In this paper, we propose a new method called \textit{Self-Training with Dynamic Weighting} (STDW), which aims to enhance robustness in Gradual Domain Adaptation (GDA) by addressing the challenge of smooth knowledge migration from the source to the target domain. Traditional GDA methods mitigate dom…

Cited by 0SourceScholar
2025

Sharpness-aware Zeroth-order Optimization for Graph Transformers

IJCAI 2025

Graph Transformers (GTs) have emerged as powerful tools for handling graph-structured data through global attention mechanisms. While GTs can effectively capture long-range dependencies, they introduce difficulties in optimization due to their complex, non-differentiable operators, which cannot be d

2025

Swarm Navigation Based on Smoothed Particle Hydrodynamics in Complex Obstacle Environments

RA-L 2025

In this letter, we propose a method for the navigation of swarm unmanned aerial vehicles (UAVs) in complex environments with obstacles. We propose an algorithmic framework based on Smoothed Particle Hydrodynamics (SPH). In this framework, each UAV is considered a particle, computing its motion infor

Cited by 1SourcecodeScholar
2025

S²MILE: Semantic-and-Structure-Aware Music-Driven Lyric Generation

AAAI 2025technical

The task of music-to-lyric generation aims to create lyrics that can be sung in harmony with the music while capturing the music’s intrinsic meaning. Previous efforts in this area have struggled to effectively handle both the structural and semantic alignments of music and lyrics, often relying on r…

Cited by 0SourcePDFScholar
2025

UACOF: A USV-AUV Collaboration Framework for Underwater Tasks Under Extreme Sea Conditions (Student Abstract)

AAAI 2025technical

Ocean exploration requires effective collaboration between the unmanned surface vehicle (USV) and autonomous underwater vehicles (AUVs). We propose UACOF, a USV-AUV collaboration framework that enhances multi-AUV performance under extreme sea conditions. The framework includes high-precision multi-A…

2025

USV-AUV Collaboration Framework for Underwater Tasks under Extreme Sea Conditions

ICASSP 2025accepted

Autonomous underwater vehicles (AUVs) are valuable for ocean exploration due to their flexibility and ability to carry communication and detection units. Nevertheless, AUVs alone often face challenges in harsh and extreme sea conditions. This study introduces a unmanned surface vehicle (USV)–AUV col…

Cited by 0SourceScholar
2025

Unlearning through Knowledge Overwriting: Reversible Federated Unlearning via Selective Sparse Adapter

CVPR 2025poster

Federated Learning is a promising paradigm for privacy-preserving collaborative model training. In practice, it is essential not only to continuously train the model to acquire new knowledge but also to guarantee old knowledge the right to be forgotten (i.e., federated unlearning), especially for pr…

2025

When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers

ICLR 2025oral

Task arithmetic refers to editing the pre-trained model by adding a weighted sum of task vectors, each of which is the weight update from the pre-trained model to fine-tuned models for certain tasks. This approach recently gained attention as a computationally efficient inference method for model ed…

Cited by 0SourcePDFScholar
2024

A Global-Local Graph Attention Network for Deformable Linear Objects Dynamic Interaction With Environment

RA-L 2024

Accurately modeling the interactions between deformable linear objects (DLOs) and their environments is crucial for active deformation control by robot manipulators. Graph Neural Networks (GNNs) have shown immense potential in the particle-based dynamics of DLOs. However, most existing studies propa

Cited by 0SourceScholar
2024

Bilateral Masking with prompt for Knowledge Graph Completion

NAACL 2024findings

The pre-trained language model (PLM) has achieved significant success in the field of knowledge graph completion (KGC) by effectively modeling entity and relation descriptions. In recent studies, the research in this field has been categorized into methods based on word matching and sentence matchin…

Cited by 1SourcePDFScholar
2024

Bridging Remote Sensors with Multisensor Geospatial Foundation Models

CVPR 2024poster

In the realm of geospatial analysis the diversity of remote sensors encompassing both optical and microwave technologies offers a wealth of distinct observational capabilities. Recognizing this we present msGFM a multisensor geospatial foundation model that effectively unifies data from four key sen…

2024

CoMM: Collaborative Multi-Agent, Multi-Reasoning-Path Prompting for Complex Problem Solving

NAACL 2024findings

Large Language Models (LLMs) have shown great ability in solving traditional natural language tasks and elementary reasoning tasks with appropriate prompting techniques. However, their ability is still limited in solving complicated science problems. In this work, we aim to push the upper bound of t…

2024

Combating Data Imbalances in Federated Semi-supervised Learning with Dual Regulators

AAAI 2024technical

Federated learning has become a popular method to learn from decentralized heterogeneous data. Federated semi-supervised learning (FSSL) emerges to train models from a small fraction of labeled data due to label scarcity on decentralized clients. Existing FSSL methods assume independent and identica…

Cited by 8SourcePDFScholar
2024

Discovering Bias in Latent Space: An Unsupervised Debiasing Approach

ICML 2024poster

The question-answering (QA) capabilities of foundation models are highly sensitive to prompt variations, rendering their performance susceptible to superficial, non-meaning-altering changes. This vulnerability often stems from the model's preference or bias towards specific input characteristics, su…

Cited by 8SourcePDFScholar
2024

FedSC: Provable Federated Self-supervised Learning with Spectral Contrastive Objective over Non-i.i.d. Data

ICML 2024poster

Recent efforts have been made to integrate self-supervised learning (SSL) with the framework of federated learning (FL). One unique challenge of federated self-supervised learning (FedSSL) is that the global objective of FedSSL usually does not equal the weighted sum of local SSL objectives. Consequ…

Cited by 4SourcePDFScholar
2024

How to Trade Off the Quantity and Capacity of Teacher Ensemble: Learning Categorical Distribution to Stochastically Employ a Teacher for Distillation

AAAI 2024technical

We observe two phenomenons with respect to quantity and capacity: 1) more teacher is not always better for multi-teacher knowledge distillation, and 2) stronger teacher is not always better for single-teacher knowledge distillation. To trade off the quantity and capacity of teacher ensemble, in this…

Cited by 1SourcePDFScholar
2024

MMGNN: A Molecular Merged Graph Neural Network for Explainable Solvation Free Energy Prediction

IJCAI 2024poster

In this paper, we address the challenge of accurately modeling and predicting Gibbs free energy in solute-solvent interactions, a pivotal yet complex aspect in the field of chemical modeling. Traditional approaches, primarily relying on deep learning models, face limitations in capturing the intrica…

Cited by 5SourcePDFScholar
2024

MobileInst: Video Instance Segmentation on the Mobile

AAAI 2024technical

Video instance segmentation on mobile devices is an important yet very challenging edge AI problem. It mainly suffers from (1) heavy computation and memory costs for frame-by-frame pixel-level instance perception and (2) complicated heuristics for tracking objects. To address these issues, we presen…

Cited by 8SourcePDFScholar
2024

MolTC: Towards Molecular Relational Modeling In Language Models

ACL 2024findings

Molecular Relational Learning (MRL), aiming to understand interactions between molecular pairs, plays a pivotal role in advancing biochemical research. Recently, the adoption of large language models (LLMs), known for their vast knowledge repositories and advanced logical inference capabilities, has…

2024

Neural Jump-Diffusion Temporal Point Processes

ICML 2024spotlight

We present a novel perspective on temporal point processes (TPPs) by reformulating their intensity processes as solutions to stochastic differential equations (SDEs). In particular, we first prove the equivalent SDE formulations of several classical TPPs, including Poisson processes, Hawkes processe…

Cited by 5SourcePDFScholar
2024

SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning

ICML 2024poster

This paper studies the transfer reinforcement learning (RL) problem where multiple RL problems have different reward functions but share the same underlying transition dynamics. In this setting, the Q-function of each RL problem (task) can be decomposed into a successor feature (SF) and a reward map…

Cited by 2SourcePDFScholar
2024

SMILE: Single-turn to Multi-turn Inclusive Language Expansion via ChatGPT for Mental Health Support

EMNLP 2024finding

Developing specialized dialogue systems for mental health support requires multi-turn conversation data, which has recently garnered increasing attention. However, gathering and releasing large-scale, real-life multi-turn conversations that could facilitate advancements in mental health support pres…

2024

Transferring Knowledge From Large Foundation Models to Small Downstream Models

ICML 2024poster

How do we transfer the relevant knowledge from ever larger foundation models into small, task-specific downstream models that can run at much lower costs? Standard transfer learning using pre-trained weights as the initialization transfers limited information and commits us to often massive pre-trai…

Cited by 2SourcePDFScholar
2024

Understanding the Therapeutic Relationship between Counselors and Clients in Online Text-based Counseling using LLMs

EMNLP 2024finding

Robust therapeutic relationships between counselors and clients are fundamental to counseling effectiveness. The assessment of therapeutic alliance is well-established in traditional face-to-face therapy but may not directly translate to text-based settings. With millions of individuals seeking supp…

Cited by 4SourcePDFScholar
2024

Unraveling the Gradient Descent Dynamics of Transformers

NeurIPS 2024poster

While the Transformer architecture has achieved remarkable success across various domains, a thorough theoretical foundation explaining its optimization dynamics is yet to be fully developed. In this study, we aim to bridge this understanding gap by answering the following two core questions: (1) Wh…

Cited by 1SourcePDFScholar
2023

3D Reconstruction of Tibia and Fibula using One General Model and Two X-ray Images

ICRA 2023poster

The 3D reconstruction of patient specific bone models plays a crucial role in orthopaedic surgery for clinical evaluation, surgical planning and precise implant design or selection. This paper considers the problem of reconstructing a patient-specific 3D tibia and fibula model from only two 2D X-ray…

Cited by 2SourceScholar
2023

Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning

NeurIPS 2023poster

In this paper, we propose a Disentangled Counterfactual Learning (DCL) approach for physical audiovisual commonsense reasoning. The task aims to infer objects’ physics commonsense based on both video and audio input, with the main challenge is how to imitate the reasoning ability of humans. Most of…

2023

Joint Edge-Model Sparse Learning is Provably Efficient for Graph Neural Networks

ICLR 2023poster

Due to the significant computational challenge of training large-scale graph neural networks (GNNs), various sparse learning techniques have been exploited to reduce memory and storage costs. Examples include graph sparsification that samples a subgraph to reduce the amount of data aggregation and m…

Cited by 20SourcePDFScholar
2023

MAS: Towards Resource-Efficient Federated Multiple-Task Learning

ICCV 2023poster

Federated learning (FL) is an emerging distributed machine learning method that empowers in-situ model training on decentralized edge devices. However, multiple simultaneous FL tasks could overload resource-constrained devices. In this work, we propose the first FL system to effectively coordinate a…

Cited by 24PDFcodeScholar
2023

Offline Imitation Learning with Variational Counterfactual Reasoning

NeurIPS 2023poster

In offline imitation learning (IL), an agent aims to learn an optimal expert behavior policy without additional online environment interactions. However, in many real-world scenarios, such as robotics manipulation, the offline dataset is collected from suboptimal behaviors without rewards. Due to th…

2023

On the Convergence and Sample Complexity Analysis of Deep Q-Networks with $\epsilon$-Greedy Exploration

NeurIPS 2023poster

This paper provides a theoretical understanding of deep Q-Network (DQN) with the $\varepsilon$-greedy exploration in deep reinforcement learning. Despite the tremendous empirical achievement of the DQN, its theoretical characterization remains underexplored. First, the exploration strategy is either…

Cited by 27SourcePDFScholar
2023

Patch-level Routing in Mixture-of-Experts is Provably Sample-efficient for Convolutional Neural Networks

ICML 2023oral

In deep learning, mixture-of-experts (MoE) activates one or few experts (sub-networks) on a per-sample or per-token basis, resulting in significant computation reduction. The recently proposed patch-level routing in MoE (pMoE) divides each input into $n$ patches (or tokens) and sends $l$ patches ($l…

2023

Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual Recognition

NeurIPS 2023poster

This work proposes POMP, a prompt pre-training method for vision-language models. Being memory and computation efficient, POMP enables the learned prompt to condense semantic information for a rich set of visual concepts with over twenty-thousand classes. Once pre-trained, the prompt with a strong t…

2023

Rethinking Document-Level Relation Extraction: A Reality Check

ACL 2023findings

Recently, numerous efforts have continued to push up performance boundaries of document-level relation extraction (DocRE) and have claimed significant progress in DocRE. In this paper, we do not aim at proposing a novel model for DocRE. Instead, we take a closer look at the field to see if these per…

2023

SKDBERT: Compressing BERT via Stochastic Knowledge Distillation

AAAI 2023technical

In this paper, we propose Stochastic Knowledge Distillation (SKD) to obtain compact BERT-style language model dubbed SKDBERT. In each distillation iteration, SKD samples a teacher model from a pre-defined teacher team, which consists of multiple teacher models with multi-level capacities, to transfe…

Cited by 16SourcePDFScholar
2023

Understanding Client Reactions in Online Mental Health Counseling

ACL 2023long

Communication success relies heavily on reading participants’ reactions. Such feedback is especially important for mental health counselors, who must carefully consider the client’s progress and adjust their approach accordingly. However, previous NLP research on counseling has mainly focused on stu…

2022

ADD 2022: the first Audio Deep Synthesis Detection Challenge

ICASSP 2022accepted

Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three t…

Cited by 0SourceScholar
2022

AutoST: Towards the Universal Modeling of Spatio-temporal Sequences

NeurIPS 2022accept

The analysis of spatio-temporal sequences plays an important role in many real-world applications, demanding a high model capacity to capture the interdependence among spatial and temporal dimensions. Previous studies provided separated network design in three categories: spatial first, temporal fir…

Cited by 10SourcePDFScholar
2022

ClusterFormer: Neural Clustering Attention for Efficient and Effective Transformer

ACL 2022long

Recently, a lot of research has been carried out to improve the efficiency of Transformer. Among them, the sparse pattern-based method is an important branch of efficient Transformers. However, some existing sparse methods usually use fixed patterns to select words, without considering similarities…

2022

How unlabeled data improve generalization in self-training? A one-hidden-layer theoretical analysis

ICLR 2022poster

Self-training, a semi-supervised learning algorithm, leverages a large amount of unlabeled data to improve learning when the labeled data are limited. Despite empirical successes, its theoretical characterization remains elusive. To the best of our knowledge, this work establishes the first theoreti…

Cited by 32SourcePDFScholar
2022

Jump Self-attention: Capturing High-order Statistics in Transformers

NeurIPS 2022accept

The recent success of Transformer has benefited many real-world applications, with its capability of building long dependency through pairwise dot-products. However, the strong assumption that elements are directly attentive to each other limits the performance of tasks with high-order dependencies…

Cited by 3SourcePDFScholar
2022

Learning Music Sequence Representation From Text Supervision

ICASSP 2022accepted

Music representation learning is notoriously difficult for its complex human-related concepts contained in the sequence of numerical signals. To excavate better MUsic SEquence Representation from labeled audio, we propose a novel text-supervision pre-training method, namely MUSER. MUSER adopts an au…

Cited by 0SourceScholar
2022

Neural Methods for Logical Reasoning over Knowledge Graphs

ICLR 2022poster

Reasoning is a fundamental problem for computers and deeply studied in Artificial Intelligence. In this paper, we specifically focus on answering multi-hop logical queries on Knowledge Graphs (KGs). This is a complicated task because, in real world scenarios, the graphs tend to be large and incomple…

2022

Syntax-guided Contrastive Learning for Pre-trained Language Model

ACL 2022findings

Syntactic information has been proved to be useful for transformer-based pre-trained language models. Previous studies often rely on additional syntax-guided attention components to enhance the transformer, which require more parameters and additional syntactic parsing in downstream tasks. This incr…

2022

Wavelet-Based Unsupervised Label-to-Image Translation

ICASSP 2022accepted

Semantic Image Synthesis (SIS) is a subclass of image-to-image translation where a semantic layout is used to generate a photorealistic image. State-of-the-art conditional Generative Adversarial Networks (GANs) need a huge amount of paired data to accomplish this task while generic un-paired image-t…

Cited by 0SourceScholar
2021

3D Reconstruction of Deformable Colon Structures based on Preoperative Model and Deep Neural Network

ICRA 2021poster

In colonoscopy procedures, it is important to rebuild and visualize the colonic surface to minimize the missing regions and reinspect for abnormalities. Due to the fast camera motion and deformation of the colon in standard forward-viewing colonoscopies, traditional simultaneous localization and map…

Cited by 10SourceScholar
2021

A Sequence-to-Set Network for Nested Named Entity Recognition

IJCAI 2021poster

Named entity recognition (NER) is a widely studied task in natural language processing. Recently, a growing number of studies have focused on the nested NER. The span-based methods, considering the entity recognition as a span classification task, can deal with nested entities naturally. But they su…

2021

Beyond Fully-Connected Layers with Quaternions: Parameterization of Hypercomplex Multiplications with $1/n$ Parameters

ICLR 2021spotlight

Recent works have demonstrated reasonable success of representation learning in hypercomplex space. Specifically, “fully-connected layers with quaternions” (quaternions are 4D hypercomplex numbers), which replace real-valued matrix multiplications in fully-connected layers with Hamilton products of…

2021

Collaborative Unsupervised Visual Representation Learning From Decentralized Data

ICCV 2021poster

Unsupervised representation learning has achieved outstanding performances using centralized data available on the Internet. However, the increasing awareness of privacy protection limits sharing of decentralized unlabeled image data that grows explosively in multiple parties (e.g. mobile phones and…

Cited by 127PDFcodeScholar
2021

Decoupling Pronunciation and Language for End-to-End Code-Switching Automatic Speech Recognition

ICASSP 2021accepted

Despite the recent significant advances witnessed in end-to-end (E2E) ASR system for code-switching, hunger for audio-text paired data limits the further improvement of the models’ performance. In this paper, we propose a decoupled transformer model to use mono-lingual paired data and unpaired text…

Cited by 0SourceScholar
2021

Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting

AAAI 2021technical

Many real-world applications require the prediction of long sequence time-series, such as electricity consumption planning. Long sequence time-series forecasting (LSTF) demands a high prediction capacity of the model, which is the ability to capture precise long-range dependency coupling between out…

2021

Knowledge Router: Learning Disentangled Representations for Knowledge Graphs

NAACL 2021long

The design of expressive representations of entities and relations in a knowledge graph is an important endeavor. While many of the existing approaches have primarily focused on learning from relational patterns and structural information, the intrinsic complexity of KG entities has been more or les…

Cited by 8SourcePDFScholar
2021

Locate and Label: A Two-stage Identifier for Nested Named Entity Recognition

ACL 2021long

Named entity recognition (NER) is a well-studied task in natural language processing. Traditional NER research only deals with flat entities and ignores nested entities. The span-based methods treat entity recognition as a span classification task. Although these methods have the innate ability to h…

2021

On Orthogonality Constraints for Transformers

ACL 2021short

Orthogonality constraints encourage matrices to be orthogonal for numerical stability. These plug-and-play constraints, which can be conveniently incorporated into model training, have been studied for popular architectures in natural language processing, such as convolutional neural networks and re…

Cited by 24SourcePDFScholar
2021

Self-Instantiated Recurrent Units with Dynamic Soft Recursion

NeurIPS 2021poster

While standard recurrent neural networks explicitly impose a chain structure on different forms of data, they do not have an explicit bias towards recursive self-instantiation where the extent of recursion is dynamic. Given diverse and even growing data modalities (e.g., logic, algorithmic input an…

Cited by 5SourcePDFScholar
2021

Why Lottery Ticket Wins? A Theoretical Perspective of Sample Complexity on Sparse Neural Networks

NeurIPS 2021poster

The lottery ticket hypothesis (LTH) states that learning on a properly pruned network (the winning ticket) has improved test accuracy over the original unpruned network. Although LTH has been justified empirically in a broad range of deep neural network (DNN) involved applications like computer visi…

Cited by 38SourcePDFScholar
2020

Fast Learning of Graph Neural Networks with Guaranteed Generalizability: One-hidden-layer Case

ICML 2020poster

Although graph neural networks (GNNs) have made great progress recently on learning from graph-structured data in practice, their theoretical guarantee on generalizability remains elusive in the literature. In this paper, we provide a theoretically-grounded generalizability analysis of GNNs with one…

Cited by 39SourcePDFScholar
2020

Synchronous Transformers for end-to-end Speech Recognition

ICASSP 2020accepted

For most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchronous problem between the encoding and decoding makes these models difficult to be applied for online speech recognition…

Cited by 0SourceScholar
2020

TRP: Trained Rank Pruning for Efficient Deep Neural Networks

IJCAI 2020poster

To enable DNNs on edge devices like mobile phones, low-rank approximation has been widely adopted because of its solid theoretical rationale and efficient implementations. Several previous works attempted to directly approximate a pre-trained model by low-rank decomposition; however, small approxima…

Cited by 0SourcePDFScholar
2019

A Tensorized Transformer for Language Modeling

NeurIPS 2019poster

Latest development of neural models has connected the encoder and decoder through a self-attention mechanism. In particular, Transformer, which is solely based on self-attention, has led to breakthroughs in Natural Language Processing (NLP) tasks. However, the multi-head attention mechanism, as a ke…

2019

Understanding Straight-Through Estimator in Training Activation Quantized Neural Nets

ICLR 2019poster

Training activation quantized neural networks involves minimizing a piecewise constant training loss whose gradient vanishes almost everywhere, which is undesirable for the standard back-propagation or chain rule. An empirical way around this issue is to use a straight-through estimator (STE) (Bengi…

Cited by 382SourcePDFScholar