← Search

Shuai Wang

194 accepted papers

2026

AHAMask: Reliable Task Specification for Large Audio Language Models Without Instructions

AAAI 2026technical

Although current large audio language models (LALMs) extend text large language models (LLMs) with generic acoustic understanding abilities, they usually suffer from prompt sensitivity, where different instructions of the same intention can yield drastically different outcomes. In this work, we pro

Cited by 0SourcePDFScholar
2026

AdaS: Adaptive Gradient Descent for Spiking Transformers

ICML 2026poster

Transformer-based Spiking Neural Networks (SNNs) combine Transformer performance with SNN energy efficiency through an event-driven self-attention mechanism. However, Spiking Transformers still lag behind their Artificial Neural Network (ANN) counterparts. Most existing studies address this issue th…

Cited by 0SourceScholar
2026

Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks

ICRA 2026poster

A major bottleneck in off-road autonomous driving research lies in the scarcity of large-scale, high-quality datasets and benchmarks. To bridge this gap, we present ORAD-3D, which, to the best of our knowledge, is the largest dataset specifically curated for off-road autonomous driving. ORAD-3D cove…

2026

Behavior Tokens Speak Louder: Disentangled Explainable Recommendation with Behavior Vocabulary

AAAI 2026technical

Recent advances in explainable recommendation have explored the integration of language models to analyze natural language rationales for user–item interactions. Despite their potential, existing methods often rely on ID-based representations that obscure semantic meaning and impose structural const

Cited by 0SourcePDFScholar
2026

COMET: A Dual Swashplate Autonomous Coaxial Bi-Copter AAV With High-Maneuverability and Long-Endurance

RA-L 2026

Coaxial bi-copter autonomous aerial vehicles (AAVs) have garnered attention due to their potential for improved rotor system efficiency and compact form factor. However, balancing efficiency, maneuverability, and compactness in coaxial bi-copter systems remains a key design challenge, limiting their

Cited by 0SourceScholar
2026

COMET: A Dual Swashplate Autonomous Coaxial Bi-Copter AAV with High-Maneuverability and Long-Endurance

ICRA 2026poster

Coaxial bi-copter unmanned aerial vehicles (UAVs) have garnered attention due to their potential for improved rotor system efficiency and compact form factor. However, balancing efficiency, maneuverability, and compactness in coaxial bi-copter systems remains a key design challenge, limiting their p…

2026

CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data

ICASSP 2026poster

Accent normalization (AN) systems often struggle with unnatural outputs and undesired content distortion, stemming from both suboptimal training data and rigid duration modeling. In this paper, we propose a "source-synthesis" methodology for training data construction. By generating source L2 speech…

Cited by 0SourcePDFScholar
2026

Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens

CVPR 2026

Visual generation with discrete tokens has gained significant attention as it enables a unified token prediction paradigm shared with language models, promising seamless multimodal architectures. However, current discrete generation methods remain limited to low-dimensional latent tokens (typically

Cited by 1SourcecodeScholar
2026

DPNet: Doppler LiDAR Motion Planning for Highly-Dynamic Environments

RA-L 2026

Existing motion planning methods often struggle with rapid-motion obstacles due to an insufficient understanding of environmental changes. To address this, we propose integrating motion planners with Doppler LiDARs, which provide not only ranging measurements but also instantaneous point velocities.

Cited by 0SourcecodeScholar
2026

DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation

CVPR 2026

Pixel diffusion aims to generate images directly in pixel space in an end-to-end fashion. This approach avoids the limitations of VAE in the two-stage latent diffusion, offering higher model capacity. Existing pixel diffusion models suffer from slow training and inference, as they usually model both

Cited by 0SourcecodeScholar
2026

DigimonGPT: An Evolvable Agent with Hierarchical Human-like Memory for Video Question Answering

AAAI 2026technical

Video question answering (VideoQA), whose goal is to produce answers through the integration of linguistic and visual understanding, has emerged as a significant research focus. Although Large Multimodal Models (LMMs) and autonomous agent methods have achieved notable advances in VideoQA, excessive

Cited by 0SourcePDFScholar
2026

Direct Preference Optimization for Speech Autoregressive Diffusion Models

ICASSP 2026poster

Autoregressive diffusion models (ARDMs) have recently been applied to speech generation, achieving state-of-the-art (SOTA) performance in zero-shot text-to-speech. By autoregressively generating continuous speech tokens with next-token diffusion, these models offer a promising alternative to next-to…

Cited by 0SourcePDFScholar
2026

EAMET: ROBUST MASSIVE MODEL EDITING VIA EMBEDDING ALIGNMENT OPTIMIZATION

ICLR 2026poster

Model editing techniques are essential for efficiently updating knowledge in large language models (LLMs). However, the effectiveness of existing approaches degrades in massive editing scenarios, particularly when evaluated with practical metrics. Their robustness is also limited in context-rich set…

Cited by 0SourcecodeScholar
2026

FALCO: Foundation Model Guided Active Learning for Cost-Effective Off-Road Freespace Detection

ICRA 2026poster

Freespace detection in unstructured off-road environments is critical for safe autonomous navigation but remains highly challenging due to ambiguous boundaries, diverse terrains, and long-tail safety-critical cases. Constructing large annotated datasets in such environments is prohibitively costly, …

Cited by 0Scholar
2026

Flowing Backwards: Improving Normalizing Flows via Reverse Representation Alignment

AAAI 2026technical

Normalizing Flows (NFs) are a class of generative models distinguished by a mathematically invertible architecture, where the forward pass transforms data into a latent space for density estimation, and the reverse pass generates new samples from this space. This characteristic creates an intrinsic

Cited by 0SourcePDFScholar
2026

GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods

ICLR 2026poster

Despite the growing interest in jailbreaks as an effective red-teaming tool for building safe and responsible large language models (LLMs), flawed evaluation system designs have led to significant discrepancies in their effectiveness assessments. With a systematic measurement study based on 37 jailb…

Cited by 0SourcecodeScholar
2026

HATS: Hardness-Aware Trajectory Synthesis for GUI Agents

CVPR 2026

Graphical user interface (GUI) agents powered by large vision-language models (VLMs) have shown remarkable potential in automating digital tasks, highlighting the need for high-quality trajectory data to support effective agent training. Yet existing trajectory synthesis pipelines often yield agents

Cited by 0SourcecodeScholar
2026

HPTune: Hierarchical Proactive Tuning for Collision-Free Model Predictive Control

ICASSP 2026poster

Parameter tuning is a powerful approach to enhance adaptability in model predictive control (MPC) motion planners. However, existing methods typically operate in a myopic fashion that only evaluates executed actions, leading to inefficient parameter updates due to the sparsity of failure events (e.g…

Cited by 0SourcePDFScholar
2026

HYBRID PRUNING: IN-SITU COMPRESSION OF SELF-SUPERVISED SPEECH MODELS FOR SPEAKER VERIFICATION AND ANTI-SPOOFING

ICASSP 2026oral

Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on resource-constrained devices. While structured pruning is a key technique for model compression, existing methods typica…

Cited by 0SourcePDFScholar
2026

HiconAgent: History Context-aware Policy Optimization for GUI Agents

CVPR 2026

Graphical User Interface (GUI) agents require effective utilization of historical context to perform sequential navigation tasks. While incorporating past actions and observations can significantly improve decision-making, naively using full history leads to excessive computational overhead and pote

Cited by 0SourcecodeScholar
2026

LLM-Driven Scenario-Aware Planning for Autonomous Driving

ICASSP 2026poster

Hybrid planner switching framework (HPSF) for autonomous driving needs to reconcile high-speed driving efficiency with safe maneuvering in dense traffic. Existing HPSF methods often fail to make reliable mode transitions or sustain efficient driving in congested environments, owing to heuristic scen…

Cited by 0SourcePDFScholar
2026

LMGL-WD: LLM-Guided Multi-Task Graph Learning for Category-Level Warehouse Demand Prediction in E-Commerce

AAAI 2026technical

In warehouse-based e-commerce, accurate category-level warehouse demand prediction is essential to ensure effective inventory management. Existing works mainly explore advanced time series models to capture the temporal dynamics, failing to mine cross-category and cross-warehouse correlations effect

Cited by 0SourcePDFScholar
2026

MOSA: Mixtures of Simple Adapters Outperform Monolithic Approaches in LLM-based Multilingual ASR

ICASSP 2026poster

LLM-based ASR overcomes multilingual data scarcity by projecting speech representations into the LLM space to leverage its robust semantic and reasoning capabilities. However, while previous approaches typically enhance performance by scaling data or model parameters, a single projector often strugg…

Cited by 0SourcePDFScholar
2026

NeuPAN: Direct Point Robot Navigation with End-to-End Model-Based Learning (Abstract Reprint)

AAAI 2026technical

Navigating a nonholonomic robot in a cluttered, unknown environment requires accurate perception and precise motion control for real-time collision avoidance. This article presents neural proximal alternating-minimization network (NeuPAN): a real-time, highly accurate, map-free, easy-to-deploy, and

Cited by 0SourcePDFScholar
2026

Neural Dynamics Self-Attention for Spiking Transformers

ICLR 2026poster

Integrating Spiking Neural Networks (SNNs) with Transformer architectures offers a promising pathway to balance energy efficiency and performance, particularly for edge vision applications. However, existing Spiking Transformers face two critical challenges: i) a substantial performance gap relative…

Cited by 0SourceScholar
2026

ORTCL: Towards Continual Learning of Time Series Foundation Models on Streaming Data via Orthogonal Rotation

AAAI 2026technical

Time Series Foundation Models (TSFMs) have emerged as a promising approach in time series analysis. Due to the large-scale parameters of TSFMs and pretraining cost, how to adapt TDFMs in streaming data is always the key factor constraining their application effectiveness. Because streaming data ofte

Cited by 0SourcePDFScholar
2026

Planar-Sector LOS Guidance for Interception of Agile Targets with Lifting-Wing Quadcopters

ICRA 2026poster

This paper proposes a Planar-Sector Line-of-Sight (PS-LOS) guidance law and an accompanying control stack for lifting-wing quadcopters, enabling robust image-based interception of agile targets. The PS-LOS relaxes conventional conical constraints, preserving maneuverability while reducing aerodynami…

2026

Position: Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

ICML 2026poster

Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the com- munity interprets model capabilities. In the past few years, awareness of benchmark quality has grown. Yet, after a decade-scale (2014 - 2025) survey over 572 …

Cited by 0SourceScholar
2026

Positional Encoding for Spiking Transformers

ICML 2026poster

Spiking Neural Networks (SNNs) demonstrate superior energy efficiency over conventional Artificial Neural Networks (ANNs). Recent advances in Transformer-based SNNs have shown encouraging performance by seamlessly integrating spike-driven computation with Transformer architectures. Positional inform…

Cited by 0SourceScholar
2026

ReDi-FM: Frozen Foundation Model for Continual Test-Time Adaptation in Medical Image Segmentation

IJCAI 2026

Continual test-time adaptation (CTTA) adapts a pre-trained medical segmentation model online to an unlabeled target stream whose distribution changes over time. However, most existing CTTA methods rely on pseudo-labeling and self-supervised objectives, which inevitably yield noisy supervision under

Cited by 0Scholar
2026

RepetitionCurse: Measuring and Understanding Router Imbalance in Mixture-of-Experts LLMs under DoS Stress

ICML 2026poster

Mixture-of-Experts architectures have become the standard for efficient LLM scaling, typically employing expert parallelism to distribute experts across devices. However, the absence of explicit load balancing constraints during inference allows adversarial inputs to trigger severe routing concentra…

Cited by 0SourceScholar
2026

Robust Spiking Neural Networks Against Adversarial Attacks

ICLR 2026poster

Spiking Neural Networks (SNNs) represent a promising paradigm for energy-efficient neuromorphic computing due to their bio-plausible and spike-driven characteristics. However, the robustness of SNNs in complex adversarial environments remains significantly constrained. In this study, we theoretical…

Cited by 0SourceScholar
2026

SDTrack: A Baseline for Event-based Tracking via Spiking Neural Networks

CVPR 2026

Event cameras provide superior temporal resolution, dynamic range, energy efficiency, and pixel bandwidth. Spiking Neural Networks (SNNs) naturally complement event data through discrete spike signals, making them ideal for event-based tracking. However, current approaches combining Artificial Neura

Cited by 0SourcecodeScholar
2026

SGERA: Stein-Guided ECG-Report Alignment for ECG Representation Learning

ICML 2026poster

Electrocardiogram (ECG) representation learning via ECG-report alignment is often hindered by the inherent structural and statistical divergence between signals and natural language. Existing methods struggle to bridge this gap with simple contrastive objectives, but struggle with distribution depen…

Cited by 0SourceScholar
2026

SUMMARY ON THE MULTILINGUAL CONVERSATIONAL SPEECH LANGUAGE MODEL CHALLENGE: DATASETS, TASKS, BASELINES, AND METHODS

ICASSP 2026poster

This paper summarizes the Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) challenge, which aims to advance the exploration of building effective multilingual conversational speech LLMs (SLLMs). We provide a detailed description of the task settings for the MLC-SLM challen…

Cited by 0SourcePDFScholar
2026

Scaling Law Analysis in Federated Learning: How to Select the Optimal Model Size?

AAAI 2026technical

The recent success of large language models (LLMs) has sparked a growing interest in training large-scale models. As the model size continues to scale, concerns are growing about the depletion of high-quality, well-curated training data. This has led practitioners to explore training approaches like

Cited by 0SourcePDFScholar
2026

SmoothSpike: Spiking Transformer with Learnable Hadamard Transformation

ICML 2026spotlight

Spiking Neural Networks (SNNs) that leverage sparse binary spikes and temporal dynamics have emerged as energy-efficient alternatives to Artificial Neural Networks (ANNs). However, SNNs suffer from limited representational capacity due to the discrete nature of spikes. Existing solutions extending s…

Cited by 0SourceScholar
2026

Spectral-Progressive Thought Flow for Lightweight Multimodal Reasoning

ICML 2026poster

Multimodal reasoning often relies on long chains of intermediate textual and visual thoughts, where accumulating visual tokens and dense cross-modal attention incur substantial computation and memory overhead. To address this challenge, we propose Spectral-Progressive Thought Flow (*SpecFlow*), a *n…

Cited by 0SourceScholar
2026

SpikingLM: Towards Fully Spiking Language Model

ICML 2026poster

Spiking Neural Networks (SNNs) offer a promising avenue toward energy-efficient language modeling by replacing multiply-accumulate operations with sparse, event-driven computation. However, constructing fully spiking language models reveals two fundamental challenges: (1) gradient degradation from d…

Cited by 0SourceScholar
2026

THE ICASSP 2026 AUTOMATIC SONG AESTHETICS EVALUATION CHALLENGE

ICASSP 2026poster

This paper summarizes the ICASSP 2026 Automatic Song Aesthetics Evaluation (ASAE) Challenge, which focuses on predicting the subjective aesthetic scores of AI-generated songs. The challenge consists of two tracks: Track 1 targets the prediction of the overall musicality score, while Track 2 focuses…

Cited by 0SourcePDFScholar
2026

THE ICASSP 2026 HUMDIAL CHALLENGE: BENCHMARKING HUMAN-LIKE SPOKEN DIALOGUE SYSTEMS IN THE LLM ERA

ICASSP 2026poster

Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowing the gap between human-machine and human-human interactions. Achieving truly ``human-like'' communication necessitates…

Cited by 0SourcePDFScholar
2026

Temporal Interaction in Spiking Transformers with Multi-Delay Mixer

CVPR 2026

Spiking Neural Networks (SNNs) have gained significant attention due to their event-driven computational paradigm, making them promising for neuromorphic computing. In recent years, the integration of SNNs and Transformer architectures has made remarkable progress in various tasks. However, existing

Cited by 0SourceScholar
2026

Towards Training-Free and Accurate ANN-to-SNN Conversion via Activation-Aware Redistribution

AAAI 2026technical

Conversion represents an effective approach for obtaining low-power models by transforming Artificial Neural Networks (ANNs) into event-driven Spiking Neural Networks (SNNs) without additional training. However, existing training-free conversion methods often incur substantial conversion errors. Her

Cited by 0SourcePDFScholar
2026

Training-Free ANN-to-SNN Conversion for High-Performance Spiking Transformers

AAAI 2026technical

Leveraging the event-driven paradigm, Spiking Neural Networks (SNNs) offer a promising approach for constructing energy-efficient Transformer architectures. Compared to directly trained Spiking Transformers, ANN-to-SNN conversion methods bypass the high training costs. However, existing methods stil

Cited by 0SourcePDFScholar
2026

USE: A Unified Model for Universal Sound Separation and Extraction

AAAI 2026technical

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require precisely specified clues to achieve optimal performance. This

Cited by 0SourcePDFScholar
2026

Understanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID Data

ICLR 2026poster

Recent research has introduced distributed self-supervised learning (D-SSL) approaches to leverage vast amounts of unlabeled decentralized data. However, D-SSL faces the critical challenge of data heterogeneity, and there is limited theoretical understanding of how different D-SSL frameworks respond…

Cited by 0SourceScholar
2026

WENETSPEECH-CHUAN: A LARGE-SCALE SICHUANESE CORPUS WITH RICH ANNOTATION FOR DIALECTAL SPEECH PROCESSING

ICASSP 2026poster

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce WenetSpeech-Chuan, a 10,000-hour, richly annotated corpus constru…

Cited by 0SourcePDFScholar
2026

WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning

ICLR 2026poster

To significantly advance the capabilities of open-source web agents, we present WebSailor-V2, a complete post-training pipeline encompassing data construction, Supervised Fine-Tuning (SFT), and Reinforcement Learning (RL). Our methodology features two key innovations: (1) On the data front, we devel…

Cited by 0SourceScholar
2026

Whole-Body Impedance Coordinative Control for a Wheel-Legged Robot on Uncertain Terrain

RA-L 2026

This article proposes a whole-body impedance coordinative control framework for a wheel-legged humanoid robot to achieve adaptability on complex terrains while maintaining the robot's upper body stability. The framework contains a bi-level control strategy. The outer level is a variable-damping impe

Cited by 0SourceScholar
2025

A Data-Efficient Progressive Learning Framework for Robot Scooping Task

ICRA 2025

Robot scooping is a challenging and important task in robotic tool manipulation research due to the complex relationship between the robot, the tool, and target objects/environment. Taking into account different tools, different target objects and varying environments, the required scooping manipula

Cited by 0SourceScholar
2025

A Systematic Survey of Automatic Prompt Optimization Techniques

EMNLP 2025

Since the advent of large language models (LLMs), prompt engineering has been a crucial step for eliciting desired responses for various Natural Language Processing (NLP) tasks. However, prompt engineering remains an impediment for end users due to rapid advances in models, tasks, and associated bes

Cited by 0SourcePDFScholar
2025

Aligning to Constraints for Data-Efficient Language Model Customization

NAACL 2025findings

General-purpose language models (LMs) are aligned to diverse user intents, but fall short when it comes to specific applications. While finetuning is the default method for customized alignment, human annotations are often unavailable in various customization scenarios. Based on the observation that…

Cited by 0SourcePDFScholar
2025

BSO: Binary Spiking Online Optimization Algorithm

ICML 2025poster

Binary Spiking Neural Networks (BSNNs) offer promising efficiency advantages for resource-constrained computing. However, their training algorithms often require substantial memory overhead due to latent weights storage and temporal processing requirements. To address this issue, we propose Binary S…

2025

Bipolar Self-attention for Spiking Transformers

NeurIPS 2025spotlight

Harnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through…

Cited by 0SourceScholar
2025

Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models

NAACL 2025short

Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple object-based visual prompting—overlaying visual cues (e.g., bounding box, circle) on images—can significantly mitigate such hallucination; however, diffe…

Cited by 0SourcePDFScholar
2025

Can’t See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs

ACL 2025long

Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. However, ensuring the safety of these models remains a significant challenge, particularly in accurately identifying whether multimodal content…

2025

Chain-of-Jailbreak Attack for Image Generation Models via Step by Step Editing

ACL 2025finding

Text-based image generation models, such as Stable Diffusion and DALL-E 3, hold significant potential in content creation and publishing workflows, making them the focus in recent years. Despite their remarkable capability to generate diverse and vivid images, considerable efforts are being made to…

2025

Clutter Resilient Occlusion Avoidance for Tightly-Coupled Motion-Assisted Detection

ICASSP 2025accepted

Occlusion is a key factor leading to detection failures. This paper proposes a motion-assisted detection (MAD) method that actively plans an executable path, for the robot to observe the target at a new viewpoint with potentially reduced occlusion. In contrast to existing MAD approaches that may fai…

Cited by 0SourceScholar
2025

Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence Modeling

NeurIPS 2025poster

The explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiote…

Cited by 0SourceScholar
2025

Differentiable Solver Search for Fast Diffusion Sampling

ICML 2025poster

Diffusion models have demonstrated remarkable generation quality but at the cost of numerous function evaluations. Recently, advanced ODE-based solvers have been developed to mitigate the substantial computational demands of reverse-diffusion solving under limited sampling steps. However, these solv…

Cited by 0SourcePDFScholar
2025

Drop the Beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation

AAAI 2025technical

Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In contrast, rap typically features simpler melodies, with a core…

2025

Enhancing Large Vision Model in Street Scene Semantic Understanding through Leveraging Posterior Optimization Trajectory

IROS 2025

To improve the generalization of the autonomous driving (AD) perception model, vehicles need to update the model over time based on the continuously collected data. As time progresses, the amount of data fitted by the AD model expands, which helps to improve the AD model generalization substantially

Cited by 6SourceScholar
2025

Fault-Tolerant Control of Lifting-Wing Multicopter Based on Nonlinear MPC

RA-L 2025

This letter proposes a fault-tolerant control (FTC) framework for a novel type of aircraft—lifting-wing multicopters. The core of the framework is an attitude controller based on nonlinear model predictive control (NMPC), where the objective function of the NMPC is designed on the based of the relax

Cited by 3SourceScholar
2025

FedEMA: Federated Exponential Moving Averaging with Negative Entropy Regularizer in Autonomous Driving

IROS 2025

Street Scene Semantic Understanding (denoted as S3U) is a crucial but complex task for autonomous driving (AD) vehicles. Their inference models typically face poor generalization due to domain-shift. Federated Learning (FL) has emerged as a promising paradigm for enhancing the generalization of AD m

Cited by 5SourceScholar
2025

Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching for Speaker Diarization

ICASSP 2025accepted

Speaker diarization is typically considered as a discriminative task, using discriminative approaches to produce fixed diarization results. In this paper, we explore for the first time the use of neural network-based generative methods for speaker diarization. We implement a Flow-Matching (FM) based…

Cited by 0SourceScholar
2025

GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents

NeurIPS 2025poster

Recent Graphical User Interface (GUI) agents replicate the R1-Zero paradigm, coupling online Reinforcement Learning (RL) with explicit chain-of-thought reasoning prior to object grounding and thereby achieving substantial performance gains. In this paper, we first conduct extensive analysis experime…

Cited by 0SourcecodeScholar
2025

Heading Adjustment by Admittance Control for Lifting-Wing Quadcopters in Strong Winds

RA-L 2025

Wind disturbance is a critical challenge for uninhabited aerial vehicles (UAVs), particularly for hybrid vertical takeoff and landing (VTOL) UAVs such as lifting-wing quadcopters, which are prone to aerodynamic disturbances like gusts and turbulence due to their wing structure. Existing control stra

Cited by 6SourceScholar
2025

LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs – No Silver Bullet for LC or RAG Routing

ICML 2025poster

As Large Language Model (LLM) context windows expand, the necessity of Retrieval-Augmented Generation (RAG) for integrating external knowledge is debated. Existing RAG vs. long-context (LC) LLM comparisons are often inconclusive due to benchmark limitations. We introduce LaRA, a novel benchmark with…

2025

Label Anything: An Interpretable, High-Fidelity and Prompt-Free Annotator

ICRA 2025

Learning-based street scene semantic understanding in autonomous driving (AD) has advanced significantly recently, but the performance of the AD model is heavily dependent on the quantity and quality of the annotated training data. However, traditional manual labeling involves high cost to annotate

Cited by 3SourceScholar
2025

LeVo: High-Quality Song Generation with Multi-Preference Alignment

NeurIPS 2025poster

Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing approaches still struggle with the complex composition of songs and the scarcity of high-quality data, leading to limit…

Cited by 0SourcecodeScholar
2025

Less is More: Empowering GUI Agent with Context-Aware Simplification

ICCV 2025poster

The research focus of GUI agents is shifting from text-dependent to pure-vision-based approaches, which, though promising, prioritize comprehensive pre-training data collection while neglecting contextual modeling challenges. We probe the characteristics of element and history contextual modeling in…

2025

MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion

ICASSP 2025accepted

In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same…

Cited by 0SourceScholar
2025

Memory-Free and Parallel Computation for Quantized Spiking Neural Networks

ICASSP 2025accepted

Quantized Spiking Neural Networks (QSNNs) offer superior energy efficiency and are well-suited for deployment on resource-limited edge devices. However, limited bit-width weight and membrane potential result in a notable performance decline. In this study, we first identify a new underlying cause fo…

Cited by 0SourceScholar
2025

MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache Optimization

ACL 2025long

Deploying large language models (LLMs) with low-rank adaptation (LoRA) on mobile devices is promising due to their capability to complete diverse domain-specific tasks while ensuring privacy and accessibility. In this paper, we introduce MobiLoRA to accelerate LoRA-based LLM inference on mobile devi…

Cited by 0SourcePDFScholar
2025

MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation

NeurIPS 2025poster

Image-to-video generation has made remarkable progress with the advancements in diffusion models, yet generating videos with realistic motion remains highly challenging. This difficulty arises from the complexity of accurately modeling motion, which involves capturing physical constraints, object in…

Cited by 0SourceScholar
2025

Multi-Level Speaker Representation for Target Speaker Extraction

ICASSP 2025accepted

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of speakers may suffer from confusion of speaker identity. In th…

Cited by 0SourceScholar
2025

NovelCR: A Large-Scale Bilingual Dataset Tailored for Long-Span Coreference Resolution

ACL 2025finding

Coreference resolution (CR) endeavors to match pronouns, noun phrases, etc. with their referent entities, acting as an important step for deep text understanding. Presently available CR datasets are either small in scale or restrict coreference resolution to a limited text span. In this paper, we pr…

2025

Opportunistic Collaborative Planning with Large Vision Model Guided Control and Joint Query-Service Optimization

IROS 2025

Navigating autonomous vehicles in open scenarios is a challenge due to the difficulties in handling unseen objects. Existing solutions either rely on small models that struggle with generalization or large models that are resource-intensive. While collaboration between the two offers a promising sol

Cited by 1SourceScholar
2025

Plugging Schema Graph into Multi-Table QA: A Human-Guided Framework for Reducing LLM Reliance

EMNLP 2025

Large language models (LLMs) have shown promise in table Question Answering (Table QA). However, extending these capabilities to multi-table QA remains challenging due to unreliable schema linking across complex tables. Existing methods based on semantic similarity work well only on simplified hand-

2025

Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking Transformers

CVPR 2025poster

Transformers significantly raise the performance limits across various tasks, spurring research into integrating them into spiking neural networks. However, a notable performance gap remains between existing spiking Transformers and their artificial neural network counterparts. Here, we first analyz…

Cited by 0SourcePDFScholar
2025

Robotic Hand Tool Use with Contact-Based Demonstration: The Case of Cucumber Peeling

IROS 2025

Robotic hand tool use has garnered significant attention from robotics researchers, because it enhances dexterity beyond the limitations imposed by manipulators with fixed tool configurations and human-involved manual tool changes. Despite extensive research, current methodologies predominantly focu

Cited by 0SourceScholar
2025

S$^2$NN: Sub-bit Spiking Neural Networks

NeurIPS 2025poster

Spiking Neural Networks (SNNs) offer an energy-efficient paradigm for machine intelligence, but their continued scaling poses challenges for resource-limited deployment. Despite recent advances in binary SNNs, the storage and computational demands remain substantial for large-scale networks. To furt…

Cited by 0SourceScholar
2025

SPA-BENCH: A COMPREHENSIVE BENCHMARK FOR SMARTPHONE AGENT EVALUATION

ICLR 2025spotlight

Smartphone agents are increasingly important for helping users control devices efficiently, with (Multimodal) Large Language Model (MLLM)-based approaches emerging as key contenders. Fairly comparing these agents is essential but challenging, requiring a varied task scope, the integration of agents…

2025

Self-Evolving Pseudo-Rehearsal for Catastrophic Forgetting with Task Similarity in LLMs

NeurIPS 2025poster

Continual learning for large language models (LLMs) demands a precise balance between $\textbf{plasticity}$ - the ability to absorb new tasks - and $\textbf{stability}$ - the preservation of previously learned knowledge. Conventional rehearsal methods, which replay stored examples, are limited by lo…

Cited by 0SourcecodeScholar
2025

SocialEval: Evaluating Social Intelligence of Large Language Models

ACL 2025long

LLMs exhibit promising Social Intelligence (SI) in modeling human behavior, raising the need to evaluate LLMs’ SI and their discrepancy with humans. SI equips humans with interpersonal abilities to behave wisely in navigating social interactions to achieve social goals. This presents an operational…

2025

SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement

NeurIPS 2025poster

Generating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-based methods often struggle to balance global coherence with local fidelity, resulting in outputs that lack musicality or s…

Cited by 0SourcecodeScholar
2025

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor

AAAI 2025technical

The emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models are capable of synthesizing both vocals and accompaniment tracks up to several minutes long concurrently, research about…

2025

Spiking Vision Transformer with Saccadic Attention

ICLR 2025poster

The combination of Spiking Neural Networks (SNNs) and Vision Transformers (ViTs) holds potential for achieving both energy efficiency and high performance, particularly suitable for edge vision applications. However, a significant performance gap still exists between SNN-based ViTs and their ANN cou…

Cited by 1SourcePDFScholar
2025

Tactile-Driven Dexterous In-Hand Writing via Extrinsic Contact Sensing

RA-L 2025

Dexterous in-hand manipulation, especially involving interactions between grasped objects and external environments, remains a formidable challenge in robotics. This study tackles the complexities of in-hand manipulation under extrinsic contact through a representative three-finger handwriting task.

Cited by 0SourcecodeScholar
2025

ToolACE: Winning the Points of LLM Function Calling

ICLR 2025poster

Function calling significantly extends the application boundary of large language models (LLMs), where high-quality and diverse training data is critical for unlocking this capability. However, collecting and annotating real function-calling data is challenging, while synthetic data from existing pi…

Cited by 23SourcePDFScholar
2025

Towards Accurate Binary Spiking Neural Networks: Learning with Adaptive Gradient Modulation Mechanism

AAAI 2025technical

Binary Spiking Neural Networks (BSNNs) inherit the event-driven paradigm of SNNs, while also adopting the reduced storage burden of binarization techniques. These distinct advantages grant BSNNs lightweight and energy-efficient characteristics, rendering them ideal for deployment on resource-constra…

2025

Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural Networks

NeurIPS 2025poster

Spiking Neural Networks (SNNs) demonstrate significant potential for energy-efficient neuromorphic computing through an event-driven paradigm. While training methods and computational models have greatly advanced, SNNs struggle to achieve competitive performance in visual long-sequence modeling task…

Cited by 0SourcecodeScholar
2025

VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

ICASSP 2025accepted

Recent TTS models with decoder-only Transformer architecture, such as SPEAR-TTS and VALL-E, achieve impressive naturalness and demonstrate the ability for zero-shot adaptation given a speech prompt. However, such decoder-only TTS models lack monotonic alignment constraints, sometimes leading to hall…

Cited by 0SourceScholar
2025

VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis

AAAI 2025technical

This paper develops a Versatile and Honest vision language Model (VHM) for remote sensing image analysis. VHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest instruction dataset comprising both factual and deceptive questions (HnstD). Un…

2024

A High-Performance Anthropomorphic Robotic Arm for Household Applications

IROS 2024poster

Anthropomorphic robotic arms, mimicking the structure and function of human arms, show great potential for helping people in various tedious and repetitive household tasks. However, such arms mostly consist of multiple serial links controlled independently by actuators at joints with high reduction…

Cited by 0SourceScholar
2024

A Robust Model Predictive Controller for Tactile Servoing

ICRA 2024poster

Tactile servoing is an effective approach to enabling robots to safely interact with unknown environments. One of the core problems in tactile servoing is to robustly converge the contact features to the desired ones via a dedicated controller. This paper proposes a Data-Driven Model Predictive Cont…

Cited by 2SourceScholar
2024

Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-Talker Speech

ICASSP 2024accepted

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the interfering speech. However, this scenario only accounts for a small…

Cited by 0SourceScholar
2024

AutoPrep: An Automatic Preprocessing Framework for In-The-Wild Speech Data

ICASSP 2024accepted

Recently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs). However, the use of in-the-wild large-scale speech data in the speech technology community remains constrained. One reason for this limitation is that…

Cited by 0SourceScholar
2024

BERGEN: A Benchmarking Library for Retrieval-Augmented Generation

EMNLP 2024finding

Retrieval-Augmented Generation allows to enhance Large Language Models with external knowledge. In response to the recent popularity of generative LLMs, many RAG approaches have been proposed, which involve an intricate number of different configurations such as evaluation datasets, collections, met…

2024

BEVLOC: End-to-End 6-DoF Localization Via Cross-Modality Correlation Under Bird's Eye View

ICASSP 2024accepted

Accurate ego-centric localization assumes a paramount significance in the domain of autonomous driving. However, traditional methods for camera-LiDAR map localization rely on perspective projection to create a unified representation, which often falls short due to challenges such as occlusion and th…

Cited by 0SourceScholar
2024

Co-Axial Slender Tubular robot (CAST): Towards Robotized Operation for Transorbital Neurosurgery with Minimal Invasiveness

ICRA 2024poster

Transorbital Neuro Surgery (TNS) offers a novel treatment towards the lesion inside skull pursuing minimal invasiveness. Most conventional TNS tools are rigid and straight, limiting the dexterity and accessibility in passing a small port. Bendable and steerable surgical tools provides an alternative…

Cited by 0SourceScholar
2024

Design and Visual Servoing Control of a Hybrid Dual-Segment Flexible Neurosurgical Robot for Intraventricular Biopsy

ICRA 2024poster

Traditional rigid endoscopes have challenges in flexibly treating tumors located deep in the brain, and low operability and fixed viewing angles limit its development. This study introduces a novel dual-segment flexible robotic endoscope MicroNeuro, designed to perform biopsies with dexterous surgic…

Cited by 3SourceScholar
2024

Dualvc 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion

ICASSP 2024accepted

Voice conversion is becoming increasingly popular, and a growing number of application scenarios require models with streaming inference capabilities. The recently proposed DualVC attempts to achieve this objective through streaming model architecture design and intra-model knowledge distillation al…

Cited by 0SourceScholar
2024

ESP-PCT: Enhanced VR Semantic Performance through Efficient Compression of Temporal and Spatial Redundancies in Point Cloud Transformers

IJCAI 2024poster

Semantic recognition is pivotal in virtual reality (VR) applications, enabling immersive and interactive experiences. A promising approach is utilizing millimeter-wave (mmWave) signals to generate point clouds. However, the high computational and memory demands of current mmWave point cloud models h…

2024

Enhancing mmWave Radar Point Cloud via Visual-inertial Supervision

ICRA 2024poster

Complementary to prevalent LiDAR and camera systems, millimeter-wave (mmWave) radar is robust to adverse weather conditions like fog, rainstorms, and blizzards but offers sparse point clouds. Current techniques enhance the point cloud by the supervision of LiDAR’s data. However, high-performance LiD…

Cited by 0SourcecodeScholar
2024

Exploring DCN-like architecture for fast image generation with arbitrary resolution

NeurIPS 2024poster

Arbitrary-resolution image generation still remains a challenging task in AIGC, as it requires handling varying resolutions and aspect ratios while maintaining high visual quality. Existing transformer-based diffusion methods suffer from quadratic computation cost and limited resolution extrapolatio…

Cited by 0SourcePDFScholar
2024

FedRC: A Rapid-Converged Hierarchical Federated Learning Framework in Street Scene Semantic Understanding

IROS 2024poster

Street Scene Semantic Understanding (denoted as TriSU) is a crucial but complex task for world-wide distributed autonomous driving (AD) vehicles (e.g., Tesla). Its inference model faces poor generalization issue due to inter-city domain-shift. Hierarchical Federated Learning (HFL) offers a potential…

Cited by 7SourceScholar
2024

Joint Input and Output Coordination for Class-Incremental Learning

IJCAI 2024poster

Incremental learning is nontrivial due to severe catastrophic forgetting. Although storing a small amount of data on old tasks during incremental learning is a feasible solution, current strategies still do not 1) adequately address the class bias problem, and 2) alleviate the mutual interference be…

Cited by 2SourcePDFScholar
2024

Leveraging in-the-wild Data for Effective Self-supervised Pretraining in Speaker Recognition

ICASSP 2024accepted

Current speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to transfer learned high-level features to the downstream speaker recognition task. H…

Cited by 2SourceScholar
2024

Multi-Uncertainty Aware Autonomous Cooperative Planning

IROS 2024poster

Autonomous cooperative planning (ACP) is a promising technique to improve the efficiency and safety of multi-vehicle interactions for future intelligent transportation systems. However, realizing robust ACP is a challenge due to the aggregation of perception, motion, and communication uncertainties.…

Cited by 1SourceScholar
2024

OOP: Object-Oriented Programming Evaluation Benchmark for Large Language Models

ACL 2024findings

Advancing automated programming necessitates robust and comprehensive code generation benchmarks, yet current evaluation frameworks largely neglect object-oriented programming (OOP) in favour of functional programming (FP), e.g., HumanEval and MBPP. To address this, our study introduces a pioneering…

2024

Robust Cross-Domain Speaker Verification with Multi-Level Domain Adapters

ICASSP 2024accepted

Speaker verification encounters significant challenges when confronted with diverse domain data, often resulting in performance degradation due to domain mismatch. To enhance performance in cross-domain scenarios, we introduce the Domain Adapter, an adaptable module designed for specific domains. Th…

Cited by 0SourceScholar
2024

Seamless Virtual Reality With Integrated Synchronizer and Synthesizer for Autonomous Driving

RA-L 2024

Virtual reality (VR) is a promising data engine for autonomous driving (AD). However, data fidelity in this paradigm is often degraded by VR inconsistency, for which the existing VR approaches become ineffective, as they ignore the inter-dependency between low-level VR synchronizer designs (i.e., da

Cited by 8SourceScholar
2024

Spike-based Neuromorphic Model for Sound Source Localization

NeurIPS 2024poster

Biological systems possess remarkable sound source localization (SSL) capabilities that are critical for survival in complex environments. This ability arises from the collaboration between the auditory periphery, which encodes sound as precisely timed spikes, and the auditory cortex, which performs…

Cited by 6SourcePDFScholar
2024

Split and Merge: Aligning Position Biases in LLM-based Evaluators

EMNLP 2024main

Large language models (LLMs) have shown promise as automated evaluators for assessing the quality of answers generated by AI systems. However, LLM-based evaluators exhibit position bias, or inconsistency, when used to evaluate candidate answers in pairwise comparisons, favoring either the first or s…

2024

Uncertainty-aware Deep Imitation Learning and Deployment for Autonomous Navigation through Crowded Intersections

IROS 2024

Navigation through crowded intersections is a challenge for autonomous vehicles, where uncertainty arises from interaction with other road users, encountering new scenes and weathers, etc. Recent end-to-end autonomous control deep models learned from human drivers have shown promising driving perfor

Cited by 2SourceScholar
2024

UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding

AAAI 2024technical

The utilization of discrete speech tokens, divided into semantic tokens and acoustic tokens, has been proven superior to traditional acoustic feature mel-spectrograms in terms of naturalness and robustness for text-to-speech (TTS) synthesis. Recent popular models, such as VALL-E and SPEAR-TTS, allow…

2023

Beyond ADMM: A Unified Client-Variance-Reduced Adaptive Federated Learning Framework

AAAI 2023technical

As a novel distributed learning paradigm, federated learning (FL) faces serious challenges in dealing with massive clients with heterogeneous data distribution and computation and communication resources. Various client-variance-reduction schemes and client sampling strategies have been respectively…

Cited by 12SourcePDFScholar
2023

Boosting Semi-Supervised Federated Learning with Model Personalization and Client-Variance-Reduction

ICASSP 2023accepted

Recently, federated learning (FL) has been increasingly appealing in distributed signal processing and machine learning. Nevertheless, the practical challenges of label deficiency and client heterogeneity form a bottleneck to its wide adoption. Although numerous efforts have been devoted to semi- su…

Cited by 0SourceScholar
2023

Byzantine-Robust Federated Learning with Optimal Statistical Rates

AISTATS 2023poster

We propose Byzantine-robust federated learning protocols with nearly optimal statistical rates based on recent progress in high dimensional robust statistics. In contrast to prior work, our proposed protocols improve the dimension dependence and achieve a near-optimal statistical rate for strongly c…

Cited by 35SourcePDFScholar
2023

Communication Resources Constrained Hierarchical Federated Learning for End-to-End Autonomous Driving

IROS 2023poster

While federated learning (FL) improves the generalization of end-to-end autonomous driving by model aggregation, the conventional single-hop FL (SFL) suffers from slow convergence rate due to long-range communications among vehicles and cloud server. Hierarchical federated learning (HFL) overcomes s…

Cited by 20SourcecodeScholar
2023

Contrastive Training Improves Zero-Shot Classification of Semi-structured Documents

ACL 2023findings

We investigate semi-structured document classification in a zero-shot setting. Classification of semi-structured documents is more challenging than that of standard unstructured documents, as positional, layout, and style information play a vital role in interpreting such documents. The standard cla…

2023

DFRD: Data-Free Robustness Distillation for Heterogeneous Federated Learning

NeurIPS 2023poster

Federated Learning (FL) is a privacy-constrained decentralized machine learning paradigm in which clients enable collaborative training without compromising private data. However, how to learn a robust global model in the data-heterogeneous and model-heterogeneous FL scenarios is challenging. To add…

Cited by 18SourcePDFScholar
2023

Explain Any Concept: Segment Anything Meets Concept-Based Explanation

NeurIPS 2023poster

EXplainable AI (XAI) is an essential topic to improve human understanding of deep neural networks (DNNs) given their black-box internals. For computer vision tasks, mainstream pixel-based XAI methods explain DNN decisions by identifying important pixels, and emerging concept-based XAI explore formin…

2023

Feature Alignment and Uniformity for Test Time Adaptation

CVPR 2023poster

Test time adaptation (TTA) aims to adapt deep neural networks when receiving out of distribution test domain samples. In this setting, the model can only access online unlabeled test samples and pre-trained models on the training domains. We first address TTA as a feature revision problem due to the…

2023

Meta-Reinforcement Learning Based on Self-Supervised Task Representation Learning

AAAI 2023technical

Meta-reinforcement learning enables artificial agents to learn from related training tasks and adapt to new tasks efficiently with minimal interaction data. However, most existing research is still limited to narrow task distributions that are parametric and stationary, and does not consider out-of-…

Cited by 16SourcePDFScholar
2023

Prototype Knowledge Distillation for Medical Segmentation with Missing Modality

ICASSP 2023accepted

Multi-modality medical imaging is crucial in clinical treatment as it can provide complementary information for medical image segmentation. However, collecting multi-modal data in clinical is difficult due to the limitation of the scan time and other clinical situations. As such, it is clinically me…

Cited by 0SourceScholar
2023

RDA: An Accelerated Collision Free Motion Planner for Autonomous Navigation in Cluttered Environments

RA-L 2023

Autonomous motion planning is challenging in multi-obstacle environments due to nonconvex collision avoidance constraints. Directly applying numerical solvers to these nonconvex formulations fails to exploit the constraint structures, resulting in excessive computation time. In this letter, we prese

Cited by 49SourcecodeScholar
2023

Relative Roughness Measurement Based Real-Time Speed Planning for Autonomous Vehicles on Rugged Road

IROS 2023poster

In order to guarantee autonomous vehicles' autonomy, mobility, and ride quality in rugged environments, a real-time speed planning method based on the time-frequency transformation of terrain characteristics is designed to achieve adaptive speed planning of autonomous vehicles in rough ground. On th…

Cited by 1SourceScholar
2023

Taxonomy Expansion for Named Entity Recognition

EMNLP 2023long main

Training a Named Entity Recognition (NER) model often involves fixing a taxonomy of entity types. However, requirements evolve and we might need the NER model to recognize additional entity types. A simple approach is to re-annotate entire dataset with both existing and additional entity types and t…

Cited by 0SourceScholar
2023

Teaching What You Should Teach: A Data-Based Distillation Method

IJCAI 2023poster

In real teaching scenarios, an excellent teacher always teaches what he (or she) is good at but the student is not. This gives the student the best assistance in making up for his (or her) weaknesses and becoming a good one overall. Enlightened by this, we introduce the "Teaching what you Should Tea…

Cited by 4SourcePDFScholar
2023

Towards Open-Vocabulary Video Instance Segmentation

ICCV 2023oral

Video Instance Segmentation (VIS) aims at segmenting and categorizing objects in videos from a closed set of training categories, lacking the generalization ability to handle novel categories in real-world videos. To address this limitation, we make the following three contributions. First, we intro…

Cited by 36PDFcodeScholar
2023

Wespeaker: A Research and Production Oriented Speaker Embedding Learning Toolkit

ICASSP 2023accepted

Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research and production oriented speaker embedding learning toolkit,…

Cited by 0SourceScholar
2022

Adversarial Examples Detection Based on Error Level Analysis and Space Mapping

ICASSP 2022accepted

Deep neural network (DNN) shows impressive performance on many tasks but they usually suffer from adversarial examples with human eyes invisible slight perturbation. Such examples can not be distinguished by human but can mislead DNN classifiers leading to its important role in DNN attack and defens…

Cited by 0SourceScholar
2022

An Adaptive Approach to Whole-Body Balance Control of Wheel-Bipedal Robot Ollie

IROS 2022poster

The wheel-bipedal robot has the advantages of both wheeled robots and legged robots, but as a cost, it is more challenging to perform flexible movements in various surroundings while keeping it balanced. The inaccurate dynamics of the robot makes the balance problem even more intractable. To solve t…

Cited by 28SourceScholar
2022

DocEE: A Large-Scale and Fine-grained Benchmark for Document-level Event Extraction

NAACL 2022long

Event extraction aims to identify an event and then extract the arguments participating in the event. Despite the great success in sentence-level event extraction, events are more naturally presented in the form of documents, with event arguments scattered in multiple sentences. However, a major bar…

2022

On the Importance of Different Frequency Bins for Speaker Verification

ICASSP 2022accepted

The majority of modern speaker verification systems take spectral analysis-based features as input, which contains multiple frequency bins. Naturally, there would be a question of whether all different frequency bins contribute equally to the speaker verification system performance? In this paper, w…

Cited by 0SourceScholar
2022

Rethinking Video Rain Streak Removal: A New Synthesis Model and a Deraining Network with Video Rain Prior

ECCV 2022poster

"Existing video synthetic models and deraining methods are mostly built on a simplified video rain model assuming that rain streak layers of different video frames are uncorrelated, thereby producing degraded performance on real-world rainy videos. To address this problem, we devise a new video rain…

2022

SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous Vehicles

NeurIPS 2022accept

As shown by recent studies, machine intelligence-enabled systems are vulnerable to test cases resulting from either adversarial manipulation or natural distribution shifts. This has raised great concerns about deploying machine learning algorithms for real-world applications, especially in safety-cr…

2022

Self-Knowledge Distillation via Feature Enhancement for Speaker Verification

ICASSP 2022accepted

As the most widely used technique, deep speaker embedding learning has become predominant in speaker verification task recently. Very large neural networks such as ECAPA-TDNN and ResNet can achieve the state-of-the-art performance. However, large models are computationally unfriendly in general, whi…

Cited by 0SourceScholar
2022

VIP-SLAM: An Efficient Tightly-Coupled RGB-D Visual Inertial Planar SLAM

ICRA 2022poster

In this paper, we propose a tightly-coupled SLAM system fused with RGB, Depth, IMU and structured plane information. Traditional sparse points based SLAM systems always maintain a mass of map points to model the environment. Huge number of map points bring us a high computational complexity, making…

Cited by 30SourceScholar
2021

Balance Control of a Novel Wheel-legged Robot: Design and Experiments

ICRA 2021poster

This paper presents a balance control technique for a novel wheel-legged robot. We first derive a dynamic model of the robot and then apply a linear feedback controller based on output regulation and linear quadratic regulator (LQR) methods to maintain the standing of the robot on the ground without…

Cited by 94SourceScholar
2021

Demystifying Model Averaging for Communication-Efficient Federated Matrix Factorization

ICASSP 2021accepted

Federated learning (FL) is encountered with the challenge of training a model in massive and heterogeneous networks. Model averaging (MA) has become a popular FL paradigm where parallel (stochastic) gradient descent (GD) is run on a small sampled subset of clients multiple times before uploading the…

Cited by 0SourceScholar
2021

Distributed Dynamic Map Fusion via Federated Learning for Intelligent Networked Vehicles

ICRA 2021poster

The technology of dynamic map fusion among networked vehicles has been developed to enlarge sensing ranges and improve sensing accuracies for individual vehicles. This paper proposes a federated learning (FL) based dynamic map fusion framework to achieve high map quality despite unknown numbers of o…

Cited by 87SourcecodeScholar
2021

FlowDriveNet: An End-to-End Network for Learning Driving Policies from Image Optical Flow and LiDAR Point Flow

ICRA 2021poster

Learning driving policies using an end-to-end network has been proved a promising solution for autonomous driving. Due to the lack of a benchmark driver behavior dataset that contains both the visual and the LiDAR data, existing works solely focus on learning driving from visual sensors. Besides, mo…

Cited by 4SourceScholar
2021

Learning from Miscellaneous Other-Class Words for Few-shot Named Entity Recognition

ACL 2021long

Few-shot Named Entity Recognition (NER) exploits only a handful of annotations to iden- tify and classify named entity mentions. Pro- totypical network shows superior performance on few-shot NER. However, existing prototyp- ical methods fail to differentiate rich seman- tics in other-class words, wh…

2021

Learning-Based Balance Control of Wheel-Legged Robots

RA-L 2021

This letter studies the adaptive optimal control problem for a wheel-legged robot in the absence of an accurate dynamic model. A crucial strategy is to exploit recent advances in reinforcement learning (RL) and adaptive dynamic programming (ADP) to derive a learning-based solution to adaptive optima

Cited by 94SourceScholar
2021

Multi-Domain Multi-Task Rehearsal for Lifelong Learning

AAAI 2021technical

Rehearsal, seeking to remind the model by storing old knowledge in lifelong learning, is one of the most effective ways to mitigate catastrophic forgetting, i.e., biased forgetting of previous knowledge when moving to new tasks. However, the old tasks of the most previous rehearsal-based methods suf…

Cited by 31SourcePDFScholar
2021

Perception Matters: Detecting Perception Failures of VQA Models Using Metamorphic Testing

CVPR 2021poster

Visual question answering (VQA) takes an image and a natural-language question as input and returns a natural-language answer. To date, VQA models are primarily assessed by their accuracy on high-level reasoning questions. Nevertheless, Given that perception tasks (e.g., recognizing objects) are the…

Cited by 51PDFcodeScholar
2021

Private Image Reconstruction from System Side Channels Using Generative Models

ICLR 2021poster

System side channels denote effects imposed on the underlying system and hardware when running a program, such as its accessed CPU cache lines. Side channel analysis (SCA) allows attackers to infer program secrets based on observed side channel signals. Given the ever-growing adoption of machine lea…

2021

Self-Supervised Learning Based Domain Adaptation for Robust Speaker Verification

ICASSP 2021accepted

Large performance degradation is often observed for speaker verification systems when applied to a new domain dataset. Given an unlabeled target-domain dataset, unsupervised domain adaptation (UDA) methods, which usually leverage adversarial training strategies, are commonly used to bridge the perfo…

Cited by 0SourceScholar
2021

SynAug: Synthesis-Based Data Augmentation for Text-Dependent Speaker Verification

ICASSP 2021accepted

Text-dependent speaker verification systems trained on large amount of labelled data exhibit remarkable performance. However, collecting the speech from a lot of speakers with target transcript is a lengthy and expensive process. In this work, we propose a synthesis based data augmentation method (S…

Cited by 0SourceScholar
2021

Unit Selection Synthesis Based Data Augmentation for Fixed Phrase Speaker Verification

ICASSP 2021accepted

Data augmentation is commonly used to help build a robust speaker verification system, especially in limited-resource case. However, conventional data augmentation methods usually focus on the diversity of acoustic environment, leaving the lexicon variation neglected. For text dependent speaker veri…

Cited by 0SourceScholar
2020

Bayes-enhanced Lifelong Attention Networks for Sentiment Classification

COLING 2020main

The classic deep learning paradigm learns a model from the training data of a single task and the learned model is also tested on the same task. This paper studies the problem of learning a sequence of tasks (sentiment classification tasks in our case). After each sentiment classification task is le…

Cited by 8SourcePDFScholar
2020

But System for the Second Dihard Speech Diarization Challenge

ICASSP 2020accepted

This paper describes the winning systems developed by the BUT team for the four tracks of the Second DIHARD Speech Diarization Challenge. For tracks 1 and 2 the systems were mainly based on performing agglomerative hierarchical clustering (AHC) of x-vectors, followed by another x-vector clustering b…

Cited by 60SourceScholar
2020

Channel Invariant Speaker Embedding Learning with Joint Multi-Task and Adversarial Training

ICASSP 2020accepted

Using deep neural network to extract speaker embedding has significantly improved the speaker verification task. However, such embeddings are still vulnerable to channel variability. Previous works have used adversarial training to suppress channel information to extract channel-invariant embedding…

Cited by 0SourceScholar
2020

Gain Scheduled Controller Design for Balancing an Autonomous Bicycle

IROS 2020poster

In this paper, the gain scheduling technique is applied to design a balance controller for an autonomous bicycle with an inertia wheel. Previously, two different balance controllers are needed depending on whether the bicycle is stationary or dynamic. The switch between the two different controllers…

Cited by 16SourceScholar
2020

Intelligent Home 3D: Automatic 3D-House Design From Linguistic Descriptions Only

CVPR 2020poster

Home design is a complex task that normally requires architects to finish with their professional skills and tools. It will be fascinating that if one can produce a house plan intuitively without knowing much knowledge about home design and experience of using complex designing tools, for example, v…

Cited by 47PDFcodeScholar
2020

Investigation of Specaugment for Deep Speaker Embedding Learning

ICASSP 2020accepted

SpecAugment is a newly proposed data augmentation method for speech recognition. By randomly masking bands in the log Mel spectogram this method leads to impressive performance improvements. In this paper, we investigate the usage of SpecAugment for speaker verification tasks. Two different models,…

Cited by 0SourceScholar
2020

Metamorphic Testing and Certified Mitigation of Fairness Violations in NLP Models

IJCAI 2020poster

Natural language processing (NLP) models have been increasingly used in sensitive application domains including credit scoring, insurance, and loan assessment. Hence, it is critical to know that the decisions made by NLP models are free of unfair bias toward certain subpopulation groups. In this pap…

Cited by 0SourcePDFScholar
2020

Nonlinear Balance Control of an Unmanned Bicycle: Design and Experiments

IROS 2020poster

In this paper, nonlinear control techniques are exploited to balance an unmanned bicycle with enlarged stability domain. We consider two cases. For the first case when the autonomous bicycle is balanced by the flywheel, the steering angle is set to zero, and the torque of the flywheel is used as the…

Cited by 26SourceScholar
2020

Optimizing Bayesian Hmm Based X-Vector Clustering for the Second Dihard Speech Diarization Challenge

ICASSP 2020accepted

This paper presents an analysis of our diarization system winning the second DIHARD speech diarization challenge, track 1. This system is based on clustering x-vector speaker embeddings extracted every 0.25s from short segments of the input recording. In this paper, we focus on the two x-vector clus…

Cited by 0SourceScholar
2020

Spatio-Temporal Ultrasonic Dataset: Learning Driving from Spatial and Temporal Ultrasonic Cues

IROS 2020poster

Recent works have proved that combining spatial and temporal visual cues can significantly improve the performance of various vision-based robotic systems. However, for the ultrasonic sensors used in most robotic tasks (e.g. collision avoidance, localization and navigation), there is a lack of bench…

Cited by 0SourceScholar
2020

Text Adaptation for Speaker Verification with Speaker-Text Factorized Embeddings

ICASSP 2020accepted

Text mismatch between pre-collected data, either training data or enrollment data, and the actual test data can significantly hurt text-dependent speaker verification (SV) system performance. Although this problem can be solved by carefully collecting data with the target speech content, such data c…

Cited by 0SourceScholar
2019

Clustering by Orthogonal Non-negative Matrix Factorization: A Sequential Non-convex Penalty Approach

ICASSP 2019accepted

The non-negative matrix factorization (NMF) model with an additional orthogonality constraint on one of the factor matrices, called the orthogonal NMF (ONMF), has been found to provide improved clustering performance over the K-means. The ONMF model is a challenging optimization problem due to the o…

Cited by 0SourceScholar
2019

Knowledge Distillation for Small Foot-print Deep Speaker Embedding

ICASSP 2019accepted

Deep speaker embedding learning is an effective method for speaker identity modelling. Very deep models such as ResNet can achieve remarkable results but are usually too computationally expensive for real applications with limited resources. On the other hand, simply reducing model size is likely to…

Cited by 0SourceScholar
2019

Massive MIMO Multicast Beamforming via Accelerated Random Coordinate Descent

ICASSP 2019accepted

One key feature of massive multiple-input multiple-output systems is the large number of antennas and users. As a result, reducing the computational complexity of beamforming design becomes imperative. To this end, the goal of this paper is to achieve a lower complexity order than that of existing b…

Cited by 0SourceScholar
2018

BML: A High-performance, Low-cost Gradient Synchronization Algorithm for DML Training

NeurIPS 2018poster

In distributed machine learning (DML), the network performance between machines significantly impacts the speed of iterative training. In this paper we propose BML, a new gradient synchronization algorithm with higher network performance and lower network cost than the current practice. BML runs on…

Cited by 39SourcePDFScholar
2018

Cell Subclass Identification in Single-Cell RNA-Sequencing Data Using Orthogonal Nonnegative Matrix Factorization

ICASSP 2018accepted

Identification of cell subclasses using single-cell RNA-Sequencing (scRNA-Seq) data is of paramount importance since it uncovers the hidden biological processes within the cell population. While the nonnegative matrix factorization (NMF) model has been reported to be effective in various unsupervise…

Cited by 0SourceScholar
2018

Focal Kl-Divergence Based Dilated Convolutional Neural Networks for Co-Channel Speaker Identification

ICASSP 2018accepted

Recognizing the identities of multiple talkers via their overlapped speech is a challenging task, it is also one main difficulty for the “cocktail party problem”. In this paper, a novel dilated convolutional neural network with a focal KL-divergence loss function is proposed to tackle this problem.…

Cited by 0SourceScholar
2018

Joint I-Vector with End-to-End System for Short Duration Text-Independent Speaker Verification

ICASSP 2018accepted

Factor analysis based i-vector has been the state-of-the-art method for speaker verification. Recently, researchers propose to build DNN based end-to-end speaker verification systems and achieve comparable performance with <i xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.…

Cited by 0SourceScholar
2017

A friction model with velocity, temperature and load torque effects for collaborative industrial robot joints

IROS 2017poster

In this paper, a comprehensive friction model for collaborative industrial robot joints is proposed which takes into account the velocity, temperature and load torque effects. The model indicates that the velocity and temperature have a strong influence on viscous friction nonlinearly, whereas load…

Cited by 40SourceScholar
2017

Deep learning of directional truncated signed distance function for robust 3D object recognition

IROS 2017poster

In this paper, we develop a novel 3D object recognition algorithm to perform detection and pose estimation jointly. We focus on analyzing the advantages of the 3D point cloud relative to the RGB-D image and try to eliminate the unpredictability of output values that inevitably occurs in regression t…

Cited by 14SourceScholar
2016

Achieving global optimality for wirelessly-powered multi-antenna TWRC with lattice codes

ICASSP 2016accepted

In this paper, we consider the joint optimization of relay transmit-receive beamformers, users' transmit powers, and users' power splitting ratios in wirelessly-powered two-way relay channel under data-rate quality-of-service constraints. In order to solve the problem, we first establish that the up…

Cited by 0SourceScholar
2016

Control and modeling for direct teaching of industrial articulated robotic arms

IROS 2016poster

This paper presents an improved force-free control method based on current, which can be applied to industrial articulated robotic arms with large mass and large friction torque for direct teaching. Three kinds of torques that influence direct teaching are analyzed, and thus a calibration method and…

Cited by 8SourceScholar