← Search

Chen Zhang

136 accepted papers

2026

Automated High-Precision Control of Twisted and Coiled Polymers Under Parameter Variability

RA-L 2026

Twisted and coiled polymer (TCP) artificial muscles offer high energy density and large deformation, but achieving reliable closed-loop control remains difficult due to time-temperature-dependent parameter drift, structural variability, and labor-intensive control parameter tuning, which accelerate

Cited by 0SourceScholar
2026

CIA: Cluster-Instance Alignment for Unsupervised Day-Night Vehicle Re-Identification

AAAI 2026technical

Cross-time vehicle re-identification (Re-ID), especially across day and night conditions, remains a challenging problem due to drastic illumination variations that lead to significant domain shifts. While existing methods perform well under daytime scenarios, their effectiveness degrades severely in

Cited by 0SourcePDFScholar
2026

CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards

ICLR 2026poster

Self-evolution is a central research topic in enabling large language model (LLM)-based agents to continually improve their capabilities after pretraining. Recent research has witnessed a transition from reinforcement learning (RL)-free to RL-based methods. Current RL-based methods either rely on de…

Cited by 0SourcecodeScholar
2026

Combinatorial Bandit Bayesian Optimization for Tensor Outputs

ICLR 2026poster

Bayesian optimization (BO) has been widely used to optimize expensive and black-box functions across various domains. Existing BO methods have not addressed tensor-output functions. To fill this gap, we propose a novel tensor-output BO method. Specifically, we first introduce a tensor-output Gaussia…

Cited by 0SourceScholar
2026

DMP-TTS: DISENTANGLED MULTI-MODAL PROMPTING FOR CONTROLLABLE TEXT-TO-SPEECH WITH CHAINED GUIDANCE

ICASSP 2026oral

Controllable text-to-speech (TTS) systems face significant challenges in achieving independent manipulation of speaker timbre and speaking style, often suffering from entanglement between these attributes. We present DMP-TTS, a latent Diffusion Transformer (DiT) framework with explicit disentangleme…

Cited by 0SourcePDFScholar
2026

DRL-STAF: A DRL Framework for State-aware Forecasting of Complex Multivariate Hidden Markov Process

ICML 2026poster

Forecasting multivariate hidden Markov processes is challenging due to nonlinear and nonstationary observations, latent state transitions, and cross-sequence dependencies. While deep learning methods achieve strong predictive accuracy, they typically lack explicit state modeling, whereas Hidden Mark…

Cited by 0SourceScholar
2026

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

IJCAI 2026

Large language model (LLM)-based text-to-speech (TTS) systems enable prompt-conditioned emotional control but struggle with fine-grained emotion intensity due to the semantic--acoustic gap between text and speech. To address this challenge, we formulate emotion intensity control in LLM-based TTS as

Cited by 0Scholar
2026

Empowering DINO Representations for Underwater Instance Segmentation via Aligner and Prompter

AAAI 2026technical

Underwater Instance Segmentation (UIS), integrating pixel-level understanding and instance-level discrimination, is a pivotal technology in marine resource exploration and ecological protection. In recent years, large-scale pretrained visual foundation models, exemplified by DINO, have advanced rapi

Cited by 0SourcePDFScholar
2026

HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics Prediction

CVPR 2026

Spatial Transcriptomics (ST) merges the benefits of pathology images and gene expression, linking molecular profiles with tissue structure to analyze spot-level function comprehensively. Predicting gene expression from histology images is a cost-effective alternative to expensive ST technologies. Ho

Cited by 0SourcecodeScholar
2026

InfinityHuman: Towards Long-Term Audio-Driven Human Animation

CVPR 2026

Audio-driven human animation has attracted wide attention thanks to its practical applications. However, critical challenges remain in generating high-resolution, long-duration videos with consistent appearance and natural hand motions. Existing methods extend videos using overlapping motion frames

Cited by 0SourcecodeScholar
2026

InstructAudio: Unified speech and music generation with natural language instruction

ICASSP 2026poster

Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only limited text-level attribute control, and rarely support dialogue generation. TTM systems are constrained by input condi…

Cited by 0SourcePDFScholar
2026

NeuroBridge: Bio-Inspired Self-Supervised EEG-to-Image Decoding via Cognitive Priors and Bidirectional Semantic Alignment

AAAI 2026technical

Visual neural decoding seeks to reconstruct or infer perceived visual stimuli from brain activity patterns, providing critical insights into human cognition and enabling transformative applications in brain-computer interfaces and artificial intelligence. Current approaches, however, remain constrai

Cited by 0SourcePDFScholar
2026

Nonparametric Teaching of Attention Learners

ICLR 2026poster

Attention learners, neural networks built on the attention mechanism, e.g., transformers, excel at learning the implicit relationships that relate sequences to their corresponding properties, e.g., mapping a given sequence of tokens to the probability of the next token. However, the learning process…

Cited by 0SourcecodeScholar
2026

Revisiting Hypernetwork in Model Heterogeneous Personalized Federated Learning

IJCAI 2026

Recent personalized federated learning research focuses on heterogeneous models across clients. However, existing methods often rely on external data, model decoupling, and partial learning, which makes them sensitive to settings. In contrast, we revisit hypernetworks and leverage their strong gener

Cited by 0Scholar
2026

SciTS: Scientific Time Series Understanding and Generation with LLMs

ICLR 2026poster

The scientific reasoning ability of large language models (LLMs) has recently attracted significant attention. Time series, as a fundamental modality in scientific data, presents unique challenges that are often overlooked in current multimodal LLMs, which either encode numerical sequences as text o…

Cited by 0SourceScholar
2026

SpatialJB: How Text Distribution Art Becomes The "Jailbreak Key" for LLM Guardrails

ICML 2026poster

While Large Language Models (LLMs) have achieved remarkable success across diverse tasks, they remain vulnerable to jailbreak attacks, which pose significant risks to their secure deployment. Current safetymechanisms primarily rely on output guardrails to filter harmful outputs, yet these defenses a…

Cited by 0SourceScholar
2026

Towards Foundation Models for Zero-Shot Time Series Anomaly Detection: Leveraging Synthetic Data and Relative Context Discrepancy

ICML 2026poster

TSAD is a critical task, but developing models that generalize to unseen data in a zero-shot manner remains a major challenge. Prevailing foundation models for TSAD predominantly rely on reconstruction-based objectives, which suffer from a fundamental objective mismatch and representation conflict: …

Cited by 0SourceScholar
2026

VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range Consistency

ICML 2026spotlight

Despite recent progress in calibration-free monocular SLAM via 3D vision foundation models, scale drift remains severe on long sequences. Motion-agnostic partitioning breaks contextual coherence and causes zero-motion drift, while conventional geometric alignment is computationally expensive. To add…

Cited by 0SourceScholar
2026

VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference

CVPR 2026

Long-context video understanding and generation pose a significant computational challenge for Transformer-based video models due to the quadratic complexity of self-attention. While existing sparse attention methods employ coarse-grained patterns to improve efficiency, they typically incur redundan

Cited by 0SourcecodeScholar
2025

A Compressive Memory-based Retrieval Approach for Event Argument Extraction

COLING 2025main

Recent works have demonstrated the effectiveness of retrieval augmentation in the Event Argument Extraction (EAE) task. However, existing retrieval-based EAE methods have two main limitations: (1) input length constraints and (2) the gap between the retriever and the inference model. These issues li…

Cited by 3SourcePDFScholar
2025

Achieving Lightweight Super-Resolution for Real-Time Computer Graphics

AAAI 2025technical

Image super-resolution (SR) is essential for bridging the gap between modern hardware and real-time computer graphics (CG) applications. It reduces CG workload by allowing low-resolution rendering, with original quality restored later via mathematical operations or machine learning. However, recent…

2025

Aligning Language Models Using Follow-up Likelihood as Reward Signal

AAAI 2025technical

In natural human-to-human conversations, participants often receive feedback signals from one another based on their follow-up reactions. These reactions can include verbal responses, facial expressions, changes in emotional state, and other non-verbal cues. Similarly, in human-machine interactions,…

2025

Bipedal Balance Control with Whole-body Musculoskeletal Standing and Falling Simulations

CoRL 2025poster

Balance control is important for human and bipedal robotic systems. While dynamic balance during locomotion has received considerable attention, quantitative understanding of static balance and falling remains limited. This work presents a hierarchical control pipeline for simulating human balance v…

Cited by 0SourceScholar
2025

Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream Knowledge

EMNLP 2025

Retrieval-augmented generation (RAG) aims to mitigate the hallucination of Large Language Models (LLMs) by retrieving and incorporating relevant external knowledge into the generation process. However, the external knowledge may contain noise and conflict with the parametric knowledge of LLMs, leadi

Cited by 0SourcePDFScholar
2025

COFlowNet: Conservative Constraints on Flows Enable High-Quality Candidate Generation

ICLR 2025poster

Generative flow networks (GFlowNets) have been considered as powerful tools for generating candidates with desired properties. Given that evaluating the property of candidates can be complex and time-consuming, existing GFlowNets train proxy models for efficient online evaluation. However, the perfo…

2025

Can We Steer Reasoning Direction by Thinking Intervention?

EMNLP 2025

Large Reason Models (LRMs) extend long reasoning process to solve complex tasks. However, due to the lack of fine-grained control, they often suffer from overthinking and erroneous reasoning problems, risking accuracy loss. To address this issue, we introduce Reasoning Direction Steering (RDS) to en

Cited by 0SourcePDFScholar
2025

ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective Primitive

NeurIPS 2025poster

Large language model (LLM) decoding suffers from high latency due to fragmented execution across operators and heavy reliance on off-chip memory for data exchange and reduction. This execution model limits opportunities for fusion and incurs significant memory traffic and kernel launch overhead. Wh…

Cited by 0SourceScholar
2025

Cross-Lingual Transfer of Cultural Knowledge: An Asymmetric Phenomenon

ACL 2025short

Despite substantial research efforts evaluating how well large language models (LLMs) handle global cultural diversity, the mechanisms behind their cultural knowledge acquisition, particularly in multilingual settings, remain unclear. We study this question by investigating how cultural knowledge tr…

2025

Cypher-RI: Reinforcement Learning for Integrating Schema Selection into Cypher Generation

NeurIPS 2025poster

The increasing utilization of graph databases across various fields stems from their capacity to represent intricate interconnections. Nonetheless, exploiting the full capabilities of graph databases continues to be a significant hurdle, largely because of the inherent difficulty in translating natu…

Cited by 0SourceScholar
2025

Debate on Graph: A Flexible and Reliable Reasoning Framework for Large Language Models

AAAI 2025technical

Large Language Models (LLMs) may suffer from hallucinations in real-world applications due to the lack of relevant knowledge. In contrast, knowledge graphs encompass extensive, multi-relational structures that store a vast array of symbolic facts. Consequently, integrating LLMs with knowledge graphs…

2025

Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

ICASSP 2025accepted

Current emotional text-to-speech (TTS) models pre-dominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending…

Cited by 0SourceScholar
2025

Enhancing Robustness of Implicit Neural Representations Against Weight Perturbations

ICASSP 2025accepted

Implicit Neural Representations (INRs) encode discrete signals in a continuous manner using neural networks, demonstrating significant value across various multimedia applications. However, the vulnerability of INRs presents a critical challenge for their real-world deployments, as the network weigh…

Cited by 0SourceScholar
2025

Exploring Intrinsic Alignments Within Text Corpus

AAAI 2025technical

Recent years have witnessed rapid advancements in the safety alignments of large language models (LLMs). Methods such as supervised instruction fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) have thus emerged as vital components in constructing LLMs. While these methods achi…

2025

FAME: Adaptive Functional Attention with Expert Routing for Function-on-Function Regression

NeurIPS 2025poster

Functional data play a pivotal role across science and engineering, yet their infinite-dimensional nature makes representation learning challenging. Conventional statistical models depend on pre-chosen basis expansions or kernels, limiting the flexibility of data-driven discovery, while many deep-le…

Cited by 0SourceScholar
2025

High-Fidelity Lightweight Mesh Reconstruction from Point Clouds

CVPR 2025highlight

Recently, learning signed distance functions (SDFs) from point clouds has become popular for reconstruction. To ensure accuracy, most methods require using high-resolution Marching Cubes for surface extraction. However, this results in redundant mesh elements, making the mesh inconvenient to use. To…

Cited by 0SourcePDFScholar
2025

LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture

EMNLP 2025

Expanding the long-context capabilities of Multi-modal Large Language Models (MLLMs) is critical for advancing video understanding and high-resolution image analysis. Achieving this requires systematic improvements in model architecture, data construction, and training strategies, particularly to ad

2025

MINR: Efficient Implicit Neural Representations for Multi-Image Encoding

ICASSP 2025accepted

Implicit Neural Representations (INRs) aim to parameterize discrete signals through implicit continuous functions. However, formulating each image with a separate neural network (typically, a Multi-Layer Perceptron (MLP)) leads to computational and storage inefficiencies when encoding multi-images.…

Cited by 0SourceScholar
2025

MiLiC-Eval: Benchmarking Multilingual LLMs for China’s Minority Languages

ACL 2025finding

Large language models (LLMs) excel in high-resource languages but struggle with low-resource languages (LRLs), particularly those spoken by minority communities in China, such as Tibetan, Uyghur, Kazakh, and Mongolian. To systematically track the progress in these languages, we introduce MiLiC-Eval,…

2025

MoDification: Mixture of Depths Made Easy

NAACL 2025long

Long-context efficiency has recently become a trending topic in serving large language models (LLMs). And mixture of depths (MoD) is proposed as a perfect fit to bring down both latency and memory. In this paper, however, we discover that MoD can barely transform existing LLMs without costly trainin…

Cited by 2SourcePDFScholar
2025

Multi-Robot Cooperative Transportation of Irregular Objects by Multi-Objective Optimization With Distributed Control

RA-L 2025

To enhance the efficiency of cooperative transportation by multiple mobile robots, we propose a transportation strategy based on multi-objective optimization of robot configurations. The system consists of multiple omnidirectional robots equipped with passively rotatable linkages, enabling the appli

Cited by 3SourceScholar
2025

MutationGuard: A Graph and Temporal-Spatial Neural Method for Detecting Mutation Telecommunication Fraud

IJCAI 2025

Telecommunication fraud refers to deceptive activities in the field of communication services. This research focuses on a category of fraud identified as ''mutation telecommunication fraud". There is currently a lack of research on mutation telecommunication fraud detection, allowing this type of fr

2025

N-ForGOT: Towards Not-forgetting and Generalization of Open Temporal Graph Learning

ICLR 2025poster

Temporal Graph Neural Networks (TGNNs) lay emphasis on capturing node interactions over time but often overlook evolution in node classes and dynamic data distributions triggered by the continuous emergence of new class labels, known as the open-set problem. This problem poses challenges for existin…

Cited by 0SourcePDFScholar
2025

Nonparametric Teaching for Graph Property Learners

ICML 2025spotlight

Inferring properties of graph-structured data, *e.g.*, the solubility of molecules, essentially involves learning the implicit mapping from graphs to their properties. This learning process is often costly for graph property learners like Graph Convolutional Networks (GCNs). To address this, we prop…

2025

PwnGPT: Automatic Exploit Generation Based on Large Language Models

ACL 2025long

Automatic exploit generation (AEG) refers to the automatic discovery and exploitation of vulnerabilities against unknown targets. Traditional AEG often targets a single type of vulnerability and still relies on templates built from expert experience. To achieve intelligent exploit generation, we est…

2025

Read it in Two Steps: Translating Extremely Low-Resource Languages with Code-Augmented Grammar Books

ACL 2025long

While large language models (LLMs) have shown promise in translating extremely low-resource languages using resources like dictionaries, the effectiveness of grammar books remains debated. This paper investigates the role of grammar books in translating extremely low-resource languages by decomposin…

2025

Towards the Law of Capacity Gap in Distilling Language Models

ACL 2025long

Language model (LM) distillation aims at distilling the knowledge in a large teacher LM to a small student one. As a critical issue facing LM distillation, a superior student often arises from a teacher of a relatively small scale instead of a larger one, especially in the presence of substantial ca…

2025

TraffiDent: A Dataset for Understanding the Interplay Between Traffic Dynamics and Incidents

NeurIPS 2025poster

Long-separated research has been conducted on two highly correlated tracks: traffic and incidents. Traffic track witnesses complicating deep learning models, e.g., to push the prediction a few percent more accurate, and the incident track only studies the incidents alone, e.g., to infer the incident…

Cited by 0SourcecodeScholar
2025

UCL-Bench: A Chinese User-Centric Legal Benchmark for Large Language Models

NAACL 2025findings

Existing legal benchmarks focusing on knowledge and logic effectively evaluate LLMs on various tasks in legal domain. However, few have explored the practical application of LLMs by actual users. To further assess whether LLMs meet the specific needs of legal practitioners in real-world scenarios, w…

2025

Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective

COLING 2025main

Enabling LLMs to handle lengthy context is currently a research hotspot. Most LLMs are built upon rotary position embedding (RoPE), a popular position encoding method. Therefore, a prominent path is to extrapolate the RoPE trained on comparably short texts to far longer texts. A heavy bunch of effor…

Cited by 7SourcePDFScholar
2025

ZigZagKV: Dynamic KV Cache Compression for Long-context Modeling based on Layer Uncertainty

COLING 2025main

Large Language models (LLMs) have become a research hotspot. To accelerate the inference of LLMs, storing computed caches in memory has become the standard technique. However, as the inference length increases, growing KV caches might lead to out-of-memory issues. Many existing methods address this…

Cited by 0SourcePDFScholar
2024

A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators

AAAI 2024technical

Automatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique, reference-free neural metrics that better align with human eva…

2024

Advancing DRL Agents in Commercial Fighting Games: Training, Integration, and Agent-Human Alignment

ICML 2024poster

Deep Reinforcement Learning (DRL) agents have demonstrated impressive success in a wide range of game genres. However, existing research primarily focuses on optimizing DRL competence rather than addressing the challenge of prolonged player interaction. In this paper, we propose a practical DRL agen…

Cited by 3SourcePDFScholar
2024

BAE-Net: a Low Complexity and High Fidelity Bandwidth-Adaptive Neural Network for Speech Super-Resolution

ICASSP 2024accepted

Speech bandwidth extension (BWE) has demonstrated promising performance in enhancing the perceptual speech quality in real communication systems. Most existing BWE researches primarily focus on fixed upsampling ratios, disregarding the fact that the effective bandwidth of captured audio may fluctuat…

Cited by 0SourceScholar
2024

Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models

EMNLP 2024finding

Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often involve processing multiple audio streams simultaneously. T…

2024

Beyond Single-Event Extraction: Towards Efficient Document-Level Multi-Event Argument Extraction

ACL 2024findings

Recent mainstream event argument extraction methods process each event in isolation, resulting in inefficient inference and ignoring the correlations among multiple events. To address these limitations, here we propose a multiple-event argument extraction model DEEIA (Dependency-guided Encoding and…

2024

CrossTune: Black-Box Few-Shot Classification with Label Enhancement

COLING 2024main

Training or finetuning large-scale language models (LLMs) requires substantial computation resources, motivating recent efforts to explore parameter-efficient adaptation to downstream tasks. One approach is to treat these models as black boxes and use forward passes (Inference APIs) to interact with…

Cited by 3SourcePDFScholar
2024

Deep Reinforcement Learning for Modelling Protein Complexes

ICLR 2024poster

Structure prediction of large protein complexes (a.k.a., protein multimer mod- elling, PMM) can be achieved through the one-by-one assembly using provided dimer structures and predicted docking paths. However, existing PMM methods struggle with vast search spaces and generalization challenges: (1) T…

Cited by 1SourcePDFScholar
2024

DynaThink: Fast or Slow? A Dynamic Decision-Making Framework for Large Language Models

EMNLP 2024main

Large language models (LLMs) have demonstrated emergent capabilities across diverse reasoning tasks via popular Chains-of-Thought (COT) prompting. However, such a simple and fast COT approach often encounters limitations in dealing with complicated problems, while a thorough method, which considers…

2024

Fast Graph Sharpness-Aware Minimization for Enhancing and Accelerating Few-Shot Node Classification

NeurIPS 2024poster

Graph Neural Networks (GNNs) have shown superior performance in node classification. However, GNNs perform poorly in the Few-Shot Node Classification (FSNC) task that requires robust generalization to make accurate predictions for unseen classes with limited labels. To tackle the challenge, we propo…

2024

FreqMark: Invisible Image Watermarking via Frequency Based Optimization in Latent Space

NeurIPS 2024poster

Invisible watermarking is essential for safeguarding digital content, enabling copyright protection and content authentication. However, existing watermarking methods fall short in robustness against regeneration attacks. In this paper, we propose a novel method called FreqMark that involves uncons…

Cited by 0SourcePDFScholar
2024

Harder Task Needs More Experts: Dynamic Routing in MoE Models

ACL 2024long

In this paper, we introduce a novel dynamic expert selection framework for Mixture of Experts (MoE) models, aiming to enhance computational efficiency and model performance by adjusting the number of activated experts based on input difficulty. Unlike existing MoE approaches that rely on fixed TopK…

2024

IINet: Implicit Intra-inter Information Fusion for Real-Time Stereo Matching

AAAI 2024technical

Recently, there has been a growing interest in 3D CNN-based stereo matching methods due to their remarkable accuracy. However, the high complexity of 3D convolution makes it challenging to strike a balance between accuracy and speed. Notably, explicit 3D volumes contain considerable redundancy. In t…

Cited by 8SourcePDFScholar
2024

Leveraging Frame Affinity for sRGB-to-RAW Video De-rendering

CVPR 2024poster

Unprocessed RAW video has shown distinct advantages over sRGB video in video editing and computer vision tasks. However capturing RAW video is challenging due to limitations in bandwidth and storage. Various methods have been proposed to address similar issues in single image RAW capture through de-…

Cited by 2SourcePDFScholar
2024

MC2: Towards Transparent and Culturally-Aware NLP for Minority Languages in China

ACL 2024long

Current large language models demonstrate deficiencies in understanding low-resource languages, particularly the minority languages in China. This limitation stems from the scarcity of available pre-training data. To address this accessibility challenge, we present MC2, a Multilingual Corpus of Mino…

2024

Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

ICLR 2024poster

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms of zero-shot TTS still face challenges in the following aspe…

2024

MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes

NeurIPS 2024poster

Talking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of appearance and talking style). While previous works typically…

2024

NondBREM: Nondeterministic Offline Reinforcement Learning for Large-Scale Order Dispatching

AAAI 2024technical

One of the most important tasks in ride-hailing is order dispatching, i.e., assigning unserved orders to available drivers. Recent order dispatching has achieved a significant improvement due to the advance of reinforcement learning, which has been approved to be able to effectively address sequenti…

Cited by 6SourcePDFScholar
2024

Nonparametric Teaching of Implicit Neural Representations

ICML 2024poster

We investigate the learning of implicit neural representation (INR) using an overparameterized multilayer perceptron (MLP) via a novel nonparametric teaching perspective. The latter offers an efficient example selection framework for teaching nonparametrically defined (viz. non-closed-form) target f…

2024

Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

ICLR 2024spotlight

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the goals of accurate 3D avatar reconstruction and stable talkin…

2024

TS-Align: A Teacher-Student Collaborative Framework for Scalable Iterative Finetuning of Large Language Models

EMNLP 2024finding

Mainstream approaches to aligning large language models (LLMs) heavily rely on human preference data, particularly when models require periodic updates. The standard process for iterative alignment of LLMs involves collecting new human feedback for each update. However, the data collection process i…

2024

Task-agnostic Distillation of Encoder-Decoder Language Models

COLING 2024main

Finetuning pretrained language models (LMs) have enabled appealing performance on a diverse array of tasks. The intriguing task-agnostic property has driven a shifted focus from task-specific to task-agnostic distillation of LMs. While task-agnostic, compute-efficient, performance-preserved LMs can…

Cited by 2SourcePDFScholar
2024

Teaching Large Language Models an Unseen Language on the Fly

ACL 2024findings

Existing large language models struggle to support numerous low-resource languages, particularly the extremely low-resource ones, for which there is minimal training data available for effective parameter updating. We thus investigate whether LLMs can learn a new language on the fly solely through p…

2024

Unlocking the Potential of Model Merging for Low-Resource Languages

EMNLP 2024finding

Adapting large language models (LLMs) to new languages typically involves continual pre-training (CT) followed by supervised fine-tuning (SFT). However, this CT-then-SFT approach struggles with limited data in the context of low-resource languages, failing to balance language modeling and task-solvi…

2024

Unveiling the Achilles’ Heel of NLG Evaluators: A Unified Adversarial Framework Driven by Large Language Models

ACL 2024findings

The automatic evaluation of natural language generation (NLG) systems presents a long-lasting challenge. Recent studies have highlighted various neural metrics that align well with human evaluations. Yet, the robustness of these evaluators against adversarial perturbations remains largely under-expl…

2023

A Low-Latency Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation

ICASSP 2023accepted

This paper describes our submission to the fourth Acoustic Echo Cancellation (AEC) Challenge, which is part of ICASSP 2023 Signal Processing Grand Challenge. The proposed system is developed based on our earlier system submitted to the ICASSP 2022 AEC challenge with significant latency and network s…

Cited by 3SourceScholar
2023

Exploiting Completeness and Uncertainty of Pseudo Labels for Weakly Supervised Video Anomaly Detection

CVPR 2023poster

Weakly supervised video anomaly detection aims to identify abnormal events in videos using only video-level labels. Recently, two-stage self-training methods have achieved significant improvements by self-generating pseudo labels and self-refining anomaly scores with these labels. As the pseudo labe…

Cited by 92SourcePDFScholar
2023

How Many Answers Should I Give? An Empirical Study of Multi-Answer Reading Comprehension

ACL 2023findings

The multi-answer phenomenon, where a question may have multiple answers scattered in the document, can be well handled by humans but is challenging enough for machine reading comprehension (MRC) systems. Despite recent progress in multi-answer MRC, there lacks a systematic analysis of how this pheno…

2023

LeanSpeech: The Microsoft Lightweight Speech Synthesis System for Limmits Challenge 2023

ICASSP 2023accepted

This paper describes the Microsoft Text-to-Speech (TTS) system: LeanSpeech for LIMMITS (Lightweight, Multi-speaker, Multi-lingual Indic TTS) Challenge 2023<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>, which is part of ICASSP2023 to encourage…

Cited by 0SourceScholar
2023

Lifting the Curse of Capacity Gap in Distilling Language Models

ACL 2023long

Pretrained language models (LMs) have shown compelling performance on various downstream tasks, but unfortunately they require a tremendous amount of inference compute. Knowledge distillation finds a path to compress LMs to small ones with a teacher-student paradigm. However, when the capacity gap b…

2023

Nonparametric Teaching for Multiple Learners

NeurIPS 2023poster

We study the problem of teaching multiple learners simultaneously in the nonparametric iterative teaching setting, where the teacher iteratively provides examples to the learner for accelerating the acquisition of a target concept. This problem is motivated by the gap between current single-learner…

2023

Relation-Aware Question Answering for Heterogeneous Knowledge Graphs

EMNLP 2023long findings

Multi-hop Knowledge Base Question Answering(KBQA) aims to find the answer entity in a knowledge graph (KG), which requires multiple steps of reasoning. Existing retrieval-based approaches solve this task by concentrating on the specific relation at different hops and predicting the intermediate enti…

Cited by 0SourcecodeScholar
2023

SMART-Degradation: A Dataset for LiDAR Degradation Evaluation in Rain

IROS 2023poster

Sensor degradation is one of the major challenges for autonomous driving. During the rain, the interference from raindrops can negatively influence LiDAR measurements. For example, valid measurements could be reduced during the rain, and some measurements may become noisy. Unreliable measurements ca…

Cited by 3SourcecodeScholar
2023

SMART-Rain: A Degradation Evaluation Dataset for Autonomous Driving in Rain

IROS 2023poster

Autonomous driving in the rain remains a challenge. One main problem is performance degradation caused by rain. This work introduces a new dataset to study this problem. Our dataset is collected from a full-scale vehicle equipped with a 3D LiDAR sensor and multiple forward-facing cameras under vario…

Cited by 5SourcecodeScholar
2023

SmartRainNet: Uncertainty Estimation For Laser Measurement in Rain

ICRA 2023poster

Adverse weather has raised a big challenge for autonomous vehicles. Unreliable measurements due to sensor degradation could seriously affect the performance of autonomous driving tasks, such as perception and localization. In this work, we study sensor degradation in rainy weather and present a nove…

Cited by 5SourceScholar
2023

The Magic of IF: Investigating Causal Reasoning Abilities in Large Language Models of Code

ACL 2023findings

Causal reasoning, the ability to identify cause-and-effect relationship, is crucial in human thinking. Although large language models (LLMs) succeed in many NLP tasks, it is still challenging for them to conduct complex causal reasoning like abductive reasoning and counterfactual reasoning. Given th…

2023

Underwater and Surface Aquatic Locomotion of Soft Biomimetic Robot Based on Bending Rolled Dielectric Elastomer Actuators

IROS 2023poster

All-around, real-time navigation and sensing across the water environments by miniature soft robotics are promising, for their merits of small size, high agility and good compliance to the unstructured surroundings. In this paper, we propose and demonstrate a mantas-like soft aquatic robot which pro…

Cited by 5SourceScholar
2023

VideoDubber: Machine Translation with Speech-Aware Length Control for Video Dubbing

AAAI 2023technical

Video dubbing aims to translate the original speech in a film or television program into the speech in a target language, which can be achieved with a cascaded system consisting of speech recognition, machine translation and speech synthesis. To ensure the translated speech to be well aligned with t…

2023

xDial-Eval: A Multilingual Open-Domain Dialogue Evaluation Benchmark

EMNLP 2023long findings

Recent advancements in reference-free learned metrics for open-domain dialogue evaluation have been driven by the progress in pre-trained language models and the availability of dialogue data with high-quality human annotations. However, current studies predominantly concentrate on English dialogues…

Cited by 0SourcecodeScholar
2022

A Deep Hierarchical Fusion Network for Fullband Acoustic Echo Cancellation

ICASSP 2022accepted

Deep learning based wideband (16kHz) acoustic echo cancellation (AEC) approaches have surpassed traditional methods. This work proposes a deep hierarchical fusion (DHF) network with intra-network and inter-network fusion to further improve the wideband AEC performance. Meanwhile, this work extends t…

Cited by 0SourceScholar
2022

A Two-Step Backward Compatible Fullband Speech Enhancement System

ICASSP 2022accepted

Speech enhancement methods based on deep learning have surpassed traditional methods. While many of these new approaches are operating on the wideband (16kHz) sample rate, a new fullband (48kHz) speech enhancement system is proposed in this paper. Compared to the existing full-band systems that util…

Cited by 0SourceScholar
2022

Analyzing and Evaluating Faithfulness in Dialogue Summarization

EMNLP 2022main

Dialogue summarization is abstractive in nature, making it suffer from factual errors. The factual correctness of summaries has the highest priority before practical applications. Many efforts have been made to improve faithfulness in text summarization. However, there is a lack of systematic study…

2022

Automatic Song Translation for Tonal Languages

ACL 2022findings

This paper develops automatic song translation (AST) for tonal languages and addresses the unique challenge of aligning words’ tones with melody of a song in addition to conveying the original meaning. We propose three criteria for effective AST—preserving meaning, singability and intelligibility—an…

Cited by 16SourcePDFScholar
2022

FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation

EMNLP 2022main

Recent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment. However, they either perform turn-level evaluation or look at a single dialogue quality dimension. One would expect a good evaluation metric to assess multiple quality di…

2022

GRELEN: Multivariate Time Series Anomaly Detection from the Perspective of Graph Relational Learning

IJCAI 2022poster

System monitoring and anomaly detection is a crucial task in daily operation. With the rapid development of cyber-physical systems and IT systems, multiple sensors get involved to represent the system state from different perspectives, which inspires us to detect anomalies considering feature depend…

Cited by 78SourcePDFScholar
2022

L3DAS22 Challenge: Learning 3D Audio Sources in a Real Office Environment

ICASSP 2022accepted

The L3DAS22 Challenge is aimed at encouraging the development of machine learning strategies for 3D speech enhancement and 3D sound localization and detection in office-like environments. This challenge improves and extends the tasks of the L3DAS21 edition <sup xmlns:mml="http://www.w3.org/1998/Math…

Cited by 62SourceScholar
2022

MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue Evaluation

AAAI 2022technical

Chatbots are designed to carry out human-like conversations across different domains, such as general chit-chat, knowledge exchange, and persona-grounded conversations. To measure the quality of such conversational agents, a dialogue evaluator is expected to conduct assessment across domains as well…

2022

Making Pretrained Language Models Good Long-tailed Learners

EMNLP 2022main

Prompt-tuning has shown appealing performance in few-shot classification by virtue of its capability in effectively exploiting pre-trained knowledge. This motivates us to check the hypothesis that prompt-tuning is also a promising choice for long-tailed classification, since the tail classes are int…

2022

Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement

ICASSP 2022accepted

Deep learning-based wideband (16kHz) speech enhancement approaches have surpassed traditional methods. This work further extends the existing wideband systems to enable full-band (48kHz) speech enhancement while simultaneously ensuring automatic speech recognition compatibility and optionally, perso…

Cited by 0SourceScholar
2022

S3T: Self-Supervised Pre-Training with Swin Transformer For Music Classification

ICASSP 2022accepted

In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive easily accessible unlabeled music data. S3T introduces a momentum-based paradigm, MoCo, with Swin Transformer as its feat…

Cited by 0SourceScholar
2022

SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian Approximation

ICLR 2022poster

Quantization of deep neural networks (DNN) has been proven effective for compressing and accelerating DNN models. Data-free quantization (DFQ) is a promising approach without the original datasets under privacy-sensitive and confidential scenarios. However, current DFQ solutions degrade accuracy, ne…

2022

Structural Bias for Aspect Sentiment Triplet Extraction

COLING 2022main

Structural bias has recently been exploited for aspect sentiment triplet extraction (ASTE) and led to improved performance. On the other hand, it is recognized that explicitly incorporating structural bias would have a negative impact on efficiency, whereas pretrained language models (PLMs) can alre…

2022

TeleMelody: Lyric-to-Melody Generation with a Template-Based Two-Stage Method

EMNLP 2022main

Lyric-to-melody generation is an important task in automatic songwriting. Previous lyric-to-melody generation systems usually adopt end-to-end models that directly generate melodies from lyrics, which suffer from several issues: 1) lack of paired lyric-melody training data; 2) lack of control on gen…

2022

Towards Effective Multi-Modal Interchanges in Zero-Resource Sounding Object Localization

NeurIPS 2022accept

Aiming to locate the object that emits a specified sound in complex scenes, the task of sounding object localization bridges two perception-oriented modalities of vision and acoustics, and brings enormous research value to the comprehensive perceptual understanding of machine intelligence. Although…

Cited by 8SourcePDFScholar
2022

XPrompt: Exploring the Extreme of Prompt Tuning

EMNLP 2022main

Prompt tuning learns soft prompts to condition the frozen Pre-trained Language Models (PLMs) for performing downstream tasks in a parameter-efficient manner. While prompt tuning has gradually reached the performance level of fine-tuning as the model scale increases, there is still a large performanc…

Cited by 39SourcePDFScholar
2021

CARE: Commonsense-Aware Emotional Response Generation with Latent Concepts

AAAI 2021technical

Rationality and emotion are two fundamental elements of humans. Endowing agents with rationality and emotion has been one of the major milestones in AI. However, in the field of conversational AI, most existing models only specialize in one aspect and neglect the other, which often leads to dull or…

2021

Deep Imitation Learning for Autonomous Navigation in Dynamic Pedestrian Environments

ICRA 2021poster

Navigation through dynamic pedestrian environments in a socially compliant manner is still a challenging task for autonomous vehicles. Classical methods usually lead to unnatural vehicle behaviours for pedestrian navigation due to the difficulty in modeling social conventions mathematically. This pa…

Cited by 19SourceScholar
2021

Denoispeech: Denoising Text to Speech with Frame-Level Noise Modeling

ICASSP 2021accepted

While neural-based text to speech (TTS) models can synthesize natural and intelligible voice, they usually require high-quality speech data, which is costly to collect. In many scenarios, only noisy speech of a target speaker is available, which presents challenges for TTS model training for this sp…

Cited by 0SourceScholar
2021

DynaEval: Unifying Turn and Dialogue Level Evaluation

ACL 2021long

A dialogue is essentially a multi-turn interaction among interlocutors. Effective evaluation metrics should reflect the dynamics of such interaction. Existing automatic metrics are focused very much on the turn-level quality, while ignoring such dynamics. To this end, we propose DynaEval, a unified…

2021

Extract, Integrate, Compete: Towards Verification Style Reading Comprehension

EMNLP 2021finding

In this paper, we present a new verification style reading comprehension dataset named VGaokao from Chinese Language tests of Gaokao. Different from existing efforts, the new dataset is originally designed for native speakers’ evaluation, thus requiring more advanced language understanding skills. T…

2021

Interactive Video Acquisition and Learning System for Motor Assessment of Parkinson's Disease

IJCAI 2021poster

Diagnosis and treatment for Parkinson's disease rely on the evaluation of motor functions, which is expensive and time consuming when performing at clinics. It is also difficult for patients to record correct movements at home without the guidance from experienced physicians. To help patients with P…

Cited by 6SourcePDFScholar
2021

OSOA: One-Shot Online Adaptation of Deep Generative Models for Lossless Compression

NeurIPS 2021poster

Explicit deep generative models (DGMs), e.g., VAEs and Normalizing Flows, have shown to offer an effective data modelling alternative for lossless compression. However, DGMs themselves normally require large storage space and thus contaminate the advantage brought by accurate data density estimatio…

Cited by 5SourcePDFScholar
2021

Revisiting Self-training for Few-shot Learning of Language Model

EMNLP 2021main

As unlabeled data carry rich task-relevant information, they are proven useful for few-shot learning of language model. The question is how to effectively make use of such data. In this work, we revisit the self-training technique for language model fine-tuning and present a state-of-the-art prompt-…

2021

UWSpeech: Speech to Speech Translation for Unwritten Languages

AAAI 2021technical

Existing speech to speech translation systems heavily rely on the text of target language: they usually translate source language either to target text and then synthesize target speech from text, or directly to target speech with target text for auxiliary training. However, those methods cannot be…

2021

iVPF: Numerical Invertible Volume Preserving Flow for Efficient Lossless Compression

CVPR 2021poster

It is nontrivial to store rapidly growing big data nowadays, which demands high-performance lossless compression techniques. Likelihood-based generative models have witnessed their success on lossless compression, where flow based models are desirable in allowing exact data likelihood optimisation w…

Cited by 46PDFScholar
2020

LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression

COLING 2020main

BERT is a cutting-edge language representation model pre-trained by a large corpus, which achieves superior performances on various natural language understanding tasks. However, a major blocking issue of applying BERT to online services is that it is memory-intensive and leads to unsatisfactory lat…

2020

Task-Level Curriculum Learning for Non-Autoregressive Neural Machine Translation

IJCAI 2020poster

Non-autoregressive translation (NAT) achieves faster inference speed but at the cost of worse accuracy compared with autoregressive translation (AT). Since AT and NAT can share model structure and AT is an easier task than NAT due to the explicit dependency on previous target-side tokens, a natural…

2019

SeerNet: Predicting Convolutional Neural Network Feature-Map Sparsity Through Low-Bit Quantization

CVPR 2019poster

In this paper we present a novel and general method to accelerate convolutional neural network (CNN) inference by taking advantage of feature map sparsity. We experimentally demonstrate that a highly quantized version of the original network is sufficient in predicting the output sparsity accurately…

Cited by 100PDFScholar
2018

Vehicle Detection, Tracking and Behavior Analysis in Urban Driving Environments Using Road Context

ICRA 2018poster

We present a real-time vehicle detection and tracking system to accomplish the complex task of driving behavior analysis in urban environments. We propose a robust fusion system that combines a monocular camera and a 2D Lidar. This system takes advantage of three key components: robust vehicle detec…

Cited by 25SourceScholar
2015

Disparity-compensated total-variation minimization for compressed-sensed multiview image reconstruction

ICASSP 2015accepted

Compressed sensing (CS) is the theory and practice of sub-Nyquist sampling of sparse signals of interest. Perfect reconstruction may then be possible with much fewer than the Nyquist required number of data. In this paper, we consider a distributed multi-view imaging system where each camera at a di…

Cited by 0SourceScholar