← Search

Jing Chen

46 accepted papers

2026

Decentralized Bandits without Global Clock for Dynamic Matching Market

ICML 2026poster

Two-sided matching markets are pervasive in numerous real-world applications, ranging from labor markets to online advertising. Recently, a rich line of research has studied the matching bandit problem, where participants learn their preferences through iterative interactions. However, existing work…

Cited by 0SourceScholar
2026

DeepBooTS: Dual-Stream Residual Boosting for Drift-Resilient Time-Series Forecasting

AAAI 2026technical

Time-Series (TS) exhibits pronounced non-stationarity. Consequently, most forecasting methods display compromised robustness to concept drift, despite the prevalent application of instance normalization. We tackle this challenge by first analysing concept drift through a bias-variance lens and provi

Cited by 0SourcePDFScholar
2026

GeoProblem Factory: A Visual Interaction System for Solvable and Controllable Geometric Problem Generation by Leveraging Symbolic Deduction Engine

AAAI 2026technical

We propose a novel system, GeoProblem Factory, designed to effectively generate high-quality geometry problems for intelligent education. The system enables to efficiently produce batches of geometry problems for teachers and students, either to save time and manual effort or to support personalized

Cited by 0SourcePDFScholar
2026

Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis

AAAI 2026technical

Large Language Models (LLMs) excel in reasoning and generation across domains, but still struggle with identifying and diagnosing complex errors. This stems mainly from training objectives that prioritize correct answers, limiting exposure to and learning from errors. While recent studies have begun

Cited by 0SourcePDFScholar
2026

The Forecast After the Forecast: A Post-Processing Shift in Time Series

ICLR 2026poster

Time series forecasting has long been dominated by advances in model architecture, with recent progress driven by deep learning and hybrid statistical techniques. However, as forecasting models approach diminishing returns in accuracy, a critical yet underexplored opportunity emerges: the strategic…

Cited by 0SourcecodeScholar
2025

A Novel Multimodal Method for Decoding Speech Perception from Brain Activities

ICASSP 2025accepted

Decoding speech from neural recordings has critical importance in application and scientific research. However, this task is still challenging with non-invasive recordings. Previous research has shown significant improvement in speech perception decoding task by leveraging wav2vec vectors and gives…

Cited by 0SourceScholar
2025

Adversarial Training and Cross-modal Feature Fusion in Multimodal Sentiment Analysis

ICASSP 2025accepted

Multimodal sentiment analysis recognizes emotions through text, audio, and visual modalities, but data incompleteness is a major challenge. Existing methods often focus on specific types of deficiencies and perform poorly when multiple types of noise are present simultaneously. To address this issue…

Cited by 0SourceScholar
2025

CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models

NAACL 2025findings

Challenges in managing linguistic diversity and integrating various musical modalities are faced by current music information retrieval systems. These limitations reduce their effectiveness in a global, multimodal music environment. To address these issues, we introduce CLaMP 2, a system compatible…

2025

FedFLD: Heterogeneous Federated Learning via Forget-Less Distillation

ICASSP 2025accepted

Federated learning, as a distributed machine learning paradigm, enhances privacy protection but faces the challenge of heterogeneity. Data-free knowledge distillation (DFKD) methods attempt to overcome this challenge by using a generator to synthesize samples for fine-tuning a global model. However,…

Cited by 0SourceScholar
2025

GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification

NeurIPS 2025poster

Aerial-Ground person re-identification (AG-ReID) is an emerging yet challenging task that aims to match pedestrian images captured from drastically different viewpoints, typically from unmanned aerial vehicles (UAVs) and ground-based surveillance cameras. The task poses significant challenges due to…

Cited by 0SourceScholar
2025

Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question Answering

CVPR 2025poster

The knowledge-based visual question answering (KB-VQA) task involves using external knowledge about the image to assist reasoning. Building on the impressive performance of multimodal large language model (MLLM), recent methods have commenced leveraging MLLM as an implicit knowledge base for reasoni…

Cited by 0SourcePDFScholar
2025

Pre-trained Semantic Interaction based Inductive Graph Neural Networks for Text Classification

COLING 2025main

Nowadays, research of Text Classification (TC) based on graph neural networks (GNNs) is on the rise. Both inductive methods and transductive methods have made significant progress. For transductive methods, the semantic interaction between texts plays a crucial role in the learning of effective text…

2025

SEI3D: CPU-only 3D Object Tracking Fusing Sparse-flow-filtered Edge and Interior Alignment

IROS 2025

Monocular 3D object tracking methods are widely employed in robotic applications, however, they often struggle with low-contrast image sequences. In this paper, we introduce a novel approach to filtering redundant edges in images by leveraging sparse interior correspondences. Our method features a s

Cited by 0SourceScholar
2025

SFADNet: Spatio-temporal Fused Graph based on Attention Decoupling Network for Traffic Prediction

ICASSP 2025accepted

In recent years, traffic flow prediction has played a crucial role in the management of intelligent transportation systems. However, traditional prediction methods are often limited by static spatial modeling, making it difficult to accurately capture the dynamic and complex relationships between ti…

Cited by 0SourceScholar
2025

Using Ear-EEG to Decode Auditory Attention in Multiple-speaker Environment

ICASSP 2025accepted

Auditory Attention Decoding (AAD) can help to determine the identity of the attended speaker during an auditory selective attention task, by analyzing and processing measurements of electroencephalography (EEG) data. Most studies on AAD are based on scalp-EEG signals in two-speaker scenarios, which…

Cited by 7SourceScholar
2025

Which Tasks Should Be Compressed Together? A Causal Discovery Approach for Efficient Multi-Task Representation Compression

ICLR 2025poster

Conventional image compression methods are inadequate for intelligent analysis, as they overemphasize pixel-level precision while neglecting semantic significance and the interaction among multiple tasks. This paper introduces a Taskonomy-Aware Multi-Task Compression framework comprising (1) inter-…

Cited by 0SourcePDFScholar
2024

A DenseNet-Based Method for Decoding Auditory Spatial Attention with EEG

ICASSP 2024accepted

Auditory spatial attention detection (ASAD) aims to decode the attended spatial location with EEG in a multiple-speaker setting. ASAD methods are inspired by the brain lateralization of cortical neural responses during the processing of auditory spatial attention, and show promising performance for…

Cited by 0SourceScholar
2024

A Fine-Grained Tri-Modal Interaction Model for Multimodal Sentiment Analysis

ICASSP 2024accepted

The methods based on multimodal representation learning enhance discriminable sentiment expression for multimodal sentiment analysis(MSA). The modal invariant and specific features serve different purposes in sentiment learning and the diversity of inter-sample and inter-category relationships takes…

Cited by 0SourceScholar
2024

HoLLMwood: Unleashing the Creativity of Large Language Models in Screenwriting via Role Playing

EMNLP 2024finding

Generative AI has demonstrated unprecedented creativity in the field of computer vision, yet such phenomena have not been observed in natural language processing. In particular, large language models (LLMs) can hardly produce written works at the level of human experts due to the extremely high comp…

Cited by 7SourcePDFScholar
2024

Semantic Reconstruction of Continuous Language from Meg Signals

ICASSP 2024accepted

Decoding language from neural signals holds considerable theoretical and practical importance. Previous research has indicated the feasibility of decoding text or speech from invasive neural signals. However, when using non-invasive neural signals, significant challenges are encountered due to their…

Cited by 0SourceScholar
2024

ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language Models

EMNLP 2024main

Tool-augmented large language models (LLMs) are rapidly being integrated into real-world applications. Due to the lack of benchmarks, the community has yet to fully understand the hallucination issues within these models. To address this challenge, we introduce a comprehensive diagnostic benchmark,…

2024

Wavelet-Decoupling Contrastive Enhancement Network for Fine-Grained Skeleton-Based Action Recognition

ICASSP 2024accepted

Skeleton-based action recognition has attracted much attention, benefiting from its succinctness and robustness. However, the minimal inter-class variation in similar action sequences often leads to confusion. The inherent spatiotemporal coupling characteristics make it challenging to mine the subtl…

Cited by 0SourceScholar
2023

A Model-Based Hearing Compensation Method Using a Self-Supervised Framework

ICASSP 2023accepted

Hearing aids can improve auditory perception for hearing-impaired (HI) listeners, but even state-of-art devices provide only limited benefits if not configured correctly for the listeners. The prescriptive fittings of hearing aids ignore the individual difference among HI listeners with identical he…

Cited by 0SourceScholar
2023

What can Discriminator do? Towards Box-free Ownership Verification of Generative Adversarial Networks

ICCV 2023poster

In recent decades, Generative Adversarial Network (GAN) and its variants have achieved unprecedented success in image synthesis. However, well-trained GANs are under the threat of illegal steal or leakage. The prior studies on remote ownership verification assume a black-box setting where the defend…

Cited by 15PDFcodeScholar
2022

Anti-Forgery: Towards a Stealthy and Robust DeepFake Disruption Attack via Adversarial Perceptual-aware Perturbations

IJCAI 2022poster

DeepFake is becoming a real risk to society and brings potential threats to both individual privacy and political security due to the DeepFaked multimedia are realistic and convincing. However, the popular DeepFake passive detection is an ex-post forensics countermeasure and failed in blocking the d…

2022

GBA: A Tuning-free Approach to Switch between Synchronous and Asynchronous Training for Recommendation Models

NeurIPS 2022accept

High-concurrency asynchronous training upon parameter server (PS) architecture and high-performance synchronous training upon all-reduce (AR) architecture are the most commonly deployed distributed training modes for recommendation models. Although synchronous AR training is designed to have higher…

Cited by 3SourcePDFScholar
2020

Efficient Uncertainty-aware Decision-making for Automated Driving Using Guided Branching

ICRA 2020poster

Decision-making in dense traffic scenarios is challenging for automated vehicles (AVs) due to potentially stochastic behaviors of other traffic participants and perception uncertainties (e.g., tracking noise and prediction errors, etc.). Although the partially observable Markov decision process (POM…

Cited by 60SourcecodeScholar
2020

Learning Event-Driven Video Deblurring and Interpolation

ECCV 2020poster

Event-based sensors, which have a response if the change of pixel intensity exceeds a triggering threshold, can capture high-speed motion with microsecond accuracy. Assisted by an event camera, we can generate high frame-rate sharp videos from low frame-rate blurry ones captured by an intensity came…

Cited by 154SourcePDFScholar
2020

Single-Channel Speech Separation Integrating Pitch Information Based on a Multi Task Learning Framework

ICASSP 2020accepted

Pitch is a critical cue for speech separation in humans' auditory perception. Although the technology of tracking pitch in single-talker speech succeeds in many applications, it's still a challenging problem to extract pitch information from speech mixtures in machine perception. In this paper, we a…

Cited by 0SourceScholar
2019

Deep Surface Normal Estimation With Hierarchical RGB-D Fusion

CVPR 2019poster

The growing availability of commodity RGB-D cameras has boosted the applications in the field of scene understanding. However, as a fundamental scene understanding task, surface normal estimation from RGB-D data lacks thorough investigation. In this paper, a hierarchical fusion network with adaptive…

Cited by 86PDFcodeScholar
2019

Integrating Spectrotemporal Context into Features Based on Auditory Perception for Classification-based Speech Separation

ICASSP 2019accepted

Speech separation, which has been a challenging task for decades, especially at low signal-to-noise ratios (SNRs), can be cast as a classification problem. In such adverse acoustic environment, extracting robust features from noisy mixtures is crucial for successful classification. In the past studi…

Cited by 0SourceScholar
2019

Predicting Vehicle Behaviors Over An Extended Horizon Using Behavior Interaction Network

ICRA 2019poster

Anticipating possible behaviors of traffic participants is an essential capability of autonomous vehicles. Many behavior detection and maneuver recognition methods only have a very limited prediction horizon that leaves inadequate time and space for planning. To avoid unsatisfactory reactive decisio…

Cited by 120SourceScholar
2019

Safe Trajectory Generation for Complex Urban Environments Using Spatio-Temporal Semantic Corridor

RA-L 2019

Planning safe trajectories for autonomous vehicles in complex urban environments is challenging since there are numerous semantic elements (such as dynamic agents, traffic lights, and speed limits) to consider. These semantic elements may have different mathematical descriptions, such as obstacle, c

Cited by 134SourcecodeScholar
2018

A Time-Weighted Method for Predicting the Intelligibility of Speech in the Presence of Interfering Sounds

ICASSP 2018accepted

The speech intelligibility index (SII) has been widely used as an objective method of predicting speech intelligibility, but its traditional form is most effective predicting speech intelligibility scores under stationary noise but not more challenging conditions (e.g., competing noise interference)…

Cited by 0SourceScholar
2017

Improving octree-based occupancy maps using environment sparsity with application to aerial robot navigation

ICRA 2017poster

In this paper, we present an improved octree-based mapping framework for autonomous navigation of mobile robots. Octree is best known for its memory efficiency for representing large-scale environments. However, existing implementations, including the state-of-the-art OctoMap [1], are computationall…

Cited by 27SourceScholar
2016

Online generation of collision-free trajectories for quadrotor flight in unknown cluttered environments

ICRA 2016

We present an online method for generating collision-free trajectories for autonomous quadrotor flight through cluttered environments. We consider the real-world scenario that the quadrotor aerial robot is equipped with limited sensing and operates in initially unknown environments. During flight, a

Cited by 211SourceScholar