← Search

Chen Yang

41 accepted papers

2026

Decentralized Attention Fails Centralized Signals: Rethinking Transformers for Medical Time Series

ICLR 2026oral

Accurate analysis of Medical time series (MedTS) data, such as Electroencephalography (EEG) and Electrocardiography (ECG), plays a pivotal role in healthcare applications, including the diagnosis of brain and heart diseases. MedTS data typically exhibits two critical patterns: **temporal dependencie…

Cited by 0SourcecodeScholar
2026

Deep Research Arena: The First Exam of LLMs’ Research Abilities via Seminar-Grounded Tasks

AAAI 2026technical

Deep research agents have attracted growing attention for their potential to orchestrate multi-stage research workflows, spanning literature synthesis, methodological design, and empirical verification. Despite these strides, evaluating their research capability faithfully is rather challenging due

Cited by 0SourcePDFScholar
2026

Dereflection Any Image with Diffusion Priors and Diversified Data

AAAI 2026technical

Reflection removal of a single image remains a highly challenging task due to the complex entanglement between target scenes and unwanted reflections. Despite significant progress, existing methods are hindered by the scarcity of high-quality, diverse data and insufficient restoration priors, result

Cited by 0SourcePDFScholar
2026

Escaping the CAM Shadow: Uncertainty-Guided Reliable Learning for Weakly Supervised Semantic Segmentation

AAAI 2026technical

Weakly supervised semantic segmentation (WSSS) suffers from an inherent mismatch between coarse image-level annotations and dense pixel-level predictions. To bridge this gap, existing methods primarily focus on generating refined class activation maps (CAM) as pseudo-labels. However, we argue that t

Cited by 0SourcePDFScholar
2026

Few-step Flow for 3D Generation via Marginal-Data Transport Distillation

AAAI 2026technical

Flow-based 3D generation models typically require dozens of sampling steps during inference. Though few-step distillation methods, particularly Consistency Models (CMs), have achieved substantial advancements in accelerating 2D diffusion models, they remain under-explored for more complex 3D generat

Cited by 0SourcePDFScholar
2026

Learning from Human Gaze: Human-like Robot Social Navigation in Dense Crowds

AAAI 2026technical

Robot navigation in dense crowds requires understanding social cues that humans naturally use, yet existing methods struggle with real-world complexity. We investigate two questions: (1) Where do pedestrians look when navigating crowds? and (2) Can eye tracking improve robot navigation? To answer, w

Cited by 0SourcePDFScholar
2026

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

ICASSP 2026poster

Mainstream Automatic Speech Recognition (ASR) systems excel at transcribing lexical content, but largely fail to recognize nonverbal vocalizations (NVs) embedded in speech, such as sighs, laughs, and coughs. This capability is important for a comprehensive understanding of human communication, as NV…

Cited by 0SourcePDFScholar
2026

Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation

ICML 2026poster

Reverse Chain-of-Thought Generation (RCG) synthesizes reasoning traces from query-answer pairs, but runs the risk of producing post-hoc rationalizations: when models can see the answer during generation, the answer serves as a cognitive anchor that shapes the entire explanation. We formalize this ph…

Cited by 0SourceScholar
2026

SMAP: Semantic Route Planning with Map-Grounded Multimodal Alignment

CVPR 2026

Semantic route planning involves generating itineraries that align with user intent while respecting real-world spatial constraints. However, text-only large language models (LLMs) often hallucinate geographically implausible routes due to poor spatial grounding. Inspired by how humans use maps for

Cited by 0SourcecodeScholar
2026

Thinking in Structures: Evaluating Spatial Intelligence through Reasoning on Constrained Manifolds

ICML 2026poster

Spatial intelligence is crucial for vision--language models (VLMs) in the physical world, yet many benchmarks evaluate largely unconstrained scenes where models can exploit 2D shortcuts. We introduce SSI-Bench, a VQA benchmark for spatial reasoning on constrained manifolds, built from complex real-w…

Cited by 0SourceScholar
2025

Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing

NeurIPS 2025poster

Large vision-language models (LVLMs) derive their capabilities from extensive training on vast corpora of visual and textual data. Empowered by large-scale parameters, these models often exhibit strong memorization of their training data, rendering them susceptible to membership inference attacks (…

Cited by 0SourcecodeScholar
2025

Geometric Imbalance in Semi-Supervised Node Classification

NeurIPS 2025poster

Class imbalance in graph data presents a significant challenge for effective node classification, particularly in semi-supervised scenarios. In this work, we formally introduce the concept of geometric imbalance, which captures how message passing on class-imbalanced graphs leads to geometric ambigu…

Cited by 0SourceScholar
2025

Latent Imputation before Prediction: A New Computational Paradigm for De Novo Peptide Sequencing

ICML 2025poster

*De novo* peptide sequencing is a fundamental computational technique for ascertaining amino acid sequences of peptides directly from tandem mass spectrometry data, eliminating the need for reference databases. Cutting-edge models encode the observed mass spectra into latent representations from whi…

2025

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

NeurIPS 2025poster

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through ite…

Cited by 0SourcecodeScholar
2025

MncCap: Mining Neural Composition for Zero-shot Image Captioning via Text-only Training

ICASSP 2025accepted

Current text-only image captioning methods leverage the shared feature space of CLIP to train zero-shot image captioning using text data only, leaving feature associations and contextual understanding not fully explored. Neurological studies have revealed that the anterior temporal lobes of the brai…

Cited by 0SourceScholar
2025

Provable Zero-Shot Generalization in Offline Reinforcement Learning

ICML 2025poster

In this work, we study offline reinforcement learning (RL) with zero-shot generalization property (ZSG), where the agent has access to an offline dataset including experiences from different environments, and the goal of the agent is to train a policy over the training environments which performs we…

Cited by 0SourcePDFScholar
2025

RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging

EMNLP 2025

We unveil that internal representations in large language models (LLMs) serve as reliable proxies of learned knowledge, and propose **RECALL**, a novel representation-aware model merging framework for continual learning without access to historical data. RECALL computes inter-model similarity from l

2025

Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation

UAI 2025

Continuous-time reinforcement learning (CTRL) provides a principled framework for sequential decision-making in environments where interactions evolve continuously over time. Despite its empirical success, the theoretical understanding of CTRL remains limited, especially in settings with general fun

2025

Segment Any 3D Gaussians

AAAI 2025technical

This paper presents SAGA (Segment Any 3D GAussians), a highly efficient 3D promptable segmentation method based on 3D Gaussian Splatting (3D-GS). Given 2D visual prompts as input, SAGA can segment the corresponding 3D target represented by 3D Gaussians within 4 ms. This is achieved by attaching a sc…

2025

URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

EMNLP 2025

Recent advances in large language models (LLMs) have driven significant progress in end-to-end spoken dialogue models (SDMs). In contrast to text-based LLMs, the evaluation framework for SDMs should encompass both cognitive dimensions (e.g., logical reasoning, knowledge) and speech-related aspects (

2025

Uncertain Pushing Adaptive Coordinated Control for the Human-Exoskeleton-Walker System

RA-L 2025

Lower Limb Exoskeletons are potential in the gait training for patients with gait disorders. For patients in the early rehabilitation stages with weak upper limb strength, it is challenge to keep balance by themselves only. A mobile robotic walker is helpful to maintain the walking balance, with the

Cited by 0SourceScholar
2025

Wcdt: World-Centric Diffusion Transformer for Traffic Scene Generation

ICRA 2025

In this paper, we introduce a novel approach for autonomous driving trajectory generation by harnessing the complementary strengths of diffusion probabilistic models (a.k.a., diffusion models) and transformers. Our proposed framework, termed the “World-centric Diffusion Transformer” (WcDT), optimize

Cited by 40SourcecodeScholar
2024

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

NeurIPS 2024spotlight

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and real…

2024

HS-GC: Holistic Semantic Embedding and Global Contrast for Effective Text Clustering

COLING 2024main

In this paper, we introduce Holistic Semantic Embedding and Global Contrast (HS-GC), an end-to-end approach to learn the instance- and cluster-level representation. Specifically, for instance-level representation learning, we introduce a new loss function that exploits different layers of semantic i…

Cited by 0SourcePDFScholar
2024

MVITP: Multi-View Image-Text Perception for Few-Shot Remote Sensing Image Classification

ICASSP 2024accepted

Few-shot learning has been extensively applied in current remote sensing image classification, enabling rapid identification of new classes by leveraging prior knowledge effectively. However, current methods mainly rely on image modality to address the issue of low intra-class similarity and high in…

Cited by 0SourceScholar
2023

NeRF-MS: Neural Radiance Fields with Multi-Sequence

ICCV 2023poster

Neural radiance fields (NeRF) achieve impressive performance in novel view synthesis when trained on only single sequence data. However, leveraging multiple sequences captured by different cameras at different times is essential for better reconstruction performance. Multi-sequence data takes two ma…

Cited by 24PDFcodeScholar
2023

NeRFVS: Neural Radiance Fields for Free View Synthesis via Geometry Scaffolds

CVPR 2023poster

We present NeRFVS, a novel neural radiance fields (NeRF) based method to enable free navigation in a room. NeRF achieves impressive performance in rendering images for novel views similar to the input views while suffering for novel views that are significantly different from the training views. To…

Cited by 13SourcePDFScholar
2023

Rethinking Semi-Supervised Imbalanced Node Classification from Bias-Variance Decomposition

NeurIPS 2023poster

This paper introduces a new approach to address the issue of class imbalance in graph neural networks (GNNs) for learning on graph-structured data. Our approach integrates imbalanced node classification and Bias-Variance Decomposition, establishing a theoretical framework that closely relates data i…

2022

A Bidirectional Soft Biomimetic Hand Driven by Water Hydraulic for Dexterous Underwater Grasping

RA-L 2022

Soft robotics shows considerable promise for various underwater applications. Soft grippers as end-effectors are particularly useful for compliant and robust grasping compared to rigid mechanisms. In this work, we describe the design, fabrication and operation of a soft robotic hand driven by water

Cited by 34SourceScholar
2022

TDv2: A Novel Tree-Structured Decoder for Offline Mathematical Expression Recognition

AAAI 2022technical

In recent years, tree decoders become more popular than LaTeX string decoders in the field of handwritten mathematical expression recognition (HMER) as they can capture the hierarchical tree structure of mathematical expressions. However previous tree decoders converted the tree structure labels int…

2022

Video Interpolation by Event-Driven Anisotropic Adjustment of Optical Flow

ECCV 2022poster

"Video frame interpolation is a challenging task due to the ever-changing real-world scene. Previous methods often calculate the bi-directional optical flows and then predict the intermediate optical flows under the linear motion assumptions, leading to isotropic intermediate flow generation. Follow…

Cited by 15SourcePDFScholar
2021

MetaCorrection: Domain-Aware Meta Loss Correction for Unsupervised Domain Adaptation in Semantic Segmentation

CVPR 2021poster

Unsupervised domain adaptation (UDA) aims to transfer the knowledge from the labeled source domain to the unlabeled target domain. Existing self-training based UDA approaches assign pseudo labels for target data and treat them as ground truth labels to fully leverage unlabeled target data for model…

Cited by 119PDFcodeScholar
2020

Modeling and Experiments on the Swallowing and Disgorging Characteristics of an Underwater Continuum Manipulator

ICRA 2020poster

Soft robots apply compliant materials to perform motions and behaviors not typically achievable by rigid robots. An underwater, compliant, multi-segment continuum manipulator that can bend, swallow, disgorge is developed in this study. The manipulator is driven by McKibben water hydraulic artificial…

Cited by 18SourceScholar
2017

Development of an inexpensive tri-axial force sensor for minimally invasive surgery

IROS 2017poster

This work presents the design and evaluation of a low-cost tri-axial force sensor, that has been developed to regain the sense of touch in minimally invasive surgeries (MIS). The force sensor uses an array of force sensitive resistors (FSR) with a mechanically pre-loaded structure to perform the for…

Cited by 35SourceScholar