← Search

SIYU CHEN

41 accepted papers

2026

Cross Domain Test Time Scaling: Scale Knowledge and Reasoning on Cross Domains

IJCAI 2026

Test-time scaling (TTS) has demonstrated remarkable potential in enhancing the reasoning capabilities of Large Language Models (LLMs) and Large Vision-Language Models (LVLMs). However, its application has primarily been limited to domains such as mathematics and programming, owing to their reasoning

Cited by 0Scholar
2026

Differentiable JPEG-based Input Perturbation for Knowledge Distillation Amplification via Conditional Mutual Information Maximization

ICLR 2026poster

Maximizing conditional mutual information (CMI) has recently been shown to enhance the effectiveness of teacher networks in knowledge distillation (KD). Prior work achieves this by fine-tuning a pretrained teacher to maximize a proxy of its CMI. However, fine-tuning large-scale teachers is often imp…

Cited by 0SourceScholar
2026

Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec

ICASSP 2026poster

Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity reconstruction under ultra-low bitrate constraints remains…

Cited by 0SourcePDFScholar
2026

How Transformers Learn Causal Structures In-Context: Explainable Mechanism Meets Theoretical Guarantee

ICLR 2026poster

Transformers have demonstrated remarkable in-context learning abilities, adapting to new tasks from just a few examples without parameter updates. However, theoretical understanding of this phenomenon typically assumes fixed dependency structures, while real-world sequences exhibit flexible, context…

Cited by 0SourceScholar
2026

ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation

AAAI 2026technical

Enabling multi-task adaptation in pre-trained Low-Rank Adaptation (LoRA) models is crucial for enhancing their generalization capabilities. Most existing pre-trained LoRA fusion methods decompose weight matrices, sharing similar parameters, while fusion divergent ones. However, this paradigm inevit

Cited by 0SourcePDFScholar
2026

Piercing the Fog: Disentangling Key Features for Vision Models in Multi-Degradation Scenarios

AAAI 2026technical

In natural scenarios, vision models often encounter the challenge of complex degradation scenarios(e.g., rain, snow, fog, or motion blur). These degradations severely corrupt image features, causing existing models to treat rarely seen or unseen degraded images as “unfamiliar”, thereby losing their

Cited by 0SourcePDFScholar
2026

SGS-3D: High-Fidelity 3D Instance Segmentation via Reliable Semantic Mask Splitting and Growing

AAAI 2026technical

Accurate 3D instance segmentation is crucial for high-quality scene understanding in the 3D vision domain. However, 3D instance segmentation based on 2D-to-3D lifting approaches struggle to produce precise instance-level segmentation, due to accumulated errors introduced during the lifting process f

Cited by 0SourcePDFScholar
2026

SpatialHand: Generative Object Manipulation from 3D Prespective

ICLR 2026poster

We introduce SpatialHand, a novel framework for generative object insertion with precise 3D control. Current generative object manipulation methods primarily operate within the 2D image plane, but often fail to grasp 3D scene complexities, leading to ambiguities in an object's 3D position, orientati…

Cited by 0SourceScholar
2026

TR-DQ: Time-Rotation Diffusion Quantization

AAAI 2026technical

Diffusion models have been widely adopted in image and video generation. However, their complex network architecture leads to high inference overhead for its generation process. Existing diffusion quantization methods primarily focus on the quantization of the model structure while ignoring the impa

Cited by 0SourcePDFScholar
2026

Taming Polysemanticity in LLMs: Theory-Grounded Feature Recovery via Sparse Autoencoders

ICLR 2026poster

We study the challenge of achieving theoretically grounded feature recovery using Sparse Autoencoders (SAEs) for the interpretation of Large Language Models. Existing SAE training algorithms often lack rigorous mathematical guarantees and suffer from practical limitations such as hyperparameter sen…

Cited by 0SourcecodeScholar
2025

An Optimized Franz-Parisi Criterion and its Equivalence with SQ Lower Bounds

NeurIPS 2025oral

Bandeira et al. (2022) introduced the Franz-Parisi (FP) criterion for characterizing the computational hard phases in statistical detection problems. The FP criterion, based on an annealed version of the celebrated Franz-Parisi potential from statistical physics, was shown to be equivalent to low-de…

Cited by 0SourceScholar
2025

Can Neural Networks Achieve Optimal Computational-statistical Tradeoff? An Analysis on Single-Index Model

ICLR 2025oral

In this work, we tackle the following question: Can neural networks trained with gradient-based methods achieve the optimal statistical-computational tradeoff in learning Gaussian single-index models? Prior research has shown that any polynomial-time algorithm under the statistical query (SQ) frame…

Cited by 0SourcePDFScholar
2025

Depth Matters: Exploring Deep Interactions of RGB-D for Semantic Segmentation in Traffic Scenes

IROS 2025

RGB-D has gradually become a crucial data source for understanding complex scenes in assisted driving. However, existing studies have paid insufficient attention to the intrinsic spatial properties of depth maps. This oversight significantly impacts the attention representation, leading to predictio

Cited by 6SourceScholar
2025

EGS-SLAM: RGB-D Gaussian Splatting SLAM With Events

RA-L 2025

Gaussian Splatting SLAM (GS-SLAM) offers a notable improvement over traditional SLAM methods, in enabling photorealistic 3D reconstruction that conventional approaches often struggle to achieve. However, existing GS-SLAM systems perform poorly under persistent and severe motion blur commonly encount

Cited by 3SourceScholar
2025

FGS-SLAM: Fourier-based Gaussian Splatting for Real-time SLAM with Sparse and Dense Map Fusion

IROS 2025

3D gaussian splatting has advanced simultaneous localization and mapping (SLAM) technology by enabling realtime positioning and the construction of high-fidelity maps. However, the uncertainty in gaussian position and initialization parameters introduces challenges, often requiring extensive iterati

Cited by 3SourcecodeScholar
2025

In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention

ICML 2025poster

We study how multi-head softmax attention models are trained to perform in-context learning on linear data. Through extensive empirical experiments and rigorous theoretical analysis, we demystify the emergence of elegant attention patterns: a diagonal and homogeneous pattern in the key-query weight…

2025

Large-Scale UWB Anchor Calibration and One-Shot Localization Using Gaussian Process

ICRA 2025

Ultra-wideband (UWB) is gaining popularity with devices like AirTags for precise home item localization but faces significant challenges when scaled to large environments like seaports. The main challenges are calibration and localization under obstructed conditions, which are common in logistics en

Cited by 15SourceScholar
2025

Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic Segmentation

NeurIPS 2025poster

Open-Vocabulary semantic segmentation (OVSS) and domain generalization in semantic segmentation (DGSS) highlight a subtle complementarity that motivates Open-Vocabulary Domain-Generalized Semantic Segmentation (OV-DGSS). OV-DGSS aims to generate pixel-level masks for unseen categories while maintain…

Cited by 0SourcecodeScholar
2025

MambaIC: State Space Models for High-Performance Learned Image Compression

CVPR 2025poster

A high-performance image compression algorithm is crucial for real-time information transmission across numerous fields. Despite rapid progress in image compression, computational inefficiency and poor redundancy modeling still pose significant bottlenecks, limiting practical applications. Inspired…

2025

Multi-Scale Temporal Neural Network for Stock Trend Prediction Enhanced by Temporal Hyepredge Learning

IJCAI 2025

Existing research in Stock Trend Prediction (STP) focuses on temporal features extracted from a temporal sequence of stock data with a look-back window, which frequently leads to the omission of important periodic patterns, such as weekly and monthly variations in stock prices. Furthermore, these me

2025

Pushing Wi-Fi Towards Fine-Grained Sensing Via Spectrogram Enhancement

ICASSP 2025accepted

In recent years, Wi-Fi sensing has attracted much attention due to the widespread deployment of communication devices. Due to advancements in signal processing algorithms, contactless sensing technology based on Wi-Fi signals has now been widely applied. However, the limited bandwidth of Wi-Fi syste…

Cited by 0SourceScholar
2025

Stronger, Steadier & Superior: Geometric Consistency in Depth VFM Forges Domain Generalized Semantic Segmentation

ICCV 2025poster

Vision Foundation Models (VFMs) have delivered remarkable performance in Domain Generalized Semantic Segmentation (DGSS). However, recent methods often overlook the fact that visual cues are susceptible, whereas the underlying geometry remains stable, rendering depth information more robust. In this…

2025

Unlocking Dark Vision Potential for Medical Image Segmentation

IJCAI 2025

Accurate segmentation of lesions is crucial for disease diagnosis and treatment planning. However, blurring and low contrast in the imaging process can affect segmentation results. We have observed that noninvasive medical imaging shares considerable similarities with natural images under low light

Cited by 0SourcePDFScholar
2025

WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models

ACL 2025long

Retrieval Augmented Generation (RAG) has gained widespread adoption owing to its capacity to empower large language models (LLMs) to integrate external knowledge. However, existing RAG frameworks are primarily designed for text-based LLMs and rely on Automatic Speech Recognition to process speech in…

Cited by 0SourcePDFScholar
2024

From Words to Actions: Unveiling the Theoretical Underpinnings of LLM-Driven Autonomous Systems

ICML 2024poster

In this work, from a theoretical lens, we aim to understand why large language model (LLM) empowered agents are able to solve decision-making problems in the physical world. To this end, consider a hierarchical reinforcement learning (RL) model where the LLM Planner and the Actor perform high-level…

Cited by 8SourcePDFScholar
2024

Salient Sparse Visual Odometry With Pose-Only Supervision

RA-L 2024

Visual Odometry (VO) is vital for the navigation of autonomous systems, providing accurate position and orientation estimates at reasonable costs. While traditional VO methods excel in some conditions, they struggle with challenges like variable lighting and motion blur. Deep learning-based VO, thou

Cited by 14SourceScholar
2024

Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers

NeurIPS 2024poster

In-context learning (ICL) is a cornerstone of large language model (LLM) functionality, yet its theoretical foundations remain elusive due to the complexity of transformer architectures. In particular, most existing work only theoretically explains how the attention mechanism facilitates ICL under c…

Cited by 11SourcePDFScholar
2023

Learning to Incentivize Information Acquisition: Proper Scoring Rules Meet Principal-Agent Model

ICML 2023poster

We study the incentivized information acquisition problem, where a principal hires an agent to gather information on her behalf. Such a problem is modeled as a Stackelberg game between the principal and the agent, where the principal announces a scoring rule that specifies the payment, and then the…

Cited by 8SourcePDFScholar
2023

PyPose: A Library for Robot Learning With Physics-Based Optimization

CVPR 2023poster

Deep learning has had remarkable success in robotic perception, but its data-centric nature suffers when it comes to generalizing to ever-changing environments. By contrast, physics-based optimization generalizes better, but it does not perform as well in complicated tasks due to the lack of high-le…

2022

Adaptive Model Design for Markov Decision Process

ICML 2022spotlight

In a Markov decision process (MDP), an agent interacts with the environment via perceptions and actions. During this process, the agent aims to maximize its own gain. Hence, appropriate regulations are often required, if we hope to take the external costs/benefits of its actions into consideration.…

Cited by 14SourcePDFScholar
2022

Continuum Manipulator With Rigid-Flexible Coupling Structure

RA-L 2022

Pneumatic-driven soft manipulator has great flexibility and adaptability, which makes it can do some complex tasks that traditional robots can't do such as minimally invasive surgery, detection, search and rescue, et al. However, due to the low stiffness of pneumatic-drive soft manipulator, it faces

Cited by 20SourceScholar
2022

DocEE: A Large-Scale and Fine-grained Benchmark for Document-level Event Extraction

NAACL 2022long

Event extraction aims to identify an event and then extract the arguments participating in the event. Despite the great success in sentence-level event extraction, events are more naturally presented in the form of documents, with event arguments scattered in multiple sentences. However, a major bar…

2022

X-Learner: Learning Cross Sources and Tasks for Universal Visual Representation

ECCV 2022poster

"In computer vision, pre-training models based on large-scale supervised learning have been proven effective over the past few years. However, existing works mostly focus on learning from the individual tasks with the single data source e.g., ImageNet for classification or COCO for detection). This…

Cited by 10SourcePDFScholar
2021

Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization

CVPR 2021poster

Localizing persons and recognizing their actions from videos is a challenging task towards high-level video under-standing. Recent advances have been achieved by modeling direct pairwise relations between entities. In this paper, we take one step further, not only model direct relations between pair…

Cited by 204PDFcodeScholar
2021

ForgeryNet: A Versatile Benchmark for Comprehensive Forgery Analysis

CVPR 2021poster

The rapid progress of photorealistic synthesis techniques has reached at a critical point where the boundary between real and manipulated images starts to blur. Thus, benchmarking and advancing digital forgery analysis have become a pressing issue. However, existing face forgery datasets either have…

Cited by 178PDFScholar
2021

Wasserstein Flow Meets Replicator Dynamics: A Mean-Field Analysis of Representation Learning in Actor-Critic

NeurIPS 2021poster

Actor-critic (AC) algorithms, empowered by neural networks, have had significant empirical success in recent years. However, most of the existing theoretical support for AC algorithms focuses on the case of linear function approximations, or linearized neural networks, where the feature representat…

Cited by 6SourcePDFScholar