← Search

Siyuan Shen

14 accepted papers

2026

Cross-Timestep: 3D Diffusion Model with Trans-temporal Memory LSTM and Adaptive Priori Decoding Strategy for Medical Segmentation

ICLR 2026poster

Diffusion models have recently demonstrated significant robustness in medical image segmentation, effectively accommodating variations across different imaging styles. However, their applications remain limited due to: (i) current successes being primarily confined to 2D segmentation tasks—we observ…

Cited by 0SourceScholar
2026

DeepSADR: Deep Transfer Learning with Subsequence Interaction and Adaptive Readout for Cancer Drug Response Prediction

ICLR 2026poster

Cancer treatment efficacy exhibits high inter-patient heterogeneity due to genomic variations. While large-scale in vitro drug response data from cancer cell lines exist, predicting patient drug responses remains challenging due to genomic distribution shifts and the scarcity of clinical response da…

Cited by 0SourcecodeScholar
2026

Hierarchical Structure-Property Alignment for Data-Efficient Molecular Generation and Editing

AAAI 2026technical

Property-constrained molecular generation and editing are crucial in AI-driven drug discovery but remain hindered by two factors: (i) capturing the complex relationships between molecular structures and multiple properties remains challenging, and (ii) the narrow coverage and incomplete annotations

Cited by 0SourcePDFScholar
2025

Enhancing Uncertainty Quantification in Large Language Models through Semantic Graph Density

UAI 2025

Large Language Models (LLMs) excel in language understanding but are susceptible to "confabulation," where they generate arbitrary, factually incorrect responses to uncertain questions. Detecting confabulation in question answering often relies on Uncertainty Quantification (UQ), which measures sema

Cited by 0SourcePDFScholar
2025

TransiT: Transient Transformer for Non-line-of-sight Videography

ICCV 2025poster

High quality and high speed videography using Non-Line-of-Sight (NLOS) imaging benefit autonomous navigation, collision prevention, and post-disaster search and rescue tasks. Current solutions have to balance between the frame rate and image quality. High frame rates, for example, can be achieved by…

Cited by 0SourcePDFScholar
2024

Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition

ICASSP 2024accepted

The mainstream paradigm of speech emotion recognition (SER) is identifying the single emotion label of the entire utterance. This line of works neglect the emotion dynamics at fine temporal granularity and mostly fail to leverage linguistic information of speech signal explicitly. In this paper, we…

Cited by 0SourceScholar
2024

EyeLS: Shadow-Guided Instrument Landing System for Target Approaching in Robotic Eye Surgery

RA-L 2024

Robotic ophthalmic surgery is an emerging technology to facilitate high-precision interventions such as subretinal injection and removing swinging tissues in retinal detachment using microscopy and iOCT. However, locating the instrument tip outside iOCT's range-limited ROI is challenging, especially

Cited by 3SourceScholar
2023

Enhancing Non-line-of-sight Imaging via Learnable Inverse Kernel and Attention Mechanisms

ICCV 2023poster

Recovering information from non-line-of-sight (NLOS) imaging is a computationally-intensive inverse problem. Most physics-based NLOS imaging methods address the complexity of this problem by assuming three-bounce reflections and no self-occlusion. However, these assumptions may break down for object…

Cited by 10PDFcodeScholar
2023

Mingling or Misalignment? Temporal Shift for Speech Emotion Recognition with Pre-Trained Representations

ICASSP 2023accepted

Fueled by recent advances of self-supervised models, pre-trained speech representations proved effective for the downstream speech emotion recognition (SER) task. Most prior works mainly focus on exploiting pre-trained representations and just adopt a linear head on top of the pre-trained model, neg…

Cited by 0SourceScholar
2023

Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition

CVPR 2023poster

Dynamic Facial Expression Recognition (DFER) is a rapidly developing field that focuses on recognizing facial expressions in video format. Previous research has considered non-target frames as noisy frames, but we propose that it should be treated as a weakly supervised problem. We also identify the…

2022

HoD-Net: High-Order Differentiable Deep Neural Networks and Applications

AAAI 2022technical

We introduce a deep architecture named HoD-Net to enable high-order differentiability for deep learning. HoD-Net is based on and generalizes the complex-step finite difference (CSFD) method. While similar to classic finite difference, CSFD approaches the derivative of a function from a higher-dimens…

Cited by 4SourcePDFScholar
2022

PhoCaL: A Multi-Modal Dataset for Category-Level Object Pose Estimation With Photometrically Challenging Objects

CVPR 2022poster

Object pose estimation is crucial for robotic applications and augmented reality. Beyond instance level 6D object pose estimation methods, estimating category-level pose and shape has become a promising trend. As such, a new research field needs to be supported by well-designed datasets. To provide…

Cited by 54PDFScholar