← Search

Shengyong Chen

19 accepted papers

2026

CodeMamba: Shifting from Target Semantics to Self-Supervised Background Manifold Learning for Singularity Detection in Infrared Sequences

ICML 2026poster

Multi-frame infrared small target detection suffers from extreme semantic paucity of targets and representation collapse due to overwhelming class imbalance, resulting in the persistent inability to accurately distinguish point-like targets from dynamic background clutter. To address these issues, w…

Cited by 0SourceScholar
2026

E²I-VRWKV: Explicit EPI-Representation and Interaction-Aware Vision-RWKV for Light Field Semantic Segmentation

ICML 2026poster

Pixel-level semantic segmentation of 4D light field (LF) data remains a considerable challenge, primarily due to the conflict between modeling complex spatial-angular dependencies and maintaining linear computational efficiency. Current linear models like VRWKV offer scalability but often fail to ca…

Cited by 0SourceScholar
2026

SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation

ICML 2026poster

Achieving pixel-level accurate segmentation of structural cracks across diverse scenarios remains a formidable challenge. Existing methods face significant bottlenecks in balancing crack topology modeling with computational efficiency, often failing to reconcile high segmentation quality with low re…

Cited by 0SourceScholar
2026

Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object Detection

CVPR 2026

Unaligned RGB-T salient object detection (SOD) remains challenging due to severe cross-modal spatial discrepancies and unreliable feature fusion. Existing methods often assume perfect alignment or rely on geometric registration, which is computationally demanding and sensitive to cross-modal inconsi

Cited by 0SourceScholar
2025

Can Students Beyond the Teacher? Distilling Knowledge from Teacher’s Bias

AAAI 2025technical

Knowledge distillation (KD) is a model compression technique that transfers knowledge from a large teacher model to a smaller student model to enhance its performance. Existing methods often assume that the student model is inherently inferior to the teacher model. However, we identify that the fund…

2025

Dust-Mamba: An Efficient Dust Storm Detection Network with Multiple Data Sources

AAAI 2025technical

Accurate detection of dust storms is challenging due to complex meteorological interactions. With the development of deep learning, deep neural networks have been increasingly applied to dust storm detection, offering better learning and generalization capabilities compared to traditional physical m…

2025

Multi-Scale Convolutional Networks with Class-Normalized Logit Clipping for Robust Sea State Estimation from Noisy Ship Motion Data

ICRA 2025

Autonomous ships utilize automation systems to achieve unmanned navigation, driving innovation in maritime transportation. However, sea conditions, influenced by dynamic factors such as wave height, wind speed, and ocean currents, present a challenge in accurately assessing these conditions. Traditi

Cited by 0SourceScholar
2025

SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in Structures

CVPR 2025poster

Pixel-level segmentation of structural cracks across various scenarios remains a considerable challenge. Current methods encounter challenges in effectively modeling crack morphology and texture, facing challenges in balancing segmentation quality with low computational resource usage. To overcome t…

2024

Intentional Evolutionary Learning for Untrimmed Videos with Long Tail Distribution

AAAI 2024technical

Human intention understanding in untrimmed videos aims to watch a natural video and predict what the person’s intention is. Currently, exploration of predicting human intentions in untrimmed videos is far from enough. On the one hand, untrimmed videos with mixed actions and backgrounds have a signif…

2023

Distilling Cross-Temporal Contexts for Continuous Sign Language Recognition

CVPR 2023poster

Continuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-the-art methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9,20,25,36] have indicated th…

Cited by 47SourcePDFScholar
2022

HMD-former: a Transformer-based Human Mesh Deformer with Inter-layer Semantic Consistency

ICRA 2022poster

We present a transformer-based network, Human Mesh Deformer (HMD-former), to tackle the problem of 3D human mesh reconstruction from a single RGB image. HMD-former applies a pre-trained CNN to extract image grid features and a transformer decoder to gradually warp the template 3D mesh to the deforme…

Cited by 1SourcecodeScholar
2022

PA-AWCNN: Two-stream Parallel Attention Adaptive Weight Network for RGB-D Action Recognition

ICRA 2022poster

Due to overly relying on appearance information or adopting direct static feature fusion, most of the existing action recognition methods based on multi-modality have poor robustness and insufficient consideration of modality differences. To address these problems, we propose a two-stream adaptive w…

Cited by 6SourcecodeScholar
2021

Consistency-Aware Graph Network for Human Interaction Understanding

ICCV 2021poster

Compared with the progress made on human activity classification, much less success has been achieved on human interaction understanding (HIU). Apart from the latter task is much more challenging, the main cause is that recent approaches learn human interactive relations via shallow graphical models…

Cited by 12PDFcodeScholar
2020

CalibRCNN: Calibrating Camera and LiDAR by Recurrent Convolutional Neural Network and Geometric Constraints

IROS 2020poster

In this paper, we present Calibration Recurrent Convolutional Neural Network (CalibRCNN) to infer a 6 degrees of freedom (DOF) rigid body transformation between 3D LiDAR and 2D camera. Different from the existing methods, our 3D-2D CalibRCNN not only uses the LSTM network to extract the temporal fea…

Cited by 72SourceScholar
2020

SiamCAR: Siamese Fully Convolutional Classification and Regression for Visual Tracking

CVPR 2020oral

By decomposing the visual tracking task into two subproblems as classification for pixel category and regression for object bounding box at this pixel, we propose a novel fully convolutional Siamese network to solve visual tracking end-to-end in a per-pixel manner. The proposed framework SiamCAR con…

Cited by 980PDFcodeScholar
2019

Modeling and Analysis of Motion Data from Dynamically Positioned Vessels for Sea State Estimation

ICRA 2019poster

Developing a reliable model to identify the sea state is significant for the autonomous ship. This paper introduces a novel deep neural network model (SeaStateNet) to estimate the sea state based on the ship motion data from dynamically positioned vessels. The SeaStateNet mainly consists of three co…

Cited by 42SourceScholar
2019

Robust High Accuracy Visual-Inertial-Laser SLAM System

IROS 2019poster

In recent years, many excellent works on visual-inertial SLAM and laser-based SLAM have been proposed. Although inertial measurement unit (IMU) significantly improve the motion estimate performance by reducing the impact of illumination variation or texture-less region on visual tracking, tracking f…

Cited by 50SourceScholar
2017

Segment-tree based cost aggregation for stereo matching with enhanced segmentation advantage

ICASSP 2017accepted

Segment-tree (ST) based cost aggregation algorithm for stereo matching successfully integrates the information of segmentation with non-local cost aggregation framework. The tree structure which is generated by the segmentation strategy directly determines the final results for this kind of algorith…

Cited by 0SourceScholar