← Search

Xu Cheng

33 accepted papers

2026

Beyond Missing Data Imputation: Information-Theoretic Coupling of Missingness and Class Imbalance for Optimal Irregular Time Series Classification

AAAI 2026technical

Irregular time series (IRTS) are prevalent in real-world applications, where uneven sampling and missing data pose fundamental challenges to deep learning-based feature modeling. Although existing methods attempt to retain timestamp information, they often overlook the structured patterns embedded w

Cited by 0SourcePDFScholar
2026

Class-Guided Network with Rare-Class Amplification for Sea State Estimation Based on Ship Motion Data

ICRA 2026poster

Accurate, real-time Sea State Estimation (SSE) is crucial for the safety and operational efficiency of Autonomous Surface Vessels (ASVs). However, existing deep learning methods for this task commonly face three major challenges: the inherent class imbalance of marine environments, the ambiguous bou…

Cited by 0Scholar
2026

CodeMamba: Shifting from Target Semantics to Self-Supervised Background Manifold Learning for Singularity Detection in Infrared Sequences

ICML 2026poster

Multi-frame infrared small target detection suffers from extreme semantic paucity of targets and representation collapse due to overwhelming class imbalance, resulting in the persistent inability to accurately distinguish point-like targets from dynamic background clutter. To address these issues, w…

Cited by 0SourceScholar
2026

E²I-VRWKV: Explicit EPI-Representation and Interaction-Aware Vision-RWKV for Light Field Semantic Segmentation

ICML 2026poster

Pixel-level semantic segmentation of 4D light field (LF) data remains a considerable challenge, primarily due to the conflict between modeling complex spatial-angular dependencies and maintaining linear computational efficiency. Current linear models like VRWKV offer scalability but often fail to ca…

Cited by 0SourceScholar
2026

Frequency-Aware Augmentation and Alignment for Time Series Contrastive Learning

IJCAI 2026

Contrastive learning has become a dominant paradigm for learning time series representations from large-scale unlabeled data. However, current methods are often adapted from computer vision and rely on random time-domain augmentations (e.g., jittering and cropping). Such augmentations can unpredicta

Cited by 0Scholar
2026

Radar-APLANC: Unsupervised Radar-based Heartbeat Sensing via Augmented Pseudo-Label and Noise Contrast

AAAI 2026technical

Frequency Modulated Continuous Wave (FMCW) radars can measure subtle chest wall oscillations to enable non-contact heartbeat sensing. However, traditional radar-based heartbeat sensing methods face performance degradation due to noise. Learning-based radar methods achieve better noise robustness but

Cited by 0SourcePDFScholar
2026

SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation

ICML 2026poster

Achieving pixel-level accurate segmentation of structural cracks across diverse scenarios remains a formidable challenge. Existing methods face significant bottlenecks in balancing crack topology modeling with computational efficiency, often failing to reconcile high segmentation quality with low re…

Cited by 0SourceScholar
2026

Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object Detection

CVPR 2026

Unaligned RGB-T salient object detection (SOD) remains challenging due to severe cross-modal spatial discrepancies and unreliable feature fusion. Existing methods often assume perfect alignment or rely on geometric registration, which is computationally demanding and sensitive to cross-modal inconsi

Cited by 0SourceScholar
2025

Can Students Beyond the Teacher? Distilling Knowledge from Teacher’s Bias

AAAI 2025technical

Knowledge distillation (KD) is a model compression technique that transfers knowledge from a large teacher model to a smaller student model to enhance its performance. Existing methods often assume that the student model is inherently inferior to the teacher model. However, we identify that the fund…

2025

Collaborative Association Network for Multi-view Multi-Human Association and Tracking using Constraint Optimization and Object Search

ICASSP 2025accepted

Multi-view multi-human association and tracking (MvMHAT) enhances scene perception using multiple cameras, crucial for applications such as surveillance and crowd analysis. Inherent feature disparities between views complicate similarity calculations. Recent works combine representation and motion i…

Cited by 0SourceScholar
2025

Efficient Large-Scale Scene Point Cloud Upsampling with Implicit Neural Networks and Spatial Hashing

ICASSP 2025accepted

Point cloud upsampling is a critical challenge in 3D vision, particularly for large-scale, real-world data. We propose ASFNet, a novel implicit neural network-based approach that uniquely combines adaptive spatial feature representation with efficient spatial hashing. This method significantly impro…

Cited by 0SourceScholar
2025

FreeNet: Liberating Depth-Wise Separable Operations for Building Faster Mobile Vision Architectures

AAAI 2025technical

In the pursuit of efficient vision architectures, substantial efforts have been devoted to optimizing operator efficiency. Depth-wise separable operators, such as DWConv, are found cheap in both FLOPs and parameters. As a result, they are increasingly incorporated into efficient backbones, trading f…

Cited by 0SourcePDFScholar
2025

From Laboratory to Real World: A New Benchmark Towards Privacy-Preserved Visible-Infrared Person Re-Identification

CVPR 2025poster

Aiming to match pedestrian images captured under varying lighting conditions, visible-infrared person re-identification (VI-ReID) has drawn intensive research attention and achieved promising results. However, in real-world surveillance contexts, data is distributed across multiple devices/entities,…

2025

FusionPhys: A Flexible Framework for Fusing Complementary Sensing Modalities in Remote Physiological Measurement

ICCV 2025poster

Remote physiological measurement using visible light cameras has emerged as a powerful tool for non-contact health monitoring, yet its reliability degrades under challenging conditions such as low-light environments or diverse skin tones. These limitations have motivated the exploration of alternati…

2025

Multi-Scale Convolutional Networks with Class-Normalized Logit Clipping for Robust Sea State Estimation from Noisy Ship Motion Data

ICRA 2025

Autonomous ships utilize automation systems to achieve unmanned navigation, driving innovation in maritime transportation. However, sea conditions, influenced by dynamic factors such as wave height, wind speed, and ocean currents, present a challenge in accurately assessing these conditions. Traditi

Cited by 0SourceScholar
2025

Multi-order Orchestrated Curriculum Distillation for Model-Heterogeneous Federated Graph Learning

NeurIPS 2025poster

Federated Graph Learning (FGL) has been shown to be particularly effective in enabling collaborative training of Graph Neural Networks (GNNs) in decentralized settings. Model-heterogeneous FGL further enhances practical applicability by accommodating client preferences for diverse model architecture…

Cited by 0SourceScholar
2025

PolypSense3D: A Multi-Source Benchmark Dataset for Depth-Aware Polyp Size Measurement in Endoscopy

NeurIPS 2025poster

Accurate polyp sizing during endoscopy is crucial for cancer risk assessment but is hindered by subjective methods and inadequate datasets lacking integrated 2D appearance, 3D structure, and real-world size information. We introduce PolypSense3D, the first multi-source benchmark dataset specifically…

Cited by 0SourcecodeScholar
2025

RankAdaptor: Hierarchical Rank Allocation for Efficient Fine-Tuning Pruned LLMs via Performance Model

NAACL 2025findings

The efficient compression of large language models (LLMs) has become increasingly popular. However, recovering the performance of compressed LLMs remains a major challenge. The current practice in LLM compression entails the implementation of structural pruning, complemented by a recovery phase that…

Cited by 0SourcePDFScholar
2025

SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in Structures

CVPR 2025poster

Pixel-level segmentation of structural cracks across various scenarios remains a considerable challenge. Current methods encounter challenges in effectively modeling crack morphology and texture, facing challenges in balancing segmentation quality with low computational resource usage. To overcome t…

2025

Serial Local Patterns and Irregular Dependencies Extract and Cascaded Fusion Network for Structural Crack Segmentation

ICASSP 2025accepted

Achieving pixel-level crack segmentation in complex scenarios is a major challenge, as current methods have difficulty effectively integrating both local features and irregular pixel dependencies. In this paper, we introduce a Cascaded Fusion Network (LICFN) specifically designed for crack segmentat…

Cited by 0SourceScholar
2024

Clarifying the Behavior and the Difficulty of Adversarial Training

AAAI 2024technical

Adversarial training is usually difficult to optimize. This paper provides conceptual and analytic insights into the difficulty of adversarial training via a simple theoretical study, where we derive an approximate dynamics of a recursive multi-step attack in a simple setting. Despite the simplicity…

Cited by 0SourcePDFScholar
2024

Differentiable Auxiliary Learning for Sketch Re-Identification

AAAI 2024technical

Sketch re-identification (Re-ID) seeks to match pedestrians' photos from surveillance videos with corresponding sketches. However, we observe that existing works still have two critical limitations: (i) cross- and intra-modality discrepancies hinder the extraction of modality-shared features, (ii) s…

Cited by 10SourcePDFScholar
2024

Layerwise Change of Knowledge in Neural Networks

ICML 2024poster

This paper aims to explain how a deep neural network (DNN) gradually extracts new knowledge and forgets noisy features through layers in forward propagation. Up to now, although how to define knowledge encoded by the DNN has not reached a consensus so far, previous studies have derived a series of m…

Cited by 5SourcePDFScholar
2023

Modality Unifying Network for Visible-Infrared Person Re-Identification

ICCV 2023poster

Visible-infrared person re-identification (VI-ReID) is a challenging task due to large cross-modality discrepancies and intra-class variations. Existing methods mainly focus on learning modality-shared representations by embedding different modalities into the same feature space. As a result, the le…

Cited by 58PDFScholar
2023

TOPLight: Lightweight Neural Networks With Task-Oriented Pretraining for Visible-Infrared Recognition

CVPR 2023poster

Visible-infrared recognition (VI recognition) is a challenging task due to the enormous visual difference across heterogeneous images. Most existing works achieve promising results by transfer learning, such as pretraining on the ImageNet, based on advanced neural architectures like ResNet and ViT.…

Cited by 14SourcePDFScholar
2023

Towards the Difficulty for a Deep Neural Network to Learn Concepts of Different Complexities

NeurIPS 2023poster

This paper theoretically explains the intuition that simple concepts are more likely to be learned by deep neural networks (DNNs) than complex concepts. In fact, recent studies have observed [24, 15] and proved [26] the emergence of interactive concepts in a DNN, i.e., it is proven that a DNN usuall…

Cited by 19SourcePDFScholar
2021

Building Interpretable Interaction Trees for Deep NLP Models

AAAI 2021technical

This paper proposes a method to disentangle and quantify interactions among words that are encoded inside a DNN for natural language processing. We construct a tree to encode salient interactions extracted by the DNN. Six metrics are proposed to analyze properties of interactions between constituent…

Cited by 43SourcePDFScholar
2021

Drop Redundant, Shrink Irrelevant: Selective Knowledge Injection for Language Pretraining

IJCAI 2021poster

Previous research has demonstrated the power of leveraging prior knowledge to improve the performance of deep models in natural language processing. However, traditional methods neglect the fact that redundant and irrelevant knowledge exists in external knowledge bases. In this study, we launched an…

Cited by 34SourcePDFScholar
2021

Towards a Unified Game-Theoretic View of Adversarial Perturbations and Robustness

NeurIPS 2021poster

This paper provides a unified view to explain different adversarial attacks and defense methods, i.e. the view of multi-order interactions between input variables of DNNs. Based on the multi-order interaction, we discover that adversarial attacks mainly affect high-order interactions to fool the DNN…

2019

Modeling and Analysis of Motion Data from Dynamically Positioned Vessels for Sea State Estimation

ICRA 2019poster

Developing a reliable model to identify the sea state is significant for the autonomous ship. This paper introduces a novel deep neural network model (SeaStateNet) to estimate the sea state based on the ship motion data from dynamically positioned vessels. The SeaStateNet mainly consists of three co…

Cited by 42SourceScholar