← Search

Shiliang Sun

23 accepted papers

2026

Incomplete Multi-View Multi-Label Classification via Shared Codebook and Fused-Teacher Self-Distillation

ICLR 2026poster

Although multi-view multi-label learning has been extensively studied, research on the dual-missing scenario, where both views and labels are incomplete, remains largely unexplored. Existing methods mainly rely on contrastive learning or information bottleneck theory to learn consistent representati…

Cited by 0SourceScholar
2026

S²-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation

IJCAI 2026

Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, but their performance degrades significantly in long-horizon tasks due to cumulative error propagation. This limitation largely arises from static feature fusion mechanisms that rely on fixed weights t

Cited by 0Scholar
2026

Temporal and Spatial Representation Learning for Multimodal Low-Beam 3D Object Detection

AAAI 2026technical

To facilitate the large-scale deployment of autonomous driving in real-world scenarios, developing low-cost and high-performance 3D object detection systems has become a critical technical challenge. Although high-beam LiDARs provide denser point cloud data, their prohibitive hardware cost and high

Cited by 0SourcePDFScholar
2025

Imagination and Contemplation: A Balanced Framework for Semantic-Augmented Multimodal Machine Translation

EMNLP 2025

Multimodal Machine Translation (MMT) enhances textual translation through auxiliary inputs such as images, which is particularly effective in resolving linguistic ambiguities. However, visual information often introduces redundancy or noise, potentially impairing translation quality. To address this

2025

Multimodal Machine Translation with Text-Image In-depth Questioning

ACL 2025finding

Multimodal machine translation (MMT) integrates visual information to address ambiguity and contextual limitations in neural machine translation (NMT). Some empirical studies have revealed that many MMT models underutilize visual data during translation. They attempt to enhance cross-modal interacti…

2025

VQA-Augmented Machine Translation with Cross-Modal Contrastive Learning

EMNLP 2025

Multimodal machine translation (MMT) aims to enhance translation quality by integrating visual information. However, existing methods often extract visual features using pre-trained models while learning text features from scratch, leading to representation imbalance. These methods are also prone to

Cited by 0SourcePDFScholar
2024

CaMIL: Causal Multiple Instance Learning for Whole Slide Image Classification

AAAI 2024technical

Whole slide image (WSI) classification is a crucial component in automated pathology analysis. Due to the inherent challenges of high-resolution WSIs and the absence of patch-level labels, most of the proposed methods follow the multiple instance learning (MIL) formulation. While MIL has been equipp…

Cited by 13SourcePDFScholar
2024

Discriminatively Fuzzy Multi-View K-means Clustering with Local Structure Preserving

AAAI 2024technical

Multi-view K-means clustering successfully generalizes K-means from single-view to multi-view, and obtains excellent clustering performance. In every view, it makes each data point close to the center of the corresponding cluster. However, multi-view K-means only considers the compactness of each cl…

Cited by 4SourcePDFScholar
2023

Structured BFGS Method for Optimal Doubly Stochastic Matrix Approximation

AAAI 2023technical

Doubly stochastic matrix plays an essential role in several areas such as statistics and machine learning. In this paper we consider the optimal approximation of a square matrix in the set of doubly stochastic matrices. A structured BFGS method is proposed to solve the dual of the primal problem. Th…

2022

Enhancing Unsupervised Domain Adaptation via Semantic Similarity Constraint for Medical Image Segmentation

IJCAI 2022poster

This work proposes a novel unsupervised cross-modality adaptive segmentation method for medical images to tackle the performance degradation caused by the severe domain shift when neural networks are being deployed to unseen modalities. The proposed method is an end-2-end framework, which conducts a…

Cited by 6SourcePDFScholar
2022

TiRGN: Time-Guided Recurrent Graph Network with Local-Global Historical Patterns for Temporal Knowledge Graph Reasoning

IJCAI 2022poster

Temporal knowledge graphs (TKGs) have been widely used in various fields that model the dynamics of facts along the timeline. In the extrapolation setting of TKG reasoning, since facts happening in the future are entirely unknowable, insight into history is the key to predicting future facts. Howeve…

2021

A Sequential Contrastive Learning Framework for Robust Dysarthric Speech Recognition

ICASSP 2021accepted

Dysarthria is a manifestation of disruption in the neuromuscular physiology resulting in uneven, slow, slurred, harsh, or quiet speech. Despite the remarkable progress of automatic speech recognition (ASR), it poses great challenges in developing stable ASR for dysarthric individuals due to the high…

Cited by 0SourceScholar
2021

ASHF-Net: Adaptive Sampling and Hierarchical Folding Network for Robust Point Cloud Completion

AAAI 2021technical

Estimating the complete 3D point cloud from an incomplete one lies at the core of many vision and robotics applications. Existing methods typically predict the complete point cloud based on the global shape representation extracted from the incomplete input. Although they could predict the overall s…

Cited by 26SourcePDFScholar
2021

Multi-Task Transformer with Input Feature Reconstruction for Dysarthric Speech Recognition

ICASSP 2021accepted

Dysarthria is a motor speech disorder caused by damage to the part of the nervous system that controls the physical production of speech. It poses great challenges in building robust dysarthric speech recognition (DSR) due to the high inter- and intra-speaker variability. To this end, we propose a m…

Cited by 0SourceScholar
2020

Semismooth Newton Algorithm for Efficient Projections onto $\ell_1, ∞$-norm Ball

ICML 2020poster

The structured sparsity-inducing $\ell_{1, \infty}$-norm, as a generalization of the classical $\ell_1$-norm, plays an important role in jointly sparse models which select or remove simultaneously all the variables forming a group. However, its resulting problem is more difficult to solve than the c…

2018

PAC-Bayes bounds for stable algorithms with instance-dependent priors

NeurIPS 2018poster

PAC-Bayes bounds have been proposed to get risk estimates based on a training sample. In this paper the PAC-Bayes approach is combined with stability of the hypothesis learned by a Hilbert space valued algorithm. The PAC-Bayes setting is used with a Gaussian prior centered at the expected output. Th…

Cited by 66SourcePDFScholar
2017

A Learning Error Analysis for Structured Prediction with Approximate Inference

NeurIPS 2017poster

In this work, we try to understand the differences between exact and approximate inference algorithms in structured prediction. We compare the estimation and approximation error of both underestimate and overestimate models. The result shows that, from the perspective of learning errors, performance…

Cited by 4SourcePDFScholar