← Search

Rui Sun

59 accepted papers

2026

Adaptive Augmentation-Aware Latent Learning for Robust LiDAR Semantic Segmentation

ICLR 2026poster

Adverse weather conditions significantly degrade the performance of LiDAR point cloud semantic segmentation networks by introducing large distribution shifts. Existing augmentation-based methods attempt to enhance robustness by simulating weather interference during training. However, they struggle…

Cited by 0SourceScholar
2026

Beyond Blind Noising: Disentangled Visual Rectification for Hallucination Mitigation in MLLMs

ICML 2026poster

Visual Contrastive Decoding (VCD) mitigates hallucinations in Multimodal Large Language Models (MLLMs) by penalizing the output shift from noise-perturbed images, assuming this shift captures the hallucination direction. We prove this assumption flawed: noise-induced drift in Language-Image Pretrain…

Cited by 0SourceScholar
2026

Beyond Logits: Coherent Hallucination Mitigation via Attention Contrastive Decoding

ICML 2026poster

Large Vision-Language Models (LVLMs) demonstrate impressive multimodal capabilities, yet suffer from hallucination—generating factually inaccurate content. Contrastive Decoding (CD) mitigates this by contrasting amateur and expert branches at the logit level. However, our investigation reveals that …

Cited by 0SourceScholar
2026

Disco: Densely-overlapping Cell Instance Segmentation via Adjacency-aware Collaborative Coloring

ICLR 2026poster

Accurate cell instance segmentation is foundational for digital pathology analysis. Existing methods based on contour detection and distance mapping still face significant challenges in processing complex and dense cellular regions. Graph coloring-based methods provide a new paradigm for this task,…

Cited by 0SourcecodeScholar
2026

From Softmax to Dirichlet: Evidential Learning for Semi-supervised Semantic Segmentation

CVPR 2026

The critical challenge of semi-supervised semantic segmentation lies in how to fully exploit a large volume of unlabeled data to improve the model's generalization performance for robust segmentation. However, existing softmax scores-based filtering methods tend to be affected by the overconfidence

Cited by 0SourceScholar
2026

MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition

AAAI 2026technical

Capsule Network (CapsNet) has demonstrated significant potential in visual recognition by capturing spatial relationships and part-whole hierarchies for learning equivariant feature representations. However, existing CapsNet and variants often rely on a single high-level feature map, overlooking the

Cited by 0SourcePDFScholar
2026

Mitigating Error Propagation in Low-Rank Approximation of Large Models via Distribution-Aware Whitening

ICML 2026poster

Low-rank approximation has emerged as a cornerstone technique for model compression and parameter-efficient fine-tuning, enabling substantial reductions in computation and memory without altering model architectures. However, existing approaches often overlook the shifts in feature distributions ind…

Cited by 0SourceScholar
2026

Pruning as a Cooperative Game: Surrogate-Assisted Layer Contribution Estimation for Large Language Models

ICLR 2026poster

While large language models (LLMs) demonstrate impressive performance across various tasks, their deployment in real-world scenarios is still constrained by high computational demands. Layer-wise pruning, a commonly employed strategy to mitigate inference costs, can partially address this challenge.…

Cited by 0SourceScholar
2026

Tracing the Heart’s Pathways: ECG Representation Learning from a Cardiac Conduction Perspective

AAAI 2026technical

The multi-lead electrocardiogram (ECG) stands as a cornerstone of cardiac diagnosis. Recent strides in electrocardiogram self-supervised learning (eSSL) have brightened prospects for enhancing representation learning without relying on high-quality annotations. Yet earlier eSSL methods suffer a key

Cited by 0SourcePDFScholar
2025

A Novel Local Search Algorithm for the Vertex Bisection Minimization Problem

IJCAI 2025

The vertex bisection minimization problem (VBMP) is a fundamental graph partitioning problem with numerous real-world applications. In this study, we propose a (k, l, S)-cluster guided local search algorithm to address this challenge. First, we propose a novel (k,l,S)-cluster enumeration procedure,

2025

Alleviate and Mining: Rethinking Unsupervised Domain Adaptation for Mitochondria Segmentation from Pseudo-Label Perspective

AAAI 2025technical

Mitochondria segmentation from electron microscopy (EM) images plays a crucial role in biological and medical research. However, models trained on source domains often suffer from performance degradation when applied to target domains due to domain shift. Unsupervised domain adaptation (UDA) methods…

Cited by 1SourcePDFScholar
2025

Balanced Learning for Domain Adaptive Semantic Segmentation

ICML 2025poster

Unsupervised domain adaptation (UDA) for semantic segmentation aims to transfer knowledge from a labeled source domain to an unlabeled target domain. Despite the effectiveness of self-training techniques in UDA, they struggle to learn each class in a balanced manner due to inherent class imbalance a…

Cited by 0SourcePDFScholar
2025

Beyond Confidence: Exploiting Homogeneous Pattern for Semi-Supervised Semantic Segmentation

ICML 2025poster

The critical challenge of semi-supervised semantic segmentation lies in how to fully exploit a large volume of unlabeled data to improve the model's generalization performance for robust segmentation. Existing methods mainly rely on confidence-based scoring functions in the prediction space to filte…

Cited by 0SourcePDFScholar
2025

BeyondMix: Leveraging Structural Priors and Long-Range Dependencies for Domain-Invariant LiDAR Segmentation

NeurIPS 2025poster

Domain adaptation for LiDAR semantic segmentation remains challenging due to the complex structural properties of point cloud data. While mix-based paradigms have shown promise, they often fail to fully leverage the rich structural priors inherent in 3D LiDAR point clouds. In this paper, we identify…

Cited by 0SourceScholar
2025

D2SA: Dual-Stage Distribution and Slice Adaptation for Efficient Test-Time Adaptation in MRI Reconstruction

NeurIPS 2025poster

Variations in Magnetic resonance imaging (MRI) scanners and acquisition protocols cause distribution shifts that degrade reconstruction performance on unseen data. Test-time adaptation (TTA) offers a promising solution to address this discrepancies. However, previous single-shot TTA approaches are…

Cited by 0SourceScholar
2025

Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence

NeurIPS 2025spotlight

AI agents today are mostly siloed — they either retrieve and reason over vast amount of digital information and knowledge obtained online; or interact with the physical world through embodied perception, planning and action — but rarely both. This separation limits their ability to solve tasks that…

Cited by 0SourceScholar
2025

InfVC: An Inference-Enhanced Local Search Algorithm for the Minimum Vertex Cover Problem in Massive Graphs

IJCAI 2025

The minimum vertex cover (MVC) problem is a classic NP-hard combinatorial optimization problem with extensive real-world applications. In this paper, we propose an efficient local search algorithm, InfVC, to solve the MVC in massive graphs, which comprises three ideas. First, we introduce an inferen

2025

LAGD: Local Topological-Alignment and Global Semantic-Deconstruction for Incremental 3D Semantic Segmentation

AAAI 2025technical

Numerous deep learning-based works focusing on 3D semantic segmentation have been proposed and have achieved impressive performance. However, due to the catastrophic forgetting, existing methods will degrade dramatically in a real-world scenario where new 3D semantic categories are arriving continua…

Cited by 0SourcePDFScholar
2025

Noise Diffusion for Enhancing Semantic Faithfulness in Text-to-Image Synthesis

CVPR 2025poster

Diffusion models have achieved impressive success in generating photorealistic images, but challenges remain in ensuring precise semantic alignment with input prompts. Optimizing the initial noisy latent offers a more efficient alternative to modifying model architectures or prompt engineering for i…

Cited by 0SourcePDFScholar
2025

NuMDS: An Efficient Local Search Algorithm for Minimum Dominating Set Problem

IJCAI 2025

The minimum dominating set (MDS) problem is a crucial NP-hard combinatorial optimization problem with wide applications in real-world scenarios. In this paper, we propose an efficient local search algorithm namely NuMDS to solve the MDS, which comprises three key ideas. First, we introduce a dominat

2025

Scaling Autonomous Agents via Automatic Reward Modeling And Planning

ICLR 2025poster

Large language models (LLMs) have demonstrated remarkable capabilities across a range of text-generation tasks. However, LLMs still struggle with problems requiring multi-step decision-making and environmental feedback, such as online shopping, scientific reasoning, and mathematical problem-solving.…

Cited by 3SourcePDFScholar
2025

Towards Robust Pseudo-Label Learning in Semantic Segmentation: An Encoding Perspective

NeurIPS 2025poster

Pseudo-label learning is widely used in semantic segmentation, particularly in label-scarce scenarios such as unsupervised domain adaptation (UDA) and semi-supervised learning (SSL). Despite its success, this paradigm can generate erroneous pseudo-labels, which are further amplified during training…

Cited by 0SourcecodeScholar
2025

Towards Unsupervised Domain Bridging via Image Degradation in Semantic Segmentation

NeurIPS 2025poster

Semantic segmentation suffers from significant performance degradation when the trained network is applied to a different domain. To address this issue, unsupervised domain adaptation (UDA) has been extensively studied. Despite the effectiveness of selftraining techniques in UDA, they still overlo…

Cited by 0SourcecodeScholar
2025

Two Losses, One Goal: Balancing Conflict Gradients for Semi-supervised Semantic Segmentation

ICCV 2025poster

Semi-supervised semantic segmentation has attracted considerable attention as it alleviates the need for extensive pixel-level annotations. However, existing methods often overlook the potential optimization conflict between supervised and unsupervised learning objectives, leading to suboptimal perf…

Cited by 0SourcePDFScholar
2024

Efficient Knowledge Infusion via KG-LLM Alignment

ACL 2024findings

To tackle the problem of domain-specific knowledge scarcity within large language models (LLMs), knowledge graph-retrievalaugmented method has been proven to be an effective and efficient technique for knowledge infusion. However, existing approaches face two primary challenges: knowledge mismatch b…

2024

Electron Microscopy Images as Set of Fragments for Mitochondrial Segmentation

AAAI 2024technical

Automatic mitochondrial segmentation enjoys great popularity with the development of deep learning. However, the coarse prediction raised by the presence of regular 3D grids in previous methods regardless of 3D CNN or the vision transformers suggest a possibly sub-optimal feature arrangement. To mit…

Cited by 8SourcePDFScholar
2024

Exploring Reliable Matching with Phase Enhancement for Night-time Semantic Segmentation

ECCV 2024poster

"Semantic segmentation of night-time images holds significant importance in computer vision, particularly for applications like night environment perception in autonomous driving systems. However, existing methods tend to parse night-time images from a day-time perspective, leaving the inherent chal…

Cited by 4SourcePDFScholar
2024

GENOME: Generative Neuro-Symbolic Visual Reasoning by Growing and Reusing Modules

ICLR 2024poster

Recent works have shown that Large Language Models (LLMs) could empower traditional neuro-symbolic models via programming capabilities to translate languages into module descriptions, thus achieving strong visual reasoning results while maintaining the model’s transparency and efficiency. However, t…

Cited by 18SourcePDFScholar
2024

GeoAB: Towards Realistic Antibody Design and Reliable Affinity Maturation

ICML 2024poster

Increasing works for antibody design are emerging to generate sequences and structures in Complementarity Determining Regions (CDRs), but problems still exist. We focus on two of them: (i) authenticity of the generated structure and (ii) rationality of the affinity maturation, and propose GeoAB as a…

Cited by 13SourcePDFScholar
2024

Image-to-Image Matching via Foundation Models: A New Perspective for Open-Vocabulary Semantic Segmentation

CVPR 2024poster

Open-vocabulary semantic segmentation (OVS) aims to segment images of arbitrary categories specified by class labels or captions. However most previous best-performing methods whether pixel grouping methods or region recognition methods suffer from false matches between image features and category l…

Cited by 15SourcePDFScholar
2024

JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images

NeurIPS 2024poster

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual understanding by relying on background language biases. Thus, strong performance on th…

2024

Learning Encodings for Constructive Neural Combinatorial Optimization Needs to Regret

AAAI 2024technical

Deep-reinforcement-learning (DRL) based neural combinatorial optimization (NCO) methods have demonstrated efficiency without relying on the guidance of optimal solutions. As the most mainstream among them, the learning constructive heuristic (LCH) achieves high-quality solutions through a rapid auto…

2024

Localization and Expansion: A Decoupled Framework for Point Cloud Few-shot Semantic Segmentation

ECCV 2024poster

"Point cloud few-shot semantic segmentation (PC-FSS) aims to segment targets of novel categories in a given query point cloud with only a few annotated support samples. The current top-performing prototypical learning methods employ prototypes originating from support samples to direct the classific…

Cited by 5SourcePDFScholar
2024

Nukplex: An Efficient Local Search Algorithm for Maximum K-Plex Problem

IJCAI 2024poster

The maximum k-plex problem (MKPP) is an significant relaxation version of the maximum clique problem with extensive applications. Recently, lots of researchers have proposed many heuristic algorithms based on various methods to solve the MKPP. In this work, to further improve the performance of solv…

2024

Pay Attention to Target: Relation-Aware Temporal Consistency for Domain Adaptive Video Semantic Segmentation

AAAI 2024technical

Video semantic segmentation has achieved conspicuous achievements attributed to the development of deep learning, but suffers from labor-intensive annotated training data gathering. To alleviate the data-hunger issue, domain adaptation approaches are developed in the hope of adapting the model train…

Cited by 14SourcePDFScholar
2024

RankMatch: Exploring the Better Consistency Regularization for Semi-supervised Semantic Segmentation

CVPR 2024poster

The key lie in semi-supervised semantic segmentation is how to fully exploit substantial unlabeled data to improve the model's generalization performance by resorting to constructing effective supervision signals. Most methods tend to directly apply contrastive learning to seek additional supervisio…

2024

Reverse Chain: A Generic-Rule for LLMs to Master Multi-API Planning

NAACL 2024findings

While enabling large language models to implement function calling (known as APIs) can greatly enhance the performance of Large Language Models (LLMs), function calling is still a challenging task due to the complicated relations between different APIs, especially in a context-learning setting witho…

2024

SafetyBench: Evaluating the Safety of Large Language Models

ACL 2024long

With the rapid development of Large Language Models (LLMs), increasing attention has been paid to their safety concerns. Consequently, evaluating the safety of LLMs has become an essential task for facilitating the broad applications of LLMs. Nevertheless, the absence of comprehensive safety evaluat…

2023

Adaptive Template Transformer for Mitochondria Segmentation in Electron Microscopy Images

ICCV 2023poster

Mitochondria, as tiny structures within the cell, are of significant importance to study cell functions for biological and clinical analysis. And exploring how to automatically segment mitochondria in electron microscopy (EM) images has attracted increasing attention. However, most of existing metho…

Cited by 19PDFScholar
2023

Alignment Before Aggregation: Trajectory Memory Retrieval Network for Video Object Segmentation

ICCV 2023poster

Memory-based methods in semi-supervised video object segmentation task achieve competitive performance by performing dense matching between query and memory frames. However, most of the existing methods neglect the fact that videos carry rich temporal information yet redundant spatial information. I…

Cited by 15PDFScholar
2023

Appearance Prompt Vision Transformer for Connectome Reconstruction

IJCAI 2023poster

Neural connectivity reconstruction aims to understand the function of biological reconstruction and promote basic scientific research. The intricate morphology and densely intertwined branches make it an extremely challenging task. Most previous best-performing methods adopt affinity learning or met…

Cited by 16SourcePDFScholar
2023

Camouflaged Instance Segmentation via Explicit De-Camouflaging

CVPR 2023highlight

Camouflaged Instance Segmentation (CIS) aims at predicting the instance-level masks of camouflaged objects, which are usually the animals in the wild adapting their appearance to match the surroundings. Previous instance segmentation methods perform poorly on this task as they are easily disturbed b…

Cited by 37SourcePDFScholar
2023

CiT-Net: Convolutional Neural Networks Hand in Hand with Vision Transformers for Medical Image Segmentation

IJCAI 2023poster

The hybrid architecture of convolutional neural networks (CNNs) and Transformer are very popular for medical image segmentation. However, it suffers from two challenges. First, although a CNNs branch can capture the local image features using vanilla convolution, it cannot achieve adaptive feature l…

2023

DAW: Exploring the Better Weighting Function for Semi-supervised Semantic Segmentation

NeurIPS 2023poster

The critical challenge of semi-supervised semantic segmentation lies in how to fully exploit a large volume of unlabeled data to improve the model’s generalization performance for robust segmentation. Existing methods tend to employ certain criteria (weighting function) to select pixel-level pseudo…

Cited by 21SourcePDFScholar
2023

DualRel: Semi-Supervised Mitochondria Segmentation From a Prototype Perspective

CVPR 2023poster

Automatic mitochondria segmentation enjoys great popularity with the development of deep learning. However, existing methods rely heavily on the labor-intensive manual gathering by experienced domain experts. And naively applying semi-supervised segmentation methods in the natural image field to mit…

Cited by 25SourcePDFScholar
2023

Fair-CDA: Continuous and Directional Augmentation for Group Fairness

AAAI 2023technical

In this work, we propose Fair-CDA, a fine-grained data augmentation strategy for imposing fairness constraints. We use a feature disentanglement method to extract the features highly related to the sensitive attributes. Then we show that group fairness can be achieved by regularizing the models on t…

Cited by 3SourcePDFScholar
2023

IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models

EMNLP 2023long findings

The field of vision-and-language (VL) understanding has made unprecedented progress with end-to-end large pre-trained VL models (VLMs). However, they still fall short in zero-shot reasoning tasks that require multi-step inferencing. To achieve this goal, previous works resort to a divide-and-conquer…

Cited by 0SourcecodeScholar
2023

Kernel-Based Tests for Likelihood-Free Hypothesis Testing

NeurIPS 2023poster

Given $n$ observations from two balanced classes, consider the task of labeling an additional $m$ inputs that are known to all belong to \emph{one} of the two classes. Special cases of this problem are well-known: with complete knowledge of class distributions ($n=\infty$) the problem is solved opt…

2023

UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding

ACL 2023findings

Vision-language tasks, such as VQA, SNLI-VE, and VCR are challenging because they require the model’s reasoning ability to understand the semantics of the visual world and natural language. Supervised methods working for vision-language tasks have been well-studied. However, solving these tasks in a…

2022

Find Someone Who: Visual Commonsense Understanding in Human-Centric Grounding

EMNLP 2022finding

From a visual scene containing multiple people, human is able to distinguish each individual given the context descriptions about what happened before, their mental/physical states or intentions, etc. Above ability heavily relies on human-centric commonsense knowledge and reasoning. For example, if…

2021

DecAug: Out-of-Distribution Generalization via Decomposed Feature Representation and Semantic Augmentation

AAAI 2021technical

While deep learning demonstrates its strong ability to handle independent and identically distributed (IID) data, it often suffers from out-of-distribution (OoD) generalization, where the test data come from another distribution (w.r.t. the training one). Designing a general OoD generalization frame…

Cited by 86SourcePDFScholar
2021

Lesion-Aware Transformers for Diabetic Retinopathy Grading

CVPR 2021poster

Diabetic retinopathy (DR) is the leading cause of permanent blindness in the working-age population. And automatic DR diagnosis can assist ophthalmologists to design tailored treatments for patients, including DR grading and lesion discovery. However, most of existing methods treat DR grading and le…

Cited by 137PDFScholar
2021

MetaAugment: Sample-Aware Data Augmentation Policy Learning

AAAI 2021technical

Automated data augmentation has shown superior performance in image recognition. Existing works search for dataset-level augmentation policies without considering individual sample variations, which are likely to be sub-optimal. On the other hand, learning different policies for different samples na…

Cited by 40SourcePDFScholar