← Search

Yujia Chen

19 accepted papers

2026

Adaptive Agent Selection and Interaction Network for Image-to-Point Cloud Registration

AAAI 2026technical

Typical detection-free methods for image-to-point cloud registration leverage transformer-based architectures to aggregate cross-modal features and establish correspondences. However, they often struggle under challenging conditions, where noise disrupts similarity computation and leads to incorrect

Cited by 0SourcePDFScholar
2026

Adaptive Augmentation-Aware Latent Learning for Robust LiDAR Semantic Segmentation

ICLR 2026poster

Adverse weather conditions significantly degrade the performance of LiDAR point cloud semantic segmentation networks by introducing large distribution shifts. Existing augmentation-based methods attempt to enhance robustness by simulating weather interference during training. However, they struggle…

Cited by 0SourceScholar
2026

Beyond Blind Noising: Disentangled Visual Rectification for Hallucination Mitigation in MLLMs

ICML 2026poster

Visual Contrastive Decoding (VCD) mitigates hallucinations in Multimodal Large Language Models (MLLMs) by penalizing the output shift from noise-perturbed images, assuming this shift captures the hallucination direction. We prove this assumption flawed: noise-induced drift in Language-Image Pretrain…

Cited by 0SourceScholar
2026

Beyond Logits: Coherent Hallucination Mitigation via Attention Contrastive Decoding

ICML 2026poster

Large Vision-Language Models (LVLMs) demonstrate impressive multimodal capabilities, yet suffer from hallucination—generating factually inaccurate content. Contrastive Decoding (CD) mitigates this by contrasting amateur and expert branches at the logit level. However, our investigation reveals that …

Cited by 0SourceScholar
2026

FS-I2P: A Hierarchical Focus–Sweep Registration Network with Dynamically Allocated Depth

ICML 2026poster

Image-to-point cloud registration is often challenged by viewpoint changes, cross-modal discrepancies, and repetitive textures, which induce scale ambiguity and consequently lead to erroneous correspondences. Recent detection-free methods alleviate this issue by leveraging multi-scale features and t…

Cited by 0SourceScholar
2026

From Softmax to Dirichlet: Evidential Learning for Semi-supervised Semantic Segmentation

CVPR 2026

The critical challenge of semi-supervised semantic segmentation lies in how to fully exploit a large volume of unlabeled data to improve the model's generalization performance for robust segmentation. However, existing softmax scores-based filtering methods tend to be affected by the overconfidence

Cited by 0SourceScholar
2026

GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation

CVPR 2026

Open-vocabulary 3D semantic segmentation aims to segment arbitrary categories beyond the training set. Existing methods predominantly rely on distilling knowledge from 2D open-vocabulary models. However, aligning 3D features to the 2D representation space restricts intrinsic 3D geometric learning an

Cited by 0SourceScholar
2026

Rethinking 2D-3D Registration: A Novel Network for High-Value Zone Selection and Representation Consistency Alignment

CVPR 2026

Both detection-then-match and detection-free methods have been extensively studied for image-to-point cloud registration, yet they still face significant challenges. The detection-then-match approach emphasizes high-quality correspondences but is limited by the availability of repeatable keypoints,

Cited by 0SourceScholar
2026

Spectral Heat Flow for Conservative Token Condensation in Vision-Language Models

ICML 2026poster

Vision-Language Models (VLMs) are costly at inference time because they must process long sequences of visual tokens. Existing token pruning methods often degrade under high compression by blindly discarding information, breaking spatial structure or collapsing diversity. We propose SpecFlow, a trai…

Cited by 0SourceScholar
2025

Alleviate and Mining: Rethinking Unsupervised Domain Adaptation for Mitochondria Segmentation from Pseudo-Label Perspective

AAAI 2025technical

Mitochondria segmentation from electron microscopy (EM) images plays a crucial role in biological and medical research. However, models trained on source domains often suffer from performance degradation when applied to target domains due to domain shift. Unsupervised domain adaptation (UDA) methods…

Cited by 1SourcePDFScholar
2025

Beyond Confidence: Exploiting Homogeneous Pattern for Semi-Supervised Semantic Segmentation

ICML 2025poster

The critical challenge of semi-supervised semantic segmentation lies in how to fully exploit a large volume of unlabeled data to improve the model's generalization performance for robust segmentation. Existing methods mainly rely on confidence-based scoring functions in the prediction space to filte…

Cited by 0SourcePDFScholar
2025

BeyondMix: Leveraging Structural Priors and Long-Range Dependencies for Domain-Invariant LiDAR Segmentation

NeurIPS 2025poster

Domain adaptation for LiDAR semantic segmentation remains challenging due to the complex structural properties of point cloud data. While mix-based paradigms have shown promise, they often fail to fully leverage the rich structural priors inherent in 3D LiDAR point clouds. In this paper, we identify…

Cited by 0SourceScholar
2025

MIND: Towards Immersive Psychological Healing with Multi-Agent Inner Dialogue

EMNLP 2025

Mental health issues are worsening in today’s competitive society, such as depression and anxiety. Traditional healings like counseling and chatbots fail to engage effectively, they often provide generic responses lacking emotional depth. Although large language models (LLMs) have the potential to c

Cited by 0SourcePDFScholar
2025

RB-Modulation: Training-Free Stylization using Reference-Based Modulation

ICLR 2025oral

We propose Reference-Based Modulation (RB-Modulation), a new plug-and-play solution for training-free personalization of diffusion models. Existing training-free approaches exhibit difficulties in (a) style extraction from reference images in the absence of additional style or content text descripti…

2025

SEAL: Semantic Attention Learning for Long Video Representation

CVPR 2025poster

Long video understanding presents challenges due to the inherent high computational complexity and redundant temporal information. An effective representation for long videos must efficiently process such redundancy while preserving essential contents for downstream tasks. This paper introduces **S…

Cited by 0SourcePDFScholar
2025

Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations

ICLR 2025poster

Generative models transform random noise into images, while their inversion aims to reconstruct structured noise for recovery and editing. This paper addresses two key tasks: (i) *inversion* and (ii) *editing* of real images using stochastic equivalents of rectified flow models (e.g., Flux). While D…

2025

Two Losses, One Goal: Balancing Conflict Gradients for Semi-supervised Semantic Segmentation

ICCV 2025poster

Semi-supervised semantic segmentation has attracted considerable attention as it alleviates the need for extensive pixel-level annotations. However, existing methods often overlook the potential optimization conflict between supervised and unsupervised learning objectives, leading to suboptimal perf…

Cited by 0SourcePDFScholar
2024

Beyond First-Order Tweedie: Solving Inverse Problems using Latent Diffusion

CVPR 2024poster

Sampling from the posterior distribution in latent diffusion models for inverse problems is computationally challenging. Existing methods often rely on Tweedie's first-order moments that tend to induce biased results. Second-order approximations are computationally prohibitive making standard revers…

Cited by 26SourcePDFScholar