← Search

Yulin He

11 accepted papers

2026

Attention to Threat-Relevant Objects: Reasoning Detection in Autonomous Driving via Multimodal Large Language Models

AAAI 2026technical

Perceiving threats is an innate human instinct. During driving, humans naturally focus their attention on objects that pose real potential risks. Motivated by this observation, we shift the focus from traditional class-based detection to a novel task termed threat-oriented reasoning detection in aut

Cited by 0SourcePDFScholar
2026

DR$^2$Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models

ICML 2026poster

Reasoning segmentation is an emerging vision-language task that requires reasoning over intricate text queries to precisely segment objects. However, existing methods typically suffer from overthinking, generating verbose reasoning chains that interfere with object localization in multimodal large l…

Cited by 0SourceScholar
2025

A Cost-effective Solution for Remote Sensing Image Segmentation via Train/Test-Time Adaptation

ICASSP 2025accepted

Remote Sensing Image (RSI) segmentation has made significant strides, emerging as a leading solution for interpreting remote sensing data. However, due to the substantial domain gap between different remote sensors and limited computational resources, existing RSI segmentation methods often suffer f…

Cited by 0SourceScholar
2025

Achieving Speed-Accuracy Balance in Vision-based 3D Occupancy Prediction via Geometric-Semantic Disentanglement

AAAI 2025technical

Occupancy prediction plays a pivotal role in autonomous driving (AD) due to its capabilities of fine-grained 3D perception and general object recognition. However, existing methods often incur high computational costs, which conflict with AD's real-time demand. To this end, we redirect the focus fro…

2025

End-to-End Low-Light Enhancement for Object Detection with Learned Metadata from RAWs

NeurIPS 2025poster

Although RAW images offer advantages over sRGB by avoiding ISP-induced distortion and preserving more information in low-light conditions, their widespread use is limited due to high storage costs, transmission burdens, and the need for significant architectural changes for downstream tasks. To addr…

Cited by 0SourceScholar
2025

Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision

AAAI 2025technical

The image compression model has long struggled with adaptability and generalization, as the decoded bitstream typically serves only human or machine needs and fails to preserve information for unseen visual tasks. Therefore, this paper innovatively introduces supervision obtained from multimodal pre…

2024

Image Coding for Analytics via Adversarially Augmented Adaptation

ICASSP 2024accepted

Image Coding for Machine (ICM) aims to compress an image so that the reconstructed one can meet the requirements of both human vision and machine vision. Existing methods apply the constraint from the downstream models to improve machine analytics performance while compromising the visual quality. T…

Cited by 0SourceScholar
2023

Reconstruction-Aware Prior Distillation for Semi-supervised Point Cloud Completion

IJCAI 2023poster

Real-world sensors often produce incomplete, irregular, and noisy point clouds, making point cloud completion increasingly important. However, most existing completion methods rely on large paired datasets for training, which is labor-intensive. This paper proposes RaPD, a novel semi-supervised poin…

Cited by 14SourcePDFScholar
2022

Adaptive Pseudo Labeling for Source-Free Domain Adaptation in Medical Image Segmentation

ICASSP 2022accepted

Domain adaptation is common but challenging in signal processing tasks due to the intrinsic discrepancy, especially in difficult-to-label medical image segmentation application scenarios. Pseudo labeling methods are widely utilized to compensate for the scarcity of annotation. However, most existing…

Cited by 0SourceScholar
2022

SCMT: Self-Correction Mean Teacher for Semi-supervised Object Detection

IJCAI 2022poster

Semi-Supervised Object Detection (SSOD) aims to improve performance by leveraging a large amount of unlabeled data. Existing works usually adopt the teacher-student framework to enforce student to learn consistent predictions over the pseudo-labels generated by teacher. However, the performance of t…

Cited by 8SourcePDFScholar