← Search

Junhao Xiao

11 accepted papers

2026

Not Just What’s There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-Tuning

AAAI 2026technical

Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing methods refine negation understanding via fine-tuning CLIP’s text encoder, risking overfitting. In this work, we propose C

Cited by 0SourcePDFScholar
2026

SurfAAV: Design and Implementation of a Novel Multimodal Surfing Aquatic-Aerial Vehicle

ICRA 2026poster

Despite significant advancements in the research of aquatic-aerial robots, existing configurations struggle to efficiently perform underwater, surface, and aerial movement. In this paper, we propose a novel multimodal surfing aquatic-aerial vehicle, SurfAAV, which efficiently integrates underwater n…

2025

A Novel Decomposed Feature-Oriented Framework for Open-Set Semantic Segmentation on LiDAR Data

ICRA 2025

Semantic segmentation is a key technique that enables mobile robots to understand and navigate surrounding environments autonomously. However, most existing works focus on segmenting known objects, overlooking the identification of unknown classes, which is common in real-world applications. In this

Cited by 1SourcecodeScholar
2025

Breaking Grid Constraints: Dynamic Graph Reconstruction Network for Multi-organ Segmentation

ICCV 2025poster

Morphological differences and dense spatial relations of organs make multi-organ segmentation challenging. Current segmentation networks, primarily based on CNNs and Transformers, represent organs by aggregating information within fixed regions. However, aggregated representations often fail to accu…

2025

Efficient Multimodal 3D Object Detector via Instance-Level Contrastive Distillation

IROS 2025

Multimodal 3D object detectors leverage the strengths of both geometry-aware LiDAR point clouds and semantically rich RGB images to enhance detection performance. However, the inherent heterogeneity between these modalities, including unbalanced convergence and modal misalignment, poses significant

Cited by 1SourcecodeScholar
2025

RecNet: Optimization for Dense Object Detection in Retail Scenarios Based on View Rectification

ICASSP 2025accepted

High-precision dense object detection in retail is crucial for automation, inventory management, and sales optimization. Our experiments revealed that detection models perform significantly better with frontal views than with oblique views, motivating the development of RecNet. RecNet utilizes a Rec…

Cited by 0SourceScholar
2025

SurfAAV: Design and Implementation of a Novel Multimodal Surfing Aquatic-Aerial Vehicle

RA-L 2025

Despite significant advancements in the research of aquatic-aerial robots, existing configurations struggle to efficiently perform underwater, surface, and aerial movement. In this paper, we propose a novel multimodal surfing aquaticaerial vehicle, SurfAAV, which efficiently integrates underwater na

Cited by 1SourceScholar
2025

Who Controls the Authorization? Invertible Networks for Copyright Protection in Text-to-Image Synthesis

ICCV 2025poster

To defend against personalized generation, a new form of infringement that is more concealed and destructive, the existing copyright protection methods is to add adversarial perturbations in images. However, these methods focus solely on countering illegal personalization, neglecting the requirement…

Cited by 0SourcePDFScholar
2023

Hybrid Map-Based Path Planning for Robot Navigation in Unstructured Environments

IROS 2023poster

Fast and accurate path planning is important for ground robots to achieve safe and efficient autonomous navigation in unstructured outdoor environments. However, most existing methods exploiting either 2D or 2.5D maps struggle to balance the efficiency and safety for ground robots navigating in such…

Cited by 14SourcecodeScholar
2021

Keypoint Matching for Point Cloud Registration Using Multiplex Dynamic Graph Attention Networks

RA-L 2021

The registration of point clouds is a key ingredient of LiDAR-based SLAM systems and mapping approaches. A challenging task in this context is finding the right data association between 3D points. This paper proposes a novel and flexible graph network architecture to tackle the keypoint matching pro

Cited by 54SourceScholar
2019

Accurate Direct Visual-Laser Odometry with Explicit Occlusion Handling and Plane Detection

ICRA 2019poster

In this paper, we address the problem of combining 3D laser scanner and camera information to estimate the motion of a mobile platform. We propose a direct laser-visual odometry approach building upon photometric image alignment. Our approach is designed to maximize the information usage of both, th…

Cited by 24SourceScholar