← Search

Xiaoqing Guo

12 accepted papers

2026

U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding

ICLR 2026poster

Ultrasound is a widely-used imaging modality critical to global healthcare, yet its interpretation remains challenging due to its varying image quality on operators, noises, and anatomical structures. Although large vision-language models (LVLMs) have demonstrated impressive multimodal capabilities…

Cited by 0SourceScholar
2026

Ultrasound-CLIP: Semantic-Aware Contrastive Pre-training for Ultrasound Image-Text Understanding

CVPR 2026

Ultrasound imaging is widely used in clinical diagnostics due to its real-time capability and radiation-free nature. However, existing vision-language pre-training models, such as CLIP, are primarily designed for other modalities, and are difficult to directly apply to ultrasound data, which exhibit

Cited by 0SourcecodeScholar
2026

Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation

CVPR 2026

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in aligning visual inputs with natural language outputs. Yet, the extent to which generated tokens depend on visual modalities remains poorly understood, limiting interpretability and reliability. In this work, we pre

Cited by 0SourcecodeScholar
2025

GaussianReg: Rapid 2D/3D Registration for Emergency Surgery via Explicit 3D Modeling with Gaussian Primitives

ICCV 2025poster

Intraoperative 2D/3D registration, which aligns preoperative CT scans with intraoperative X-ray images, is critical for surgical navigation. However, existing methods require extensive preoperative training (several hours), making them unsuitable for emergency surgeries where minutes significantly i…

2024

Diversified and Personalized Multi-rater Medical Image Segmentation

CVPR 2024highlight

Annotation ambiguity due to inherent data uncertainties such as blurred boundaries in medical scans and different observer expertise and preferences has become a major obstacle for training deep-learning based medical image segmentation models. To address it the common practice is to gather multiple…

2023

Digital Twin-Driven Mixed Reality Framework for Immersive Teleoperation With Haptic Rendering

RA-L 2023

Teleoperation has widely contributed to many applications. Consequently, the design of intuitive and ergonomic control interfaces for teleoperation has become crucial. The rapid advancement of Mixed Reality (MR) has yielded tangible benefits in human-robot interaction. MR provides an immersive envir

Cited by 21SourceScholar
2023

TIMS: A Tactile Internet-Based Micromanipulation System with Haptic Guidance for Surgical Training

IROS 2023poster

Microsurgery involves the dexterous manipulation of delicate tissue or fragile structures, such as small blood vessels and nerves, under a microscope. To address the limitations of imprecise manipulation of human hands, robotic systems have been developed to assist surgeons in performing complex mic…

Cited by 8SourcecodeScholar
2022

SimT: Handling Open-Set Noise for Domain Adaptive Semantic Segmentation

CVPR 2022poster

This paper studies a practical domain adaptative (DA) semantic segmentation problem where only pseudo-labeled target data is accessible through a black-box model. Due to the domain gap and label shift between two domains, pseudo-labeled target data contains mixed closed-set and open-set label noises…

Cited by 34PDFcodeScholar
2021

COINet: Adaptive Segmentation with Co-Interactive Network for Autonomous Driving

IROS 2021poster

Semantic segmentation serves as a cornerstone for safety autonomous driving and has been achieved remarkable progress at the price of dense annotations. Unsupervised domain adaptation was widely utilized to addresses this labor-intensive problem, which transfers the knowledge learned from labeled sy…

Cited by 6SourceScholar
2021

MetaCorrection: Domain-Aware Meta Loss Correction for Unsupervised Domain Adaptation in Semantic Segmentation

CVPR 2021poster

Unsupervised domain adaptation (UDA) aims to transfer the knowledge from the labeled source domain to the unlabeled target domain. Existing self-training based UDA approaches assign pseudo labels for target data and treat them as ground truth labels to fully leverage unlabeled target data for model…

Cited by 119PDFcodeScholar