← Search

Xiaolong Liu

20 accepted papers

2026

Active Reasoning Vision-Language Model via Sequential Experimental Design

ICML 2026poster

Visual perception in modern Vision-Language Models (VLM) is constrained by a fundamental perceptual bandwidth bottleneck: a broad field-of-view inevitably sacrifices the fine-grained details necessary for complex reasoning. Inspired by the classical paradigms of active vision and information foragin…

Cited by 0SourceScholar
2026

Chain-of-Models Pre-Training: Rethinking Training Acceleration of Vision Foundation Models

CVPR 2026

In this paper, we present Chain-of-Models Pre-Training (CoM-PT), a novel performance-lossless training acceleration method for vision foundation models (VFMs). This approach fundamentally differs from existing acceleration methods in its core motivation: rather than optimizing each model individuall

Cited by 0SourcecodeScholar
2025

KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA

ICLR 2025poster

Biomedical reasoning integrates structured, codified knowledge with tacit, experience-driven insights. Depending on the context, quantity, and nature of available evidence, researchers and clinicians use diverse strategies, including rule-based, prototype-based, and case-based reasoning. Effective m…

Cited by 0SourcePDFScholar
2025

Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking

AAAI 2025technical

Visual Entity Linking (VEL) is a crucial task for achieving fine-grained visual understanding, matching objects within images (visual mentions) to entities in a knowledge base. Previous VEL tasks rely on textual inputs, but writing queries for complex scenes can be challenging. Visual inputs like cl…

2024

ScaleKD: Strong Vision Transformers Could Be Excellent Teachers

NeurIPS 2024poster

In this paper, we question if well pre-trained vision transformer (ViT) models could be used as teachers that exhibit scalable properties to advance cross architecture knowledge distillation research, in the context of adopting mainstream large-scale visual recognition datasets for evaluation. To ma…

2023

Augmentation-Free Dense Contrastive Knowledge Distillation for Efficient Semantic Segmentation

NeurIPS 2023poster

In recent years, knowledge distillation methods based on contrastive learning have achieved promising results on image classification and object detection tasks. However, in this line of research, we note that less attention is paid to semantic segmentation. Existing methods heavily rely on data aug…

2023

Class Lifelong Learning for Intent Detection via Structure Consolidation Networks

ACL 2023findings

Intent detection, which estimates diverse intents behind user utterances, is an essential component of task-oriented dialogue systems. Previous intent detection models are usually trained offline, which can only handle predefined intent classes. In the real world, new intents may keep challenging de…

Cited by 3SourcePDFScholar
2023

NORM: Knowledge Distillation via N-to-One Representation Matching

ICLR 2023poster

Existing feature distillation methods commonly adopt the One-to-one Representation Matching between any pre-selected teacher-student layer pair. In this paper, we present $N$-to-$O$ne $R$epresentation $M$atching (NORM), a new two-stage knowledge distillation method, which relies on a simpleFeature T…

2023

SOOD: Towards Semi-Supervised Oriented Object Detection

CVPR 2023poster

Semi-Supervised Object Detection (SSOD), aiming to explore unlabeled data for boosting object detectors, has become an active task in recent years. However, existing SSOD approaches mainly focus on horizontal objects, leaving multi-oriented objects that are common in aerial images unexplored. This p…

2022

Enhanced Accuracy in Magnetic Actuation: Closed-Loop Control of a Magnetic Agent With Low-Error Numerical Magnetic Model Estimation

RA-L 2022

Magnetic actuation holds promise for wirelessly controlling small, magnetic surgical tools and may enable the next generation of ultra minimally invasive surgical robotic systems. Precise torque and force exertion are required for safe surgical operations and accurate state control. Dipole field est

Cited by 10SourceScholar
2021

Localization and Control of Magnetic Suture Needles in Cluttered Surgical Site with Blood and Tissue

IROS 2021poster

Real-time visual localization of needles is necessary for various surgical applications, including surgical automation and visual feedback. In this study we investigate localization and autonomous robotic control of needles in the context of our magneto-suturing system. Our system holds the potentia…

Cited by 8SourceScholar
2021

Multi-Shot Temporal Event Localization: A Benchmark

CVPR 2021poster

Current developments in temporal event or action localization usually target actions captured by a single camera. However, extensive events or actions in the wild may be captured as a sequence of shots by multiple cameras at different positions. In this paper, we propose a new and challenging task c…

Cited by 109PDFcodeScholar
2020

Towards Autonomous Control of Magnetic Suture Needles

IROS 2020poster

This paper proposes a magnetic needle steering controller to manipulate mesoscale magnetic suture needles for executing planned suturing motion. This is an initial step towards our research objective: enabling autonomous control of magnetic suture needles for suturing tasks in minimally invasive sur…

Cited by 12SourceScholar
2019

Towards A Generic In Vivo In Situ Camera Lens Cleaning Module for Laparoscopic Surgery

IROS 2019poster

This paper proposes a generic cleaning module to address lens fogging and soiling problems for insertable robotic cameras in laparoscopic surgery. The proposed lens cleaning module features minimal intraoperative interruption for surgeons to maintain clear visual field. The technical challenges for…

Cited by 2SourceScholar
2018

Design and Test of an In-Vivo Robotic Camera Integrated with Optimized Illumination System for Single-port Laparoscopic Surgery

ICRA 2018poster

This paper proposes a novel in-vivo robotic laparo-scopic camera design with an optimized illumination system, which is a crucial component for achieving high imaging quality. The robotic camera design with three extendable wings can reserve sufficient on-board space to harbor the optimized illumina…

Cited by 3SourceScholar
2017

A novel laparoscopic camera robot with in-vivo lens cleaning and debris prevention modules

IROS 2017poster

Robotic systems have recently drawn attention in minimally invasive surgeries due to their increased dexterity feature. A major drawback of these systems is image blurring due to lens contamination which cause imaging impairment during up to 40% of surgery time. This paper demonstrates a novel lapar…

Cited by 15SourceScholar
2017

Modeling and analysis of a laparoscopic camera's interaction with abdomen tissue

ICRA 2017poster

Robotic camera systems have recently drawn attention in minimally invasive surgeries(MIS). Control and manipulation of these systems during traversing abdominal cavity is associated with camera-tissue interaction. This paper demonstrates a theoretical and experimental analysis of a wireless laparosc…

Cited by 14SourceScholar
2016

The design and experiments of a small wheel-legged mobile robot system with two robotic arms

IROS 2016poster

In this paper, we developed a small wheel-legged mobile robot system, which could walk on different road environments using wheels or legs. It is composed of mechanical, sensor and control subsystems. The mechanical subsystem includes a wheel-legged mobile platform, a rigid robotic arm and a flexibl…

Cited by 11SourceScholar
2015

Design and analysis of a magnetic actuated capsule camera robot for single incision laparoscopic surgery

IROS 2015poster

This paper presents the design of a novel insertable robotic capsule camera system for single incision laparoscopic surgery. This design features a unified mechanism for anchoring, navigating, and rotating an insertable camera by externally generated rotational magnetic field. The design is inspired…

Cited by 24SourceScholar