← Search

Lihao Liu

14 accepted papers

2026

ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction

AAAI 2026technical

Large language models (LLMs) are increasingly used in scientific domains. While they can produce reasoning-like content via methods such as chain-of-thought prompting, these outputs are typically unstructured and informal, obscuring whether models truly understand the fundamental reasoning paradigms

Cited by 0SourcePDFScholar
2026

EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval

CVPR 2026

Retrieval-augmented generation (RAG) has emerged as a critical paradigm for grounding Multimodal Large Language Models (MLLMs) in external knowledge. Recent GraphRAG methods introduce structured entity-relation graphs to improve retrieval and reasoning. However, they remain limited by treating knowl

Cited by 0SourceScholar
2026

HyperST: Hierarchical Hyperbolic Learning for Spatial Transcriptomics Prediction

CVPR 2026

Spatial Transcriptomics (ST) merges the benefits of pathology images and gene expression, linking molecular profiles with tissue structure to analyze spot-level function comprehensively. Predicting gene expression from histology images is a cost-effective alternative to expensive ST technologies. Ho

Cited by 0SourcecodeScholar
2026

IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation

AAAI 2026technical

Recent visual generative models enable story generation with consistent characters from text, but human-centric story generation faces additional challenges, such as maintaining detailed and diverse human face consistency and coordinating multiple characters across different images. This paper prese

Cited by 0SourcePDFScholar
2026

S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything Without Supervision

AAAI 2026technical

Recent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining schedule is multi-stage, requiring a time-consuming pseudo-masks generation process between each training epoch. This

Cited by 0SourcePDFScholar
2026

UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis

ICML 2026poster

Medical diagnosis demands models that can process multimodal medical inputs, such as medical images and patient histories, and generate diverse outputs including textual reports and visual content, such as annotations or segmentation masks. Despite this need, existing medical AI models disrupt this …

Cited by 0SourceScholar
2025

A Magnetically-Actuated Ultrasound Capsule Endoscope (MUSCE) for Endoluminal Imaging in Tubular Environments

RA-L 2025

Endoscopic ultrasound (EUS) has the ability to image tissue in and beyond the wall of the gastrointestinal (GI) tract, assisting in the early diagnosis of digestive diseases. However, traditional EUS based on flexible endoscopes could make the operation procedure traumatic and intolerable to patient

Cited by 9SourceScholar
2025

Detect Any Mirrors: Boosting Learning Reliability on Large-Scale Unlabeled Data with an Iterative Data Engine

CVPR 2025poster

Mirror detection is a challenging task because a mirror's visual appearance varies depending on the reflected content. Due to limited annotated data, current methods failed to generalize well for detecting diverse mirror scenes. Semi-supervised learning with large-scale unlabeled data can improve ge…

2025

Simultaneous 6-DOF localization and scanning angle detection of magnetic ultrasound capsule endoscope (MUSCE) with internal sensors

IROS 2025

Localization of magnetically actuated capsule endoscope (MCE) is essential for accurate actuation. Despite extensive progress in pose estimation using internal magnetic field sensors and external magnetic sources, it remains challenging to achieve localization when a time-varying internal magnetic f

Cited by 0SourceScholar
2025

Throwing Planning Diffusion: A Solution to Learning and Planning of Robotic Throwing

IROS 2025

Dynamic manipulation enables efficient interaction tasks, such as throwing, which rely on finding one or more high-quality trajectories from the initial state to the goal state. While model-free learning methods have been used to acquire efficient robot manipulation configurations, traditional plann

Cited by 0SourceScholar
2025

Toward Fair and Accurate Cross-Domain Medical Image Segmentation: A VLM-Driven Active Domain Adaptation Paradigm

ICCV 2025poster

Fairness in AI-assisted medical image analysis is crucial for equitable healthcare, but is often neglected, especially in prevalent cross-domain scenarios (diverse demographics and imaging protocols). Effective and equitable deployment of AI models in these scenarios is critical, yet traditional Uns…

2024

YOLOv10: Real-Time End-to-End Object Detection

NeurIPS 2024poster

Over the past years, YOLOs have emerged as the predominant paradigm in the field of real-time object detection owing to their effective balance between computational cost and detection performance. Researchers have explored the architectural designs, optimization objectives, data augmentation strate…

2023

SCOTCH and SODA: A Transformer Video Shadow Detection Framework

CVPR 2023poster

Shadows in videos are difficult to detect because of the large shadow deformation between frames. In this work, we argue that accounting for shadow deformation is essential when designing a video shadow detection method. To this end, we introduce the shadow deformation attention trajectory (SODA), a…