← Search

Zhaoshuo Li

14 accepted papers

2026

DuoGen: Towards Autonomous Interleaved Multimodal Generation

CVPR 2026

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts for reasoning. However, the quality of existing interleaved generation models under general instructions remains limited

Cited by 0SourceScholar
2026

SAGE: Scalable Agentic 3D Scene Generation for Embodied AI

CVPR 2026

Real-world data collection for embodied agents remains costly and unsafe, calling for scalable, realistic, and simulator-ready 3D environments. However, existing scene-generation systems often rely on rule-based or task-specific pipelines, yielding artifacts and physically invalid scenes. We present

Cited by 0SourcecodeScholar
2025

ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary

CVPR 2025poster

Designing 3D scenes is traditionally a challenging and laborious task that demands both artistic expertise and proficiency with complex software. Recent advances in text-to-3D generation have greatly simplified this process by letting users create scenes based on simple text descriptions. However, a…

Cited by 0SourcePDFScholar
2025

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

CVPR 2025poster

Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor control. While this paradigm effectively utilizes large-scale data from both robotic and non-robotic sources, current VLA…

2025

EdgeRunner: Auto-regressive Auto-encoder for Artistic Mesh Generation

ICLR 2025poster

Current auto-regressive mesh generation methods suffer from issues such as incompleteness, insufficient detail, and poor generalization. In this paper, we propose an Auto-regressive Auto-encoder (ArAE) model capable of generating high-quality 3D meshes with up to 4,000 faces at a spatial resolution…

Cited by 22SourcePDFScholar
2024

Ada-Tracker: Soft Tissue Tracking via Inter-Frame and Adaptive-template Matching

ICRA 2024poster

Soft tissue tracking is crucial for computer-assisted interventions. Existing approaches mainly rely on extracting discriminative features from the template and videos to recover corresponding matches. However, it is difficult to adopt these techniques in surgical scenes, where tissues are changing…

Cited by 2SourcecodeScholar
2023

Improving Surgical Situational Awareness with Signed Distance Field: A Pilot Study in Virtual Reality

IROS 2023poster

The introduction of image-guided surgical navigation (IGSN) has greatly benefited technically demanding surgical procedures by providing real-time support and guidance to the surgeon during surgery. To develop effective IGSN, a careful selection of the surgical information and the medium to present…

Cited by 5SourceScholar
2023

Neuralangelo: High-Fidelity Neural Surface Reconstruction

CVPR 2023poster

Neural surface reconstruction has been shown to be powerful for recovering dense 3D surfaces via image-based neural rendering. However, current methods struggle to recover detailed structures of real-world scenes. To address the issue, we present Neuralangelo, which combines the representation power…

2022

Context-Enhanced Stereo Transformer

ECCV 2022poster

"Stereo depth estimation is of great interest for computer vision research. However, existing methods struggles to generalize and predict reliably in hazardous regions, such as large uniform regions. To overcome these limitations, we propose Context Enhanced Path (CEP). CEP improves the generalizati…

2022

SAGE: SLAM with Appearance and Geometry Prior for Endoscopy

ICRA 2022poster

In endoscopy, many applications (e.g., surgical navigation) would benefit from a real-time method that can simultaneously track the endoscope and reconstruct the dense 3D geometry of the observed anatomy from a monocular endoscopic video. To this end, we develop a Simultaneous Localization and Mappi…

Cited by 44SourcecodeScholar
2021

Revisiting Stereo Depth Estimation From a Sequence-to-Sequence Perspective With Transformers

ICCV 2021poster

Stereo depth estimation relies on optimal correspondence matching between pixels on epipolar lines in the left and right images to infer depth. In this work, we revisit the problem from a sequence-to-sequence correspondence perspective to replace cost volume construction with dense pixel matching us…

Cited by 338PDFcodeScholar
2020

Anatomical Mesh-Based Virtual Fixtures for Surgical Robots

IROS 2020poster

This paper presents a dynamic constraint formulation to provide protective virtual fixtures of 3D anatomical structures from polygon mesh representations. The proposed approach can anisotropically limit the tool motion of surgical robots without any assumption of the local anatomical shape close to…

Cited by 18SourcecodeScholar
2019

A Novel Semi-Autonomous Control Framework for Retina Confocal Endomicroscopy Scanning

IROS 2019poster

In this paper, a novel semi-autonomous control framework is presented for enabling probe-based confocal laser endomicroscopy (pCLE) scan of the retinal tissue. With pCLE, retinal layers such as nerve fiber layer (NFL) and retinal ganglion cell (RGC) can be scanned and characterized in real-time for…

Cited by 1SourceScholar
2018

Free Head Movement Eye Gaze Contingent Ultrasound Interfaces for the da Vinci Surgical System

RA-L 2018

The current practice of intraoperative ultrasound requires an assistant because the surgeon's hands are occupied with surgical tools or console instruments. This process can be tedious and prone to error. Eye gaze is a promising control modality that can help address this issue. In previous work, a

Cited by 15SourceScholar