← Search

Zongyu Li

8 accepted papers

2026

PSBench: Editing Image via GUI Agents in Photoshop

ICML 2026poster

Photoshop is a professional image editing software whose complex multi-level menus, fine-grained operations, and layer-based non-destructive editing pose substantial challenges for automated agents. Existing GUI benchmarks and methods primarily target web interfaces and short-horizon, low-complexity…

Cited by 0SourceScholar
2026

SDiD:Shared diffusion prior for efficient distributed stereo image compression

ICML 2026poster

Stereo vision is widely utilized in automotive imagery and 3D reconstruction, creating a demand for compressing stereo images. Existing methods for stereo image compression often employ VAE-like architectures based on distortion optimization, leading to subpar perceptual quality at low bitrates. Whi…

Cited by 0SourceScholar
2025

Reducing Confounding Bias without Data Splitting for Causal Inference via Optimal Transport

ICML 2025poster

Causal inference seeks to estimate the effect given a treatment such as a medicine or the dosage of a medication. To reduce the confounding bias caused by the non-randomized treatment assignment, most existing methods reduce the shift between subpopulations receiving different treatments. However, t…

Cited by 0SourcePDFScholar
2023

Evaluating the Task Generalization of Temporal Convolutional Networks for Surgical Gesture and Motion Recognition Using Kinematic Data

RA-L 2023

Fine-grained activity recognition enables explainable analysis of procedures for skill assessment, autonomy, and error detection in robot-assisted surgery. However, existing recognition models suffer from the limited availability of annotated datasets with both kinematic and video data and an inabil

Cited by 8SourcecodeScholar
2023

Robotic Scene Segmentation with Memory Network for Runtime Surgical Context Inference

IROS 2023poster

Surgical context inference has recently garnered significant attention in robot-assisted surgery as it can facilitate workflow analysis, skill assessment, and error detection. However, runtime context inference is challenging since it requires timely and accurate detection of the interactions among…

Cited by 2SourcecodeScholar
2023

Towards Surgical Context Inference and Translation to Gestures

ICRA 2023poster

Manual labeling of gestures in robot-assisted surgery is labor intensive, prone to errors, and requires expertise or training. We propose a method for automated and explainable generation of gesture transcripts that leverages the abundance of data for image segmentation. Surgical context is detected…

Cited by 4SourcecodeScholar