← Search

Yiming Zhao

20 accepted papers

2026

Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models

ICLR 2026poster

Although current large Vision-Language Models (VLMs) have advanced in multimodal understanding and reasoning, their fundamental perceptual and reasoning abilities remain limited. Specifically, even on simple jigsaw tasks, existing VLMs perform near randomly, revealing deficiencies in core perception…

Cited by 0SourcecodeScholar
2026

SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model

CVPR 2026

This paper introduces a novel architecture for trajectory-conditioned forecasting of future 3D scene occupancy. In contrast to methods that rely on variational autoencoders (VAEs) to generate discrete occupancy tokens, which inherently limit representational capacity, our approach predicts multi-fra

Cited by 0SourcecodeScholar
2026

V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction

ICLR 2026poster

Large Vision-Language Models (LVLMs) have made significant strides in the field of video understanding in recent times. Nevertheless, existing video benchmarks predominantly rely on text prompts for evaluation, which often require complex referential language and diminish both the accuracy and effic…

Cited by 0SourcecodeScholar
2025

ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation

CVPR 2025poster

Multi-layer image generation is a fundamental task that enables users to isolate, select, and edit specific image layers, thereby revolutionizing interactions with generative models. In this paper, we introduce the Anonymous Region Transformer (ART), which facilitates the direct generation of variab…

Cited by 4SourcePDFScholar
2025

CLEA: Closed-Loop Embodied Agent for Enhancing Task Execution in Dynamic Environments

IROS 2025

Large Language Models (LLMs) exhibit remarkable capabilities in the hierarchical decomposition of complex tasks through semantic reasoning. However, their application in embodied systems faces challenges in ensuring reliable execution of subtask sequences and achieving one-shot success in long-term

Cited by 5SourcecodeScholar
2025

Dual Encoder Contrastive Learning with Augmented Views for Graph Anomaly Detection

IJCAI 2025

Graph anomaly detection (GAD), which aims to identify patterns that deviate significantly from normal nodes in attributed networks, is widely used in financial fraud, cybersecurity, and bioinformatics. The paradigms of jointly optimizing contrastive learning and reconstruction learning have shown si

Cited by 0SourcePDFScholar
2025

EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision

CVPR 2025highlight

Touch contact and pressure are essential for understanding how humans interact with objects and offer insights that benefit applications in mixed reality and robotics. Estimating these interactions from an egocentric camera perspective is challenging, largely due to the lack of comprehensive dataset…

Cited by 3SourcePDFScholar
2025

Enhancing Large Vision-Language Models with Ultra-Detailed Image Caption Generation

EMNLP 2025

High-quality image captions are essential for improving modality alignment and visual understanding in Large Vision-Language Models (LVLMs). However, the scarcity of ultra-detailed image caption data limits further advancements. This paper presents a systematic pipeline for generating high-quality,

2025

GPEN: Global Position Encoding Network for Enhanced Subgraph Representation Learning

ICML 2025poster

Subgraph representation learning has attracted growing interest due to its wide applications in various domains. However, existing methods primarily focus on local neighborhood structures while overlooking the significant impact of global structural information, in particular the influence of multi-…

Cited by 0SourcePDFScholar
2025

SEEA-R1: Tree-Structured Reinforcement Fine-Tuning for Self-Evolving Embodied Agents

NeurIPS 2025poster

Self-evolution, the ability of agents to autonomously improve their reasoning and behavior, is essential for the embodied domain with long-horizon, real-world tasks. Despite current advancements in reinforcement fine-tuning (RFT) showing strong performance in enhancing reasoning in LLMs, its potenti…

Cited by 0SourcecodeScholar
2025

State Feedback Enhanced Graph Differential Equations for Multivariate Time Series Forecasting

IJCAI 2025

Multivariate time series forecasting holds significant theoretical and practical importance in various fields, including web analytics and transportation. Recently, graph neural networks and graph differential equations have shown exceptional capabilities in modeling spatio-temporal features. Howeve

2024

Multi-Level Contrastive Learning For Hybrid Cross-Modal Retrieval

ICASSP 2024accepted

Hybrid image retrieval is a significant task for a wide range of applications. In this scenario, the hybrid query for searching images consists of a reference image and a text modifier. The reference image provides a vital visual context and displays some semantic details, while the text modifier sp…

Cited by 0SourceScholar
2023

Human from Blur: Human Pose Tracking from Blurry Images

ICCV 2023poster

We propose a method to estimate 3D human poses from substantially blurred images. The key idea is to tackle the inverse problem of image deblurring by modeling the forward problem with a 3D human model, a texture map, and a sequence of poses to describe human motion. The blurring process is then mod…

Cited by 3PDFScholar
2022

A Divide-and-Merge Point Cloud Clustering Algorithm for LiDAR Panoptic Segmentation

ICRA 2022poster

Clustering objects from the LiDAR point cloud is an important research problem with many applications such as autonomous driving. To meet the real-time requirement, existing research proposed to apply the connected-component-labeling (CCL) technique on LiDAR spherical range image with a heuristic co…

Cited by 25SourcecodeScholar
2021

FIDNet: LiDAR Point Cloud Semantic Segmentation with Fully Interpolation Decoding

IROS 2021poster

Projecting the point cloud on the 2D spherical range image transforms the LiDAR semantic segmentation to a 2D segmentation task on the range image. However, the LiDAR range image is still naturally different from the regular 2D RGB image; for example, each position on the range image encodes the uni…

Cited by 78SourcecodeScholar