← Search

yuhao liu

26 accepted papers

2026

Adversarially Robust Approximate Furthest Neighbor

ICML 2026poster

We work in the adaptive query model, where one is given a point set $P \subset \mathbb{R}^d$ and seeks to construct a data structure that can answer correctly and efficiently a sequence of adaptive queries. In this model, an adversary observes the answers returned by the data structure to previous q…

Cited by 0SourceScholar
2026

DA$^{2}$: Depth Anything in Any Direction

ICLR 2026poster

Panorama has a full FoV (360$^\circ\times$180$^\circ$), offering a more complete visual description than perspective images. Thanks to this characteristic, panoramic depth estimation is gaining increasing traction in 3D vision. However, due to the scarcity of panoramic data, previous methods are oft…

Cited by 0SourcecodeScholar
2026

Finite-Time Convergence Analysis of ODE-based Generative Models for Stochastic Interpolants

ICLR 2026poster

Stochastic interpolants offer a robust framework for continuously transforming samples between arbitrary data distributions via ordinary or stochastic differential equations (ODEs/SDEs), holding significant promise for generative modeling. While previous studies have analyzed the finite-time converg…

Cited by 0SourceScholar
2026

GenSplat: Bridging the Generalization Gap in 3DGS Language Comprehension

CVPR 2026

In this paper, we propose GenSplat, a novel approach for language comprehension in 3D Gaussian Splatting (3DGS). Unlike previous methods that either achieve cross-scene generalization by being bounded to a predefined vocabulary or handle free-form language by overfitting to individual scenes, GenSpl

Cited by 0SourcecodeScholar
2026

World-Shaper: A Unified Framework for 360° Panoramic Editing

ICML 2026poster

Being able to edit panoramic images is crucial for creating realistic 360° visual experiences. However, existing perspective-based image editing methods fail to model the spatial structure of panoramas. Conventional cube-map decompositions attempt to overcome this problem but inevitably break global…

Cited by 0SourceScholar
2025

Language-Guided Salient Object Ranking

CVPR 2025poster

Salient Object Ranking (SOR) aims to study human attention shifts across different objects in the scene. It is a challenging task, as it requires comprehension of the relations among the salient objects in the scene. However, existing works often overlook such relations or model them implicitly. In…

Cited by 0SourcePDFScholar
2025

STC-Tracker: Spatiotemporal-Consistent Multi-Robot Collaboration Framework for Long-Term Dynamic Object Tracking

IROS 2025

Multi-robot cooperative tracking, as a vital sub-field of multi-robot collaboration, exhibits significant potential in areas such as military reconnaissance and emergency rescue. Conventional dynamic object tracking methods often face issues of incomplete target detection and even loss in complex sc

Cited by 0SourceScholar
2025

Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding

NeurIPS 2025poster

Spatio-temporal video grounding (STVG) aims at localizing the spatio-temporal tube of a video, as specified by the input text query. In this paper, we utilize multimodal large language models (MLLMs) to explore a zero-shot solution in STVG. We reveal two key insights about MLLMs: (1) MLLMs tend to…

Cited by 0SourcecodeScholar
2024

Boosting Weakly Supervised Referring Image Segmentation via Progressive Comprehension

NeurIPS 2024poster

This paper explores the weakly-supervised referring image segmentation (WRIS) problem, and focuses on a challenging setup where target localization is learned directly from image-text pairs. We note that the input text description typically already contains detailed information on how to localize t…

Cited by 2SourcePDFScholar
2024

Decentralizing Coherent Joint Transmission Precoding Via Deterministic Equivalents

ICASSP 2024accepted

In order to control the inter-cell interference for a multi-cell multi-user multiple-input multiple-output network, we consider the precoder design for coordinated multi-point with downlink coherent joint transmission. To avoid costly information exchange among the cooperating base stations in a cen…

Cited by 0SourceScholar
2024

Diff-Plugin: Revitalizing Details for Diffusion-based Low-level Tasks

CVPR 2024poster

Diffusion models trained on large-scale datasets have achieved remarkable progress in image synthesis. However due to the randomness in the diffusion process they often struggle with handling diverse low-level tasks that require details preservation. To overcome this limitation we present a new Diff…

Cited by 23SourcePDFScholar
2024

Multi-View Dynamic Reflection Prior for Video Glass Surface Detection

AAAI 2024technical

Recent research has shown significant interest in image-based glass surface detection (GSD). However, detecting glass surfaces in dynamic scenes remains largely unexplored due to the lack of a high-quality dataset and an effective video glass surface detection (VGSD) method. In this paper, we propos…

2024

Novel Architecture of Deep Feature-Based Gaussian Processes with an Ensemble of Kernels

ICASSP 2024accepted

The inherent adaptability and flexibility of Gaussian processes lie in the capability of their kernel functions to capture diverse data characteristics. Thus, selecting an appropriate kernel function is crucial because an improper choice can detrimentally affect the model’s performance. One way to e…

Cited by 0SourceScholar
2024

Recasting Regional Lighting for Shadow Removal

AAAI 2024technical

Removing shadows requires an understanding of both lighting conditions and object textures in a scene. Existing methods typically learn pixel-level color mappings between shadow and non-shadow images, in which the joint modeling of lighting and object textures is implicit and inadequate. We observe…

2023

Referring Image Segmentation Using Text Supervision

ICCV 2023poster

Existing Referring Image Segmentation (RIS) methods typically require expensive pixel-level or box-level annotations for supervision. In this paper, we observe that the referring texts used in RIS already provide sufficient information to localize the target object. Hence, we propose a novel weakly-…

Cited by 34PDFcodeScholar
2023

VF-Taco2: Towards Fast and Lightweight Synthesis for Autoregressive Models with Variation Autoencoder and Feature Distillation

ICASSP 2023accepted

With the development of deep learning, end-to-end neural text-to-speech (TTS) systems have achieved significant improvements in high-quality speech synthesis. However, most of these systems are attention-based autoregressive models, resulting in slow synthesis speed and large model parameter sizes.…

Cited by 0SourceScholar
2021

Tripartite Information Mining and Integration for Image Matting

ICCV 2021poster

With the development of deep convolutional neural networks, image matting has ushered in a new phase. Regarding the nature of image matting, most researches have focused on solutions for transition regions. However, we argue that many existing approaches are excessively focused on transition-dominan…

Cited by 72PDFcodeScholar
2020

Attention-Guided Hierarchical Structure Aggregation for Image Matting

CVPR 2020poster

Existing deep learning based matting algorithms primarily resort to high-level semantic features to improve the overall structure of alpha mattes. However, we argue that advanced semantics extracted from CNNs contribute unequally for alpha perception and we are supposed to reconcile advanced semanti…

Cited by 213PDFScholar