← Search

Miao Wang

14 accepted papers

2026

Image Guides Images: Consistent Video Amodal Completion with Rectified In-Context Exemplar Guidance

CVPR 2026

Video amodal completion (VAC) aims to mimic the human brain's ability to implicitly perceive the complete appearance of partially occluded objects, thereby facilitating recognition and understanding. Existing VAC methods finetune video generation models on custom datasets, yet these datasets often h

Cited by 0SourcecodeScholar
2026

Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion

ICML 2026poster

GUI grounding maps natural language instructions to the correct interface elements, serving as the perception foundation for GUI agents. Existing approaches predominantly rely on fine-tuning multimodal large language models (MLLMs) using large-scale GUI datasets to predict target element coordinates…

Cited by 0SourceScholar
2025

COLA: Collaborative Multi-Agent Framework with Dynamic Task Scheduling for GUI Automation

EMNLP 2025

With the rapid advancements in Large Language Models (LLMs), an increasing number of studies have leveraged LLMs as the cognitive core of agents to address complex task decision-making challenges. Specially, recent research has demonstrated the potential of LLM-based agents on automating GUI operati

2025

GPAvatar: High-fidelity Head Avatars by Learning Efficient Gaussian Projections

CVPR 2025poster

Existing radiance field-based head avatar methods have mostly relied on pre-computed explicit priors (e.g., mesh, point) or neural implicit representations, making it challenging to achieve high fidelity with both computational efficiency and low memory consumption. To overcome this, we present GPAv…

Cited by 0SourcePDFScholar
2025

PAT: Pruning-Aware Tuning for Large Language Models

AAAI 2025technical

Large language models (LLMs) excel in language tasks, especially with supervised fine-tuning after pre-training. However, their substantial memory and computational requirements hinder practical applications. Structural pruning, which reduces less significant weight dimensions, is one solution. Yet,…

2024

A Non-parametric Graph Clustering Framework for Multi-View Data

AAAI 2024technical

Multi-view graph clustering (MVGC) derives encouraging grouping results by seamlessly integrating abundant information inside heterogeneous data, and has captured surging focus recently. Nevertheless, the majority of current MVGC works involve at least one hyper-parameter, which not only requires…

Cited by 19SourcePDFScholar
2024

DVSAI: Diverse View-Shared Anchors Based Incomplete Multi-View Clustering

AAAI 2024technical

In numerous real-world applications, it is quite common that sample information is partially available for some views due to machine breakdown or sensor failure, causing the problem of incomplete multi-view clustering (IMVC). While several IMVC approaches using view-shared anchors have successfully…

Cited by 17SourcePDFScholar
2024

Exploring Regional Clues in CLIP for Zero-Shot Semantic Segmentation

CVPR 2024poster

CLIP has demonstrated marked progress in visual recognition due to its powerful pre-training on large-scale image-text pairs. However it still remains a critical challenge: how to transfer image-level knowledge into pixel-level understanding tasks such as semantic segmentation. In this paper to solv…

2024

Language Embedded 3D Gaussians for Open-Vocabulary Scene Understanding

CVPR 2024poster

Open-vocabulary querying in 3D space is challenging but essential for scene understanding tasks such as object localization and segmentation. Language-embedded scene representations have made progress by incorporating language features into 3D spaces. However their efficacy heavily depends on neural…

2024

Neural 3D Strokes: Creating Stylized 3D Scenes with Vectorized 3D Strokes

CVPR 2024poster

We present Neural 3D Strokes a novel technique to generate stylized images of a 3D scene at arbitrary novel views from multi-view 2D images. Different from existing methods which apply stylization to trained neural radiance fields at the voxel level our approach draws inspiration from image-to-paint…

2022

C3-STISR: Scene Text Image Super-resolution with Triple Clues

IJCAI 2022poster

Scene text image super-resolution (STISR) has been regarded as an important pre-processing task for text recognition from low-resolution scene text images. Most recent approaches use the recognizer's feedback as clues to guide super-resolution. However, directly using recognition clue has two proble…

2022

Rendering-Aware HDR Environment Map Prediction from a Single Image

AAAI 2022technical

High dynamic range (HDR) illumination estimation from a single low dynamic range (LDR) image is a significant task in computer vision, graphics, and augmented reality. We present a two-stage deep learning-based method to predict an HDR environment map from a single narrow field-of-view LDR image. We…

Cited by 13SourcePDFScholar
2019

Example-Guided Style-Consistent Image Synthesis From Semantic Labeling

CVPR 2019poster

Example-guided image synthesis aims to synthesize an image from a semantic label map and an exemplary image indicating style. We use the term "style" in this problem to refer to implicit characteristics of images, for example: in portraits "style" includes gender, racial identity, age, hairstyle;…

Cited by 102PDFcodeScholar