← Search

Wenhao Wang

23 accepted papers

2026

CustomTex: High-fidelity Indoor Scene Texturing via Multi-Reference Customization

CVPR 2026

The creation of high-fidelity, customizable 3D indoor scene textures remains a significant challenge. While text-driven methods offer flexibility, they lack the precision for fine-grained, instance-level control, and often produce textures with insufficient quality, artifacts, and baked-in shading.

Cited by 0SourceScholar
2026

InfoMosaic-Bench: Evaluating Multi-Source Information Seeking in Tool-Augmented Agents

ICLR 2026poster

Information seeking is a fundamental requirement for humans. However, existing LLM agents rely heavily on open-web search, which exposes two fundamental weaknesses: online content is noisy and unreliable, and many real-world tasks require precise, domain-specific knowledge unavailable from the web.…

Cited by 0SourceScholar
2026

LLM-Driven Corrective Robot Operation Code Generation with Static Text-Based Simulation

ICRA 2026poster

Recent advances in Large language models (LLMs) have demonstrated their promising capabilities of generating robot operation code to enable LLM-driven robots. To enhance the reliability of operation code generated by LLMs, corrective designs with feedback from the observation of executing code have …

2026

MCP-Persona: Benchmarking LLM Agents on Personalized MCP Tools and Tasks

ICML 2026poster

a transformative standard for connecting large language models (LLMs) with external data sources and tools, and has been rapidly adopted across personal applications and development platforms. However, existing benchmarks predominantly focus on generic information-seeking tools and fail to capture t…

Cited by 0SourceScholar
2025

Can Retelling Have Adequate Information for Reasoning? An Enhancement Method for Imperfect Video Understanding with Large Language Model

IJCAI 2025

Large Language Models (LLMs) demonstrate strong capabilities in video understanding. However, it exhibits hallucinations and factual errors in video description. On the one hand, existing Multimodal Large Language Models (MLLMs) are primarily trained by combining language models and vision models, w

Cited by 0SourcePDFScholar
2025

Captured by Captions: On Memorization and its Mitigation in CLIP Models

ICLR 2025poster

Multi-modal models, such as CLIP, have demonstrated strong performance in aligning visual and textual representations, excelling in tasks like image retrieval and zero-shot classification. Despite this success, the mechanisms by which these models utilize training data, particularly the role of memo…

Cited by 0SourcePDFScholar
2025

Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Prediction

EMNLP 2025

Natural Language Inference (NLI) is a fundamental task in natural language processing. While NLI has developed many subdirections such as sentence-level NLI, document-level NLI and cross-lingual NLI, Cross-Document Cross-Lingual NLI (CDCL-NLI) remains largely unexplored. In this paper, we propose a

Cited by 0SourcePDFScholar
2025

FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User Data

EMNLP 2025

Mobile GUI agents have attracted tremendous research participation recently. Traditional approaches to mobile agent training rely on centralized data collection, leading to high cost and limited scalability. Distributed training utilizing federated learning offers an alternative by harnessing real-w

2025

Generalizable Humanoid Manipulation with 3D Diffusion Policies

IROS 2025

Humanoid robots capable of autonomous operation in diverse environments have long been a goal for roboticists. However, autonomous manipulation by humanoid robots has largely been restricted to one specific scene, primarily due to the difficulty of acquiring generalizable skills and the expensivenes

Cited by 40SourcecodeScholar
2025

Origin Identification for Text-Guided Image-to-Image Diffusion Models

ICML 2025poster

Text-guided image-to-image diffusion models excel in translating images based on textual prompts, allowing for precise and creative visual modifications. However, such a powerful technique can be misused for *spreading misinformation*, *infringing on copyrights*, and *evading content tracing*. This…

2025

TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation

ICCV 2025poster

Video generation models are revolutionizing content creation, with image-to-video models drawing increasing attention due to their enhanced controllability, visual consistency, and practical applications. However, despite their popularity, these models rely on user-provided text and image prompts, a…

2024

KnowledgeSG: Privacy-Preserving Synthetic Text Generation with Knowledge Distillation from Server

EMNLP 2024main

The success of large language models (LLMs) facilitate many parties to fine-tune LLMs on their own private data. However, this practice raises privacy concerns due to the memorization of LLMs. Existing solutions, such as utilizing synthetic data for substitution, struggle to simultaneously improve p…

2024

LOFT: Latent Space Optimization and Generator Fine-Tuning for Defending Against Deepfakes

ICASSP 2024accepted

DeepFakes pose a significant threat to individual reputations and society as a whole. Existing proactive defense strategies concentrate on adding adversarial perturbations to images to disrupt or nullify the generation of DeepFakes, but these approaches are easily detectable by human perception and…

Cited by 0SourceScholar
2024

MS-DETR: Efficient DETR Training with Mixed Supervision

CVPR 2024poster

DETR accomplishes end-to-end object detection through iteratively generating multiple object candidates based on image features and promoting one candidate for each ground-truth object. The traditional training procedure using one-to-one supervision in the original DETR lacks direct supervision for…

2024

Memorization in Self-Supervised Learning Improves Downstream Generalization

ICLR 2024poster

Self-supervised learning (SSL) has recently received significant attention due to its ability to train high-performance encoders purely on unlabeled data---often scraped from the internet. This data can still be sensitive and empirical evidence suggests that SSL encoders memorize private information…

2024

VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models

NeurIPS 2024poster

The arrival of Sora marks a new era for text-to-video diffusion models, bringing significant advancements in video generation and potential applications. However, Sora, along with other text-to-video diffusion models, is highly reliant on prompts, and there is no publicly available dataset that feat…

2023

A Benchmark and Asymmetrical-Similarity Learning for Practical Image Copy Detection

AAAI 2023technical

Image copy detection (ICD) aims to determine whether a query image is an edited copy of any image from a reference set. Currently, there are very limited public benchmarks for ICD, while all overlook a critical challenge in real-world applications, i.e., the distraction from hard negative queries. S…

2021

Learning Anchored Unsigned Distance Functions With Gradient Direction Alignment for Single-View Garment Reconstruction

ICCV 2021poster

While single-view 3D reconstruction has made significant progress benefiting from deep shape representations in recent years, garment reconstruction is still not solved well due to open surfaces, diverse topologies and complex geometric details. In this paper, we propose a novel learnable Anchored U…

Cited by 55PDFcodeScholar