← Search

Wenqing Wang

9 accepted papers

2026

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

ICML 2026poster

Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improvement. However, existing embodied benchmarks fail to provide actionable insights because they focus on task-level evaluation rather than discovering capability bottlenecks. To address t…

Cited by 0SourceScholar
2026

SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States

ICML 2026poster

Long-Text Understanding (LTU) at million-token scale requires balancing reasoning fidelity with computational efficiency. Frontier long-context LLMs can process millions of token contexts end-to-end, but they suffer from high token consumption and attention dilution. In parallel, specialized LTU age…

Cited by 0SourceScholar
2025

Crab: A Novel Configurable Role-Playing LLM with Assessing Benchmark

ACL 2025long

This study introduces Crab, a novel Configurable Role-Playing (RP) LLM with Assessing Benchmark, which consists of Role-Centric Dataset Curation, Persona-Embodying LLM Construction, and Comprehensive Benchmark Creation for RP dialogue generation. Distinct from traditional RP models that employ only…

Cited by 0SourcePDFScholar
2025

Improving Gaussian Splatting with Localized Points Management

CVPR 2025highlight

Point management is critical for optimizing 3D Gaussian Splatting models, as point initiation (e.g., via structure from motion) is often distributionally inappropriate. Typically, Adaptive Density Control (ADC) algorithm is adopted, leveraging view-averaged gradient magnitude thresholding for point…

Cited by 0SourcePDFScholar
2025

Towards A “Novel” Benchmark: Evaluating Literary Fiction with Large Language Models

ACL 2025finding

Current exploration on creative generation focuses mainly on short stories, poetry, and scripts. With the expansion of Large Language Models (LLMs) context windows, “novel” avenues emerge. This study aims to extend the boundaries of Natural Language Generation (NLG) evaluation by exploring LLMs’ cap…

2024

Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation

NAACL 2024long

Human evaluation serves as the gold standard for assessing the quality of Natural Language Generation (NLG) systems. Nevertheless, the evaluation guideline, as a pivotal element ensuring reliable and reproducible human assessment, has received limited attention. Our investigation revealed that only…

2023

Boosting Face Recognition Performance with Synthetic Data and Limited Real Data

ICASSP 2023accepted

Face recognition is one of the most precise and straightforward methods to establish individual identity, and is important in our daily life. To solve the issues of privacy, bias, and collection difficulty caused by face recognition relying heavily on collecting a huge number of real face images fro…

Cited by 0SourceScholar