← Search

Xuhong Wang

6 accepted papers

2026

Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration

CVPR 2026

An ideal embodied agent should possess lifelong learning capabilities to handle long-horizon and complex tasks, enabling continuous operation in general environments. This not only requires the agent to accurately accomplish given tasks but also to leverage long-term episodic memory to optimize deci

Cited by 0SourcecodeScholar
2026

TESTAGENT: AUTOMATIC BENCHMARKING AND EXPLORATORY INTERACTION FOR EVALUATING LLMS IN VERTICAL DOMAINS

ICASSP 2026oral

As Large Language Models (LLMs) are increasingly deployed in highly specialized vertical domains, the evaluation of their domain-specific performance becomes critical. However, existing evaluations for vertical domains typically rely on the labor-intensive construction of static single-turn datasets…

Cited by 0SourcePDFScholar
2026

TPRU: Advancing Temporal and Procedural Understanding in Large Multimodal Models

ICLR 2026poster

Multimodal Large Language Models (MLLMs), particularly smaller, deployable variants, exhibit a critical deficiency in understanding temporal and procedural visual data, a bottleneck hindering their application in real-world embodied AI. This gap is largely caused by a systemic failure in training pa…

Cited by 0SourcecodeScholar
2026

World2Minecraft: Occupancy-Driven simulated scenes Construction

ICLR 2026poster

Embodied intelligence requires high-fidelity simulation environments to support perception and decision-making, yet existing platforms often suffer from data contamination and limited flexibility. To mitigate this, we propose World2Minecraft to convert real-world scenes into structured Minecraft env…

Cited by 0SourceScholar
2025

Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning

EMNLP 2025

Recent advancements in large language models (LLMs) have shifted the post-training paradigm from traditional instruction tuning and human preference alignment toward reinforcement learning (RL) focused on reasoning capabilities. However, most current methods rely on rule-based evaluations of answer

2020

Efficient Spatial-Temporal Normalization of SAE Representation for Event Camera

RA-L 2020

Event-based cameras are a new type of vision sensor that can encode spatial-temporal context in a pixel-level event stream. Its appealing properties offer great potential for applications requiring low processing latency and low power consumption. As an effective representation of events, the surfac

Cited by 13SourceScholar