← Search

Xingtao Wang

11 accepted papers

2026

MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image Compression

AAAI 2026technical

Recent advances in extreme image compression have revealed that mapping pixel data into highly compact latent representations can significantly improve coding efficiency. However, most existing methods compress images into 2-D latent spaces via convolutional neural networks (CNNs) or Swin Transforme

Cited by 0SourcePDFScholar
2026

T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates

AAAI 2026technical

Recent advances in video generation techniques have given rise to an emerging paradigm of generative video coding for Ultra-Low Bitrate (ULB) scenarios by leveraging powerful generative priors. However, most existing methods are limited by domain specificity (e.g., facial or human videos) or excessi

Cited by 0SourcePDFScholar
2025

DexFlyWheel: A Scalable and Self-improving Data Generation Framework for Dexterous Manipulation

NeurIPS 2025spotlight

Dexterous manipulation is critical for advancing robot capabilities in real-world applications, yet diverse and high-quality datasets remain scarce. Existing data collection methods either rely on human teleoperation or require significant human engineering, or generate data with limited diversity,…

Cited by 0SourceScholar
2025

Digging into Intrinsic Contextual Information for High-fidelity 3D Point Cloud Completion

AAAI 2025technical

The common occurrence of occlusion-induced incompleteness in point clouds has made point cloud completion (PCC) a highly-concerned task in the field of geometric processing. Existing PCC methods typically produce complete point clouds from partial point clouds in a coarse-to-fine paradigm, with the…

2025

Hyperbolic-Constraint Point Cloud Reconstruction from Single RGB-D Images

AAAI 2025technical

Reconstructing desired objects and scenes has long been a primary goal in 3D computer vision. Single-view point cloud reconstruction has become a popular technique due to its low cost and accurate results. However, single-view reconstruction methods often rely on expensive CAD models and complex geo…

Cited by 0SourcePDFScholar
2025

Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval

AAAI 2025technical

Due to the challenges in acquiring paired Text-3D data and the inherent irregularity of 3D data structures, combined representation learning of 3D point clouds and text remains unexplored. In this paper, we propose a novel Riemann-based Multi-scale Attention Reasoning Network (RMARN) for text-3D ret…

2025

SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action Recognition

ICCV 2025poster

Open-vocabulary action recognition (OVAR) extends recognition systems to identify unseen action categories. While large-scale vision-language models (VLMs) like CLIP have enabled OVAR in image domains, their adaptation to event data remains underexplored. Event cameras offer high temporal resolution…

2025

TS-Net: Assembling Task-specific Features from Multiple Feature Levels for Multi-task Learning

ICASSP 2025accepted

Multi-task learning (MTL) has become an attractive topic that leverages shared knowledge to improve performance and enhance generalization. However, most existing works neglect the varying contribution of multi-level features to sub-task representations. In this paper, we explore the impact of multi…

Cited by 0SourceScholar
2025

Training-Free Task Planning by Parsing Language Signals With Common Sense

ICASSP 2025accepted

Task planning refers to autonomously organizing actions in response to instruction signals, especially language signals. Previous reinforcement learning and imitation learning methods always require a large amount of task-related data (data interacting with the environment or expert demonstrations)…

Cited by 0SourceScholar
2024

Hyper-MD: Mesh Denoising with Customized Parameters Aware of Noise Intensity and Geometric Characteristics

CVPR 2024poster

Mesh denoising (MD) is a critical task in geometry processing as meshes from scanning or AIGC techniques are susceptible to noise contamination. The challenge of MD lies in the diverse nature of mesh facets in terms of geometric characteristics and noise distributions. Despite recent advancements in…

Cited by 0SourcePDFScholar
2024

Toward a Stable, Fair, and Comprehensive Evaluation of Object Hallucination in Large Vision-Language Models

NeurIPS 2024poster

Given different instructions, large vision-language models (LVLMs) exhibit different degrees of object hallucinations, posing a significant challenge to the evaluation of object hallucinations. Overcoming this challenge, existing object hallucination evaluation methods average the results obtained f…

Cited by 4SourcePDFScholar