← Search

Wenlong Wang

11 accepted papers

2026

Self-Supervised Adaptive Transformer for Surgical Step Recognition in Robotic-Assisted Radical Prostatectomy

RA-L 2026

The automatic recognition of surgical steps is essential for enhancing situational awareness and workflow automation in robotic-assisted surgery. However, existing vision-based approaches exhibit limitations in effectively leveraging rich spatial-temporal information from surgical videos, particular

Cited by 0SourceScholar
2025

AVS3P10 Standard for Real-time Speech Coding

ICASSP 2025accepted

As the tenth part of the third-generation AVS standard series for real-time speech coding, AVS3P10 is the recent standard completed in the Audio Video Coding Standards Workgroup of China (AVS). Combining the state-of-the-art deep generative networks and signal processing methods, AVS3P10 targets def…

Cited by 0SourceScholar
2025

Drama: Mamba-Enabled Model-Based Reinforcement Learning Is Sample and Parameter Efficient

ICLR 2025poster

Model-based reinforcement learning (RL) offers a solution to the data inefficiency that plagues most model-free RL algorithms. However, learning a robust world model often requires complex and deep architectures, which are computationally expensive and challenging to train. Within the world model, s…

2025

FreeAlign: Superior Text-Image Alignment by Modulating Prompt Attention

ICASSP 2025accepted

In recent years, Text-to-Image (T2I) models have made remarkable advancements, yet accurate accurate association of attributes remains a key challenge. This paper presents FreeAlign, a novel training-free framework designed to enhance attribute alignment in T2I generation. By modulating attention an…

Cited by 0SourceScholar
2025

Learning Statistical and Physical Modeling for Consistency Human Motion Prediction

ICASSP 2025accepted

Diffusion denoising models have great potential in generating diverse and realistic human motions. However, despite the impressive performance of existing methods, they still face some issues. The diffusion process often significantly overlooks physical laws, leading to physically implausible motion…

Cited by 0SourceScholar
2025

Stable Control Visual AutoRegressive Model: Precise and Efficient Image Generation via Scale Alignment

ICASSP 2025accepted

Although diffusion models advance condition-based visual generation, they suffer from speed and cost issues, unlike faster AutoRegressive methods that are limited in performance. To address these, we introduce the Stable Control Visual AutoRegressive Model (SCVAR). SCVAR ensures stable control by al…

Cited by 0SourceScholar
2025

Synthesizing Efficient Trajectory-Controllable Co-Speech Gesture with Latent Consistency Model

ICASSP 2025accepted

Diffusion models excel at audio-driven gesture generation. However, the reverse denoising process is computationally intensive and time-redundant. Additionally, existing methods predominantly focus on the quality and diversity of upper-body gestures, without considering the movement trajectories of…

Cited by 0SourceScholar
2024

Applying Neural Monte Carlo Tree Search to Unsignalized Multi-intersection Scheduling for Autonomous Vehicles

IROS 2024poster

Dynamic scheduling of access to shared resources by autonomous systems is a challenging problem, characterized as being NP-hard. The complexity of this task leads to a combinatorial explosion of possibilities in highly dynamic systems where arriving requests must be continuously scheduled subject to…

Cited by 0SourceScholar
2024

On Unique Localization of Uncorrelated Constant-Modulus Sources Using Sparse Linear Arrays

ICASSP 2024accepted

In direction-of-arrival (DOA) estimation, both stochastic and deterministic priors of source signals have been used to localize more sources than sensors. The stochastic prior refers to uncorrelateness of sources, while the deterministic one includes the constant modulus (CM) property of sources. Ex…

Cited by 0SourceScholar