← Search

Yonghao Dang

6 accepted papers

2026

Exploring Position Encoding Mechanism in Diffusion U-Net for Training-free High-resolution Image Generation

AAAI 2026technical

Denoising higher-resolution latents using a pre-trained U-Net often results in repetitive and disordered image patterns. In this work, we are motivated to reveal the intrinsic cause of such pattern disruption in high-resolution image generation. Through theoretical analysis and empirical studies, we

Cited by 0SourcePDFScholar
2026

ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers

CVPR 2026

Leveraging pre-trained Diffusion Transformers (DiTs) for high-resolution (HR) image synthesis often leads to spatial layout collapse and degraded texture fidelity. Prior work mitigates these issues with complex pipelines that first perform a base-resolution (i.e., training-resolution) denoising proc

Cited by 0SourceScholar
2025

3DWSNet: A Novel 3D Wavelet Spiking Neural Network for Event-based Action Recognition

IROS 2025

In robotics applications, event cameras provide low-latency and high-dynamic-range sensing by asynchronously detecting brightness changes, making them well-suited for capturing fast motions and subtle cues in dynamic environments. However, most existing Spiking Neural Network (SNN)-based methods enh

Cited by 1SourceScholar
2025

MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation

IROS 2025

Human action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action recognition, the mask-based reconstruction paradigm learns the spatial structure and motion patterns of the skeleton by

Cited by 0SourcecodeScholar
2025

Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning

ICRA 2025

This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the l

Cited by 1SourcecodeScholar
2025

Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation

IROS 2025

The significant advancements in embodied vision navigation have raised concerns about its susceptibility to adversarial attacks exploiting deep neural networks. Investigating the adversarial robustness of embodied vision navigation is crucial, especially given the threat of 3D physical attacks that

Cited by 7SourcecodeScholar