← Search

Jianqin Yin

15 accepted papers

2026

Exploring Position Encoding Mechanism in Diffusion U-Net for Training-free High-resolution Image Generation

AAAI 2026technical

Denoising higher-resolution latents using a pre-trained U-Net often results in repetitive and disordered image patterns. In this work, we are motivated to reveal the intrinsic cause of such pattern disruption in high-resolution image generation. Through theoretical analysis and empirical studies, we

Cited by 0SourcePDFScholar
2026

ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers

CVPR 2026

Leveraging pre-trained Diffusion Transformers (DiTs) for high-resolution (HR) image synthesis often leads to spatial layout collapse and degraded texture fidelity. Prior work mitigates these issues with complex pipelines that first perform a base-resolution (i.e., training-resolution) denoising proc

Cited by 0SourceScholar
2025

3DWSNet: A Novel 3D Wavelet Spiking Neural Network for Event-based Action Recognition

IROS 2025

In robotics applications, event cameras provide low-latency and high-dynamic-range sensing by asynchronously detecting brightness changes, making them well-suited for capturing fast motions and subtle cues in dynamic environments. However, most existing Spiking Neural Network (SNN)-based methods enh

Cited by 1SourceScholar
2025

AlignCAPE: Support and Query Feature Aligning for Category-Agnostic Pose Estimation

IROS 2025

Recent advancements in category-agnostic pose estimation have focused on developing a unified model capable of localizing keypoint coordinates across arbitrary categories, which enables robots to accurately interact with diverse objects by understanding their poses. While existing methods predominan

Cited by 0SourceScholar
2025

Learning a Unified Policy for Position and Force Control in Legged Loco-Manipulation

CoRL 2025oral

Robotic loco-manipulation tasks often involve contact-rich interactions with the environment, requiring the joint modeling of contact force and robot position. However, recent visuomotor policies often focus solely on position or force control, overlooking their integration. In this work, we propose…

Cited by 0SourceScholar
2025

MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection

ICASSP 2025accepted

Sound Event Detection (SED) plays a vital role in comprehending and perceiving acoustic scenes. Previous methods have demonstrated impressive capabilities. However, they are deficient in learning features of complex scenes from heterogeneous dataset. In this paper, we introduce a novel dual-branch a…

Cited by 0SourceScholar
2025

MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation

IROS 2025

Human action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action recognition, the mask-based reconstruction paradigm learns the spatial structure and motion patterns of the skeleton by

Cited by 0SourcecodeScholar
2025

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

ICCV 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated significant success in visual understanding tasks. However, challenges persist in adapting these models for video comprehension due to the large volume of data and temporal complexity. Existing Video-LLMs using uniform frame sampling often s…

Cited by 0SourcePDFScholar
2025

Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation

IROS 2025

The significant advancements in embodied vision navigation have raised concerns about its susceptibility to adversarial attacks exploiting deep neural networks. Investigating the adversarial robustness of embodied vision navigation is crucial, especially given the threat of 3D physical attacks that

Cited by 7SourcecodeScholar
2024

Lifting by Image – Leveraging Image Cues for Accurate 3D Human Pose Estimation

AAAI 2024technical

The "lifting from 2D pose" method has been the dominant approach to 3D Human Pose Estimation (3DHPE) due to the powerful visual analysis ability of 2D pose estimators. Widely known, there exists a depth ambiguity problem when estimating solely from 2D pose, where one 2D pose can be mapped to multipl…

Cited by 13SourcePDFScholar
2023

Target-Aware Spatio-Temporal Reasoning via Answering Questions in Dynamic Audio-Visual Scenarios

EMNLP 2023long findings

Audio-visual question answering (AVQA) is a challenging task that requires multistep spatio-temporal reasoning over multimodal contexts. Recent works rely on elaborate target-agnostic parsing of audio-visual scenes for spatial grounding while mistreating audio and video as separate entities for temp…

Cited by 0SourceScholar
2022

Deep Tri-Training for Semi-Supervised Image Segmentation

RA-L 2022

Semantic segmentation is of great value to autonomous driving and many robotic applications, while it highly depends on costly and time-consuming pixel-level annotation. To make full use of unlabeled data, this work proposes a deep tri-training framework (dubbed DTT) to utilize labeled along with un

Cited by 11SourceScholar
2021

Dynamic tracking for microrobot with active magnetic sensor array

ICRA 2021poster

Accurate position feedback in a wide range is critical for medical microrobotics and robot-assisted examinations, such as colonoscopy, bronchoscopy and capsule endoscopy examination. Among the many modalities of positioning feedback, magnetic tracking is a preferable method due to the unique advanta…

Cited by 11SourceScholar
2021

Neighborhood Spatial Aggregation based Efficient Uncertainty Estimation for Point Cloud Semantic Segmentation

ICRA 2021poster

Uncertainty estimation for point cloud semantic segmentation is to quantify the confidence degree for the predicted label of points, which is essential for decision-making tasks. This paper proposes a neighborhood spatial aggregation based method, NSA-MC dropout, to achieve efficient uncertainty est…

Cited by 4SourcecodeScholar