← Search

Zemin Yang

7 accepted papers

2026

Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Models

AAAI 2026technical

Affordance grounding focuses on predicting the specific regions of objects that are associated with the actions to be performed by robots. It plays a vital role in the fields of human-robot interaction, human-object interaction, embodied manipulation, and embodied perception. Existing models often n

Cited by 20SourcePDFScholar
2026

From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges

ICML 2026poster

Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemming from the fundamental spatiotemporal scale mismatch between cognition and action. Existing generative policies typically adopt a "Generation-from-Noise" paradig…

Cited by 0SourceScholar
2026

HUMOF: Human Motion Forecasting in Interactive Social Scenes

ICLR 2026poster

Complex dynamic scenes present significant challenges for predicting human behavior due to the abundance of interaction information, such as human-human and human-environment interactions. These factors complicate the analysis and understanding of human behavior, thereby increasing the uncertainty i…

Cited by 0SourceScholar
2025

EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild

CVPR 2025poster

Our work aims to reconstruct hand-object interactions from a single-view image, which is a fundamental but ill-posed task.Unlike methods that reconstruct from videos, multi-view images, or predefined 3D templates, single-view reconstruction faces significant challenges due to inherent ambiguities an…

2025

EvolvingGrasp: Evolutionary Grasp Generation via Efficient Preference Alignment

ICCV 2025poster

Dexterous robotic hands often struggle to generalize effectively in complex environments due to models trained on low-diversity data. However, the real world presents an inherently unbounded range of scenarios. A natural solution is to enable robots learning from experience in complex environments--…

Cited by 0SourcePDFScholar
2025

FreqPolicy: Frequency Autoregressive Visuomotor Policy with Continuous Tokens

NeurIPS 2025poster

Learning effective visuomotor policies for robotic manipulation is challenging, as it requires generating precise actions while maintaining computational efficiency. Existing methods remain unsatisfactory due to inherent limitations in the essential action representation and the basic network archit…

Cited by 0SourcecodeScholar
2025

UniDemoiré: Towards Universal Image Demoiréing with Data Generation and Synthesis

AAAI 2025technical

Image demoiréing poses one of the most formidable challenges in image restoration, primarily due to the unpredictable and anisotropic nature of moiré patterns. Limited by the quantity and diversity of training data, current methods tend to overfit to a single moiré domain, resulting in performance d…