← Search

Jiahui Yang

8 accepted papers

2025

BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent

NeurIPS 2025poster

In the field of AI-driven human-GUI interaction automation, while rapid advances in multimodal large language models and reinforcement fine-tuning techniques have yielded remarkable progress, a fundamental challenge persists: their interaction logic significantly deviates from natural human-GUI comm…

Cited by 0SourceScholar
2025

Deep Reactive Policy: Learning Reactive Manipulator Motion Planning for Dynamic Environments

CoRL 2025poster

Generating collision-free motion in dynamic, partially observable environments is a fundamental challenge for robotic manipulators. Classical motion planners can compute globally optimal trajectories but require full environment knowledge and are typically too slow for dynamic scenes. Neural motion…

Cited by 0SourceScholar
2025

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation

CVPR 2025poster

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current methods face fundamental challenges: a) the absence of frame…

Cited by 1SourcePDFScholar
2025

Neural MP: A Neural Motion Planner

IROS 2025

The current paradigm for motion planning generates solutions from scratch for every new problem, which consumes significant amounts of time and computational resources. For complex, cluttered scenes, motion planning approaches can often take minutes to produce a solution, while humans are able to ac

Cited by 0SourcecodeScholar
2025

Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

ICCV 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated significant success in visual understanding tasks. However, challenges persist in adapting these models for video comprehension due to the large volume of data and temporal complexity. Existing Video-LLMs using uniform frame sampling often s…

Cited by 0SourcePDFScholar
2025

QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized Generation

ICCV 2025poster

Existing text-to-image models often rely on parame- ter fine-tuning techniques such as Low-Rank Adaptation (LoRA) to customize visual attributes. However, when com- bining multiple LoRA models for content-style fusion tasks, unstructured modifications of weight matrices often lead to undesired featu…

Cited by 0SourcePDFScholar
2024

Bimanual Dexterity for Complex Tasks

CoRL 2024poster

To train generalist robot policies, machine learning methods often require a substantial amount of expert human teleoperation data. An ideal robot for humans collecting data is one that closely mimics them: bimanual arms and dexterous hands. However, creating such a bimanual teleoperation system wit…

Cited by 21SourcecodeScholar
2023

A lightweight high-voltage boost circuit for soft-actuated micro-aerial-robots

ICRA 2023poster

Flight is an energetically expensive task. While aerial insects can effortlessly fly through natural environments, achieving power autonomous flights in insect-scale robots remains a major challenge. In prior works, we developed soft-actuated insect-scale aerial robots that demonstrated unique capab…

Cited by 5SourceScholar