← Search

Runhao Zeng

14 accepted papers

2026

Master Skill Learning with Policy-Grounded Synergy of LLM-based Reward Shaping and Exploring

ICLR 2026poster

The acquisition of robotic skills via reinforcement learning (RL) is crucial for advancing embodied intelligence, but designing effective reward functions for complex tasks remains challenging. Recent methods using large language models (LLMs) can generate reward functions from language instructions…

Cited by 0SourceScholar
2026

Revisiting Cross-Architecture Distillation: Adaptive Dual-Teacher Transfer for Lightweight Video Models

AAAI 2026technical

Vision Transformers (ViTs) have achieved strong performance in video action recognition, but their high computational cost limits their practicality. Lightweight CNNs are more efficient but suffer from accuracy gaps. Cross-Architecture Knowledge Distillation (CAKD) addresses this by transferring kno

Cited by 0SourcePDFScholar
2026

UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model

AAAI 2026technical

Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions—remains highly challenging. Recent research on enhancing language-guided navigation reasoning using pre-trained large language models (LLMs) has show

Cited by 0SourcePDFScholar
2026

Whole-Body Coordination for Dynamic Object Grasping with Legged Manipulators

AAAI 2026technical

Quadrupedal robots with manipulators offer strong mobility and adaptability for grasping in unstructured, dynamic environments through coordinated whole-body control. However, existing research has predominantly focused on static-object grasping, neglecting the challenges posed by dynamic targets an

Cited by 0SourcePDFScholar
2025

Efficient Language-instructed Skill Acquisition via Reward-Policy Co-Evolution

AAAI 2025technical

The ability to autonomously explore and resolve tasks with minimal human guidance is crucial for the self-development of embodied intelligence. Although reinforcement learning methods can largely ease human effort, it's challenging to design reward functions for real-world tasks, especially for hig…

2025

OVG-HQ: Online Video Grounding with Hybrid-modal Queries

ICCV 2025poster

Video grounding (VG) task focuses on locating specific moments in a video based on a query, usually in text form. However, traditional VG struggles with some scenarios like streaming video or queries using visual cues. To fill this gap, we present a new task named Online Video Grounding with Hybrid-…

Cited by 0SourcePDFScholar
2025

Temporal Action Detection Model Compression by Progressive Block Drop

CVPR 2025poster

Temporal action detection (TAD) aims to identify and localize action instances in untrimmed videos, which is essential for various video understanding tasks. However, recent improvements in model performance, driven by larger feature extractors and datasets, have led to increased computational deman…

Cited by 0SourcePDFScholar
2025

Understanding Emotional Body Expressions via Large Language Models

AAAI 2025technical

Emotion recognition based on body movements is vital in human-computer interaction. However, existing emotion recognition methods predominantly focus on enhancing classification accuracy, often neglecting the provision of textual explanations to justify their classifications. In this paper, we propo…

2024

Benchmarking the Robustness of Temporal Action Detection Models Against Temporal Corruptions

CVPR 2024poster

Temporal action detection (TAD) aims to locate action positions and recognize action categories in long-term untrimmed videos. Although many methods have achieved promising results their robustness has not been thoroughly studied. In practice we observe that temporal information in videos can be occ…

2024

MemSAM: Taming Segment Anything Model for Echocardiography Video Segmentation

CVPR 2024poster

We propose a novel echocardiographical video segmentation model by adapting SAM to medical videos to address some long-standing challenges in ultrasound video segmentation including (1) massive speckle noise and artifacts (2) extremely ambiguous boundaries and (3) large variations of targeting objec…

2022

Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation

NeurIPS 2022accept

We address a practical yet challenging problem of training robot agents to navigate in an environment following a path described by some language instructions. The instructions often contain descriptions of objects in the environment. To achieve accurate and efficient navigation, it is critical to b…

2021

RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning

AAAI 2021technical

We study unsupervised video representation learning that seeks to learn both motion and appearance features from unlabeled video only, which can be reused for downstream tasks such as action recognition. This task, however, is extremely challenging due to 1) the highly complex spatial-temporal infor…

2019

Graph Convolutional Networks for Temporal Action Localization

ICCV 2019poster

Most state-of-the-art action localization systems process each action proposal individually, without explicitly exploiting their relations during learning. However, the relations between proposals actually play an important role in action localization, since a meaningful action always consists of mu…

Cited by 640PDFcodeScholar