← Search

Min-Chun Hu

5 accepted papers

2026

VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning

IJCAI 2026

Designing effective reward functions remains a major challenge in reinforcement learning (RL), particularly in open-ended environments where task goals are abstract and difficult to quantify. In this work, we present VLM-AR3L, a framework that leverages Vision-Language Models (VLMs) to provide both

Cited by 0Scholar
2025

CoachMe: Decoding Sport Elements with a Reference-Based Coaching Instruction Generation Model

ACL 2025long

Motion instruction is a crucial task that helps athletes refine their technique by analyzing movements and providing corrective guidance. Although recent advances in multimodal models have improved motion understanding,generating precise and sport-specific instruction remains challenging due to the…

2025

Cross-Layer Cache Aggregation for Token Reduction in Ultra-Fine-Grained Image Recognition

ICASSP 2025accepted

Ultra-fine-grained image recognition (UFGIR) is a challenging task that involves classifying images within a macro-category. While traditional FGIR deals with classifying different species, UFGIR goes beyond by classifying sub-categories within a species such as cultivars of a plant. In recent times…

Cited by 0SourceScholar
2023

Scalable Spatial Memory for Scene Rendering and Navigation

AAAI 2023technical

Neural scene representation and rendering methods have shown promise in learning the implicit form of scene structure without supervision. However, the implicit representation learned in most existing methods is non-expandable and cannot be inferred online for novel scenes, which makes the learned r…

Cited by 1SourcePDFScholar
2021

STR-GQN: Scene Representation and Rendering for Unknown Cameras Based on Spatial Transformation Routing

ICCV 2021poster

Geometry-aware modules are widely applied in recent deep learning architectures for scene representation and rendering. However, these modules require intrinsic camera information that might not be obtained accurately. In this paper, we propose a Spatial Transformation Routing (STR) mechanism to mod…

Cited by 5PDFScholar