← Search

Marcelo H. Ang Jr.

13 accepted papers

2026

FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation

ICRA 2026poster

Force sensing is a crucial modality for Vision-Language-Action (VLA) frameworks, as it enables fine-grained perception and dexterous manipulation in contact-rich tasks. We present Force-Distilled VLA (FD-VLA), a novel framework that integrates force awareness into contact-rich manipulation without r…

2026

From Dream to Action: Hierarchical Policy Learning with 3D World Imagination for Robotic Manipulation

ICRA 2026poster

Recent advancements in robotics have focused on developing foundation models capable of generating both actions and future states. Typically, these policies leverage world models to depict human-like imagination. However, most methods remain confined to the 2D domain, where they forecast only the fi…

Cited by 0Scholar
2026

IMPACT: Behavioral Intention-Aware Multimodal Trajectory Prediction With Adaptive Context Trimming

RA-L 2026

This paper presents a unified framework that jointly predicts behavioral intentions and vectorized occupancy, leveraging them as priors to dynamically prune context information during trajectory decoding, thereby enhancing prediction accuracy, interpretability, and efficiency. While most prior work

Cited by 4SourceScholar
2026

IMPACT: Behavioral Intention-Aware Multimodal Trajectory Prediction with Adaptive Context Trimming

ICRA 2026poster

This paper presents a unified framework that jointly predicts behavioral intentions and vectorized occupancy, leveraging them as priors to dynamically prune context information during trajectory decoding, thereby enhancing prediction accuracy, interpretability, and efficiency. While most prior work …

2026

RoboMT: Human-Like Compliance Control for Assembly Via a Bilateral Robotic Teleoperation and Hybrid Mamba-Transformer Framework

ICRA 2026poster

Robotic compliance control is critical for delicate tasks such as electronic connector assembly, where precise force regulation and adaptability are paramount. However, traditional methods often struggle with modeling inaccuracies and sensor noise. Inspired by human adaptability in complex assembly …

Cited by 0Scholar
2026

URPlanner: A Universal Paradigm for Collision-Free Robotic Motion Planning Based on Deep Reinforcement Learning

ICRA 2026poster

Collision-free motion planning for redundant robot manipulators in complex environments is yet to be explored. Although recent advancements at the intersection of deep reinforcement learning (DRL) and robotics have highlighted its potential to handle versatile robotic tasks, current DRL-based collis…

2025

AGI-Elo: How Far Are We From Mastering A Task?

NeurIPS 2025poster

As the field progresses toward Artificial General Intelligence (AGI), there is a pressing need for more comprehensive and insightful evaluation frameworks that go beyond aggregate performance metrics. This paper introduces a unified rating system that jointly models the difficulty of individual test…

Cited by 0SourcecodeScholar
2022

TAda! Temporally-Adaptive Convolutions for Video Understanding

ICLR 2022poster

Spatial convolutions are widely used in numerous deep video models. It fundamentally assumes spatio-temporal invariance, i.e., using shared weights for every location in different frames. This work presents Temporally-Adaptive Convolutions (TAdaConv) for video understanding, which shows that adaptiv…

2021

Multi-Scale Feature Aggregation by Cross-Scale Pixel-to-Region Relation Operation for Semantic Segmentation

RA-L 2021

Exploiting multi-scale features has shown great potential in tackling semantic segmentation problems. The aggregation is commonly done with sum or concatenation (concat) followed by convolutional (conv) layers. However, it fully passes down the high-level context to the following hierarchy without c

Cited by 4SourceScholar
2020

Shape Prior Deformation for Categorical 6D Object Pose and Size Estimation

ECCV 2020poster

We present a novel learning approach to recover the 6D poses and sizes of unseen object instances from an RGB-D image. To handle the intra-class shape variation, we propose a deep network to reconstruct the 3D object model by explicitly modeling the deformation from a pre-learned categorical shape p…

2017

A Two-Stage Optimized Next-View Planning Framework for 3-D Unknown Environment Exploration, and Structural Reconstruction

RA-L 2017

In this paper, we present a solution for autonomous exploration and reconstruction in 3-D unknown environments without a priori knowledge of the environments. In our framework, a two-stage heuristic information gain-based next-view planning algorithm is performed to dynamically select and update can

Cited by 120SourceScholar