← Search

Teli Ma

14 accepted papers

2026

PA-BiCoop: A Primary-Auxiliary Cooperative Framework for General Bimanual Manipulation

ICRA 2026poster

Bimanual manipulation is essential for advanced robotic systems because it offers higher efficiency and flexibility compared to single-arm configurations. However, existing approaches either lack inter-arm interaction or ignore the need for a dynamic division of labor, treating the arms as functiona…

2025

DynaMind: Reasoning over Abstract Video Dynamics for Embodied Decision-Making

ICML 2025poster

Integrating natural language instructions and visual perception with decision-making is a critical challenge for embodied agents. Existing methods often struggle to balance the conciseness of language commands with the richness of video content. To bridge the gap between modalities, we propose extra…

Cited by 0SourcePDFScholar
2025

Exploring the Limits of Vision-Language-Action Manipulation in Cross-task Generalization

NeurIPS 2025poster

The generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings. However, the cross-task generalization capabilities of existing VLA models remain significantly underexplored. To address this…

Cited by 0SourceScholar
2025

GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation

CoRL 2025poster

Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence—particularly through the lens of *actionable affordances*. However, transferring such knowledge remains challenging due to: 1) a lack of large-scale data…

Cited by 0SourceScholar
2025

Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation

CVPR 2025poster

Learning generalizable visual representations across different embodied environments is essential for effective robotic manipulation in real-world scenarios. However, the limited scale and diversity of robot demonstration data pose a significant challenge. Recent research has explored leveraging lar…

Cited by 8SourcePDFScholar
2025

Omni-Perception: Omnidirectional Collision Avoidance of Legged Robots in Dynamic Environments

CoRL 2025oral

Agile locomotion in complex 3D environments requires robust spatial awareness to safely avoid diverse obstacles such as aerial clutter, uneven terrain, and dynamic agents. Depth-based perception approaches often struggle with sensor noise, lighting variability, computational overhead from intermedia…

Cited by 0SourceScholar
2024

An Examination of the Compositionality of Large Generative Vision-Language Models

NAACL 2024long

With the success of Large Language Models (LLMs), many Generative Vision-Language Models (GVLMs) have been constructed via multimodal instruction tuning. However, the performance of GVLMs in multimodal compositional reasoning remains under-explored. In this paper, we examine both the evaluation metr…

2024

Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation

CoRL 2024poster

Developing robots capable of executing various manipulation tasks, guided by natural language instructions and visual observations of intricate real-world environments, remains a significant challenge in robotics. Such robot agents need to understand linguistic commands and distinguish between the…

Cited by 11SourceScholar
2023

Resilient Binary Neural Network

AAAI 2023technical

Binary neural networks (BNNs) have received ever-increasing popularity for their great capability of reducing storage burden as well as quickening inference time. However, there is a severe performance drop compared with {real-valued} networks, due to its intrinsic frequent weight oscillation during…

2023

Synchronize Feature Extracting and Matching: A Single Branch Framework for 3D Object Tracking

ICCV 2023poster

Siamese network has been a de facto benchmark framework for 3D LiDAR object tracking with a shared-parametric encoder extracting features from template and search region, respectively. This paradigm relies heavily on an additional matching network to model the cross-correlation/similarity of the tem…

Cited by 18PDFScholar
2022

IDa-Det: An Information Discrepancy-Aware Distillation for 1-Bit Detectors

ECCV 2022poster

"Knowledge distillation (KD) has been proven to be useful for training compact object detection models. However, we observe that KD is often effective when the teacher model and student counterpart share similar proposal information. This explains why existing KD methods are less effective for 1-bit…

2022

MCMAE: Masked Convolution Meets Masked Autoencoders

NeurIPS 2022accept

Vision Transformers (ViT) become widely-adopted architectures for various vision tasks. Masked auto-encoding for feature pretraining and multi-scale hybrid convolution-transformer architectures can further unleash the potentials of ViT, leading to state-of-the-art performances on image classificatio…

2022

Recurrent Bilinear Optimization for Binary Neural Networks

ECCV 2022poster

"Binary Neural Networks (BNNs) show great promise for real-world embedded devices. As one of the critical steps to achieve a powerful BNN, the scale factor calculation plays an essential role in reducing the performance gap to their real-valued counterparts. However, existing BNNs neglect the intrin…

2021

Dual-stream Network for Visual Recognition

NeurIPS 2021poster

Transformers with remarkable global representation capacities achieve competitive results for visual tasks, but fail to consider high-level local pattern information in input images. In this paper, we present a generic Dual-stream Network (DS-Net) to fully explore the representation capacity of loc…

Cited by 71SourcePDFScholar