← Search

Mingtao Feng

19 accepted papers

2026

Learning Hierarchical Hyperbolic Mixture Model for Part-aware 3D Generation

CVPR 2026

3D shape generation has become increasingly important for graphics and vision applications. Current part-aware 3D generation usually overlooks hierarchical part relations or inefficiently encodes multi-level semantics in Euclidean space. Thus we propose a novel framework for hierarchical and efficie

Cited by 0SourceScholar
2025

Beyond Human Perception: Understanding Multi-Object World from Monocular View

CVPR 2025poster

Language and binocular vision play a crucial role in human understanding of the world. Advancements in artificial intelligence have also made it possible for machines to develop 3D perception capabilities essential for high-level scene understanding. However, only monocular cameras are often availab…

2025

Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object Detection

CVPR 2025poster

Tiny object detection remains challenging in spite of the success of generic detectors. The dramatic performance degradation of generic detectors on tiny objects is mainly due to the the weak representations of extremely limited pixels. To address this issue, we propose a plug-and-play architecture…

Cited by 0SourcePDFScholar
2025

Gain from Neighbors: Boosting Model Robustness in the Wild via Adversarial Perturbations Toward Neighboring Classes

CVPR 2025poster

Recent approaches, such as data augmentation, adversarial training, and transfer learning, have shown potential in addressing the issue of performance degradation caused by distributional shifts. However, they typically demand careful design in terms of data or models and lack awareness of the impac…

Cited by 0SourcePDFScholar
2025

Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generation

CVPR 2025poster

3D content creation has achieved significant progress in terms of both quality and speed. Although current Gaussian Splatting-based methods can produce 3D objects within seconds, they are still limited by complex preprocessing or low controllability. In this paper, we introduce a novel framework des…

Cited by 0SourcePDFScholar
2025

Mono3DVLT: Monocular-Video-Based 3D Visual Language Tracking

CVPR 2025poster

Visual-Language Tracking (VLT) is emerging as a promising paradigm to bridge the human-machine performance gap. For single objects, VLT broadens the problem scope to text-driven video comprehension. Yet, this direction is still confined to 2D spatial extents, currently lacking the ability to deal wi…

2025

Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud Localization

ICCV 2025poster

Text to point cloud cross-modal localization is a crucial vision-language task for future human-robot collaboration. Existing coarse-to-fine frameworks assume that each query text precisely corresponds to the center area of a submap, limiting their applicability in real-world scenarios. This work re…

2025

Semantic Ambiguity Modeling and Propagation for Fine-Grained Visual Cross View Geo-Localization

AAAI 2025technical

Visual cross view geo-localization is generally approached within a joint retrieval-and-calibration framework. However, existing methods overlook semantic ambiguities arising from query and reference images characterized by low overlap, dynamic foregrounds, viewpoint changes, and perceptual aliasing…

2024

Decentralized Multi-Robot Navigation Coupled with Spatial-Temporal RetNet Based on Deep Reinforcement Learning

IROS 2024poster

Navigating robots through dynamic multi-robot environments, avoiding collisions with both other robots and obstacles, has emerged as a central challenge in robotics. The existing approaches fall short in allowing the policy network to effectively capture spatial-temporal reciprocal collision avoidan…

Cited by 0SourceScholar
2024

Domain Adaptation in Visual Reinforcement Learning via Self-Expert Imitation with Purifying Latent Feature

IROS 2024poster

Generalizing visual reinforcement learning is fundamental to robot visual navigation, involving the acquisition of a policy from interactions with source environments to facilitate adaptation to analogous, yet unfamiliar target environments. Recent advancements capitalize on data augmentation techni…

Cited by 0SourceScholar
2024

L4D-Track: Language-to-4D Modeling Towards 6-DoF Tracking and Shape Reconstruction in 3D Point Cloud Stream

CVPR 2024poster

3D visual language multi-modal modeling plays an important role in actual human-computer interaction. However the inaccessibility of large-scale 3D-language pairs restricts their applicability in real-world scenarios. In this paper we aim to handle a real-time multi-task for 6-DoF pose tracking of u…

Cited by 0SourcePDFScholar
2024

Referring Human Pose and Mask Estimation In the Wild

NeurIPS 2024poster

We introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to pr…

2023

3D Spatial Multimodal Knowledge Accumulation for Scene Graph Prediction in Point Cloud

CVPR 2023poster

In-depth understanding of a 3D scene not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since 3D scenes contain partially scanned objects with physical connections, dense placement, changing sizes, and a wide…

2023

DANet: Density Adaptive Convolutional Network With Interactive Attention for 3D Point Clouds

RA-L 2023

Local features and contextual dependencies are crucial for 3D point cloud analysis. Many works have been devoted to designing better local convolutional kernels that exploit the contextual dependencies. However, current point convolutions lack robustness to varying point cloud density. Moreover, con

Cited by 7SourceScholar
2023

Sketch and Text Guided Diffusion Model for Colored Point Cloud Generation

ICCV 2023poster

Diffusion probabilistic models have achieved remarkable success in text guided image generation. However, generating 3D shapes is still challenging due to the lack of sufficient data containing 3D models along with their descriptions. Moreover, text based descriptions of 3D shapes are inherently amb…

Cited by 32PDFScholar
2022

ICK-Track: A Category-Level 6-DoF Pose Tracker Using Inter-Frame Consistent Keypoints for Aerial Manipulation

IROS 2022poster

Robots that are supposed to interact with or manipulate objects in the world must be able to track the poses of objects in their sensor data. Thus, Detecting and tracking the 6-DoF poses of targeted objects is important for aerial manipulation and is still in the early stage due to the high dynamics…

Cited by 8SourcecodeScholar
2022

Learning From Pixel-Level Noisy Label: A New Perspective for Light Field Saliency Detection

CVPR 2022poster

Saliency detection with light field images is becoming attractive given the abundant cues available, however, this comes at the expense of large-scale pixel level annotated data which is expensive to generate. In this paper, we propose to learn light field saliency from pixel-level noisy labels obta…

Cited by 25PDFcodeScholar
2021

Free-Form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloud

ICCV 2021poster

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of…

Cited by 100PDFcodeScholar
2018

3D Face Reconstruction from Light Field Images: A Model-free Approach

ECCV 2018poster

Reconstructing 3D facial geometry from a single RGB image has recently instigated wide research interest. However, it is still an ill-posed problem and most methods rely on prior models hence undermining the accuracy of the recovered 3D faces. In this paper, we exploit the Epipolar Plane Images (EPI…

Cited by 50SourcePDFScholar