← Search

Minghao Chen

35 accepted papers

2026

AutoRPA: Efficient GUI Automation through LLM-Driven Code Synthesis from Interactions

ICML 2026poster

Large Language Model (LLM) based agents have demonstrated proficiency in multi-step interactions with graphical user interfaces (GUIs). While most research focuses on improving single-task performance, practical scenarios often involve repetitive GUI tasks for which invoking LLM reasoning repeatedly…

Cited by 0SourceScholar
2026

LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis

CVPR 2026

Recent work has shown that neural networks can perform 3D tasks such as Novel View Synthesis (NVS) without explicit 3D reconstruction. Even so, we argue that strong 3D inductive biases are still helpful in the design of such networks. We show this point by introducing LagerNVS, an encoder-decoder ne

Cited by 0SourcecodeScholar
2026

REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting

CVPR 2026

Bridging the gap between complex human instructions and precise 3D object grounding remains a significant challenge in vision and robotics. Existing 3D segmentation methods often struggle to interpret ambiguous, reasoning-based instructions, while 2D vision-language models that excel at such reasoni

Cited by 0SourcecodeScholar
2026

SRSplat: Feed-Forward Super-Resolution Gaussian Splatting from Sparse Multi-View Images

AAAI 2026technical

Feed-forward 3D reconstruction from sparse, low-resolution (LR) images is a crucial capability for real-world applications, such as autonomous driving and embodied AI. However, existing methods often fail to recover fine texture details. This limitation stems from the inherent lack of high-frequency

Cited by 0SourcePDFScholar
2026

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

CVPR 2026

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D understanding or video generation, a significant gap remains in

Cited by 0SourceScholar
2026

Sparse4DGS: 4D Gaussian Splatting for Sparse-Frame Dynamic Scene Reconstruction

AAAI 2026technical

Dynamic Gaussian Splatting approaches have achieved remarkable performance for 4D scene reconstruction. However, these approaches rely on dense-frame video sequences for photorealistic reconstruction. In real-world scenarios, due to equipment constraints, sometimes only sparse frames are accessible.

Cited by 0SourcePDFScholar
2026

VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding

CVPR 2026

Long-form video understanding remains challenging due to the extended temporal structure and dense multimodal cues. Despite recent progress, many existing approaches still rely on hand-crafted reasoning pipelines or employ token-consuming video preprocessing to guide MLLMs in autonomous reasoning. T

Cited by 0SourcecodeScholar
2026

WorldGen: From Text to Traversable and Interactive 3D Worlds

CVPR 2026

We introduce WorldGen, a method for generating large, fully formed, navigable 3D worlds from a single text prompt. Existing approaches to 3D scene generation often trade off scene diversity, completeness, and correctness in different ways. We push this envelope by producing large scenes explicitly d

Cited by 0SourceScholar
2025

AutoPartGen: Autoregressive 3D Part Generation and Discovery

NeurIPS 2025poster

We introduce AutoPartGen, a model that generates objects composed of 3D parts in an autoregressive manner. This model can take as input an image of an object, 2D masks of the object's parts, or an existing 3D object, and generate a corresponding compositional 3D reconstruction. Our approach builds…

Cited by 0SourceScholar
2025

PartGen: Part-level 3D Generation and Reconstruction with Multi-view Diffusion Models

CVPR 2025highlight

Text- or image-to-3D generators and 3D scanners can now produce 3D assets with high-quality shapes and textures, but as single, fused entities lacking meaningful structure. In contrast, most applications and creative workflows require 3D assets to be composed of distinct, meaningful parts that can b…

Cited by 5SourcePDFScholar
2025

STraj: Self-training for Bridging the Cross-Geography Gap in Trajectory Prediction

AAAI 2025technical

Accurate trajectory prediction has prominent significance in autonomous driving scenarios. Most existing methods predict the trajectory of an agent by learning its interaction with other agents and the map within the scenario. However, the heterogeneous distribution of these elements across differe…

2025

VGGT: Visual Geometry Grounded Transformer

CVPR 2025award

We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. This approach is a step forward in 3D computer vision, where models have typicall…

2024

A Constrained Path Following Method for Snake-like Manipulators via Controlled Winding Uncoiling Strategy

ICRA 2024poster

Benefiting from its hyper-redundant structure, the biomimetic snake-like manipulator retains its remarkable flexibility even within confined spaces. However, its motion planning and control pose significant challenges. This paper imitates the winding uncoiling behavior of snakes to achieve controlla…

Cited by 1SourceScholar
2024

AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning

NeurIPS 2024poster

Large Language Models (LLM) based agents have shown promise in autonomously completing tasks across various domains, e.g., robotics, games, and web navigation. However, these agents typically require elaborate design and expert prompts to solve tasks in specific domains, which limits their adaptabil…

2024

G2LTraj: A Global-to-Local Generation Approach for Trajectory Prediction

IJCAI 2024poster

Predicting future trajectories of traffic agents accurately holds substantial importance in various applications such as autonomous driving. Previous methods commonly infer all future steps of an agent either recursively or simultaneously. However, the recursive strategy suffers from the accumulated…

2024

Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-dataset 3D Object Detection

CVPR 2024poster

Recent self-training techniques have shown notable improvements in unsupervised domain adaptation for 3D object detection (3D UDA). These techniques typically select pseudo labels i.e. 3D boxes to supervise models for the target domain. However this selection process inevitably introduces unreliable…

2024

TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without Training

AAAI 2024technical

Contrastive Language-Image Pre-training (CLIP) has demonstrated impressive capabilities in open-vocabulary classification. The class token in the image encoder is trained to capture the global features to distinguish different text descriptions supervised by contrastive loss, making it highly effect…

2023

A Novel Coiled Cable-Conduit-Driven Hyper-Redundant Manipulator for Remote Operating in Narrow Spaces

IROS 2023poster

Operating in narrow spaces is an important challenge in the development of robots. Redundant manipulators are one way to solve this problem, but their mechanism design and control method still have much room for improvement. In this paper, we propose a coiled cable-conduit-driven hyper-redundant man…

Cited by 4SourceScholar
2023

A Passive Compliance Obstacle-Crossing Robot for Power Line Inspection and Maintenance

RA-L 2023

In scenarios of the overhead power line system, manual methods are inefficient and unsafe. Meanwhile, the majority of cantilevered robots have poor efficiency when crossing obstacles. This letter proposes a novel power line inspection and maintenance robot to solve these problems. The robot employs

Cited by 15SourceScholar
2023

CLIP Is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation

CVPR 2023poster

Weakly supervised semantic segmentation (WSSS) with image-level labels is a challenging task. Mainstream approaches follow a multi-stage framework and suffer from high training costs. In this paper, we explore the potential of Contrastive Language-Image Pre-training models (CLIP) to localize differe…

2023

PVT-SSD: Single-Stage 3D Object Detector With Point-Voxel Transformer

CVPR 2023poster

Recent Transformer-based 3D object detectors learn point cloud features either from point- or voxel-based representations. However, the former requires time-consuming sampling while the latter introduces quantization errors. In this paper, we present a novel Point-Voxel Transformer for single-stage…

2022

Expanding Language-Image Pretrained Models for General Video Recognition

ECCV 2022poster

"Contrastive language-image pretraining has shown great success in learning visual-textual joint representation from web-scale data, demonstrating remarkable “zero-shot” generalization ability for various image tasks. However, how to effectively expand such new language-image pretraining methods to…

2022

Frame-Wise Action Representations for Long Videos via Sequence Contrastive Learning

CVPR 2022poster

Prior works on action representation learning mainly focus on designing various architectures to extract the global representations for short video clips. In contrast, many practical applications such as video alignment have strong demand for learning dense representations for long videos. In this p…

Cited by 53PDFcodeScholar
2021

Accelerate CNNs from Three Dimensions: A Comprehensive Pruning Framework

ICML 2021spotlight

Most neural network pruning methods, such as filter-level and layer-level prunings, prune the network model along one dimension (depth, width, or resolution) solely to meet a computational budget. However, such a pruning policy often leads to excessive reduction of that dimension, thus inducing a hu…

Cited by 77SourcePDFScholar
2021

CRFL: Certifiably Robust Federated Learning against Backdoor Attacks

ICML 2021spotlight

Federated Learning (FL) as a distributed learning paradigm that aggregates information from diverse clients to train a shared global model, has demonstrated great success. However, malicious clients can perform poisoning attacks and model replacement to introduce backdoors into the trained global mo…

2021

KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge Distillation

ICML 2021spotlight

Conventional unsupervised multi-source domain adaptation (UMDA) methods assume all source domains can be accessed directly. However, this assumption neglects the privacy-preserving policy, where all the data and computations must be kept decentralized. There exist three challenges in this scenario:…

2021

One-Shot Neural Ensemble Architecture Search by Diversity-Guided Search Space Shrinking

CVPR 2021poster

Despite remarkable progress achieved, most neural architecture search (NAS) methods focus on searching for one single accurate and robust architecture. To further build models with better generalization capability and performance, model ensemble is usually adopted and performs better than stand-alon…

Cited by 37PDFcodeScholar
2021

Rethinking and Improving Relative Position Encoding for Vision Transformer

ICCV 2021poster

Relative position encoding (RPE) is important for transformer to capture sequence ordering of input tokens. General efficacy has been proven in natural language processing. However, in computer vision, its efficacy is not well studied and even remains controversial, e.g., whether relative position e…

Cited by 451PDFcodeScholar
2021

SHOT-VAE: Semi-supervised Deep Generative Models With Label-aware ELBO Approximations

AAAI 2021technical

Semi-supervised variational autoencoders (VAEs) have obtained strong results, but have also encountered the challenge that good ELBO values do not always imply accurate inference results.In this paper, we investigate and propose two causes of this problem: (1) The ELBO objective cannot utilize the l…

2021

Salient Object Ranking With Position-Preserved Attention

ICCV 2021poster

Instance segmentation can detect where the objects are in an image, but hard to understand the relationship between them. We pay attention to a typical relationship, relative saliency. A closely related task, salient object detection, predicts a binary map highlighting a visually salient region whil…

Cited by 31PDFcodeScholar
2021

Searching the Search Space of Vision Transformer

NeurIPS 2021poster

Vision Transformer has shown great visual representation power in substantial vision tasks such as recognition and detection, and thus been attracting fast-growing efforts on manually designing more effective architectures. In this paper, we propose to use neural architecture search to automate this…