← Search

Xuming He

64 accepted papers

2026

AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis

CVPR 2026

Generating human grasping poses that accurately reflect both object geometry and user-specified interaction semantics is essential for natural hand-object interactions in AR/VR and embodied AI. However, existing semantic grasping approaches struggle with the large modality gap between 3D object repr

Cited by 0SourceScholar
2026

CryoKRAQEN: Kernel-Regularized Annealing for Quantized Embedding Networks in Cryo-EM Heterogeneous Reconstruction

CVPR 2026

Heterogeneous reconstruction in cryo-electron microscopy (Cryo-EM) is fundamental for understanding macromolecular structural diversity, yet remains challenging due to extreme noise, continuous conformational changes, and ambiguous image-to-structure mappings. Existing neural approaches often rely o

Cited by 0SourceScholar
2026

Omni-Weather: Unified Multimodal Foundation Model for Weather Generation and Understanding

ICLR 2026poster

Weather modeling requires both accurate prediction and mechanistic interpretation, yet existing methods treat these goals in isolation, separating generation from understanding. To address this gap, we present Omni-Weather, the first multimodal foundation model that unifies weather generation and un…

Cited by 0SourcecodeScholar
2026

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum

ICLR 2026poster

Knowledge-Based Visual Question Answering (KB-VQA) requires models to answer questions about an image by integrating external knowledge, posing significant challenges due to noisy retrieval and the structured, encyclopedic nature of the knowledge base. These characteristics create a distributional g…

Cited by 0SourceScholar
2026

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition

CVPR 2026

Open-domain visual entity recognition (VER) seeks to associate images with entities in encyclopedic knowledge bases such as Wikipedia. Recent generative methods tailored for VER demonstrate strong performance but incur high computational costs, limiting their scalability and practical deployment. In

Cited by 0SourcecodeScholar
2025

DiffSR: Learning Radar Reflectivity Synthesis via Diffusion Model from Satellite Observations

ICASSP 2025accepted

Weather radar data synthesis can fill in data for areas where ground observations are missing. Existing methods often employ reconstruction-based approaches with MSE loss to reconstruct radar data from satellite observation. However, such methods lead to over-smoothing, which hinders the generation…

Cited by 0SourceScholar
2025

GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation

NeurIPS 2025poster

While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utilization. We present a reasoning-enhanced framework that systematically integrates structured reasoning, action prediction,…

Cited by 0SourceScholar
2025

GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View Localization

ICCV 2025poster

Cross-view localization, the task of estimating a camera's 3-degrees-of-freedom (3-DoF) pose by aligning ground-level images with aerial images, is crucial for large-scale outdoor applications like autonomous navigation and augmented reality. Existing methods often rely on fully supervised learning,…

2025

LithoSim: A Large, Holistic Lithography Simulation Benchmark for AI-Driven Semiconductor Manufacturing

NeurIPS 2025poster

Lithography orchestrates a symphony of light, mask and photochemicals to transfer the integrated circuit patterns onto the wafer. Lithography simulation serves as the critical nexus between circuit design and manufacturing, where its speed and accuracy fundamentally govern the optimization quality o…

Cited by 0SourcecodeScholar
2025

NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation

NeurIPS 2025poster

Reinforcement learning (RL) has shown promise in enhancing the general Chain-of-Thought (CoT) reasoning capabilities of multimodal large language models (MLLMs). However, when applied to improve general CoT reasoning, existing RL frameworks often struggle to generalize beyond the training distributi…

Cited by 0SourceScholar
2025

RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts

NeurIPS 2025poster

Quality analysis of weather forecasts is an essential topic in meteorology. Although traditional score-based evaluation metrics can quantify certain forecast errors, they are still far from meteorological experts in terms of descriptive capability, interpretability, and understanding of dynamic evol…

Cited by 0SourceScholar
2025

Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph Generation

AAAI 2025technical

Open-vocabulary Scene Graph Generation (OV-SGG) overcomes the limitations of the closed-set assumption by aligning visual relationship representations with open-vocabulary textual representations. This enables the identification of novel visual relationships, making it applicable to real-world scena…

Cited by 1SourcePDFScholar
2025

Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning

NeurIPS 2025poster

Scientific discoveries increasingly rely on complex multimodal reasoning based on information-intensive scientific data and domain-specific expertise. Empowered by expert-level scientific benchmarks, scientific Multimodal Large Language Models (MLLMs) hold the potential to significantly enhance this…

Cited by 0SourceScholar
2025

TokMan:Tokenize Manhattan Mask Optimization for Inverse Lithography

NeurIPS 2025poster

Manhattan representations, defined by axis-aligned, orthogonal structures, are widely used in vision, robotics, and semiconductor design for their geometric regularity and algorithmic simplicity. In integrated circuit (IC) design, Manhattan geometry is key for routing, design rule checking, and lith…

Cited by 0SourceScholar
2024

"SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models"

ECCV 2024poster

"We present , a versatile multi-modal large language model (MLLM) with a joint mixing of model weights, visual embeddings and image scales. First, for stronger vision-language alignment, we unfreeze the large language model (LLM) during pre-training, and introduce a weight mix strategy between LLMs…

2024

CryoGEM: Physics-Informed Generative Cryo-Electron Microscopy

NeurIPS 2024poster

In the past decade, deep conditional generative models have revolutionized the generation of realistic images, extending their application from entertainment to scientific domains. Single-particle cryo-electron microscopy (cryo-EM) is crucial in resolving near-atomic resolution 3D structures of prot…

Cited by 0SourcePDFScholar
2024

Dual-level Adaptive Self-Labeling for Novel Class Discovery in Point Cloud Segmentation

ECCV 2024poster

"We tackle the novel class discovery in point cloud segmentation, which discovers novel classes based on the semantic knowledge of seen classes. Existing work proposes an online point-wise clustering method with a simplified equal class-size constraint on the novel classes to avoid degenerate soluti…

2024

From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models

CVPR 2024poster

Scene graph generation (SGG) aims to parse a visual scene into an intermediate graph representation for downstream reasoning tasks. Despite recent advancements existing methods struggle to generate scene graphs with novel visual relation concepts. To address this challenge we introduce a new open-vo…

2024

Generalize or Detect? Towards Robust Semantic Segmentation Under Multiple Distribution Shifts

NeurIPS 2024poster

In open-world scenarios, where both novel classes and domains may exist, an ideal segmentation model should detect anomaly classes for safety and generalize to new domains. However, existing methods often struggle to distinguish between domain-level and semantic-level distribution shifts, leading to…

2024

Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language Reasoning

CVPR 2024poster

Generative vision-language models (VLMs) have shown impressive performance in zero-shot vision-language tasks like image captioning and visual question answering.However improving their zero-shot reasoning typically requires second-stage instruction tuning which relies heavily on human-labeled or la…

2024

Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training

AAAI 2024technical

Image captioning aims at generating descriptive and meaningful textual descriptions of images, enabling a broad range of vision-language applications. Prior works have demonstrated that harnessing the power of Contrastive Image Language Pre-training (CLIP) offers a promising approach to achieving ze…

2024

Multi-Level Progressive Reinforcement Learning for Control Policy in Physical Simulations

ICRA 2024poster

Training model-free intelligent agents in complex real-world scenarios using reinforcement learning (RL) often necessitates simulation-based environments due to high physical expenses. However, when simulation takes a long time, e.g., in an unsteady 3D fluid simulation with interactions to the contr…

Cited by 0SourceScholar
2024

P$^2$OT: Progressive Partial Optimal Transport for Deep Imbalanced Clustering

ICLR 2024poster

Deep clustering, which learns representation and semantic clustering without labels information, poses a great challenge for deep learning-based approaches. Despite significant progress in recent years, most existing methods focus on uniformly distributed datasets, significantly limiting the practic…

2024

RealDex: Towards Human-like Grasping for Robotic Dexterous Hand

IJCAI 2024poster

In this paper, we introduce RealDex, a pioneering dataset capturing authentic dexterous hand grasping motions infused with human behavioral patterns, enriched by multi-view and multimodal visual data. Utilizing a teleoperation system, we seamlessly synchronize human-robot hand poses in real time. Th…

2023

ATTA: Anomaly-aware Test-Time Adaptation for Out-of-Distribution Detection in Segmentation

NeurIPS 2023poster

Recent advancements in dense out-of-distribution (OOD) detection have primarily focused on scenarios where the training and testing datasets share a similar domain, with the assumption that no domain shift exists between them. However, in real-world situations, domain shift often exits and significa…

2023

CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free Attention

AAAI 2023technical

Contrastive Language-Image Pre-training (CLIP) has been shown to learn visual representations with promising zero-shot performance. To further improve its downstream accuracy, existing works propose additional learnable modules upon CLIP and fine-tune them by few-shot training sets. However, the res…

2023

Exploring Learning-Based Control Policy for Fish-Like Robots in Altered Background Flows

IROS 2023poster

The study of motion control for the fish-like robots in complex fluid fields is of great importance in improving the performance of underwater vehicles, due to its strong maneuverability, propulsion efficiency, and deceptive visual appearance. In this article, a novel learning-based control framewor…

Cited by 2SourceScholar
2023

Grounded Image Text Matching with Mismatched Relation Reasoning

ICCV 2023poster

This paper introduces Grounded Image Text Matching with Mismatched Relation (GITM-MR), a novel visual-linguistic joint task that evaluates the relation understanding capabilities of transformer-based pre-trained models. GITM-MR requires a model to first determine if an expression describes an image,…

Cited by 8PDFcodeScholar
2023

HOICLIP: Efficient Knowledge Transfer for HOI Detection With Vision-Language Models

CVPR 2023poster

Human-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Recently, Contrastive Language-Image Pre-training (CLIP) has shown great potential in providing interaction prior for HOI detectors via knowledge distillation. However, such approaches ofte…

2023

Human-centric Scene Understanding for 3D Large-scale Scenarios

ICCV 2023poster

Human-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human-environment interactions, severe occlusions in crowds, etc. In this paper, we present a large-scale multi-modal dataset…

Cited by 26PDFcodeScholar
2023

MILD: Modeling the Instance Learning Dynamics for Learning with Noisy Labels

IJCAI 2023poster

Despite deep learning has achieved great success, it often relies on a large amount of training data with accurate labels, which are expensive and time-consuming to collect. A prominent direction to reduce the cost is to learn with noisy labels, which are ubiquitous in the real-world applications. A…

2023

Modeling Multimodal Aleatoric Uncertainty in Segmentation with Mixture of Stochastic Experts

ICLR 2023poster

Equipping predicted segmentation with calibrated uncertainty is essential for safety-critical applications. In this work, we focus on capturing the data-inherent uncertainty (aka aleatoric uncertainty) in segmentation, typically when ambiguities exist in input images. Due to the high-dimensional out…

2023

Weakly-supervised HOI Detection via Prior-guided Bi-level Representation Learning

ICLR 2023poster

Human object interaction (HOI) detection plays a crucial role in human-centric scene understanding and serves as a fundamental building block for many vision tasks. One generalizable and scalable strategy for HOI detection is to use weak supervision, learning from image-level annotations only. This…

Cited by 15SourcePDFScholar
2022

FishGym: A High-Performance Physics-based Simulation Framework for Underwater Robot Learning

ICRA 2022poster

Bionic underwater robots have demonstrated their superiority in many applications. Yet, training their intelligence for a variety of tasks that mimic the behavior of underwater creatures poses a number of challenges in practice, mainly due to lack of a large amount of available training data as well…

Cited by 14SourceScholar
2022

Generative Negative Text Replay for Continual Vision-Language Pretraining

ECCV 2022poster

"Vision-language pre-training (VLP) has attracted increasing attention recently. With a large amount of image-text pairs, VLP models trained with contrastive loss have achieved impressive performance in various tasks, especially the zero-shot generalization on downstream datasets. In practical appli…

Cited by 26SourcePDFScholar
2022

KD-VLP: Improving End-to-End Vision-and-Language Pretraining with Object Knowledge Distillation

NAACL 2022findings

Self-supervised vision-and-language pretraining (VLP) aims to learn transferable multi-modal representations from large-scale image-text data and to achieve strong performances on a broad scope of vision-language tasks after finetuning. Previous mainstream VLP approaches typically adopt a two-step s…

Cited by 32SourcePDFScholar
2022

Learning Semantic Correspondence with Sparse Annotations

ECCV 2022poster

"Finding dense semantic correspondence is a fundamental problem in computer vision, which remains challenging in complex scenes due to background clutter, extreme intra-class variation, and a severe lack of ground truth. In this paper, we aim to address the challenge of label sparsity in semantic co…

2021

Bipartite Graph Network With Adaptive Message Passing for Unbiased Scene Graph Generation

CVPR 2021poster

Scene graph generation is an important visual understanding task with a broad range of vision applications. Despite recent tremendous progress, it remains challenging due to the intrinsic long-tailed class distribution and large intra-class variation. To address these issues, we introduce a novel co…

Cited by 281PDFcodeScholar
2021

Distribution Alignment: A Unified Framework for Long-Tail Visual Recognition

CVPR 2021poster

Despite the success of the deep neural networks, it remains challenging to effectively build a system for long-tail visual recognition tasks. To address this problem, we first investigate the performance bottleneck of the two-stage learning framework via ablative study. Motivated by our discovery, w…

Cited by 365PDFcodeScholar
2021

Dynamic Grained Encoder for Vision Transformers

NeurIPS 2021poster

Transformers, the de-facto standard for language modeling, have been recently applied for vision tasks. This paper introduces sparse queries for vision transformers to exploit the intrinsic spatial redundancy of natural images and save computational costs. Specifically, we propose a Dynamic Grained…

2021

GNeRF: GAN-Based Neural Radiance Field Without Posed Camera

ICCV 2021poster

We introduce GNeRF, a framework to marry Generative Adversarial Networks (GAN) with Neural Radiance Field (NeRF) reconstruction for the complex scenarios with unknown and even randomly initialized camera poses. Recent NeRF-based advances have gained popularity for remarkable realistic novel view syn…

Cited by 223PDFcodeScholar
2021

Learning Implicit Temporal Alignment for Few-shot Video Classification

IJCAI 2021poster

Few-shot video classification aims to learn new video categories with only a few labeled examples, alleviating the burden of costly annotation in real-world applications. However, it is particularly challenging to learn a class-invariant spatial-temporal representation in such a setting. To address…

2020

Part-aware Prototype Network for Few-shot Semantic Segmentation

ECCV 2020poster

Few-shot semantic segmentation aims to learn to segment new object classes with only a few annotated examples, which has a wide range of real-world applications. Most existing methods either focus on the restrictive setting of one-way few-shot segmentation or suffer from incomplete coverage of objec…

2019

Dynamic Context Correspondence Network for Semantic Alignment

ICCV 2019poster

Establishing semantic correspondence is a core problem in computer vision and remains challenging due to large intra-class variations and lack of annotated data. In this paper, we aim to incorporate global semantic context in a flexible manner to overcome the limitations of prior work that relies on…

Cited by 107PDFScholar
2019

LatentGNN: Learning Efficient Non-local Relations for Visual Recognition

ICML 2019oral

Capturing long-range dependencies in feature representations is crucial for many visual recognition tasks. Despite recent successes of deep convolutional networks, it remains challenging to model non-local context relations between visual features. A promising strategy is to model the feature contex…

2019

Pose-Aware Multi-Level Feature Network for Human Object Interaction Detection

ICCV 2019oral

Reasoning human object interactions is a core problem in human-centric scene understanding and detecting such relations poses a unique challenge to vision systems due to large variations in human-object configurations, multiple co-occurring relation instances and subtle visual difference between rel…

Cited by 276PDFcodeScholar
2018

SemStyle: Learning to Generate Stylised Image Captions Using Unaligned Text

CVPR 2018poster

Linguistic style is an essential part of written communication, with the power to affect both clarity and attractiveness. With recent advances in vision and language, we can start to tackle the problem of generating image captions that are both visually grounded and appropriately styled. Existing ap…

2017

Indoor Scene Parsing With Instance Segmentation, Semantic Labeling and Support Relationship Inference

CVPR 2017poster

Over the years, indoor scene parsing has attracted a growing interest in the computer vision community. Existing methods have typically focused on diverse subtasks of this challenging problem. In particular, while some of them aim at segmenting the image into regions, such as object or surface insta…

Cited by 39PDFScholar
2015

Indoor Scene Structure Analysis for Single Image Depth Estimation

CVPR 2015poster

We tackle the problem of single image depth estimation, which, without additional knowledge, suffers from many ambiguities. Unlike previous approaches that only reason locally, we propose to exploit the global structure of the scene to estimate its depth. To this end, we introduce a hierarchical rep…

Cited by 143SourcePDFScholar
2015

Separating Objects and Clutter in Indoor Scenes

CVPR 2015poster

Objects' spatial layout estimation and clutter identification are two important tasks to understand indoor scenes. We propose to solve both of these problems in a joint framework using RGBD images of indoor scenes. In contrast to recent approaches which focus on either one of these two problems, we…

Cited by 26SourcePDFScholar