← Search

Hulingxiao He

6 accepted papers

2026

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning

ICLR 2026poster

Any entity in the visual world can be hierarchically grouped based on shared characteristics and mapped to fine-grained sub-categories. While Multi-modal Large Language Models (MLLMs) achieve strong performance on coarse-grained visual tasks, they often struggle with Fine-Grained Visual Recognition…

Cited by 0SourcecodeScholar
2026

Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models

ICML 2026poster

Taxonomies provide key information about the semantic relationships between concepts and the inherent organization of vision and language. Despite their impressive capabilities, large multimodal models (LMMs) often lack taxonomic knowledge, leading to low hierarchical visual recognition (HVR) consis…

Cited by 0SourceScholar
2026

Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models

CVPR 2026

A high-performing, general-purpose visual understanding model should map visual inputs to a taxonomic tree of labels, identify novel categories beyond the training set for which few or no publicly available images exist. Large Multimodal Models (LMMs) have achieved remarkable progress in fine-graine

Cited by 0SourcecodeScholar
2026

Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping

CVPR 2026

The widespread use of smartphones has made photography ubiquitous, yet a clear gap remains between ordinary users and professional photographers, who can identify aesthetic issues and provide actionable shooting guidance during capture. We define this capability as aesthetic guidance (AG) --- an ess

Cited by 0SourcecodeScholar
2025

Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models

ICLR 2025poster

Multi-modal large language models (MLLMs) have shown remarkable abilities in various visual understanding tasks. However, MLLMs still struggle with fine-grained visual recognition (FGVR), which aims to identify subordinate-level categories from images. This can negatively impact more advanced capabi…