← Search

Hong Li

20 accepted papers

2026

Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models

ICLR 2026poster

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in image and video understanding, their ability to comprehend the physical world has become an increasingly important research focus. Despite their improvements, current MLLMs struggle significantly with high-le…

Cited by 0SourceScholar
2026

Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation

ICRA 2026poster

Transparent objects remain notoriously hard for perception systems: refraction, reflection and transmission break the assumptions behind stereo, ToF and purely discriminative monocular depth, causing holes and temporally unstable estimates. Our key observation is that modern video diffusion models a…

2026

Light of Normals: Unified Feature Representation for Universal Photometric Stereo

ICLR 2026poster

Universal photometric stereo (PS) is defined by two factors: it must (i) operate under arbitrary, unknown lighting conditions and (ii) avoid reliance on specific illumination models. Despite progress (e.g., SDM UniPS), two challenges remain. First, current encoders cannot guarantee that illumination…

Cited by 0SourcecodeScholar
2026

PartDiffuser: Part-wise 3D Mesh Generation via Discrete Diffusion

CVPR 2026

Existing autoregressive (AR) methods for generating artist-designed meshes struggle to balance global structural consistency with high-fidelity local details, and are susceptible to error accumulation. To address this, we propose PartDiffuser, a novel semi-autoregressive diffusion framework for poin

Cited by 0SourceScholar
2025

AnimateAnything: Consistent and Controllable Animation for Video Generation

CVPR 2025poster

We propose a unified approach for video-controlled generation, enabling text-based guidance and manual annotations to control the generation of videos, similar to camera direction guidance. Specifically, we designed a two-stage algorithm. In the first stage, we convert all control information into f…

Cited by 10SourcePDFScholar
2025

IPDreamer: Appearance-Controllable 3D Object Generation with Complex Image Prompts

ICLR 2025poster

Recent advances in 3D generation have been remarkable, with methods such as DreamFusion leveraging large-scale text-to-image diffusion-based models to guide 3D object generation. These methods enable the synthesis of detailed and photorealistic textured objects. However, the appearance of 3D objects…

2025

The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs

ICLR 2025poster

Multi-modal Large Language Models (MLLMs) have exhibited impressive capability. However, recently many deficiencies of MLLMs have been found compared to human intelligence, $\textit{e.g.}$, hallucination. To drive the MLLMs study, the community dedicated efforts to building larger benchmarks with co…

Cited by 0SourcePDFScholar
2025

UniTransfer: Video Concept Transfer via Progressive Spatio-Temporal Decomposition

NeurIPS 2025poster

Recent advancements in video generation models have enabled the creation of diverse and realistic videos, with promising applications in advertising and film production. However, as one of the essential tasks of video generation models, video concept transfer remains significantly challenging. Exist…

Cited by 0SourceScholar
2025

WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection

ICCV 2025poster

Leveraging the complementary characteristics of visible (RGB) and infrared (IR) imagery offers significant potential for improving object detection. In this paper, we propose WaveMamba, a cross-modality fusion method that efficiently integrates the unique and complementary frequency features of RGB…

Cited by 0SourcePDFScholar
2024

A Relation-Aware Heterogeneous Graph Transformer on Dynamic Fusion for Multimodal Classification Tasks

ICASSP 2024accepted

Multimodal fusion aims to improve the performance of models for applications by extracting and fusing information in different modalities, including texts, images or others. Recent researches have shown that multimodal fusion is beneficial in many multimedia tasks. In this paper, we study typical mu…

Cited by 0SourceScholar
2024

DiffuX2CT: Diffusion Learning to Reconstruct CT Images from Biplanar X-Rays

ECCV 2024poster

"Computed tomography (CT) is widely utilized in clinical settings because it delivers detailed 3D images of the human body. However, performing CT scans is not always feasible due to radiation exposure and limitations in certain surgical environments. As an alternative, reconstructing CT images from…

Cited by 3SourcePDFScholar
2024

Hierarchical Aligned Multimodal Learning for NER on Tweet Posts

AAAI 2024technical

Mining structured knowledge from tweets using named entity recognition (NER) can be beneficial for many downstream applications such as recommendation and intention under standing. With tweet posts tending to be multimodal, multimodal named entity recognition (MNER) has attracted more attention. In…

Cited by 3SourcePDFScholar
2024

UV-IDM: Identity-Conditioned Latent Diffusion Model for Face UV-Texture Generation

CVPR 2024poster

3D face reconstruction aims at generating high-fidelity 3D face shapes and textures from single-view or multi-view images. However current prevailing facial texture generation methods generally suffer from low-quality texture identity information loss and inadequate handling of occlusions. To solve…

2023

Boosting Multi-modal Model Performance with Adaptive Gradient Modulation

ICCV 2023poster

While the field of multi-modal learning keeps growing fast, the deficiency of the standard joint training paradigm has become clear through recent studies. They attribute the sub-optimal performance of the jointly trained model to the modality competition phenomenon. Existing works attempt to improv…

Cited by 29PDFcodeScholar
2023

Improving the Modality Representation with multi-view Contrastive Learning for Multimodal Sentiment Analysis

ICASSP 2023accepted

Modality representation learning is an important problem for multimodal sentiment analysis (MSA), since the highly distinguishable representations can contribute to improving the analysis effect. Previous works of MSA have usually focused on internal fusion strategies for different modalities within…

Cited by 0SourceScholar
2022

FNeVR: Neural Volume Rendering for Face Animation

NeurIPS 2022accept

Face animation, one of the hottest topics in computer vision, has achieved a promising performance with the help of generative models. However, it remains a critical challenge to generate identity preserving and photo-realistic images due to the sophisticated motion deformation and complex facial de…

2022

Rethinking Spatial Invariance of Convolutional Networks for Object Counting

CVPR 2022poster

Previous work generally believes that improving the spatial invariance of convolutional networks is the key to object counting. However, after verifying several mainstream counting networks, we surprisingly found too strict pixel-level spatial invariance would cause overfit noise in the density map…

Cited by 124PDFcodeScholar