← Search

Haodong LI

20 accepted papers

2026

AD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured Reasoning

ICML 2026poster

Multimodal understanding of advertising videos is essential for interpreting the intricate relationship between visual storytelling and abstract persuasion strategies. However, despite excelling at general search, existing agents often struggle to bridge the cognitive gap between pixel-level percept…

Cited by 0SourceScholar
2026

D$^2$GS: Depth-and-Density Guided Gaussian Splatting for Stable and Accurate Sparse-View Reconstruction

ICLR 2026poster

Recent advances in 3D Gaussian Splatting (3DGS) enable real-time, high-fidelity novel view synthesis (NVS) with explicit 3D representations. However, performance degradation and instability remain significant under sparse-view conditions. In this work, we identify two key failure modes under sparse-…

Cited by 0SourcecodeScholar
2026

DA$^{2}$: Depth Anything in Any Direction

ICLR 2026poster

Panorama has a full FoV (360$^\circ\times$180$^\circ$), offering a more complete visual description than perspective images. Thanks to this characteristic, panoramic depth estimation is gaining increasing traction in 3D vision. However, due to the scarcity of panoramic data, previous methods are oft…

Cited by 0SourcecodeScholar
2026

Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation

CVPR 2026

In this work, we present a panoramic metric depth foundation model that generalizes across diverse scene distances. We explore a data-in-the-loop paradigm from the view of both data construction and framework design. We collect a large-scale dataset by combining public datasets, high-quality synthet

Cited by 0SourcecodeScholar
2026

Enabling Your Forensic Detector Know How Well It Performs on Distorted Samples

ICLR 2026poster

Generative AI has substantially facilitated realistic image synthesizing, posing great challenges for reliable forensics. When image forensic detectors are deployed in the wild, the inputs usually undergone various distortions including compression, rescaling, and lossy transmission. Such distortion…

Cited by 0SourceScholar
2026

From Parameters to Feature Space: Task Arithmetic for Backdoor Mitigation in Model Merging

ICML 2026poster

Model merging (MM) has gained significant attention as a cost-effective approach to integrate multiple task-specific models into a unified model. However, recent work reveals that MM is highly susceptible to backdoor attacks. Existing defenses based on task arithmetic often fail to eliminate backdoo…

Cited by 0SourceScholar
2026

PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception

ICML 2026poster

We introduce the Perception Rubric Benchmark (PRB), a rubric-based evaluation framework for Multimodal Large Language Models (MLLMs) that addresses the growing gap between benchmark scores and human-perceived quality. While standard perception metrics approach saturation, they produce compressed ran…

Cited by 0SourceScholar
2026

Proact-VL: A Proactive VideoLLM for Real-Time AI Companions

ICML 2026poster

Proactive and real-time interactive experiences are essential for human-like AI companions, yet face three key challenges: (1) achieving low-latency inference under continuous streaming inputs, (2) autonomously deciding when to respond, and (3) controlling both quality and quantity of generated cont…

Cited by 0SourceScholar
2026

Subspace-Aware Graph Construction and Contrastive Alignment for Multimodal Recommendation with Large Language Models

AAAI 2026technical

Multimedia content offers additional context for recommender systems to better understand user interests. Existing studies on multimodal recommendation primarily focus on constructing item-item semantic graphs. However, most of these methods capture only shallow semantic structures based on feature

Cited by 0SourcePDFScholar
2026

TraceTrans: Translation and Spatial Tracing for Surgical Prediction

AAAI 2026technical

Image-to-image translation models have achieved notable success in converting images across visual domains and are increasingly used for medical tasks such as predicting post-operative outcomes and modeling disease progression. However, most existing methods primarily aim to match the target distrib

Cited by 0SourcePDFScholar
2026

UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark

CVPR 2026

In real-world multimodal applications, systems usually need to comprehend arbitrarily combined and interleaved multimodal inputs from users, while also generating outputs in any interleaved multimedia form. This capability defines the goal of any-to-any interleaved multimodal learning under a unifie

Cited by 0SourceScholar
2025

Balancing User-Item Structure and Interaction with Large Language Models and Optimal Transport for Multimedia Recommendation

IJCAI 2025

The rapid growth of multimedia content has driven the development of recommender systems. Most previous work focuses on uncovering latent relationships among items to learn better representations. However, this approach does not sufficiently account for user affinities, potentially leading to an imb

2025

DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image Generation

ICLR 2025poster

In the realm of image generation, creating customized images from visual prompt with additional textual instruction emerges as a promising endeavor. However, existing methods, both tuning-based and tuning-free, struggle with interpreting the subject-essential attributes from the visual prompt. This…

Cited by 2SourcePDFScholar
2025

Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation

NeurIPS 2025poster

In this paper, we propose \textbf{Jasmine}, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD’s visual priors to enhance the sharpness and generalization of unsupervised prediction. Previous SD-based methods are all supervi…

Cited by 0SourceScholar
2025

Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction

ICLR 2025poster

Leveraging the visual priors of pre-trained text-to-image diffusion models offers a promising solution to enhance zero-shot generalization in dense prediction tasks. However, existing methods often uncritically use the original diffusion formulation, which may not be optimal due to the fundamental d…

Cited by 33SourcePDFScholar
2025

Multi-Scale Temporal Neural Network for Stock Trend Prediction Enhanced by Temporal Hyepredge Learning

IJCAI 2025

Existing research in Stock Trend Prediction (STP) focuses on temporal features extracted from a temporal sequence of stock data with a look-back window, which frequently leads to the omission of important periodic patterns, such as weekly and monthly variations in stock prices. Furthermore, these me

2024

LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching

CVPR 2024highlight

The recent advancements in text-to-3D generation mark a significant milestone in generative models unlocking new possibilities for creating imaginative 3D assets across various real-world scenarios. While recent advancements in text-to-3D generation have shown promise they often fall short in render…

2022

Dual Capsule Attention Mask Network with Mutual Learning for Visual Question Answering

COLING 2022main

A Visual Question Answering (VQA) model processes images and questions simultaneously with rich semantic information. The attention mechanism can highlight fine-grained features with critical information, thus ensuring that feature extraction emphasizes the objects related to the questions. However,…

Cited by 5SourcePDFScholar