← Search

Hao Lu

57 accepted papers

2026

ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Video Understanding

CVPR 2026

We revisit video hallucination in multimodal large language models (Video-MLLMs) from a semantic aggregation perspective. While prior work attributes hallucinations to language priors, missing frames, or visual encoder biases, these explanations overlook errors arising during the aggregation of corr

Cited by 0SourcecodeScholar
2026

EVA: Efficient Reinforcement Learning for End-to-End Video Agent

CVPR 2026

Video understanding with multimodal large language models (MLLMs) remains challenging due to the long token sequences of videos, which contain extensive temporal dependencies and redundant frames.Existing approaches typically treat MLLMs as passive recognizers, processing entire videos or uniformly

Cited by 0SourcecodeScholar
2026

GaussianMatch: Semi-Supervised Regression with Pseudo-Label Filtering via Multi-View Gaussian Consistency

CVPR 2026

Semi-Supervised Regression (SSR) is essential in domains like sentiment analysis and healthcare where labeled data is limited but unlabeled data is plentiful. Despite its practical importance, SSR remains underexplored due to the lack of effective pseudo-labeling strategies for continuous outputs. U

Cited by 0SourcecodeScholar
2026

IPFormer: Instance Prompt-guided Transformer for Multi-modal Multi-shot Video Understanding

AAAI 2026technical

Video Large Language Models (VideoLLMs), which adopt large language models for video understanding, have been demonstrated for single-shot videos. However, they usually struggle in multi-shot videos with frequent shot changes, varying camera angles, etc., which makes VideoLLMs hardly answer question

Cited by 0SourcePDFScholar
2026

MCTSr-Zero: Self-Reflective Psychological Counseling Dialogues Generation via Principles and Adaptive Exploration

AAAI 2026technical

The integration of Monte Carlo Tree Search (MCTS) with Large Language Models (LLMs) has demonstrated significant success in structured, problem-oriented tasks. However, applying these methods to open-ended dialogues, such as those in psychological counseling, presents unique challenges. Unlike tasks

Cited by 0SourcePDFScholar
2026

MV-FAC: Mean–Variance Value Function Factorization for Multi-Robot Mean–Standard Deviation Moving Target Search

IJCAI 2026

This paper studies a risk-sensitive formulation of the multi-robot search problem, termed multi-robot mean-standard deviation search (MuRMSS), in which a team of robots cooperatively search for a moving target by minimizing a linear combination of the mean and standard deviation of search time. Howe

Cited by 0Scholar
2026

MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis

CVPR 2026

Accurate and interpretable multi-disease diagnosis remains a critical challenge in medical research, particularly when leveraging heterogeneous multimodal medical data. Current approaches often rely on single-modal data, limiting their ability to comprehensively understand complex diseases. To addre

Cited by 0SourcecodeScholar
2026

Momentum Memory for Knowledge Distillation in Computational Pathology

CVPR 2026

Multimodal learning that integrates genomics and histopathology has shown strong potential in cancer diagnosis, yet its clinical translation is hindered by the limited availability of paired histology-genomics data. Knowledge distillation (KD) offers a practical solution by transferring genomic supe

Cited by 0SourcecodeScholar
2026

Plant Taxonomy Meets Plant Counting: A Fine-Grained, Taxonomic Dataset for Counting Hundreds of Plant Species

CVPR 2026

Visually cataloging and quantifying the natural world requires pushing the boundaries of both detailed visual classification and counting at scale. Despite significant progress, particularly in crowd and traffic analysis, the fine-grained, taxonomy-aware plant counting remains underexplored in visio

Cited by 0SourcecodeScholar
2026

PolarGuide-GSDR: 3D Gaussian Splatting Driven by Polarization Priors and Deferred Reflection for Real-World Reflective Scenes

CVPR 2026

Polarization-aware Neural Radiance Fields (NeRF) enables novel view synthesis of specular scenes but suffers from slow training, inefficient rendering, and material/viewpoint assumptions. 3D Gaussian Splatting (3DGS) supports real-time rendering but struggles with reflection reconstruction due to re

Cited by 0SourceScholar
2025

DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving

NeurIPS 2025poster

Large reconstruction model has remarkable progress, which can directly predict 3D or 4D representations for unseen scenes and objects. However, current work has not systematically explored the potential of large reconstruction models in the field of autonomous driving. To achieve this, we introduce…

Cited by 0SourcecodeScholar
2025

GDRIVE: Adaptive Object Detection in Autonomous Vehicles via Graph-Based Feature Learning

ICASSP 2025accepted

Navigating domain shifts in object detection is crucial for autonomous driving systems, particularly under varying weather conditions and diverse visual perspectives. Existing Cross-Domain Object Detection methods often struggle due to their reliance on broad semantic models, which can introduce bia…

Cited by 0SourceScholar
2025

MA-Det: A Discriminative Morphology-Aware Detector for Cervical Lesion Cell Clumps

ICASSP 2025accepted

Automated detection of cervical lesion cell clumps is crucial for cervical cancer screening. However, the dense packing and overlap of cells, caused by adhesion molecules, make detection challenging. To address this issue, we propose the Morphology-Aware Detector (MA-Det). Specifically, by innovativ…

Cited by 0SourceScholar
2025

MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models

CVPR 2025poster

Recent advances in generalizable 3D Gaussian Splatting have demonstrated promising results in real-time high-fidelity rendering without per-scene optimization, yet existing approaches still struggle to handle unfamiliar visual content during inference on novel scenes due to limited generalizability.…

2025

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

NeurIPS 2025poster

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization…

Cited by 0SourcecodeScholar
2025

Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models

ICRA 2025

Large Language Models (LLMs) have made substantial advancements in the field of robotic and autonomous driving. This study presents the first Occupancy-based Large Language Model (Occ-LLM), which represents a pioneering effort to integrate LLMs with an important representation. To effectively encode

Cited by 23SourceScholar
2025

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model

CVPR 2025poster

Periodic or quasi-periodic phenomena reveal intrinsic characteristics in various natural processes, such as weather patterns, movement behaviors, traffic flows, and biological signals. Given that these phenomena span multiple modalities, the capabilities of Multimodal Large Language Models (MLLMs) o…

2025

RhythmGuassian: Repurposing Generalizable Gaussian Model For Remote Physiological Measurement

ICCV 2025poster

Remote Photoplethysmography (rPPG) enables non-contact extraction of physiological signals, providing significant advantages in medical monitoring, emotion recognition, and face anti-spoofing. However, the extraction of reliable rPPG signals is hindered by motion variations in real-world environment…

2025

Towards Generalizable Multi-Camera 3D Object Detection via Perspective Rendering

AAAI 2025technical

Detecting and localizing objects in 3D space using multiple cameras, known as Multi-Camera 3D Object Detection (MC3D-Det), has gained prominence with the advent of bird's-eye view (BEV) approaches. However, these methods often struggle with the serious domain gaps caused by various viewpoints and en…

2024

An Incremental Unified Framework for Small Defect Inspection

ECCV 2024poster

"Artificial Intelligence (AI)-driven defect inspection is pivotal in industrial manufacturing. However, existing inspection systems are typically designed for specific industrial products and struggle with diverse product portfolios and evolving processes. Although some previous studies attempt to a…

2024

Backdoor Contrastive Learning via Bi-level Trigger Optimization

ICLR 2024poster

Contrastive Learning (CL) has attracted enormous attention due to its remarkable capability in unsupervised representation learning. However, recent works have revealed the vulnerability of CL to backdoor attacks: the feature extractor could be misled to embed backdoored data close to an attack targ…

2024

Bi-TTA: Bidirectional Test-Time Adapter for Remote Physiological Measurement

ECCV 2024poster

"Remote photoplethysmography (rPPG) is gaining prominence for its non-invasive approach to monitoring physiological signals using only cameras. Despite its promise, the adaptability of rPPG models to new, unseen domains is hindered due to the environmental sensitivity of physiological signals. To ad…

2024

HAWK: Learning to Understand Open-World Video Anomalies

NeurIPS 2024poster

Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prev…

2024

Sampling-based Motion Planning for Optimal Probability of Collision under Environment Uncertainty

IROS 2024poster

Motion planning is a fundamental capability in robotics applications. Real-world scenarios can introduce uncertainty to the motion planning problem. In this work we study environment uncertainty in general high-dimensional problems wherein the choice of appropriate metrics and formulations are shown…

Cited by 0SourceScholar
2024

Self-Motion As Supervision For Egocentric Audiovisual Localization

ICASSP 2024accepted

Sound source localization is a key requirement for many assistive applications of augmented reality, such as speech enhancement. In conversational settings, potential sources of interest may be approximated by active speaker detection. However, localizing speakers in crowded, noisy environments is c…

Cited by 0SourceScholar
2024

Unifying Automatic and Interactive Matting with Pretrained ViTs

CVPR 2024poster

Automatic and interactive matting largely improve image matting by respectively alleviating the need for auxiliary input and enabling object selection. Due to different settings on whether prompts exist they either suffer from weakness in instance completeness or region details. Also when dealing wi…

2024

Vision Transformer Off-the-Shelf: A Surprising Baseline for Few-Shot Class-Agnostic Counting

AAAI 2024technical

Class-agnostic counting (CAC) aims to count objects of interest from a query image given few exemplars. This task is typically addressed by extracting the features of query image and exemplars respectively and then matching their feature similarity, leading to an extract-then-match paradigm. In this…

2023

ESTextSpotter: Towards Better Scene Text Spotting with Explicit Synergy in Transformer

ICCV 2023poster

In recent years, end-to-end scene text spotting approaches are evolving to the Transformer-based framework. While previous studies have shown the crucial importance of the intrinsic synergy between text detection and recognition, recent advances in Transformer-based methods usually adopt an implicit…

Cited by 37PDFcodeScholar
2023

Fast Full-frame Video Stabilization with Iterative Optimization

ICCV 2023poster

Video stabilization refers to the problem of transforming a shaky video into a visually pleasing one. The question of how to strike a good trade-off between visual quality and computational speed has remained one of the open challenges in video stabilization. Inspired by the analogy between wobbly f…

Cited by 12PDFcodeScholar
2023

Find Beauty in the Rare: Contrastive Composition Feature Clustering for Nontrivial Cropping Box Regression

AAAI 2023technical

Automatic image cropping algorithms aim to recompose images like human-being photographers by generating the cropping boxes with improved composition quality. Cropping box regression approaches learn the beauty of composition from annotated cropping boxes. However, the bias of annotations leads to q…

Cited by 6SourcePDFScholar
2023

Infusing Definiteness into Randomness: Rethinking Composition Styles for Deep Image Matting

AAAI 2023technical

We study the composition style in deep image matting, a notion that characterizes a data generation flow on how to exploit limited foregrounds and random backgrounds to form a training dataset. Prior art executes this flow in a completely random manner by simply going through the foreground pool or…

2023

Learning Second-Order Attentive Context for Efficient Correspondence Pruning

AAAI 2023technical

Correspondence pruning aims to search consistent correspondences (inliers) from a set of putative correspondences. It is challenging because of the disorganized spatial distribution of numerous outliers, especially when putative correspondences are largely dominated by outliers. It's more challengin…

2022

BokehMe: When Neural Rendering Meets Classical Rendering

CVPR 2022oral

We propose BokehMe, a hybrid bokeh rendering framework that marries a neural renderer with a classical physically motivated renderer. Given a single image and a potentially imperfect disparity map, BokehMe generates high-resolution photo-realistic bokeh effects with adjustable blur size, focal plane…

Cited by 48PDFcodeScholar
2022

FADE: Fusing the Assets of Decoder and Encoder for Task-Agnostic Upsampling

ECCV 2022poster

"We consider the problem of task-agnostic feature upsampling in dense prediction where an upsampling operator is required to facilitate both region-sensitive tasks like semantic segmentation and detail-sensitive tasks such as image matting. Existing upsampling operators often can work well in either…

Cited by 48SourcePDFScholar
2022

FIORA : A Flexible Tendon-Driven Continuum Manipulator for Laparoscopic Surgery

RA-L 2022

Aiming at solving the problems of low flexibility, poor internal extension, and chopstick effect in most laparoscopic surgical robots, this letter presents a flexible tendon-driven continuum surgical manipulator with eight degrees of freedom, called FIORA. The manipulator is composed of three indepe

Cited by 29SourceScholar
2022

MPIB: An MPI-Based Bokeh Rendering Framework for Realistic Partial Occlusion Effects

ECCV 2022poster

"Partial occlusion effects are a phenomenon that blurry objects near a camera are semi-transparent, resulting in partial appearance of occluded background. However, it is challenging for existing bokeh rendering methods to simulate realistic partial occlusion effects due to the missing information o…

2022

Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic Counting

CVPR 2022poster

Class-agnostic counting (CAC) aims to count all instances in a query image given few exemplars. A standard pipeline is to extract visual features from exemplars and match them with query images to infer object counts. Two essential components in this pipeline are feature representation and similarit…

Cited by 109PDFcodeScholar
2022

Robust Object Detection with Inaccurate Bounding Boxes

ECCV 2022poster

"Learning accurate object detectors often requires large-scale training data with precise object bounding boxes. However, labeling such data is expensive and time-consuming. As the crowd-sourcing labeling process and the ambiguities of the objects may raise noisy bounding box annotations, the object…

2022

SAPA: Similarity-Aware Point Affiliation for Feature Upsampling

NeurIPS 2022accept

We introduce point affiliation into feature upsampling, a notion that describes the affiliation of each upsampled point to a semantic cluster formed by local decoder feature points with semantic similarity. By rethinking point affiliation, we present a generic formulation for generating upsampling k…

2021

ASV-SUBTOOLS: Open Source Toolkit for Automatic Speaker Verification

ICASSP 2021accepted

In this paper, we introduce a new open source toolkit for automatic speaker verification (ASV), named ASV-Subtools. Adopting PyTorch as main deep learning engine and Kaldi toolkit for data processing, ASV-Subtools allows users to develop modern speaker recognizers flexibly and efficiently. The toolk…

Cited by 0SourceScholar
2021

Bootstrapping Fitted Q-Evaluation for Off-Policy Inference

ICML 2021spotlight

Bootstrapping provides a flexible and effective approach for assessing the quality of batch reinforcement learning, yet its theoretical properties are poorly understood. In this paper, we study the use of bootstrapping in off-policy evaluation (OPE), and in particular, we focus on the fitted Q-evalu…

Cited by 52SourcePDFScholar
2021

TransView: Inside, Outside, and Across the Cropping View Boundaries

ICCV 2021poster

We show that relation modeling between visual elements matters in cropping view recommendation. Cropping view recommendation addresses the problem of image recomposition conditioned on the composition quality and the ranking of views (cropped sub-regions). This task is challenging because the visual…

Cited by 23PDFScholar
2020

Weighing Counts: Sequential Crowd Counting by Reinforcement Learning

ECCV 2020poster

We formulate counting as a sequential decision problem and present a novel crowd counting model solvable by deep reinforcement learning. In contrast to existing counting models that directly output count values, we divide one-step estimation into a sequence of much easier and more tractable sub-deci…

Cited by 97SourcePDFScholar
2019

From Open Set to Closed Set: Counting Objects by Spatial Divide-and-Conquer

ICCV 2019poster

Visual counting, a task that predicts the number of objects from an image/video, is an open-set problem by nature, i.e., the number of population can vary in [0,+[?]) in theory. However, the collected images and labeled count values are limited in reality, which means only a small closed set is obse…

Cited by 214PDFcodeScholar
2018

Monocular Relative Depth Perception With Web Stereo Data Supervision

CVPR 2018poster

In this paper we study the problem of monocular relative depth perception in the wild. We introduce a simple yet effective method to automatically generate dense relative depth annotations from web stereo images, and propose a new dataset that consists of diverse images as well as corresponding dens…

Cited by 253SourcePDFScholar
2018

The Edge Density Barrier: Computational-Statistical Tradeoffs in Combinatorial Inference

ICML 2018oral

We study the hypothesis testing problem of inferring the existence of combinatorial structures in undirected graphical models. Although there exist extensive studies on the information-theoretic limits of this problem, it remains largely unexplored whether such limits can be attained by efficient al…

Cited by 10SourcePDFScholar
2017

When Unsupervised Domain Adaptation Meets Tensor Representations

ICCV 2017poster

Domain adaption (DA) allows machine learning methods trained on data sampled from one distribution to be applied to data sampled from another. It is thus of great practical importance to the application of such methods. Despite the fact that tensor representations are widely used in Computer Vision…

Cited by 89PDFcodeScholar