← Search

Zhihui Li

16 accepted papers

2026

BiOTPrompt: Bidirectional Optimal Transport Guided Prompting for Disease Evolution-aware Radiology Report Generation

CVPR 2026

Radiology report generation (RRG) aims to automatically describe medical images via free-text reports. In clinical practice, comparing current and prior chest X-rays is essential for assessing disease progression, motivating the development of longitudinal RRG methods. However, most existing approac

Cited by 0SourcecodeScholar
2026

CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation Detection

ICML 2026poster

The rapid rise of generative AI has made multimodal fake news increasingly realistic and pervasive, posing severe threats to public trust and social stability. Existing detection methods rely heavily on manipulation-specific models and large-scale labeled data, resulting in poor generalization to em…

Cited by 0SourceScholar
2026

Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning

CVPR 2026

Human video generation has advanced rapidly with the development of diffusion models, but the high computational cost and substantial memory consumption associated with training these models on high-resolution, multi-frame data pose significant challenges. In this paper, we propose Entropy-Guided Pr

Cited by 0SourcecodeScholar
2026

Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition

CVPR 2026

Recently, masked skeleton reconstruction models have emerged as strong action representation learners, driving significant progress in self-supervised skeleton-based action recognition. However, existing state-of-the-art methods must predict an exceedingly large number of spatiotemporal patches, sig

Cited by 0SourcecodeScholar
2026

LangField4D: Learning Identity-Adaptive and Spatio-Temporal Continuous 4D Language Fields for Dynamic Scenes

CVPR 2026

Constructing a 4D language field that supports open-vocabulary queries is essential for semantic perception and interaction in dynamic environments. Existing 4D Gaussian-based approaches face two major challenges. First, the assumption of a static identity per Gaussian leads to semantic inconsistenc

Cited by 0SourceScholar
2026

ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation

CVPR 2026

The ability to extrapolate dynamic 3D scenes beyond the observed timeframe is fundamental to advancing physical world understanding and predictive modeling. Existing dynamic 3D reconstruction methods have achieved high-fidelity rendering of temporal interpolation, but typically lack physical consist

Cited by 0SourceScholar
2026

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

ICLR 2026poster

Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability that is overlooked by conventional video LLMs evaluated in offline settings. The shift to an online, streaming paradigm introduces significant challen…

Cited by 0SourceScholar
2026

RiskProp: Collision-Anchored Self-Supervised Risk Propagation For Early Accident Anticipation

CVPR 2026

Accident anticipation aims to predict impending collisions from dashcam videos and trigger early alerts. Existing methods rely on binary supervision with manually annotated "anomaly onset" frames, which are subjective and inconsistent, leading to inaccurate risk estimation. In contrast, we propose R

Cited by 0SourcecodeScholar
2026

SANER: Switchable Adapter with Non-parametric Enhanced Routing for Person De-Reidentification

CVPR 2026

Person De-Reidentification (De-ReID) is an emerging and safety-critical task that aims to selectively forget specific individuals in surveillance systems while preserving the recognition capability for others. Existing methods typically learn both forgetting and retaining objectives within a unified

Cited by 0SourcecodeScholar
2026

Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models

AAAI 2026technical

Text-guided image inpainting aims to inpaint masked image regions based on a textual prompt while preserving the background. Although diffusion-based methods have become dominant, their property of modeling the entire image in latent space makes it challenging for the results to align well with prom

Cited by 0SourcePDFScholar
2024

Masked Distillation Advances Self-Supervised Transformer Architecture Search

ICLR 2024poster

Transformer architecture search (TAS) has achieved remarkable progress in automating the neural architecture design process of vision transformers. Recent TAS advancements have discovered outstanding transformer architectures while saving tremendous labor from human experts. However, it is still cum…

Cited by 2SourcePDFScholar
2023

HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object Segmentation

ICCV 2023poster

Referring Video Object Segmentation (RVOS) is to segment the object instance from a given video, according to the textual description of this object. However, in the open world, the object descriptions are often diversified in contents and flexible in lengths. This leads to the key difficulty in RVO…

Cited by 30PDFScholar
2022

Airborne Mimo Radar Transmit-Receive Design Under Spectral Constraint in Signal-Dependent Clutter

ICASSP 2022accepted

This paper considers the joint design of the transmit waveform and receive filter for airborne multiple-input multiple-output (MIMO) radar under spectral constraint in signal-dependent clutter. The spatial-frequency spectral compatibility constraint is imposed in the joint design problem. To tackle…

Cited by 0SourceScholar
2021

Parameter Identifiability Of Spatial-Smoothing-Based Bistatic Mimo Radar

ICASSP 2021accepted

Diversity smoothing has been widely developed for angle estimation with bistatic multiple input multiple output (MIMO) radar in the presence of coherent targets, the parameter identifiability of which is an important issue. In this paper, we are devoted to establishing more accurate conditions by st…

Cited by 0SourceScholar