← Search

Wei-Hong Li

11 accepted papers

2026

3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding

CVPR 2026

This paper addresses the challenge of training a single network to jointly perform multiple dense prediction tasks, such as segmentation and depth estimation, i.e., multi-task learning (MTL). Current approaches mainly capture cross-task relations in the 2D image space, often leading to unstructured

Cited by 0SourcecodeScholar
2025

FairGen: Enhancing Fairness in Text-to-Image Diffusion Models via Self-Discovering Latent Directions

ICCV 2025poster

While Diffusion Models (DM) exhibit remarkable performance across various image generative tasks, they nonetheless reflect the inherent bias presented in the training set.As DMs are now widely used in real-world applications, these biases could perpetuate a distorted worldview and hinder opportuniti…

2025

UniSTD: Towards Unified Spatio-Temporal Learning across Diverse Disciplines

CVPR 2025poster

Traditional spatiotemporal models generally rely on task-specific architectures, which limit their generalizability and scalability across diverse tasks due to domain-specific design requirements. In this paper, we introduce UniSTD, a unified Transformer-based framework for spatiotemporal modeling,…

2024

$\textit{Bifr\"ost}$: 3D-Aware Image Compositing with Language Instructions

NeurIPS 2024poster

This paper introduces $\textit{Bifröst}$, a novel 3D-aware framework that is built upon diffusion models to perform instruction-based image composition. Previous methods concentrate on image compositing at the 2D level, which fall short in handling complex spatial relationships ($\textit{e.g.}$, occ…

Cited by 0SourcePDFScholar
2024

Multi-task Learning with 3D-Aware Regularization

ICLR 2024poster

Deep neural networks have become the standard solution for designing models that can perform multiple dense computer vision tasks such as depth estimation and semantic segmentation thanks to their ability to capture complex correlations in high dimensional feature space across tasks. However, the cr…

2020

Learning to Detect Important People in Unlabelled Images for Semi-Supervised Important People Detection

CVPR 2020poster

Important people detection is to automatically detect the individuals who play the most important roles in a social event image, which requires the designed model to understand a high-level pattern. However, existing methods rely heavily on supervised learning using large quantities of annotated ima…

Cited by 21PDFScholar
2020

MINI-Net: Multiple Instance Ranking Network for Video Highlight Detection

ECCV 2020poster

We address the weakly supervised video highlight detection problem for learning to detect segments that are more attractive in training videos given their video event label but without expensive supervision of manually annotating highlight segments. While manually averting localizing highlight segme…

Cited by 87SourcePDFScholar