← Search

Junhui Hou

57 accepted papers

2026

Consistency Geodesic Bridge: Image Restoration with Pretrained Diffusion Models

ICLR 2026poster

Bridge diffusion models have shown great promise in image restoration by constructing a direct path from degraded to clean images. However, they often rely on predefined, high-action trajectories, which limits both sampling efficiency and final restoration quality. To address this, we propose a Cons…

Cited by 0SourcecodeScholar
2026

DiCaP: Distribution-Calibrated Pseudo-labeling for Semi-Supervised Multi-Label Learning

AAAI 2026technical

Semi-supervised multi-label learning (SSMLL) aims to address the challenge of limited labeled data in multi-label learning (MLL) by leveraging unlabeled data to improve the model’s performance. While pseudo-labeling has become a dominant strategy in SSMLL, most existing methods assign equal weights

Cited by 0SourcePDFScholar
2026

ESMC: MLLM-Based Embedding Selection for Explainable Multiple Clustering

AAAI 2026technical

Typical deep clustering methods, while achieving notable progress, can only provide one clustering result per dataset. This limitation arises from their assumption of a fixed underlying data distribution, which may fail to meet user needs and provide unsatisfactory clustering outcomes. Our work inve

Cited by 0SourcePDFScholar
2026

From Extrinsic to Intrinsic: Geodesic-Guided Representation Learning for 3D Geometric Data

ICML 2026poster

Geometric analysis fundamentally distinguishes between extrinsic and intrinsic perspectives. The dominant paradigm in current 3D representation learning relies on either extrinsic spatial structures or high-level semantics, struggling to capture the essence of shape identity and underlying manifold …

Cited by 0SourceScholar
2026

Metric–-Phase Fields: Decoupling Distance and Sign for Thin-Structure Reconstruction from Unoriented Point Clouds

ICML 2026poster

Neural Signed Distance Functions (SDFs) excel at reconstructing watertight manifolds but fail on thin structures and open boundaries due to strict inside-outside constraints. Conversely, Unsigned Distance Fields (UDFs) accommodate general geometries but suffer from gradient singularities at the zero…

Cited by 0SourceScholar
2026

One Coin Has Two Sides: Single Poistive Multi Label Learning from Salient Annotations

ICML 2026poster

Single-Positive Multi-Label Learning (SPML) studies learning from incomplete supervision, where each instance is annotated with only one positive label despite potentially belonging to multiple categories. While existing methods assume the annotated labels are randomly distributed, real-world annota…

Cited by 0SourceScholar
2026

Samples Are Not Equal: A Sample Selection Approach for Deep Clustering

ICLR 2026poster

Deep clustering has recently achieved remarkable progress across various domains. However, existing clustering methods typically treat all samples equally, neglecting the inherent differences in their feature patterns and learning states. Such redundant learning often drives models to overemphasize…

Cited by 0SourcecodeScholar
2026

Scaling Dense Event-Stream Pretraining from Visual Foundation Models

CVPR 2026

Learning versatile, fine-grained representations from irregular event streams is pivotal yet nontrivial, primarily due to the heavy annotation that hinders scalability in dataset size, semantic richness, and application scope. To mitigate this dilemma, we launch a novel self-supervised pretraining m

Cited by 0SourcecodeScholar
2026

TINA: Text-Free Inversion Attack for Unlearned Text-to-Image Diffusion Models

CVPR 2026

Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for their safe deployment to prevent the creation of harmful content.This has fostered a dynamic interplay between the development of erasure defenses and the adversarial probes desi

Cited by 0SourcecodeScholar
2026

Towards Better IncomLDL: We Are Unaware of Hidden Labels in Advance

AAAI 2026technical

Label distribution learning (LDL) is a novel paradigm that describe the samples by label distribution of a sample. However, acquiring LDL dataset is costly and time-consuming, which leads to the birth of incomplete label distribution learning (IncomLDL). All the previous IncomLDL methods set the de

Cited by 0SourcePDFScholar
2025

A Lightweight UDF Learning Framework for 3D Reconstruction Based on Local Shape Functions

CVPR 2025poster

Unsigned distance fields (UDFs) provide a versatile framework for representing a diverse array of 3D shapes, encompassing both watertight and non-watertight geometries. Traditional UDF learning methods typically require extensive training on large 3D shape datasets, which is costly and necessitates…

2025

Acc3D: Accelerating Single Image to 3D Diffusion Models via Edge Consistency Guided Score Distillation

CVPR 2025poster

We present Acc3D to tackle the challenge of accelerating the diffusion process to generate 3D models from single images. To derive high-quality reconstructions through few-step inferences, we emphasize the critical issue of regularizing the learning of score function in states of random noise. To th…

Cited by 0SourcePDFScholar
2025

Generalization Performance of Ensemble Clustering: From Theory to Algorithm

ICML 2025poster

Ensemble clustering has demonstrated great success in practice; however, its theoretical foundations remain underexplored. This paper examines the generalization performance of ensemble clustering, focusing on generalization error, excess risk and consistency. We derive a convergence rate of general…

2025

Keep It on a Leash: Controllable Pseudo-label Generation Towards Realistic Long-Tailed Semi-Supervised Learning

NeurIPS 2025poster

Current long-tailed semi-supervised learning methods assume that labeled data exhibit a long-tailed distribution, and unlabeled data adhere to a typical predefined distribution (i.e., long-tailed, uniform, or inverse long-tailed). However, the distribution of the unlabeled data is generally unknown…

Cited by 0SourcecodeScholar
2025

MoDGS: Dynamic Gaussian Splatting from Casually-captured Monocular Videos with Depth Priors

ICLR 2025poster

In this paper, we propose MoDGS, a new pipeline to render novel-view images in dynamic scenes using only casually captured monocular videos. Previous monocular dynamic NeRF or Gaussian Splatting methods strongly rely on the rapid movement of input cameras to construct multiview consistency but fail…

2025

NVS-Solver: Video Diffusion Model as Zero-Shot Novel View Synthesizer

ICLR 2025poster

By harnessing the potent generative capabilities of pre-trained large video diffusion models, we propose a new novel view synthesis paradigm that operates without the need for training. The proposed method adaptively modulates the diffusion sampling process with the given views to enable the creatio…

2025

ParaSolver: A Hierarchical Parallel Integral Solver for Diffusion Models

ICLR 2025poster

This paper explores the challenge of accelerating the sequential inference process of Diffusion Probabilistic Models (DPMs). We tackle this critical issue from a dynamic systems perspective, in which the inherent sequential nature is transformed into a parallel sampling process. Specifically, we pro…

2025

You Can Trust Your Clustering Model: A Parameter-free Self-Boosting Plug-in for Deep Clustering

NeurIPS 2025poster

Recent deep clustering models have produced impressive clustering performance. However, a common issue with existing methods is the disparity between global and local feature structures. While local structures typically show strong consistency and compactness within class samples, global features o…

Cited by 0SourcecodeScholar
2024

E-Motion: Future Motion Simulation via Event Sequence Diffusion

NeurIPS 2024poster

Forecasting a typical object's future motion is a critical task for interpreting and interacting with dynamic environments in computer vision. Event-based sensors, which could capture changes in the scene with exceptional temporal granularity, may potentially offer a unique opportunity to predict fu…

2024

Fine-grained Image-to-LiDAR Contrastive Distillation with Visual Foundation Models

NeurIPS 2024poster

Contrastive image-to-LiDAR knowledge transfer, commonly used for learning 3D representations with synchronized images and point clouds, often faces a self-conflict dilemma. This issue arises as contrastive losses unintentionally dissociate features of unmatched points and pixels that share semantic…

2024

Flatten Anything: Unsupervised Neural Surface Parameterization

NeurIPS 2024poster

Surface parameterization plays an essential role in numerous computer graphics and geometry processing applications. Traditional parameterization approaches are designed for high-quality meshes laboriously created by specialized 3D modelers, thus unable to meet the processing demand for the current…

2024

PrefPaint: Aligning Image Inpainting Diffusion Model with Human Preference

NeurIPS 2024poster

In this paper, we make the first attempt to align diffusion models for image inpainting with human aesthetic standards via a reinforcement learning framework, significantly improving the quality and visual appeal of inpainted images. Specifically, instead of directly measuring the divergence with pa…

2024

Segment Any Event Streams via Weighted Adaptation of Pivotal Tokens

CVPR 2024poster

In this paper we delve into the nuanced challenge of tailoring the Segment Anything Models (SAMs) for integration with event data with the overarching objective of attaining robust and universal object segmentation within the event-centric domain. One pivotal issue at the heart of this endeavor is t…

2023

Bidirectional Propagation for Cross-Modal 3D Object Detection

ICLR 2023poster

Recent works have revealed the superiority of feature-level fusion for cross-modal 3D object detection, where fine-grained feature propagation from 2D image pixels to 3D LiDAR points has been widely adopted for performance improvement. Still, the potential of heterogeneous feature propagation betwee…

Cited by 2SourcePDFScholar
2023

Cross-Modal Orthogonal High-Rank Augmentation for RGB-Event Transformer-Trackers

ICCV 2023poster

This paper addresses the problem of cross-modal object tracking from RGB videos and event data. Rather than constructing a complex cross-modal fusion network, we explore the great potential of a pre-trained vision Transformer (ViT). Particularly, we delicately investigate plug-and-play training augm…

Cited by 36PDFcodeScholar
2023

Downstream-agnostic Adversarial Examples

ICCV 2023poster

Self-supervised learning usually uses a large amount of unlabeled data to pre-train an encoder which can be used as a general-purpose feature extractor, such that downstream users only need to perform fine-tuning operations to enjoy the benefit of "big model". Despite this promising prospect, the se…

Cited by 30PDFcodeScholar
2023

GeoUDF: Surface Reconstruction from 3D Point Clouds via Geometry-guided Distance Representation

ICCV 2023poster

We present a learning-based method, namely GeoUDF, to tackle the long-standing and challenging problem of reconstructing a discrete surface from a sparse point cloud. To be specific, we propose a geometry-guided learning method for UDF and its gradient estimation that explicitly formulates the unsig…

Cited by 27PDFcodeScholar
2023

Global Structure-Aware Diffusion Process for Low-light Image Enhancement

NeurIPS 2023poster

This paper studies a diffusion-based framework to address the low-light image enhancement problem. To harness the capabilities of diffusion models, we delve into this intricate process and advocate for the regularization of its inherent ODE-trajectory. To be specific, inspired by the recent research…

2023

NeuroGF: A Neural Representation for Fast Geodesic Distance and Path Queries

NeurIPS 2023poster

Geodesics play a critical role in many geometry processing applications. Traditional algorithms for computing geodesics on 3D mesh models are often inefficient and slow, which make them impractical for scenarios requiring extensive querying of arbitrary point-to-point geodesics. Recently, deep impli…

2023

PointCA: Evaluating the Robustness of 3D Point Cloud Completion Models against Adversarial Examples

AAAI 2023technical

Point cloud completion, as the upstream procedure of 3D recognition and segmentation, has become an essential part of many tasks such as navigation and scene understanding. While various point cloud completion models have demonstrated their powerful capabilities, their robustness against adversarial…

Cited by 13SourcePDFScholar
2023

Unleash the Potential of Image Branch for Cross-modal 3D Object Detection

NeurIPS 2023poster

To achieve reliable and precise scene understanding, autonomous vehicles typically incorporate multiple sensing modalities to capitalize on their complementary attributes. However, existing cross-modal 3D detectors do not fully utilize the image domain information to address the bottleneck issues of…

2022

IDEA-Net: Dynamic 3D Point Cloud Interpolation via Deep Embedding Alignment

CVPR 2022poster

This paper investigates the problem of temporally interpolating dynamic 3D point clouds with large non-rigid deformation. We formulate the problem as estimation of point-wise trajectories (i.e., smooth curves) and further reason that temporal irregularity and under-sampling are two major challenges.…

Cited by 23PDFcodeScholar
2022

Learning Graph-embedded Key-event Back-tracing for Object Tracking in Event Clouds

NeurIPS 2022accept

Event data-based object tracking is attracting attention increasingly. Unfortunately, the unusual data structure caused by the unique sensing mechanism poses great challenges in designing downstream algorithms. To tackle such challenges, existing methods usually re-organize raw event data (or event…

2022

WarpingGAN: Warping Multiple Uniform Priors for Adversarial 3D Point Cloud Generation

CVPR 2022poster

We propose WarpingGAN, an effective and efficient 3D point cloud generation network. Unlike existing methods that generate point clouds by directly learning the mapping functions between latent codes and 3D shapes, WarpingGAN learns a unified local-warping function to warp multiple identical pre-def…

Cited by 26PDFcodeScholar
2021

Clustering Ensemble Meets Low-rank Tensor Approximation

AAAI 2021technical

This paper explores the problem of clustering ensemble, which aims to combine multiple base clusterings to produce better performance than that of the individual one. The existing clustering ensemble methods generally construct a co-association matrix, which indicates the pairwise similarity between…

2021

CorrNet3D: Unsupervised End-to-End Learning of Dense Correspondence for 3D Point Clouds

CVPR 2021poster

Motivated by the intuition that one can transform two aligned point clouds to each other more easily and meaningfully than a misaligned pair, we propose CorrNet3D -the first unsupervised and end-to-end deep learning-based framework - to drive the learning of dense correspondence between 3D shapes by…

Cited by 97PDFcodeScholar
2021

Learning Dynamic Interpolation for Extremely Sparse Light Fields With Wide Baselines

ICCV 2021poster

In this paper, we tackle the problem of dense light field (LF) reconstruction from sparsely-sampled ones with wide baselines and propose a learnable model, namely dynamic interpolation, to replace the commonly-used geometry warping operation. Specifically, with the estimated geometric relation betwe…

Cited by 20PDFcodeScholar
2021

Recurrent Multi-View Alignment Network for Unsupervised Surface Registration

CVPR 2021poster

Learning non-rigid registration in an end-to-end manner is challenging due to the inherent high degrees of freedom and the lack of labeled training data. In this paper, we resolve these two challenges simultaneously. First, we propose to represent the non-rigid transformation with a point-wise combi…

Cited by 55PDFcodeScholar
2021

Semantic-Embedded Unsupervised Spectral Reconstruction From Single RGB Images in the Wild

ICCV 2021poster

This paper investigates the problem of reconstructing hyperspectral (HS) images from single RGB images captured by commercial cameras, without using paired HS and RGB images during training. To tackle this challenge, we propose a new lightweight and end-to-end learning-based framework. Specifically,…

Cited by 29PDFcodeScholar
2020

CoADNet: Collaborative Aggregation-and-Distribution Networks for Co-Salient Object Detection

NeurIPS 2020poster

Co-Salient Object Detection (CoSOD) aims at discovering salient objects that repeatedly appear in a given query group containing two or more relevant images. One challenging issue is how to effectively capture co-saliency cues by modeling and exploiting inter-image relationships. In this paper, we p…

2020

Deep Spatial-angular Regularization for Compressive Light Field Reconstruction over Coded Apertures

ECCV 2020poster

Coded aperture is a promising approach for capturing the 4-D light field (LF), in which the 4-D data are compressively modulated into 2-D coded measurements that are further decoded by reconstruction algorithms. The bottleneck lies in the reconstruction algorithms, resulting in rather limited recons…

2020

Light Field Spatial Super-Resolution via Deep Combinatorial Geometry Embedding and Structural Consistency Regularization

CVPR 2020poster

Light field (LF) images acquired by hand-held devices usually suffer from low spatial resolution as the limited sampling resources have to be shared with the angular dimension. LF spatial super-resolution (SR) thus becomes an indispensable part of the LF camera processing pipeline. The high-dimensio…

Cited by 188PDFcodeScholar
2020

PUGeo-Net: A Geometry-centric Network for 3D Point Cloud Upsampling

ECCV 2020poster

In this paper, we propose a novel deep neural network based method, called PUGeo-Net, for upsampling 3D point clouds. PUGeo-Net incorporates discrete differential geometry into deep learning elegantly by learning the first and second fundamental forms that are able to fully represent the local geome…

2020

Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement

CVPR 2020poster

The paper presents a novel method, Zero-Reference Deep Curve Estimation (Zero-DCE), which formulates light enhancement as a task of image-specific curve estimation with a deep network. Our method trains a lightweight deep network, DCE-Net, to estimate pixel-wise and high-order curves for dynamic ran…

Cited by 2093PDFcodeScholar
2018

Fast Light Field Reconstruction With Deep Coarse-To-Fine Modeling of Spatial-Angular Clues

ECCV 2018poster

Densely-sampled light fields (LFs) are beneficial to many applications such as depth inference and post-capture refocusing. However, it is costly and challenging to capture them. In this paper, we propose a learning based algorithm to reconstruct a densely-sampled LF fast and accurately from a spars…

2018

Robust Video Content Alignment and Compensation for Rain Removal in a CNN Framework

CVPR 2018poster

Rain removal is important for improving the robustness of outdoor vision based systems. Current rain removal methods show limitations either for complex dynamic scenes shot from fast moving cameras, or under torrential rain fall with opaque occlusions. We propose a novel derain algorithm, which appl…

Cited by 209SourcePDFScholar
2017

Sparse representation for colors of 3D point cloud via virtual adaptive sampling

ICASSP 2017accepted

Sparse signal representation has proven to be an extremely powerful tool in a wide range of engineering applications. However, most of the existing techniques are designed for regular data (such as audio signals and images/videos) that uniformly lies in regular Euclidian spaces. This paper aims at e…

Cited by 0SourceScholar