← Search

Nicu Sebe

140 accepted papers

2026

Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding

CVPR 2026

While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-explored. We address this gap by introducing Chorus, a multi-teacher pretraining framework that learns a holistic feed-forward 3D Gaussian Splatting (3DG

Cited by 0SourcecodeScholar
2026

Cross Domain Test Time Scaling: Scale Knowledge and Reasoning on Cross Domains

IJCAI 2026

Test-time scaling (TTS) has demonstrated remarkable potential in enhancing the reasoning capabilities of Large Language Models (LLMs) and Large Vision-Language Models (LVLMs). However, its application has primarily been limited to domains such as mathematics and programming, owing to their reasoning

Cited by 0Scholar
2026

Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar Diagnosis

CVPR 2026

Statistically consistent methods based on the noise transition matrix (T) offer a theoretically grounded solution to Learning with Noisy Labels (LNL), with guarantees of convergence to the optimal clean-data classifier. In practice, however, these methods are often outperformed by empirical approach

Cited by 0SourceScholar
2026

Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization

ICLR 2026poster

Degradation-agnostic image restoration aims to handle diverse corruptions with one unified model, but faces fundamental challenges in balancing efficiency and performance across different degradation types. Existing approaches either sacrifice efficiency for versatility or fail to capture the distin…

Cited by 0SourcecodeScholar
2026

Fast and Stable Riemannian Metrics on SPD Manifolds via Cholesky Product Geometry

ICLR 2026poster

Recent advances in Symmetric Positive Definite (SPD) matrix learning show that Riemannian metrics are fundamental to effective SPD neural networks. Motivated by this, we revisit the geometry of the Cholesky factors and uncover a simple product structure that enables convenient metric design. Buildin…

Cited by 0SourcecodeScholar
2026

FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy

ICML 2026poster

Visuomotor policies aim to learn complex manipulation tasks from expert demonstrations. However, generating smooth and coherent trajectories remains challenging, as it requires balancing proximal precision with distal foresight. Existing approaches typically focus on optimizing intra-chunk action di…

Cited by 0SourceScholar
2026

Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation

CVPR 2026

Knowledge distillation (KD) has been widely applied in semantic segmentation to compress large models, but conventional approaches primarily preserve in-domain accuracy while neglecting out-of-domain generalization, which is essential under distribution shifts. This limitation becomes more severe wi

Cited by 0SourcecodeScholar
2026

Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals

ICLR 2026poster

While virtual try-on (VTON) systems aim to render a garment onto a target person, this paper tackles the novel task of virtual try-off (VTOFF), which addresses the inverse problem: generating standardized product images from real-world photos of clothed individuals. Unlike VTON, which must resolve d…

Cited by 0SourcecodeScholar
2026

Masked Clustering Prediction for Unsupervised Point Cloud Pre-training

AAAI 2026technical

Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challenge of learning dense and informative semantic features from point clouds via standard ViTs remains underexplored. We pr

Cited by 0SourcePDFScholar
2026

Open-Vocabulary Domain Generalization in Urban-Scene Segmentation

CVPR 2026

Domain Generalization in Semantic Segmentation (DG-SS) aims to enable segmentation models to perform robustly in unseen environments. However, conventional DG-SS methods are restricted to a fixed set of known categories, limiting their applicability in open-world scenarios. Recent progress in Vision

Cited by 0SourcecodeScholar
2026

Open-World Deepfake Attribution via Confidence-Aware Asymmetric Learning

AAAI 2026technical

The proliferation of synthetic facial imagery has intensified the need for robust Open-World DeepFake Attribution (OW-DFA), which aims to attribute both known and unknown forgeries using labeled data for known types and unlabeled data containing a mixture of known and novel types. However, existing

Cited by 0SourcePDFScholar
2026

PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems

CVPR 2026

Poisoning input views of 3D reconstruction systems has been recently studied. However, existing studies simply backpropagate adversarial gradients through the 3D reconstruction pipeline as a whole, without uncovering the new vulnerability rooted in specific modules of the pipeline. In this paper, we

Cited by 0SourcecodeScholar
2026

Probing Newtonian Mechanics in Video Generative Models with Real Physical Systems

ICML 2026poster

Recent advances in image and video generation raise hopes that these models possess world modeling capabilities—the ability to generate realistic, physically plausible videos. This could revolutionize applications in robotics, autonomous driving, and scientific simulation. However, before treating t…

Cited by 0SourceScholar
2026

Rethinking Expressivity and Degradation-Awareness in Attention for All-in-One Blind Image Restoration

ICLR 2026poster

All-in-one image restoration (IR) aims to recover high-quality images from diverse degradations, which in real-world settings are often mixed and unknown. Unlike single-task IR, this problem requires a model to approximate a family of heterogeneous inverse functions, making it fundamentally more cha…

Cited by 0SourceScholar
2026

Riemannian Graph Convolutional Network for Skeleton-Based Two-Person Interaction Recognition

IJCAI 2026

In the field of skeleton-based human action recognition, Graph Convolutional Networks (GCNs) have become a dominant framework. However, existing GCN-based approaches often treat the sequences of two-person interaction as separate entities, ignoring the inherent semantic dependencies and spatial corr

Cited by 0Scholar
2026

Riemannian High-Order Pooling for Brain Foundation Models

ICLR 2026poster

Electroencephalography (EEG) is a noninvasive technique for measuring brain electrical activity that supports a wide range of brain-computer interaction applications. Motivated by the breakthroughs of Large Language Models (LLMs), recent efforts have begun to explore Large EEG foundation Models trai…

Cited by 0SourcecodeScholar
2026

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

CVPR 2026

Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level visual representations. To address this problem, we introduce TerraScope, a unified VLM that delivers pixel-grounded geospa

Cited by 0SourceScholar
2026

The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category Discovery

CVPR 2026

Generalized Category Discovery (GCD) leverages labeled data to categorize unlabeled samples from known or unknown classes. Most previous methods jointly optimize supervised and unsupervised objectives and achieve promising results. However, inherent optimization interference still limits their abili

Cited by 0SourcecodeScholar
2026

Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models

CVPR 2026

Video Large Language Models (VLLMs) demonstrate strong video understanding but suffer from inefficiency due to redundant visual tokens. Existing pruning primary targets intra-frame spatial redundancy or prunes inside the LLM with shallow-layer overhead, yielding suboptimal spatiotemporal reduction a

Cited by 0SourcecodeScholar
2026

Wasserstein-Aligned Hyperbolic Multi-View Clustering

AAAI 2026technical

Multi-view clustering (MVC) aims to uncover the latent structure of multi-view data by learning view-common and view-specific information. Although recent studies have explored hyperbolic representations for better tackling the representation gap between different views, they focus primarily on inst

Cited by 0SourcePDFScholar
2025

A Correlation Manifold Self-Attention Network for EEG Decoding

IJCAI 2025

Riemannian neural networks, which generalize the deep learning paradigm to non-Euclidean geometries, have garnered widespread attention across diverse applications in artificial intelligence. Among these, the representative attention models have been studied on various non-Euclidean spaces to geomet

2025

CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP

CVPR 2025poster

Despite its prevalent use in image-text matching tasks in a zero-shot manner, CLIP has been shown to be highly vulnerable to adversarial perturbations added onto images. Recent studies propose to finetune the vision encoder of CLIP with adversarial samples generated on the fly, and show improved rob…

2025

Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding

CVPR 2025poster

The lack of a large-scale 3D-text corpus has led recent works to distill open-vocabulary knowledge from vision-language models (VLMs). However, these methods typically rely on a single VLM to align the feature spaces of 3D models within a common language space, which limits the potential of 3D model…

2025

Curriculum Direct Preference Optimization for Diffusion and Consistency Models

CVPR 2025poster

Direct Preference Optimization (DPO) has been proposed as an effective and efficient alternative to reinforcement learning from human feedback (RLHF). In this paper, we propose a novel and enhanced version of DPO based on curriculum learning for text-to-image generation. Our method is divided into t…

2025

FisherTune: Fisher-Guided Robust Tuning of Vision Foundation Models for Domain Generalized Segmentation

CVPR 2025poster

Vision Foundation Models (VFMs) excel in generalization due to large-scale pretraining, but fine-tuning them for Domain Generalized Semantic Segmentation (DGSS) while maintaining this ability remains a challenge. Existing approaches either selectively fine-tune parameters or freeze the VFMs and upda…

Cited by 0SourcePDFScholar
2025

Generate, Refine, and Encode: Leveraging Synthesized Novel Samples for On-the-Fly Fine-Grained Category Discovery

ICCV 2025poster

In this paper, we investigate a practical yet challenging task: On-the-fly Category Discovery (OCD). This task focuses on the online identification of newly arriving stream data that may belong to both known and unknown categories, utilizing the category knowledge from only labeled data. Existing OC…

2025

Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation

ICCV 2025poster

Video instance segmentation (VIS) has gained significant attention for its capability in tracking and segmenting object instances across video frames. However, most of the existing VIS approaches unrealistically assume that the categories of object instances remain fixed over time. Moreover, they ex…

2025

Multi-Stage Multimodal Distillation for Audio-Visual Speaker Tracking

ICASSP 2025accepted

Speaker tracking plays a crucial role in various human-robot interaction applications. Recently, leveraging multimodal information, such as audio and visual signals, has become an important strategy for enhancing the robustness of the tracking system. However, current methods face challenges in effe…

Cited by 0SourceScholar
2025

Multi-focal Conditioned Latent Diffusion for Person Image Synthesis

CVPR 2025poster

The Latent Diffusion Model (LDM) has demonstrated strong capabilities in high-resolution image generation and has been widely employed for Pose-Guided Person Image Synthesis (PGPIS), yielding promising results. However, the compression process of LDM often results in the deterioration of details, pa…

2025

Pseudo-SD: Pseudo Controlled Stable Diffusion for Semi-Supervised and Cross-Domain Semantic Segmentation

ICCV 2025poster

Pseudo-labeling is a key technique of semi-supervised and cross-domian semantic segmentation, yet its efficacy is often hampered by the intrinsic noise of pseudo-labels. This study introduces Pseudo-SD, a novel framework that redefines the utilization of pseudo-label knowledge through Stable Diffusi…

2025

Robust Consensus Anchor Learning for Efficient Multi-view Subspace Clustering

ICML 2025poster

As a leading unsupervised classification algorithm in artificial intelligence, multi-view subspace clustering segments unlabeled data from different subspaces. Recent works based on the anchor have been proposed to decrease the computation complexity for the datasets with large scales in multi-view…

Cited by 0SourcePDFScholar
2025

SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting

NeurIPS 2025poster

3D Gaussian Splatting (3DGS) serves as a highly performant and efficient encoding of scene geometry, appearance, and semantics. Moreover, grounding language in 3D scenes has proven to be an effective strategy for 3D scene understanding. Current Language Gaussian Splatting line of work fall into thre…

Cited by 0SourceScholar
2025

SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining

ICCV 2025poster

Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training, or together at inference. This highlights a clear absence of a model capable of processing 3D data…

2025

Superpowering Open-Vocabulary Object Detectors for X-ray Vision

ICCV 2025poster

Open-vocabulary object detection (OvOD) is set to revolutionize security screening by enabling systems to recognize any item in X-ray scans. However, developing effective OvOD models for X-ray imaging presents unique challenges due to data scarcity and the modality gap that prevents direct adoption…

2025

Towards a General Attention Framework on Gyrovector Spaces for Matrix Manifolds

NeurIPS 2025poster

Deep neural networks operating on non-Euclidean geometries have recently demonstrated impressive performance across various machine-learning applications. Several studies have extended the attention mechanism to different manifolds. However, most existing non-Euclidean attention models are tailored…

Cited by 0SourceScholar
2025

Understanding Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian Geometry

ICLR 2025poster

Global Covariance Pooling (GCP) has been demonstrated to improve the performance of Deep Neural Networks (DNNs) by exploiting second-order statistics of high-level representations. GCP typically performs classification of the covariance matrices by applying matrix function normalization, such as mat…

2025

What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models

ICCV 2025poster

Instruction-based image editing models offer increased personalization opportunities in generative tasks. However, properly evaluating their results is challenging, and most of the existing metrics lag in terms of alignment with human judgment and explainability. To tackle these issues, we introduce…

2025

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

NeurIPS 2025poster

Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visual…

Cited by 0SourceScholar
2024

3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance

ECCV 2024poster

"In this paper, we propose 3DSS-VLG, a weakly supervised approach for 3D Semantic Segmentation with 2D Vision-Language Guidance, an alternative approach that a 3D model predicts dense-embedding for each point which is co-embedded with both the aligned image and text spaces from the 2D vision-languag…

2024

Connectivity-Driven Pseudo-Labeling Makes Stronger Cross-Domain Segmenters

NeurIPS 2024poster

Presently, pseudo-labeling stands as a prevailing approach in cross-domain semantic segmentation, enhancing model efficacy by training with pixels assigned with reliable pseudo-labels. However, we identify two key limitations within this paradigm: (1) under relatively severe domain shifts, most sel…

Cited by 1SourcePDFScholar
2024

Democratizing Fine-grained Visual Recognition with Large Language Models

ICLR 2024poster

Identifying subordinate-level categories from images is a longstanding task in computer vision and is referred to as fine-grained visual recognition (FGVR). It has tremendous significance in real-world applications since an average layperson does not excel at differentiating species of birds or mush…

Cited by 8SourcePDFScholar
2024

Denoising Diffusion Probabilistic Models for Action-Conditioned 3D Motion Generation

ICASSP 2024accepted

Diffusion-based generative models have proven to be highly effective in various domains of synthesis. In this work, we propose a conditional paradigm utilizing the denoising diffusion probabilistic model (DDPM) to address the challenge of realistic and diverse action-conditioned 3D skeleton-based mo…

Cited by 0SourceScholar
2024

Diversity-Authenticity Co-constrained Stylization for Federated Domain Generalization in Person Re-identification

AAAI 2024technical

This paper tackles the problem of federated domain generalization in person re-identification (FedDG re-ID), aiming to learn a model generalizable to unseen domains with decentralized source domains. Previous methods mainly focus on preventing local overfitting. However, the direction of diversifyin…

2024

Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose Estimation

CVPR 2024highlight

Transformers have been successfully applied in the field of video-based 3D human pose estimation. However the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this paper we present a plug-and-play pruning-and-recovering framew…

2024

LESS: Label-Efficient and Single-Stage Referring 3D Segmentation

NeurIPS 2024poster

Referring 3D Segmentation is a visual-language task that segments all points of the specified object from a 3D point cloud described by a sentence of query. Previous works perform a two-stage paradigm, first conducting language-agnostic instance segmentation then matching with given text query. Howe…

2024

Learning to Distinguish Samples for Generalized Category Discovery

ECCV 2024poster

"Generalized Category Discovery (GCD) utilizes labelled data from seen categories to cluster unlabelled samples from both seen and unseen categories. Previous methods have demonstrated that assigning pseudo-labels for representation learning is effective. However, these methods commonly predict pseu…

2024

Mitigating robust overfitting via self-residual-calibration regularization (Abstract Reprint)

IJCAI 2024poster

Overfitting in adversarial training has attracted the interest of researchers in the community of artificial intelligence and machine learning in recent years. To address this issue, in this paper we begin by evaluating the defense performances of several calibration methods on various robust models…

Cited by 0SourcePDFScholar
2024

OpenBias: Open-set Bias Detection in Text-to-Image Generative Models

CVPR 2024highlight

Text-to-image generative models are becoming increasingly popular and accessible to the general public. As these models see large-scale deployments it is necessary to deeply investigate their safety and fairness to not disseminate and perpetuate any kind of biases. However existing works focus on de…

2024

PAIR Diffusion: A Comprehensive Multimodal Object-Level Image Editor

CVPR 2024poster

Generative image editing has recently witnessed extremely fast-paced growth. Some works use high-level conditioning such as text while others use low-level conditioning. Nevertheless most of them lack fine-grained control over the properties of the different objects present in the image i.e. object-…

2024

Prototypical Hash Encoding for On-the-Fly Fine-Grained Category Discovery

NeurIPS 2024poster

In this paper, we study a practical yet challenging task, On-the-fly Category Discovery (OCD), aiming to online discover the newly-coming stream data that belong to both known and unknown classes, by leveraging only known category knowledge contained in labeled data. Previous OCD methods employ the…

2024

RMLR: Extending Multinomial Logistic Regression into General Geometries

NeurIPS 2024poster

Riemannian neural networks, which extend deep learning techniques to Riemannian spaces, have gained significant attention in machine learning. To better classify the manifold-valued features, researchers have started extending Euclidean multinomial logistic regression (MLR) into Riemannian manifolds…

2024

Riemannian Multinomial Logistics Regression for SPD Neural Networks

CVPR 2024poster

Deep neural networks for learning Symmetric Positive Definite (SPD) matrices are gaining increasing attention in machine learning. Despite the significant progress most existing SPD networks use traditional Euclidean classifiers on an approximated space rather than intrinsic classifiers that accurat…

2024

Sharing Key Semantics in Transformer Makes Efficient Image Restoration

NeurIPS 2024poster

Image Restoration (IR), a classic low-level vision task, has witnessed significant advancements through deep models that effectively model global information. Notably, the emergence of Vision Transformers (ViTs) has further propelled these advancements. When computing, the self-attention mechanism,…

2024

Stable Neighbor Denoising for Source-free Domain Adaptive Segmentation

CVPR 2024poster

We study source-free unsupervised domain adaptation (SFUDA) for semantic segmentation which aims to adapt a source-trained model to the target domain without accessing the source data. Many works have been proposed to address this challenging problem among which uncertainty based self-training is a…

2024

Textual Knowledge Matters: Cross-Modality Co-Teaching for Generalized Visual Class Discovery

ECCV 2024poster

"In this paper, we study the problem of Generalized Category Discovery (GCD), which aims to cluster unlabeled data from both known and unknown categories using the knowledge of labeled data from known categories. Current GCD methods rely on only visual cues, which however neglect the multi-modality…

2023

Attribute-Preserving Face Dataset Anonymization via Latent Code Optimization

CVPR 2023highlight

This work addresses the problem of anonymizing the identity of faces in a dataset of images, such that the privacy of those depicted is not violated, while at the same time the dataset is useful for downstream task such as for training machine learning models. To the best of our knowledge, we are th…

2023

Cross-Modality Earth Mover’s Distance for Visible Thermal Person Re-identification

AAAI 2023technical

Visible thermal person re-identification (VT-ReID) suffers from inter-modality discrepancy and intra-identity variations. Distribution alignment is a popular solution for VT-ReID, however, it is usually restricted to the influence of the intra-identity variations. In this paper, we propose the Cross…

Cited by 40SourcePDFScholar
2023

Discover and Align Taxonomic Context Priors for Open-world Semi-Supervised Learning

NeurIPS 2023poster

Open-world Semi-Supervised Learning (OSSL) is a realistic and challenging task, aiming to classify unlabeled samples from both seen and novel classes using partially labeled samples from the seen classes. Previous works typically explore the relationship of samples as priors on the pre-defined sing…

2023

Dynamic Conceptional Contrastive Learning for Generalized Category Discovery

CVPR 2023poster

Generalized category discovery (GCD) is a recently proposed open-world problem, which aims to automatically cluster partially labeled data. The main challenge is that the unlabeled data contain instances that are not only from known categories of the labeled data but also from novel categories. This…

2023

Dynamically Instance-Guided Adaptation: A Backward-Free Approach for Test-Time Domain Adaptive Semantic Segmentation

CVPR 2023poster

In this paper, we study the application of Test-time domain adaptation in semantic segmentation (TTDA-Seg) where both efficiency and effectiveness are crucial. Existing methods either have low efficiency (e.g., backward optimization) or ignore semantic adaptation (e.g., distribution alignment). Besi…

2023

Edge Guided GANs with Contrastive Learning for Semantic Image Synthesis

ICLR 2023poster

We propose a novel \underline{e}dge guided \underline{g}enerative \underline{a}dversarial \underline{n}etwork with \underline{c}ontrastive learning (ECGAN) for the challenging semantic image synthesis task. Although considerable improvement has been achieved, the quality of synthesized images is far…

2023

Graph Transformer GANs for Graph-Constrained House Generation

CVPR 2023poster

We present a novel graph Transformer generative adversarial network (GTGAN) to learn effective graph node relations in an end-to-end fashion for the challenging graph-constrained house generation task. The proposed graph-Transformer-based generator includes a novel graph Transformer encoder that com…

Cited by 31SourcePDFScholar
2023

Latent Traversals in Generative Models as Potential Flows

ICML 2023poster

Despite the significant recent progress in deep generative models, the underlying structure of their latent spaces is still poorly understood, thereby making the task of performing semantically meaningful latent traversals an open research challenge. Most prior work has aimed to solve this challenge…

2023

Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision Transformers

CVPR 2023poster

Position Embeddings (PEs), an arguably indispensable component in Vision Transformers (ViTs), have been shown to improve the performance of ViTs on many vision tasks. However, PEs have a potentially high risk of privacy leakage since the spatial information of the input patches is exposed. This cave…

2023

PI-Trans: Parallel-Convmlp and Implicit-Transformation Based Gan for Cross-View Image Translation

ICASSP 2023accepted

For semantic-guided cross-view image translation, it is crucial to learn where to sample pixels from the source view image and where to reallocate them guided by the target view semantic map, especially when there is little overlap or drastic view difference between the source and target images. Hen…

Cited by 0SourceScholar
2023

StylerDALLE: Language-Guided Style Transfer Using a Vector-Quantized Tokenizer of a Large-Scale Generative Model

ICCV 2023poster

Despite the progress made in the style transfer task, most previous work focus on transferring only relatively simple features like color or texture, while missing more abstract concepts such as overall art expression or painter-specific traits. However, these abstract semantics can be captured by m…

Cited by 14PDFcodeScholar
2022

3D-Aware Semantic-Guided Generative Model for Human Synthesis

ECCV 2022poster

"Generative Neural Radiance Field (GNeRF) models, which extract implicit 3D representations from 2D images, have recently been shown to produce realistic images representing rigid/semi-rigid objects, such as human faces or cars. However, they usually struggle to generate high-quality images represen…

2022

Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation

NeurIPS 2022accept

In this paper, we consider the problem of domain generalization in semantic segmentation, which aims to learn a robust model using only labeled synthetic (source) data. The model is expected to perform well on unseen real (target) domains. Our study finds that the image style variation can largely i…

Cited by 94SourcePDFScholar
2022

CoSMix: Compositional Semantic Mix for Domain Adaptation in 3D LiDAR Segmentation

ECCV 2022poster

"3D LiDAR semantic segmentation is fundamental for autonomous driving. Several Unsupervised Domain Adaptation (UDA) methods for point cloud data have been recently proposed to improve model generalization for different sensors and environments. Researchers working on UDA problems in the image domain…

2022

GIPSO: Geometrically Informed Propagation for Online Adaptation in 3D LiDAR Segmentation

ECCV 2022poster

"3D point cloud semantic segmentation is fundamental for autonomous driving. Most approaches in the literature neglect an important aspect, i.e., how to deal with domain shift when handling dynamic scenes. This can significantly hinder the navigation capabilities of self-driving vehicles. This paper…

2022

Geometry-Contrastive Transformer for Generalized 3D Pose Transfer

AAAI 2022technical

We present a customized 3D mesh Transformer model for the pose transfer task. As the 3D pose transfer essentially is a deformation procedure dependent on the given meshes, the intuition of this work is to perceive the geometric inconsistency between the given meshes with the powerful self-attention…

2022

Hyperbolic Vision Transformers: Combining Improvements in Metric Learning

CVPR 2022poster

Metric learning aims to learn a highly discriminative model encouraging the embeddings of similar classes to be close in the chosen metrics and pushed apart for dissimilar ones. The common recipe is to use an encoder to extract embeddings and a distance-based loss function to match the representatio…

Cited by 135PDFcodeScholar
2022

Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model

CVPR 2022poster

To achieve disentangled image manipulation, previous works depend heavily on manual annotation. Meanwhile, the available manipulations are limited to a pre-defined set the models were trained for. We propose a novel framework, i.e., Predict, Prevent, and Evaluate (PPE), for disentangled text-driven…

Cited by 49PDFcodeScholar
2022

Style-Hallucinated Dual Consistency Learning for Domain Generalized Semantic Segmentation

ECCV 2022poster

"In this paper, we study the task of synthetic-to-real domain generalized semantic segmentation, which aims to learn a model that is robust to unseen real-world scenes using only synthetic data. The large domain shift between synthetic and real-world data, including the limited source environmental…

2022

Uncertainty-Guided Source-Free Domain Adaptation

ECCV 2022poster

"Source-free domain adaptation (SFDA) aims to adapt a classifier to an unlabelled target data set by only using a pre-trained source model. However, the absence of the source data and the domain shift makes the predictions on the target data unreliable. We propose quantifying the uncertainty in the…

2021

Curriculum Graph Co-Teaching for Multi-Target Domain Adaptation

CVPR 2021poster

In this paper we address multi-target domain adaptation (MTDA), where given one labeled source dataset and multiple unlabeled target datasets that differ in data distributions, the task is to learn a robust predictor for all the target domains. We identify two key aspects that can help to alleviate…

Cited by 84PDFcodeScholar
2021

Efficient Training of Visual Transformers with Small Datasets

NeurIPS 2021poster

Visual Transformers (VTs) are emerging as an architectural paradigm alternative to Convolutional networks (CNNs). Differently from CNNs, VTs can capture global relations between image elements and they potentially have a larger representation capacity. However, the lack of the typical convolutional…

2021

Exploiting Sample Correlation for Crowd Counting With Multi-Expert Network

ICCV 2021poster

Crowd counting is a difficult task because of the diversity of scenes. Most of the existing crowd counting methods adopt complex structures with massive backbones to enhance the generalization ability. Unfortunately, the performance of existing methods on large-scale data sets is not satisfactory. I…

Cited by 38PDFScholar
2021

Intrinsic-Extrinsic Preserved GANs for Unsupervised 3D Pose Transfer

ICCV 2021poster

With the strength of deep generative models, 3D pose transfer regains intensive research interests in recent years. Existing methods mainly rely on a variety of constraints to achieve the pose transfer over 3D meshes, e.g., the need for manually encoding for shape and pose disentanglement. In this p…

Cited by 33PDFcodeScholar
2021

Joint Noise-Tolerant Learning and Meta Camera Shift Adaptation for Unsupervised Person Re-Identification

CVPR 2021poster

This paper considers the problem of unsupervised person re-identification (re-ID), which aims to learn discriminative models with unlabeled data. One popular method is to obtain pseudo-label by clustering and use them to optimize the model. Although this kind of approach has shown promising accuracy…

Cited by 153PDFcodeScholar
2021

Learning to Attack Real-World Models for Person Re-identification via Virtual-Guided Meta-Learning

AAAI 2021technical

Recent advances in person re-identification (re-ID) have led to impressive retrieval accuracy. However, existing re-ID models are challenged by the adversarial examples crafted by adding quasi-imperceptible perturbations. Moreover, re-ID systems face the domain shift issue that training and testing…

2021

Learning to Generalize Unseen Domains via Memory-based Multi-Source Meta-Learning for Person Re-Identification

CVPR 2021poster

Recent advances in person re-identification (ReID) obtain impressive accuracy in the supervised and unsupervised learning settings. However, most of the existing methods need to train a new model for a new domain by accessing data. Due to public privacy, the new domain data are not always accessible…

Cited by 266PDFcodeScholar
2021

Neighborhood Contrastive Learning for Novel Class Discovery

CVPR 2021poster

In this paper, we address Novel Class Discovery (NCD), the task of unveiling new classes in a set of unlabeled samples given a labeled dataset with known classes. We exploit the peculiarities of NCD to build a new framework, named Neighborhood Contrastive Learning (NCL), to learn discriminative repr…

Cited by 190PDFcodeScholar
2021

OpenMix: Reviving Known Knowledge for Discovering Novel Visual Categories in an Open World

CVPR 2021poster

In this paper, we tackle the problem of discovering new classes in unlabeled visual data given labeled data from disjoint classes. Existing methods typically first pre-train a model with labeled data, and then identify new classes in unlabeled data via unsupervised clustering. However, the labeled d…

Cited by 152PDFScholar
2021

Smoothing the Disentangled Latent Style Space for Unsupervised Image-to-Image Translation

CVPR 2021poster

Image-to-Image (I2I) multi-domain translation models are usually evaluated also using the quality of their semantic interpolation results. However, state-of-the-art models frequently show abrupt changes in the image appearance during interpolation, and usually perform poorly in interpolations across…

Cited by 58PDFScholar
2021

Transformer-Based Attention Networks for Continuous Pixel-Wise Prediction

ICCV 2021poster

While convolutional neural networks have shown a tremendous impact on various computer vision tasks, they generally demonstrate limitations in explicitly modeling long-range dependencies due to the intrinsic locality of the convolution operation. Initially designed for natural language processing ta…

Cited by 239PDFcodeScholar
2021

Whitening for Self-Supervised Representation Learning

ICML 2021spotlight

Most of the current self-supervised representation learning (SSL) methods are based on the contrastive loss and the instance-discrimination task, where augmented versions of the same image instance ("positives") are contrasted with instances extracted from other images ("negatives"). For the learnin…

2021

Why Approximate Matrix Square Root Outperforms Accurate SVD in Global Covariance Pooling?

ICCV 2021poster

Global Covariance Pooling (GCP) aims at exploiting the second-order statistics of the convolutional feature. Its effectiveness has been demonstrated in boosting the classification performance of Convolutional Neural Networks (CNNs). Singular Value Decomposition (SVD) is used in GCP to compute the ma…

Cited by 28PDFcodeScholar
2020

Local Class-Specific and Global Image-Level Generative Adversarial Networks for Semantic-Guided Scene Generation

CVPR 2020poster

In this paper, we address the task of semantic-guided scene generation. One open challenge widely observed in global image-level generation methods is the difficulty of generating small objects and detailed local texture. To tackle this issue, in this work we consider learning the scene generation i…

Cited by 192PDFcodeScholar
2020

Online Depth Learning Against Forgetting in Monocular Videos

CVPR 2020poster

Online depth learning is the problem of consistently adapting a depth estimation model to handle a continuously changing environment. This problem is challenging due to the network easily overfits on the current environment and forgets its past experiences. To address such problem, this paper presen…

Cited by 49PDFScholar
2020

Reverse Perspective Network for Perspective-Aware Object Counting

CVPR 2020poster

One of the critical challenges of object counting is the dramatic scale variations, which is introduced by arbitrary perspectives. We propose a reverse perspective network to solve the scale variations of input images, instead of generating perspective maps to smooth final outputs. The reverse persp…

Cited by 168PDFScholar
2020

Weakly-Supervised Crowd Counting Learns from Sorting rather than Locations

ECCV 2020poster

In crowd counting datasets, the location labels are costly, yet, they are not taken into the evaluation metrics. Besides, existing multi-task approaches employ high-level tasks to improve counting accuracy. This research tendency increases the demand for more annotations. In this paper, we propose a…

Cited by 108SourcePDFScholar
2019

Animating Arbitrary Objects via Deep Motion Transfer

CVPR 2019oral

This paper introduces a novel deep learning framework for image animation. Given an input image with a target object and a driving video sequence depicting a moving object, our framework generates a video in which the target object is animated according to the driving sequence. This is achieved thro…

Cited by 442PDFcodeScholar
2019

Budget-Aware Adapters for Multi-Domain Learning

ICCV 2019poster

Multi-Domain Learning (MDL) refers to the problem of learning a set of models derived from a common deep architecture, each one specialized to perform a task in a certain domain (e.g., photos, sketches, paintings). This paper tackles MDL with a particular interest in obtaining domain-specific models…

Cited by 45PDFScholar
2019

First Order Motion Model for Image Animation

NeurIPS 2019poster

Image animation consists of generating a video sequence so that an object in a source image is animated according to the motion of a driving video. Our framework addresses this problem without using any annotation or prior information about the specific object to animate. Once trained on a set of vi…

2019

Multi-Channel Attention Selection GAN With Cascaded Semantic Guidance for Cross-View Image Translation

CVPR 2019oral

Cross-view image translation is challenging because it involves images with drastically different views and severe deformation. In this paper, we propose a novel approach named Multi-Channel Attention SelectionGAN (SelectionGAN) that makes it possible to generate images of natural scenes in arbitrar…

Cited by 441PDFcodeScholar
2019

Pattern-Affinitive Propagation Across Depth, Surface Normal and Semantic Segmentation

CVPR 2019poster

In this paper, we propose a novel Pattern-Affinitive Propagation (PAP) framework to jointly predict depth, surface normal and semantic segmentation. The motivation behind it comes from the statistic observation that pattern-affinitive pairs recur much frequently across different tasks as well as wit…

Cited by 392PDFScholar
2019

Refine and Distill: Exploiting Cycle-Inconsistency and Knowledge Distillation for Unsupervised Monocular Depth Estimation

CVPR 2019poster

Nowadays, the majority of state of the art monocular depth estimation techniques are based on supervised deep learning models. However, collecting RGB images with associated depth maps is a very time consuming procedure. Therefore, recent works have proposed deep architectures for addressing the mon…

Cited by 172PDFScholar
2019

Unsupervised Domain Adaptation Using Feature-Whitening and Consensus Loss

CVPR 2019poster

A classifier trained on a dataset seldom works on other datasets obtained under different conditions due to domain shift. This problem is commonly addressed by domain adaptation methods. In this work we introduce a novel deep learning framework which unifies different paradigms in unsupervised domai…

Cited by 207PDFcodeScholar
2018

Deformable GANs for Pose-Based Human Image Generation

CVPR 2018poster

In this paper we address the problem of generating person images conditioned on a given pose. Specifically, given an image of a person and a target pose, we synthesize a new image of that person in the novel pose. In order to deal with pixel-to-pixel misalignments caused by the pose differences, w…

2018

Every Smile Is Unique: Landmark-Guided Diverse Smile Generation

CVPR 2018poster

Each smile is unique: one person surely smiles in different ways (e.g., closing/opening the eyes or mouth). Given one input image of a neutral face, can we generate multiple smile videos with distinctive characteristics? To tackle this one-to-many video generation problem, we propose a novel deep le…

Cited by 82SourcePDFScholar
2018

Group Consistent Similarity Learning via Deep CRF for Person Re-Identification

CVPR 2018poster

Person re-identification benefits greatly from deep neural networks (DNN) to learn accurate similarity metrics and robust feature embeddings. However, most of the current methods impose only local constraints for similarity learning. In this paper, we incorporate constraints on large image groups by…

Cited by 280SourcePDFScholar
2018

PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing

CVPR 2018poster

Depth estimation and scene parsing are two particularly important tasks in visual scene understanding. In this paper we tackle the problem of simultaneous depth estimation and scene parsing in a joint CNN. The task can be typically treated as a deep multi-task learning problem [42]. Different from p…

Cited by 596SourcePDFScholar
2018

Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation

CVPR 2018poster

Recent works have shown the benefit of integrating Conditional Random Fields (CRFs) models into deep architectures for improving pixel-level prediction tasks. Following this line of research, in this paper we introduce a novel approach for monocular depth estimation. Similarly to previous works, our…

2017

A cross-modal adaptation approach for brain decoding

ICASSP 2017accepted

Brain decoding has become a hot topic in many recent brain studies. In a typical neuroimaging experiment, participants are presented with different categories of stimuli while their concurrent brain activity is recorded. Then a classifier is trained on the features extracted from the recorded brain…

Cited by 0SourceScholar
2017

Learning Cross-Modal Deep Representations for Robust Pedestrian Detection

CVPR 2017poster

This paper presents a novel method for detecting pedestrians under adverse illumination conditions. Our approach relies on a novel cross-modality learning framework and it is based on two main phases. First, given a multimodal dataset, a deep convolutional network is employed to learn a non-linear m…

Cited by 257PDFScholar
2017

Learning Deep Structured Multi-Scale Features using Attention-Gated CRFs for Contour Prediction

NeurIPS 2017poster

Recent works have shown that exploiting multi-scale representations deeply learned via convolutional neural networks (CNN) is of tremendous importance for accurate contour detection. This paper presents a novel approach for predicting contours which advances the state of the art in two fundamental a…

2017

Multi-Scale Continuous CRFs as Sequential Deep Networks for Monocular Depth Estimation

CVPR 2017spotlight

This paper addresses the problem of depth estimation from a single still image. Inspired by recent works on multi-scale convolutional neural networks (CNN), we propose a deep model which fuses complementary information derived from multiple CNN side outputs. Different from previous methods, the inte…

Cited by 561PDFcodeScholar
2017

Spatio-Temporal Vector of Locally Max Pooled Features for Action Recognition in Videos

CVPR 2017poster

We introduce Spatio-Temporal Vector of Locally Max Pooled Features (ST-VLMPF), a super vector-based encoding method specifically designed for local deep features encoding. The proposed method addresses an important problem of video understanding: how to build a video representation that incorporate…

Cited by 68PDFScholar
2016

Recognizing Emotions From Abstract Paintings Using Non-Linear Matrix Completion

CVPR 2016poster

Advanced computer vision and machine learning techniques tried to automatically categorize the emotions elicited by abstract paintings with limited success. Since the annotation of the emotional content is highly resource-consuming, datasets of abstract paintings are either constrained in size or pa…

Cited by 120PDFcodeScholar
2016

Self-Adaptive Matrix Completion for Heart Rate Estimation From Face Videos Under Realistic Conditions

CVPR 2016oral

Recent studies in computer vision have shown that, while practically invisible to a human observer, skin color changes due to blood flow can be captured on face videos and, surprisingly, be used to estimate the heart rate (HR). While considerable progress has been made in the last few years, still m…

Cited by 413PDFScholar
2015

Localize Me Anywhere, Anytime: A Multi-Task Point-Retrieval Approach

ICCV 2015poster

Image-based localization is an essential complement to GPS localization. Current image-based localization methods are based on either 2D-to-3D or 3D-to-2D to find the correspondences, which ignore the real scene geometric attributes. The main contribution of our paper is that we use a 3D model recon…

Cited by 39PDFScholar
2015

Optimal Graph Learning With Partial Tags and Multiple Features for Image and Video Annotation

CVPR 2015poster

In multimedia annotation, due to the time constraints and the tediousness of manual tagging, it is quite common to utilize both tagged and untagged data to improve the performance of supervised learning when only limited tagged training data are available. This is often done by adding a geometri…

Cited by 93SourcePDFScholar
2015

The S-Hock Dataset: Analyzing Crowds at the Stadium

CVPR 2015poster

The topic of crowd modeling in computer vision usually assumes a single generic typology of crowd, which is very simplistic. In this paper we adopt a taxonomy that is widely accepted in sociology, focusing on a particular category, the spectator crowd, which is formed by people "interested in watchi…

Cited by 57SourcePDFScholar
2015

Unsupervised Tube Extraction Using Transductive Learning and Dense Trajectories

ICCV 2015poster

We address the problem of automatic extraction of foreground objects from videos. The goal is to provide a method for unsupervised collection of samples which can be further used for object detection training without any human intervention. We use the well known Selective Search approach to produce…

Cited by 40PDFcodeScholar