← Search

Luc Van Gool

339 accepted papers

2026

AR-VLA: Autoregressive Action Expert for Vision–Language–Action Models

RSS 2026poster

We propose a standalone autoregressive (AR) Action Expert that generates actions as a continuous causal sequence while conditioning on refreshable vision-language prefixes. In contrast to existing Vision-Language-Action (VLA) models and diffusion policies that reset temporal context with each new ob…

Cited by 0SourceScholar
2026

Autonomous Vehicle Path Planning by Searching with Differentiable Simulation

AAAI 2026technical

Planning allows an agent to safely refine its actions before executing them in the real world. In autonomous driving, this is crucial to avoid collisions and navigate in complex, dense traffic scenarios. One way to plan is to search for the best action sequence. However, this is challenging when all

Cited by 0SourcePDFScholar
2026

Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding

CVPR 2026

While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-explored. We address this gap by introducing Chorus, a multi-teacher pretraining framework that learns a holistic feed-forward 3D Gaussian Splatting (3DG

Cited by 0SourcecodeScholar
2026

Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization

ICLR 2026poster

Degradation-agnostic image restoration aims to handle diverse corruptions with one unified model, but faces fundamental challenges in balancing efficiency and performance across different degradation types. Existing approaches either sacrifice efficiency for versatility or fail to capture the distin…

Cited by 0SourcecodeScholar
2026

EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering

AAAI 2026technical

Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities such as cooking and cleaning. In contrast, real-world deplo

Cited by 0SourcePDFScholar
2026

EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark

ICLR 2026poster

Most existing benchmarks for egocentric vision understanding focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first comprehensive benchmark for nighttime egocentric vision, wit…

Cited by 0SourcecodeScholar
2026

FireScope: Wildfire Risk Raster Prediction With a Chain-of-Thought Oracle

CVPR 2026

Predicting wildfire risk is a reasoning-intensive spatial problem that requires the integration of visual, climatic, and geographic factors to infer continuous risk maps. Existing methods lack the causal reasoning and multimodal understanding required for reliable generalization. We introduce FireSc

Cited by 0SourcecodeScholar
2026

GaussianVLM: Scene-Centric 3D Vision-Language Models Using Language-Aligned Gaussian Splats for Embodied Reasoning and Beyond

ICRA 2026poster

As multimodal language models advance, their application to 3D scene understanding is a fast-growing frontier, driving the development of 3D Vision-Language Models (VLMs). Current methods show strong dependence on object detectors, introducing processing bottlenecks and limitations in taxonomic flex…

2026

Inferring Compositional 4D Scenes without Ever Seeing One

CVPR 2026

Scenes in the real world are often composed of several static and dynamic objects. Capturing their 4-dimensional structures, composition and spatio-temporal configuration in-the-wild, though extremely interesting, is equally hard.Therefore, existing works often focus on one object at a time, while r

Cited by 0SourcecodeScholar
2026

LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have achieved significant advances in integrating visual and linguistic information, yet their ability to reason about complex and real-world scenarios remains limited. Existing benchmarks are usually constructed in a task-oriented manner, without a guarantee…

Cited by 0SourceScholar
2026

Position: Vision encoders should be image size agnostic and task driven

ICML 2026poster

This position paper argues that the next generation of vision encoders should be image size agnostic and task driven. The source of our inspiration is biological. Not a structural aspect of biological vision, but a behavioral trait – efficiency. We focus on a couple of ways in which vision in nature…

Cited by 0SourceScholar
2026

ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy Prediction

CVPR 2026

3D semantic occupancy prediction is central to autonomous driving, yet current methods are vulnerable to long-tailed class bias and out-of-distribution (OOD) inputs, often overconfidently assigning anomalies to rare classes. We present ProOOD, a lightweight, plug-and-play method that couples prototy

Cited by 0SourcecodeScholar
2026

Rethinking Expressivity and Degradation-Awareness in Attention for All-in-One Blind Image Restoration

ICLR 2026poster

All-in-one image restoration (IR) aims to recover high-quality images from diverse degradations, which in real-world settings are often mixed and unknown. Unlike single-task IR, this problem requires a model to approximate a family of heterogeneous inverse functions, making it fundamentally more cha…

Cited by 0SourceScholar
2026

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding

CVPR 2026

Robotic Foundation Models (RFMs) hold great promise as generalist, end-to-end systems for robot control.Yet their ability to generalize across new environments, tasks, and embodiments remains limited.We argue that a major bottleneck lies in their foundations: most RFMs are built by fine-tuning inter

Cited by 0SourceScholar
2026

Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysis

CVPR 2026

Prevalent Computational Aberration Correction (CAC) methods are typically tailored to specific optical systems, leading to poor generalization and labor-intensive re-training for new lenses.Developing CAC paradigms capable of generalizing across diverse photographic lenses offers a promising solutio

Cited by 0SourcecodeScholar
2026

Unlocking Efficient Vehicle Dynamics Modeling via Analytic World Models

AAAI 2026technical

Differentiable simulators represent an environment’s dynamics as a differentiable function. Within robotics and autonomous driving, this property is used in Analytic Policy Gradients (APG), which relies on backpropagating through the dynamics to train accurate policies for diverse tasks. Here we sho

Cited by 0SourcePDFScholar
2026

V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence

CVPR 2026

Cross-view object correspondence, exemplified by the representative task of ego-exo object correspondence, aims to establish consistent associations of the same object across different viewpoints (e.g., ego-centric and exo-centric). This task poses significant challenges due to drastic viewpoint and

Cited by 0SourcecodeScholar
2025

Articulate3D: Holistic Understanding of 3D Scenes as Universal Scene Description

ICCV 2025poster

3D scene understanding is a long-standing challenge in computer vision and a key component in enabling mixed reality, wearable computing, and embodied AI. Providing a solution to these applications requires a multifaceted approach that covers scene-centric, object-centric, as well as interaction-cen…

Cited by 0SourcePDFScholar
2025

CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes

RA-L 2025

Leveraging multiple sensors is crucial for robust semantic perception in autonomous driving, as each sensor type has complementary strengths and weaknesses. However, existing sensor fusion methods often treat sensors uniformly across all conditions, leading to suboptimal performance. By contrast, we

Cited by 37SourcecodeScholar
2025

Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection

NeurIPS 2025poster

Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel objects with only a handful of labeled samples from previously unseen domains. While data augmentation and generative methods have shown promise in few-shot learning, their effectiveness for CD-FSOD remains unclear due to the need…

Cited by 0SourcecodeScholar
2025

Exploration-Driven Generative Interactive Environments

CVPR 2025poster

Modern world models require costly and time-consuming collection of large video datasets with action demonstrations by people or by environment-specific agents. To simplify training, we focus on using many virtual environments for inexpensive, automatically collected interaction data. Genie, a recen…

2025

Generalist Robot Manipulation beyond Action Labeled Data

CoRL 2025poster

Recent advances in generalist robot manipulation leverage pre-trained Vision–Language Models (VLMs) and large-scale robot demonstrations to tackle diverse tasks in a zero-shot manner. A key challenge remains: scaling high-quality, action-labeled robot demonstration data, which existing methods rely…

Cited by 0SourceScholar
2025

LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation

NeurIPS 2025poster

We propose LangHOPS, the first Multimodal Large Language Model (MLLM)-based framework for open-vocabulary object–part instance segmentation. Given an image, LangHOPS can jointly detect and segment hierarchical object and part instances from open-vocabulary candidate categories. Unlike prior approach…

Cited by 0SourceScholar
2025

Learning to Prompt with Text Only Supervision for Vision-Language Models

AAAI 2025technical

Foundational vision-language models like CLIP are emerging as a promising paradigm in vision due to their excellent generalization. However, adapting these models for downstream tasks while maintaining their generalization remains challenging. In literature, one branch of methods adapts CLIP by lear…

2025

Leveraging Driver Field-of-View for Multimodal Ego-Trajectory Prediction

ICLR 2025poster

Understanding drivers’ decision-making is crucial for road safety. Although predicting the ego-vehicle’s path is valuable for driver-assistance systems, existing methods mainly focus on external factors like other vehicles’ motions, often neglecting the driver’s attention and intent. To address this…

2025

Locate Anything on Earth: Advancing Open-Vocabulary Object Detection for Remote Sensing Community

AAAI 2025technical

Object detection, particularly open-vocabulary object detection, plays a crucial role in Earth sciences, such as environmental monitoring, natural disaster assessment, and land-use planning. However, existing open-vocabulary detectors, primarily trained on natural-world images, struggle to generaliz…

2025

Low-Light Image Enhancement Using Event-Based Illumination Estimation

ICCV 2025poster

Low-light image enhancement (LLIE) aims to improve the visibility of images captured in poorly lit environments. Prevalent event-based solutions primarily utilize events triggered by motion, i.e., "motion events" to strengthen only the edge texture, while leaving the high dynamic range and excellent…

2025

MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning

ICCV 2025poster

Scaling up model size and training data has advanced foundation models for instance-level perception, achieving state-of-the-art in-domain and zero-shot performance across object detection and segmentation. However, their high computational cost limits adoption on resource-constrained platforms. We…

2025

MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures

CVPR 2025poster

The automated analysis of chemical literature holds promise to accelerate discovery in fields such as material science and drug development. In particular, search capabilities for chemical structures and Markush structures (chemical structure templates) within patent documents are valuable, e.g., fo…

2025

ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives

ICCV 2025poster

Bridging the gap between ego-centric and exo-centric views has been a long-standing question in computer vision. In this paper, we focus on the emerging Ego-Exo object correspondence task, which aims to understand object relations across ego-exo perspectives through segmentation. While numerous segm…

Cited by 0SourcePDFScholar
2025

One2Any: One-Reference 6D Pose Estimation for Any Object

CVPR 2025poster

6D object pose estimation remains challenging for many applications due to dependencies on complete 3D models, multi-view images, or training limited to specific object categories. These requirements make generalization to novel objects difficult for which neither 3D models nor multi-view images may…

2025

PBR-NeRF: Inverse Rendering with Physics-Based Neural Fields

CVPR 2025poster

We tackle the ill-posed inverse rendering problem in 3D reconstruction with a Neural Radiance Field (NeRF) approach informed by Physics-Based Rendering (PBR) theory, named PBR-NeRF. Our method addresses a key limitation in most NeRF and 3D Gaussian Splatting approaches: they estimate view-dependent…

2025

ReVLA: Reverting Visual Domain Limitation of Robotic Foundation Models

ICRA 2025

Recent progress in large language models and access to large-scale robotic datasets has sparked a paradigm shift in robotics models transforming them into generalists able to adapt to various tasks, scenes, and robot modalities. A large step for the community are open Vision Language Action models w

Cited by 22SourceScholar
2025

Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization

ICCV 2025poster

Fusing and balancing multi-modal inputs from novel sensors for dense prediction tasks, particularly semantic segmentation, is critically important yet remains a significant challenge. One major limitation is the tendency of multi-modal frameworks to over-rely on easily learnable modalities, a phenom…

Cited by 0SourcePDFScholar
2025

Samba: Synchronized Set-of-Sequences Modeling for Multiple Object Tracking

ICLR 2025spotlight

Multiple object tracking in complex scenarios - such as coordinated dance performances, team sports, or dynamic animal groups - presents unique challenges. In these settings, objects frequently move in coordinated patterns, occlude each other, and exhibit long-term dependencies in their trajectories…

Cited by 2SourcePDFScholar
2025

SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting

NeurIPS 2025poster

3D Gaussian Splatting (3DGS) serves as a highly performant and efficient encoding of scene geometry, appearance, and semantics. Moreover, grounding language in 3D scenes has proven to be an effective strategy for 3D scene understanding. Current Language Gaussian Splatting line of work fall into thre…

Cited by 0SourceScholar
2025

SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining

ICCV 2025poster

Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training, or together at inference. This highlights a clear absence of a model capable of processing 3D data…

2025

StateSpaceDiffuser: Bringing Long Context to Diffusion World Models

NeurIPS 2025poster

World models have recently gained prominence for action-conditioned visual prediction in complex environments. However, relying on only a few recent observations causes them to lose long-term context. Consequently, within a few steps, the generated scenes drift from what was previously observed, und…

Cited by 0SourcecodeScholar
2025

Understanding Museum Exhibits using Vision-Language Reasoning

ICCV 2025poster

Museums serve as repositories of cultural heritage and historical artifacts from diverse epochs, civilizations, and regions, preserving well-documented collections that encapsulate vast knowledge, which, when systematically structured into large-scale datasets, can train specialized models. Visitors…

Cited by 0SourcePDFScholar
2025

UniK3D: Universal Camera Monocular 3D Estimation

CVPR 2025poster

Monocular 3D estimation is crucial for visual perception. However, current methods fall short by relying on oversimplified assumptions, such as pinhole camera models or rectified images. These limitations severely restrict their general applicability, causing poor performance in real-world scenarios…

2025

What You Have is What You Track: Adaptive and Robust Multimodal Tracking

ICCV 2025poster

Multimodal data is known to be helpful for visual tracking by improving robustness to appearance variations. However, sensor synchronization challenges often compromise data availability, particularly in video settings where shortages can be temporal. Despite its importance, this area remains undere…

2025

XTrack: Multimodal Training Boosts RGB-X Video Object Trackers

ICCV 2025poster

Multimodal sensing has proven valuable for visual tracking, as different sensor types offer unique strengths in handling one specific challenging scene where object appearance varies. While a generalist model capable of leveraging all modalities would be ideal, development is hindered by data sparsi…

2024

"SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking"

ECCV 2024poster

"Open-vocabulary Multiple Object Tracking (MOT) aims to generalize trackers to novel categories not in the training set. Currently, the best-performing methods are mainly based on pure appearance matching. Due to the complexity of motion patterns in the large-vocabulary scenarios and unstable classi…

2024

A Unified and Interpretable Emotion Representation and Expression Generation

CVPR 2024poster

Canonical emotions such as happy sad and fear are easy to understand and annotate. However emotions are often compound e.g. happily surprised and can be mapped to the action units (AUs) used for expressing emotions and trivially to the canonical ones. Intuitively emotions are continuous as represent…

Cited by 5SourcePDFScholar
2024

Continuous Pose for Monocular Cameras in Neural Implicit Representation

CVPR 2024poster

In this paper we showcase the effectiveness of optimizing monocular camera poses as a continuous function of time. The camera poses are represented using an implicit neural function which maps the given time to the corresponding camera pose. The mapped camera poses are then used for the downstream t…

2024

Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector

ECCV 2024poster

"This paper studies the challenging cross-domain few-shot object detection (CD-FSOD), aiming to develop an accurate object detector for novel domains with minimal labeled examples. While transformer-based open-set detectors, such as DE-ViT, show promise in traditional few-shot object detection, thei…

2024

DGInStyle: Domain-Generalizable Semantic Segmentation with Image Diffusion Models and Stylized Semantic Control

ECCV 2024poster

"Large, pretrained latent diffusion models (LDMs) have demonstrated an extraordinary ability to generate creative content, specialize to user data through few-shot fine-tuning, and condition their output on other modalities, such as semantic maps. However, are they usable as large-scale data generat…

2024

Deep Equilibrium Diffusion Restoration with Parallel Sampling

CVPR 2024poster

Diffusion model-based image restoration (IR) aims to use diffusion models to recover high-quality (HQ) images from degraded images achieving promising performance. Due to the inherent property of diffusion models most existing methods need long serial sampling chains to restore HQ images step-by-ste…

2024

Equivariant Multi-Modality Image Fusion

CVPR 2024poster

Multi-modality image fusion is a technique that combines information from different sensors or modalities enabling the fused image to retain complementary features from each modality such as functional highlights and texture details. However effective training of such fusion models is challenging du…

2024

Event-Free Moving Object Segmentation from Moving Ego Vehicle

IROS 2024poster

Moving object segmentation (MOS) in dynamic scenes is an important, challenging, but under-explored research topic for autonomous driving, especially for sequences obtained from moving ego vehicles. Most segmentation methods leverage motion cues obtained from optical flow maps. However, since these…

Cited by 5SourcecodeScholar
2024

Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding

ECCV 2024poster

"3D visual grounding is the task of localizing the object in a 3D scene which is referred by a description in natural language. With a wide range of applications ranging from autonomous indoor robotics to AR/VR, the task has recently risen in popularity. A common formulation to tackle 3D visual grou…

2024

HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point Cloud

CVPR 2024highlight

Extracting keypoint locations from input hand frames known as 3D hand pose estimation is a critical task in various human-computer interaction applications. Essentially the 3D hand pose estimation can be regarded as a 3D point subset generative problem conditioned on input frames. Thanks to the rece…

2024

Image Fusion via Vision-Language Model

ICML 2024poster

Image fusion integrates essential information from multiple images into a single composite, enhancing structures, textures, and refining imperfections. Existing methods predominantly focus on pixel-level and semantic visual features for recognition, but often overlook the deeper text-level semantic…

2024

Implicit Zoo: A Large-Scale Dataset of Neural Implicit Functions for 2D Images and 3D Scenes

NeurIPS 2024poster

Neural implicit functions have demonstrated significant importance in various areas such as computer vision, graphics. Their advantages include the ability to represent complex shapes and scenes with high fidelity, smooth interpolation capabilities, and continuous representations. Despite these bene…

2024

Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning

CVPR 2024poster

Recovering the 3D scene geometry from a single view is a fundamental yet ill-posed problem in computer vision. While classical depth estimation methods infer only a 2.5D scene representation limited to the image plane recent approaches based on radiance fields reconstruct a full 3D representation. H…

2024

Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models

ECCV 2024poster

"Text-to-Image (T2I) models excel at synthesizing concepts such as nouns, appearances, and styles. To enable customized content creation based on a few example images of a concept, methods such as Textual Inversion and DreamBooth invert the desired concept and enable synthesizing it in new scenes. H…

Cited by 11SourcePDFScholar
2024

Lightweight Image Super-Resolution via Flexible Meta Pruning

ICML 2024poster

Lightweight image super-resolution (SR) methods have obtained promising results with moderate model complexity. These approaches primarily focus on a lightweight architecture design, but neglect to further reduce network redundancy. While some model compression techniques try to achieve more lightwe…

Cited by 1SourcePDFScholar
2024

Loopy-SLAM: Dense Neural SLAM with Loop Closures

CVPR 2024poster

Neural RGBD SLAM techniques have shown promise in dense Simultaneous Localization And Mapping (SLAM) yet face challenges such as error accumulation during camera tracking resulting in distorted maps. In response we introduce Loopy-SLAM that globally optimizes poses and the dense 3D model. We use fra…

2024

MICDrop: Masking Image and Depth Features via Complementary Dropout for Domain-Adaptive Semantic Segmentation

ECCV 2024poster

"Unsupervised Domain Adaptation (UDA) is the task of bridging the domain gap between a labeled source domain, e.g., synthetic data, and an unlabeled target domain. We observe that current UDA methods show inferior results on fine structures and tend to oversegment objects with ambiguous appearance.…

2024

MUSES: The Multi-Sensor Semantic Perception Dataset for Driving under Uncertainty

ECCV 2024poster

"Achieving level-5 driving automation in autonomous vehicles necessitates a robust semantic visual perception system capable of parsing data from different sensors across diverse conditions. However, existing semantic perception datasets often lack important non-camera modalities typically used in a…

2024

Matching Anything by Segmenting Anything

CVPR 2024highlight

The robust association of the same objects across video frames in complex scenes is crucial for many applications especially object tracking. Current methods predominantly rely on labeled domain-specific video datasets which limits cross-domain generalization of learned similarity embeddings. We pro…

2024

MoVideo: Motion-Aware Video Generation with Diffusion Models

ECCV 2024poster

"While recent years have witnessed great progress on using diffusion models for video generation, most of them are simple extensions of image generation frameworks, which fail to explicitly consider one of the key differences between videos and images, , motion. In this paper, we propose a novel mot…

Cited by 20SourcePDFScholar
2024

PALM: Predicting Actions through Language Models

ECCV 2024poster

"Understanding human activity is a crucial yet intricate task in egocentric vision, a field that focuses on capturing visual perspectives from the camera wearer’s viewpoint. Traditional methods heavily rely on representation learning that is trained on a large amount of video data. However, a major…

2024

Probabilistic Sampling of Balanced K-Means using Adiabatic Quantum Computing

CVPR 2024poster

Adiabatic quantum computing (AQC) is a promising approach for discrete and often NP-hard optimization problems. Current AQCs allow to implement problems of research interest which has sparked the development of quantum representations for many computer vision tasks. Despite requiring multiple measur…

Cited by 1SourcePDFScholar
2024

Real-World Mobile Image Denoising Dataset with Efficient Baselines

CVPR 2024poster

The recently increased role of mobile photography has raised the standards of on-device photo processing tremendously. Despite the latest advancements in camera hardware the mobile camera sensor area cannot be increased significantly due to physical constraints leading to a pixel size of 0.6--2.0 \m…

2024

Rethinking Few-shot 3D Point Cloud Semantic Segmentation

CVPR 2024poster

This paper revisits few-shot 3D point cloud semantic segmentation (FS-PCS) with a focus on two significant issues in the state-of-the-art: foreground leakage and sparse point distribution. The former arises from non-uniform point sampling allowing models to distinguish the density disparities betwee…

2024

SILC: Improving Vision Language Pretraining with Self-Distillation

ECCV 2024poster

"Image-Text pretraining on web-scale image caption datasets has become the default recipe for open vocabulary classification and retrieval models thanks to the success of CLIP and its variants. Several works have also used CLIP features for dense prediction tasks and have shown the emergence of open…

Cited by 27SourcePDFScholar
2024

Self-supervised Shape Completion via Involution and Implicit Correspondences

ECCV 2024poster

"3D shape completion is traditionally solved using supervised training or by distribution learning on complete shape examples. Recently self-supervised learning approaches that do not require any complete 3D shape examples have gained more interests. In this paper, we propose a non-adversarial self-…

2024

SemiVL: Semi-Supervised Semantic Segmentation with Vision-Language Guidance

ECCV 2024poster

"In semi-supervised semantic segmentation, a model is trained with a limited number of labeled images along with a large corpus of unlabeled images to reduce the high annotation effort. While previous methods are able to learn good segmentation boundaries, they are prone to confuse classes with simi…

2024

Sharing Key Semantics in Transformer Makes Efficient Image Restoration

NeurIPS 2024poster

Image Restoration (IR), a classic low-level vision task, has witnessed significant advancements through deep models that effectively model global information. Notably, the emergence of Vision Transformers (ViTs) has further propelled these advancements. When computing, the self-attention mechanism,…

2024

Single-Model and Any-Modality for Video Object Tracking

CVPR 2024poster

In the realm of video object tracking auxiliary modalities such as depth thermal or event data have emerged as valuable assets to complement the RGB trackers. In practice most existing RGB trackers learn a single set of parameters to use them across datasets and applications. However a similar singl…

2024

Stereo Risk: A Continuous Modeling Approach to Stereo Matching

ICML 2024oral

We introduce Stereo Risk, a new deep-learning approach to solve the classical stereo-matching problem in computer vision. As it is well-known that stereo matching boils down to a per-pixel disparity estimation problem, the popular state-of-the-art stereo-matching approaches widely rely on regressing…

Cited by 8SourcePDFScholar
2024

Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation

CVPR 2024poster

We study object interaction anticipation in egocentric videos. This task requires an understanding of the spatio-temporal context formed by past actions on objects coined "action context". We propose TransFusion a multimodal transformer-based architecture for short-term object interaction anticipati…

Cited by 5SourcePDFScholar
2024

Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits

ECCV 2024poster

"CLIP is a powerful and widely used tool for understanding images in the context of natural language descriptions to perform nuanced tasks. However, it does not offer application-specific fine-grained and structured understanding, due to its generic nature. In this work, we aim to adapt CLIP for fin…

2024

Ternary-Type Opacity and Hybrid Odometry for RGB NeRF-SLAM

IROS 2024poster

In this work, we address the challenge of deploying Neural Radiance Field (NeRFs) in Simultaneous Localization and Mapping (SLAM) under the condition of lacking depth information, relying solely on RGB inputs. The key to unlocking the full potential of NeRF in such a challenging context lies in the…

Cited by 0SourceScholar
2024

UniDepth: Universal Monocular Metric Depth Estimation

CVPR 2024highlight

Accurate monocular metric depth estimation (MMDE) is crucial to solving downstream tasks in 3D perception and modeling. However the remarkable accuracy of recent MMDE methods is confined to their training domains. These methods fail to generalize to unseen domains even in the presence of moderate do…

2024

Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes

CVPR 2024highlight

The estimation of implicit cross-frame correspondences and the high computational cost have long been major challenges in video semantic segmentation (VSS) for driving scenes. Prior works utilize keyframes feature propagation or cross-frame attention to address these issues. By contrast we are the f…

2024

Walker: Self-supervised Multiple Object Tracking by Walking on Temporal Object Appearance Graphs

ECCV 2024poster

"The supervision of state-of-the-art multiple object tracking (MOT) methods requires enormous annotation efforts to provide bounding boxes for all frames of all videos, and instance IDs to associate them through time. To this end, we introduce Walker, the first self-supervised tracker that learns fr…

2023

A Multiplicative Value Function for Safe and Efficient Reinforcement Learning

IROS 2023poster

An emerging field of sequential decision problems is safe Reinforcement Learning (RL), where the objective is to maximize the reward while obeying safety constraints. Being able to handle constraints is essential for deploying RL agents in real-world environments, where constraint violations can har…

Cited by 1SourcecodeScholar
2023

Autodecoding Latent 3D Diffusion Models

NeurIPS 2023poster

Diffusion-based methods have shown impressive visual results in the text-to-image domain. They first learn a latent space using an autoencoder, then run a denoising process on the bottleneck to generate new samples. However, learning an autoencoder requires substantial data in the target domain. Suc…

2023

Basic Binary Convolution Unit for Binarized Image Restoration Network

ICLR 2023poster

Lighter and faster image restoration (IR) models are crucial for the deployment on resource-limited devices. Binary neural network (BNN), one of the most promising model compression methods, can dramatically reduce the computations and parameters of full-precision convolutional neural networks (CNN)…

2023

CDDFuse: Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image Fusion

CVPR 2023poster

Multi-modality (MM) image fusion aims to render fused images that maintain the merits of different modalities, e.g., functional highlight and detailed textures. To tackle the challenge in modeling cross-modality features and decomposing desirable modality-specific and modality-shared features, we pr…

2023

CiaoSR: Continuous Implicit Attention-in-Attention Network for Arbitrary-Scale Image Super-Resolution

CVPR 2023poster

Learning continuous image representations is recently gaining popularity for image super-resolution (SR) because of its ability to reconstruct high-resolution images with arbitrary scales from low-resolution inputs. Existing methods mostly ensemble nearby features to predict the new pixel at any que…

2023

Continuous Pseudo-Label Rectified Domain Adaptive Semantic Segmentation With Implicit Neural Representations

CVPR 2023poster

Unsupervised domain adaptation (UDA) for semantic segmentation aims at improving the model performance on the unlabeled target domain by leveraging a labeled source domain. Existing approaches have achieved impressive progress by utilizing pseudo-labels on the unlabeled target-domain images. Yet the…

2023

Contrastive Model Adaptation for Cross-Condition Robustness in Semantic Segmentation

ICCV 2023poster

Standard unsupervised domain adaptation methods adapt models from a source to a target domain using labeled source data and unlabeled target data jointly. In model adaptation, on the other hand, access to the labeled source data is prohibited, i.e., only the source-trained model and unlabeled target…

Cited by 16PDFcodeScholar
2023

DDFM: Denoising Diffusion Model for Multi-Modality Image Fusion

ICCV 2023oral

Multi-modality image fusion aims to combine different modalities to produce fused images that retain the complementary features of each modality, such as functional highlights and texture details. To leverage strong generative priors and address challenges such as unstable training and lack of inter…

Cited by 210PDFcodeScholar
2023

DREAMWALKER: Mental Planning for Continuous Vision-Language Navigation

ICCV 2023poster

VLN-CE is a recently released embodied task, where AI agents need to navigate a freely traversable environment to reach a distant target location, given language instructions. It poses great challenges due to the huge space of possible strategies. Driven by the belief that the ability to anticipate…

Cited by 38PDFcodeScholar
2023

DiffDreamer: Towards Consistent Unsupervised Single-view Scene Extrapolation with Conditional Diffusion Models

ICCV 2023poster

Scene extrapolation---the idea of generating novel views by flying into a given image---is a promising, yet challenging task. For each predicted frame, a joint inpainting and 3D refinement problem has to be solved, which is ill posed and includes a high level of ambiguity. Moreover, training data fo…

Cited by 38PDFcodeScholar
2023

DiffIR: Efficient Diffusion Model for Image Restoration

ICCV 2023poster

Diffusion model (DM) has achieved SOTA performance by modeling the image synthesis process into a sequential application of a denoising network. However, different from image synthesis generating each pixel from scratch, most pixels of image restoration (IR) are given. Thus, for IR, traditional DMs…

Cited by 291PDFcodeScholar
2023

EDAPS: Enhanced Domain-Adaptive Panoptic Segmentation

ICCV 2023poster

With autonomous industries on the rise, domain adaptation of the visual perception stack is an important research direction due to the cost savings promise. Much prior art was dedicated to domain-adaptive semantic segmentation in the synthetic-to-real context. Despite being a crucial output of the p…

Cited by 14PDFcodeScholar
2023

Edge Guided GANs with Contrastive Learning for Semantic Image Synthesis

ICLR 2023poster

We propose a novel \underline{e}dge guided \underline{g}enerative \underline{a}dversarial \underline{n}etwork with \underline{c}ontrastive learning (ECGAN) for the challenging semantic image synthesis task. Although considerable improvement has been achieved, the quality of synthesized images is far…

2023

Efficient and Explicit Modelling of Image Hierarchies for Image Restoration

CVPR 2023poster

The aim of this paper is to propose a mechanism to efficiently and explicitly model image hierarchies in the global, regional, and local range for image restoration. To achieve that, we start by analyzing two important properties of natural images including cross-scale similarity and anisotropic ima…

2023

Event-Based Frame Interpolation With Ad-Hoc Deblurring

CVPR 2023poster

The performance of video frame interpolation is inherently correlated with the ability to handle motion in the input scene. Even though previous works recognize the utility of asynchronous event information for this task, they ignore the fact that motion may or may not result in blur in the input vi…

2023

Graph Transformer GANs for Graph-Constrained House Generation

CVPR 2023poster

We present a novel graph Transformer generative adversarial network (GTGAN) to learn effective graph node relations in an end-to-end fashion for the challenging graph-constrained house generation task. The proposed graph-Transformer-based generator includes a novel graph Transformer encoder that com…

Cited by 31SourcePDFScholar
2023

I2MVFormer: Large Language Model Generated Multi-View Document Supervision for Zero-Shot Image Classification

CVPR 2023highlight

Recent works have shown that unstructured text (documents) from online sources can serve as useful auxiliary information for zero-shot image classification. However, these methods require access to a high-quality source like Wikipedia and are limited to a single source of information. Large Language…

2023

Improving Online Lane Graph Extraction by Object-Lane Clustering

ICCV 2023poster

Autonomous driving requires accurate local scene understanding information. To this end, autonomous agents deploy object detection and online BEV lane graph extraction methods as a part of their perception stack. In this work, we propose an architecture and loss formulation to improve the accuracy o…

Cited by 9PDFScholar
2023

Indiscernible Object Counting in Underwater Scenes

CVPR 2023poster

Recently, indiscernible scene understanding has attracted a lot of attention in the vision community. We further advance the frontier of this field by systematically studying a new challenge named indiscernible object counting (IOC), the goal of which is to count objects that are blended with respec…

2023

Introducing Language Guidance in Prompt-based Continual Learning

ICCV 2023poster

Continual Learning aims to learn a single model on a sequence of tasks without having access to data from previous tasks. The biggest challenge in the domain still remains catastrophic forgetting: a loss in performance on seen classes of earlier tasks. Some existing methods rely on an expensive repl…

Cited by 56PDFcodeScholar
2023

Knowledge Distillation based Degradation Estimation for Blind Super-Resolution

ICLR 2023poster

Blind image super-resolution (Blind-SR) aims to recover a high-resolution (HR) image from its corresponding low-resolution (LR) input image with unknown degradations. Most of the existing works design an explicit degradation estimator for each degradation to guide SR. However, it is infeasible to pr…

2023

L2E: Lasers to Events for 6-DoF Extrinsic Calibration of Lidars and Event Cameras

ICRA 2023poster

As neuromorphic technology is maturing, its application to robotics and autonomous vehicle systems has become an area of active research. In particular, event cameras have emerged as a compelling alternative to frame-based cameras in low-power and latency-demanding applications. To enable event came…

Cited by 13SourcecodeScholar
2023

LART: Neural Correspondence Learning with Latent Regularization Transformer for 3D Motion Transfer

NeurIPS 2023poster

3D motion transfer aims at transferring the motion from a dynamic input sequence to a static 3D object and outputs an identical motion of the target with high-fidelity and realistic visual effects. In this work, we propose a novel 3D Transformer framework called LART for 3D motion transfer. With car…

2023

MIC: Masked Image Consistency for Context-Enhanced Domain Adaptation

CVPR 2023poster

In unsupervised domain adaptation (UDA), a model trained on source data (e.g. synthetic) is adapted to target data (e.g. real-world) without access to target annotation. Most previous UDA methods struggle with classes that have a similar visual appearance on the target domain as no ground truth is a…

2023

Real-Time Motion Prediction via Heterogeneous Polyline Transformer with Relative Pose Encoding

NeurIPS 2023poster

The real-world deployment of an autonomous driving system requires its components to run on-board and in real-time, including the motion prediction module that predicts the future trajectories of surrounding traffic participants. Existing agent-centric methods have demonstrated outstanding performan…

2023

Revisiting Evaluation Metrics for Semantic Segmentation: Optimization and Evaluation of Fine-grained Intersection over Union

NeurIPS 2023poster

Semantic segmentation datasets often exhibit two types of imbalance: \textit{class imbalance}, where some classes appear more frequently than others and \textit{size imbalance}, where some objects occupy more pixels than others. This causes traditional evaluation metrics to be biased towards \textit…

2023

SMAE: Few-Shot Learning for HDR Deghosting With Saturation-Aware Masked Autoencoders

CVPR 2023poster

Generating a high-quality High Dynamic Range (HDR) image from dynamic scenes has recently been extensively studied by exploiting Deep Neural Networks (DNNs). Most DNNs-based methods require a large amount of training data with ground truth, requiring tedious and time-consuming work. Few-shot HDR ima…

Cited by 19SourcePDFScholar
2023

Single Image Depth Prediction Made Better: A Multivariate Gaussian Take

CVPR 2023poster

Neural-network-based single image depth prediction (SIDP) is a challenging task where the goal is to predict the scene's per-pixel depth at test time. Since the problem, by definition, is ill-posed, the fundamental goal is to come up with an approach that can reliably model the scene depth from a se…

Cited by 26SourcePDFScholar
2023

Source-free Depth for Object Pop-out

ICCV 2023poster

Depth cues are known to be useful for visual perception. However, direct measurement of depth is often impracticable. Fortunately, though, modern learning-based methods offer promising depth maps by inference in the wild. In this work, we adapt such depth inference models for object segmentation usi…

Cited by 70PDFcodeScholar
2023

Spherical Space Feature Decomposition for Guided Depth Map Super-Resolution

ICCV 2023poster

Guided depth map super-resolution (GDSR), as a hot topic in multi-modal image processing, aims to upsample low-resolution (LR) depth maps with additional information involved in high-resolution (HR) RGB images from the same scene. The critical step of this task is to effectively extract domain-share…

Cited by 35PDFcodeScholar
2023

Structured Sparsity Learning for Efficient Video Super-Resolution

CVPR 2023poster

The high computational costs of video super-resolution (VSR) models hinder their deployment on resource-limited devices, e.g., smartphones and drones. Existing VSR models contain considerable redundant filters, which drag down the inference efficiency. To prune these unimportant filters, we develop…

2023

Surface Normal Clustering for Implicit Representation of Manhattan Scenes

ICCV 2023poster

Novel view synthesis and 3D modeling using implicit neural field representation are shown to be very effective for calibrated multi-view cameras. Such representations are known to benefit from additional geometric and semantic supervision. Most existing methods that exploit additional supervision re…

Cited by 4PDFcodeScholar
2023

TrafficBots: Towards World Models for Autonomous Driving Simulation and Motion Prediction

ICRA 2023poster

Data-driven simulation has become a favorable way to train and test autonomous driving algorithms. The idea of replacing the actual environment with a learned simulator has also been explored in model-based reinforcement learning in the context of world models. In this work, we show data-driven traf…

Cited by 48SourceScholar
2023

Unbalanced Optimal Transport: A Unified Framework for Object Detection

CVPR 2023poster

During training, supervised object detection tries to correctly match the predicted bounding boxes and associated classification scores to the ground truth. This is essential to determine which predictions are to be pushed towards which solutions, or to be discarded. Popular matching strategies incl…

2023

Unsupervised Deep Probabilistic Approach for Partial Point Cloud Registration

CVPR 2023poster

Deep point cloud registration methods face challenges to partial overlaps and rely on labeled data. To address these issues, we propose UDPReg, an unsupervised deep probabilistic registration framework for point clouds with partial overlaps. Specifically, we first adopt a network to learn posterior…

2023

VA-DepthNet: A Variational Approach to Single Image Depth Prediction

ICLR 2023top-25%

We introduce VA-DepthNet, a simple, effective, and accurate deep neural network approach for the single-image depth prediction (SIDP) problem. The proposed approach advocates using classical first-order variational constraints for this problem. While state-of-the-art deep neural network methods for…

2023

Video Task Decathlon: Unifying Image and Video Tasks in Autonomous Driving

ICCV 2023poster

Performing multiple heterogeneous visual tasks in dynamic scenes is a hallmark of human perception capability. Despite remarkable progress in image and video recognition via representation learning, current research still focuses on designing specialized networks for singular, homogeneous, or simple…

Cited by 7PDFScholar
2022

3D Compositional Zero-Shot Learning with DeCompositional Consensus

ECCV 2022poster

"Parts represent a basic unit of geometric and semantic similarity across different objects. We argue that part knowledge should be composable beyond the observed object classes. Towards this, we present 3D Compositional Zero-shot Learning as a problem of part generalization from seen to unseen obje…

Cited by 15SourcePDFScholar
2022

A Real-Time Online Learning Framework for Joint 3D Reconstruction and Semantic Segmentation of Indoor Scenes

RA-L 2022

This letter presents a real-time online vision framework to jointly recover an indoor scene’s 3D structure and semantic label. Given noisy depth maps, a camera trajectory, and 2D semantic labels at train time, the proposed deep neural network based approach learns to fuse the depth over frames with

Cited by 26SourcecodeScholar
2022

Adiabatic Quantum Computing for Multi Object Tracking

CVPR 2022poster

Multi-Object Tracking (MOT) is most often approached in the tracking-by-detection paradigm, where object detections are associated through time. The association step naturally leads to discrete optimization problems. As these optimization problems are often NP-hard, they can only be solved exactly f…

Cited by 34PDFScholar
2022

Arbitrary-Scale Image Synthesis

CVPR 2022poster

Positional encodings have enabled recent works to train a single adversarial network that can generate images of different scales. However, these approaches are either limited to a set of discrete scales or struggle to maintain good perceptual quality at the scales for which the model is not trained…

Cited by 25PDFcodeScholar
2022

Coarse-To-Fine Feature Mining for Video Semantic Segmentation

CVPR 2022poster

The contextual information plays a core role in semantic segmentation. As for video semantic segmentation, the contexts include static contexts and motional contexts, corresponding to static content and moving content in a video clip, respectively. The static contexts are well exploited in image sem…

Cited by 75PDFcodeScholar
2022

Coarse-to-Fine Sparse Transformer for Hyperspectral Image Reconstruction

ECCV 2022poster

"Many learning-based algorithms have been developed to solve the inverse problem of coded aperture snapshot spectral imaging (CASSI). However, CNN-based methods show limitations in capturing long-range dependencies. Previous Transformer-based methods densely sample tokens, some of which are uninform…

2022

Collapse by Conditioning: Training Class-conditional GANs with Limited Data

ICLR 2022poster

Class-conditioning offers a direct means to control a Generative Adversarial Network (GAN) based on a discrete input variable. While necessary in many applications, the additional information provided by the class labels could even be expected to benefit the training of the GAN itself. On the contra…

2022

Counterfactual Cycle-Consistent Learning for Instruction Following and Generation in Vision-Language Navigation

CVPR 2022poster

Since the rise of vision-language navigation (VLN), great progress has been made in instruction following -- building a follower to navigate environments under the guidance of instructions. However, far less attention has been paid to the inverse task: instruction generation -- learning a speaker to…

Cited by 62PDFcodeScholar
2022

DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic Segmentation

CVPR 2022poster

As acquiring pixel-wise annotations of real-world images for semantic segmentation is a costly process, a model can instead be trained with more accessible synthetic data and adapted to real images without requiring their annotations. This process is studied in unsupervised domain adaptation (UDA).…

Cited by 617PDFcodeScholar
2022

Degradation-Aware Unfolding Half-Shuffle Transformer for Spectral Compressive Imaging

NeurIPS 2022accept

In coded aperture snapshot spectral compressive imaging (CASSI) systems, hyperspectral image (HSI) reconstruction methods are employed to recover the spatial-spectral signal from a compressed measurement. Among these algorithms, deep unfolding methods demonstrate promising performance but suffer fro…

2022

End-to-End Optimization of LiDAR Beam Configuration for 3D Object Detection and Localization

RA-L 2022

Existing learning methods for LiDAR-based applications use 3D points scanned under a pre-determined beam configuration, e.g., the elevation angles of beams are often evenly distributed. Those fixed configurations are task-agnostic, so simply using them can lead to sub-optimal performance. In this wo

Cited by 17SourcecodeScholar
2022

Event-Based Fusion for Motion Deblurring with Cross-Modal Attention

ECCV 2022poster

"Traditional frame-based cameras inevitably suffer from motion blur due to long exposure times. As a kind of bio-inspired camera, the event camera records the intensity changes in an asynchronous way with high temporal resolution, providing valid image degradation information within the exposure tim…

2022

Flow-Guided Sparse Transformer for Video Deblurring

ICML 2022spotlight

Exploiting similar and sharper scene patches in spatio-temporal neighborhoods is critical for video deblurring. However, CNN-based methods show limitations in capturing long-range dependencies and modeling non-local self-similarity. In this paper, we propose a novel framework, Flow-Guided Sparse Tra…

2022

Generative Flows With Invertible Attentions

CVPR 2022poster

Flow-based generative models have shown an excellent ability to explicitly learn the probability density function of data via a sequence of invertible transformations. Yet, learning attentions in generative flows remains understudied, while it has made breakthroughs in other domains. To fill the gap…

Cited by 16PDFcodeScholar
2022

HDNet: High-Resolution Dual-Domain Learning for Spectral Compressive Imaging

CVPR 2022poster

The rapid development of deep learning provides a better solution for the end-to-end reconstruction of hyperspectral image (HSI). However, existing learning-based methods have two major defects. Firstly, networks with self-attention usually sacrifice internal resolution to balance model performance…

Cited by 188PDFcodeScholar
2022

HRDA: Context-Aware High-Resolution Domain-Adaptive Semantic Segmentation

ECCV 2022poster

"Unsupervised domain adaptation (UDA) aims to adapt a model trained on the source domain (e.g. synthetic data) to the target domain (e.g. real-world data) without requiring further annotations on the target domain. This work focuses on UDA for semantic segmentation as real-world pixel-wise annotatio…

2022

Highly Accurate Dichotomous Image Segmentation

ECCV 2022poster

"We present a systematic study on a new task called dichotomous image segmentation (DIS), which aims to segment highly accurate objects from natural images. To this end, we collected the first large-scale DIS dataset, called DIS5K, which contains 5,470 high-resolution (e.g., 2K, 4K or larger) images…

2022

I2DFormer: Learning Image to Document Attention for Zero-Shot Image Classification

NeurIPS 2022accept

Despite the tremendous progress in zero-shot learning (ZSL), the majority of existing methods still rely on human-annotated attributes, which are difficult to annotate and scale. An unsupervised alternative is to represent each class using the word embedding associated with its semantic class name.…

Cited by 55SourcePDFScholar
2022

Implicit Neural Representations for Image Compression

ECCV 2022poster

"Implicit Neural Representations (INRs) gained attention as a novel and effective representation for various data types. Recently, prior work applied INRs to image compressing. Such compression algorithms are promising candidates as a general purpose approach for any coordinate-based data modality.…

2022

Learnable Online Graph Representations for 3D Multi-Object Tracking

RA-L 2022

Autonomous systems that operate in dynamic environments require robust object tracking in 3D as one of their key components. Most recent approaches for 3D multi-object tracking (MOT) from LIDAR use object dynamics together with a set of handcrafted features to match detections of objects across mult

Cited by 78SourceScholar
2022

Learning Online Multi-sensor Depth Fusion

ECCV 2022poster

"Many hand-held or mixed reality devices are used with a single sensor for 3D reconstruction, although they often comprise multiple sensors. Multi-sensor depth fusion is able to substantially improve the robustness and accuracy of 3D reconstruction methods, but existing techniques are not robust eno…

2022

LiDAR Snowfall Simulation for Robust 3D Object Detection

CVPR 2022oral

3D object detection is a central task for applications such as autonomous driving, in which the system needs to localize and classify surrounding traffic agents, even in the presence of adverse weather. In this paper, we address the problem of LiDAR-based 3D object detection under snowfall. Due to t…

Cited by 147PDFcodeScholar
2022

Lidar Line Selection with Spatially-Aware Shapley Value for Cost-Efficient Depth Completion

CoRL 2022poster

Lidar is a vital sensor for estimating the depth of a scene. Typical spinning lidars emit pulses arranged in several horizontal lines and the monetary cost of the sensor increases with the number of these lines. In this work, we present the new problem of optimizing the positioning of lidar lines to…

Cited by 2SourceScholar
2022

MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation

CVPR 2022poster

Estimating 3D human poses from monocular videos is a challenging task due to depth ambiguity and self-occlusion. Most existing works attempt to solve both issues by exploiting spatial and temporal relationships. However, those works ignore the fact that it is an inverse problem where multiple feasib…

Cited by 415PDFcodeScholar
2022

Mask-Guided Spectral-Wise Transformer for Efficient Hyperspectral Image Reconstruction

CVPR 2022poster

Hyperspectral image (HSI) reconstruction aims to recover the 3D spatial-spectral signal from a 2D measurement in the coded aperture snapshot spectral imaging (CASSI) system. The HSI representations are highly similar and correlated across the spectral dimension. Modeling the inter-spectra interactio…

Cited by 333PDFcodeScholar
2022

Mining Relations among Cross-Frame Affinities for Video Semantic Segmentation

ECCV 2022poster

"The essence of video semantic segmentation (VSS) is how to leverage temporal information for prediction. Previous efforts are mainly devoted to developing new techniques to calculate the cross-frame affinities such as optical flow and attention. Instead, this paper contributes from a different angl…

2022

OSFormer: One-Stage Camouflaged Instance Segmentation with Transformers

ECCV 2022poster

"We present OSFormer, the first one-stage transformer framework for camouflaged instance segmentation (CIS). OSFormer is based on two key designs. First, we design a location-sensing transformer (LST) to obtain the location label and instance-aware parameters by introducing the location-guided queri…

2022

On the Practicality of Deterministic Epistemic Uncertainty

ICML 2022spotlight

A set of novel approaches for estimating epistemic uncertainty in deep neural networks with a single forward pass has recently emerged as a valid alternative to Bayesian Neural Networks. On the premise of informative representations, these deterministic uncertainty methods (DUMs) achieve strong perf…

2022

P3Depth: Monocular Depth Estimation With a Piecewise Planarity Prior

CVPR 2022poster

Monocular depth estimation is vital for scene understanding and downstream tasks. We focus on the supervised setup, in which ground-truth depth is available only at training time. Based on knowledge about the high regularity of real 3D scenes, we propose a method that learns to selectively leverage…

Cited by 170PDFcodeScholar
2022

Perceptual Learned Video Compression with Recurrent Conditional GAN

IJCAI 2022poster

This paper proposes a Perceptual Learned Video Compression (PLVC) approach with recurrent conditional GAN. We employ the recurrent auto-encoder-based compression network as the generator, and most importantly, we propose a recurrent conditional discriminator, which judges raw vs. compressed video co…

2022

Pix2NeRF: Unsupervised Conditional p-GAN for Single Image to Neural Radiance Fields Translation

CVPR 2022poster

We propose a pipeline to generate Neural Radiance Fields (NeRF) of an object or a scene of a specific class, conditioned on a single input image. This is a challenging task, as training NeRF requires multiple views of the same scene, coupled with corresponding poses, which are hard to obtain. Our me…

Cited by 105PDFcodeScholar
2022

Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model

CVPR 2022poster

To achieve disentangled image manipulation, previous works depend heavily on manual annotation. Meanwhile, the available manipulations are limited to a pre-defined set the models were trained for. We propose a novel framework, i.e., Predict, Prevent, and Evaluate (PPE), for disentangled text-driven…

Cited by 49PDFcodeScholar
2022

Probabilistic Warp Consistency for Weakly-Supervised Semantic Correspondences

CVPR 2022poster

We propose Probabilistic Warp Consistency, a weakly-supervised learning objective for semantic matching. Our approach directly supervises the dense matching scores predicted by the network, encoded as a conditional probability distribution. We first construct an image triplet by applying a known war…

Cited by 38PDFcodeScholar
2022

RePaint: Inpainting Using Denoising Diffusion Probabilistic Models

CVPR 2022poster

Free-form inpainting is the task of adding new content to an image in the regions specified by an arbitrary binary mask. Most existing approaches train for a certain distribution of masks, which limits their generalization capabilities to unseen mask types. Furthermore, training with pixel-wise and…

Cited by 1817PDFcodeScholar
2022

Recurrent Video Restoration Transformer with Guided Deformable Attention

NeurIPS 2022accept

Video restoration aims at restoring multiple high-quality frames from multiple low-quality frames. Existing video restoration methods generally fall into two extreme cases, i.e., they either restore all frames in parallel or restore the video frame by frame in a recurrent way, which would result in…

2022

Reference-Based Image Super-Resolution with Deformable Attention Transformer

ECCV 2022poster

"Reference-based image super-resolution (RefSR) aims to exploit auxiliary reference (Ref) images to super-resolve low-resolution (LR) images. Recently, RefSR has been attracting great attention as it provides an alternative way to surpass single image SR. However, addressing the RefSR problem has tw…

2022

Revisiting Random Channel Pruning for Neural Network Compression

CVPR 2022poster

Channel (or 3D filter) pruning serves as an effective way to accelerate the inference of neural networks. There has been a flurry of algorithms that try to solve this practical problem, each being claimed effective in some ways. Yet, a benchmark to compare those algorithms directly is lacking, mainl…

Cited by 131PDFcodeScholar
2022

SHIFT: A Synthetic Driving Dataset for Continuous Multi-Task Domain Adaptation

CVPR 2022poster

Adapting to a continuously evolving environment is a safety-critical challenge inevitably faced by all autonomous-driving systems. Existing image- and video-based driving datasets, however, fall short of capturing the mutable nature of the real world. In this paper, we introduce the largest syntheti…

Cited by 166PDFScholar
2022

TACS: Taxonomy Adaptive Cross-Domain Semantic Segmentation

ECCV 2022poster

"Traditional domain adaptive semantic segmentation addresses the task of adapting a model to a novel target domain under limited or no additional supervision. While tackling the input domain gap, the standard domain adaptation settings assume no domain change in the output space. In semantic predict…

2022

Topology Preserving Local Road Network Estimation From Single Onboard Camera Image

CVPR 2022poster

Knowledge of the road network topology is crucial for autonomous planning and navigation. Yet, recovering such topology from a single image has only been explored in part. Furthermore, it needs to refer to the ground plane, where also the driving actions are taken. This paper aims at extracting the…

Cited by 53PDFcodeScholar
2022

Towards Interpretable Video Super-Resolution via Alternating Optimization

ECCV 2022poster

"In this paper, we study a practical space-time video super-resolution (STVSR) problem which aims at generating a high-framerate high-resolution sharp video from a low-framerate low-resolution blurry video. Such problem often occurs when recording a fast dynamic event with a low-framerate and low-re…

2022

Transform Your Smartphone into a DSLR Camera: Learning the ISP in the Wild

ECCV 2022poster

"We propose a trainable Image Signal Processing (ISP) framework that produces DSLR quality images given RAW images captured by a smartphone. To address the color misalignments between training image pairs, we employ a color-conditional ISP network and optimize a novel parametric color mapping betwee…

2022

Transforming Model Prediction for Tracking

CVPR 2022poster

Optimization based tracking methods have been widely successful by integrating a target model prediction module, providing effective global reasoning by minimizing an objective function. While this inductive bias integrates valuable domain knowledge, it limits the expressivity of the tracking networ…

Cited by 380PDFcodeScholar
2022

Uncertainty-Aware Deep Multi-View Photometric Stereo

CVPR 2022poster

This paper presents a simple and effective solution to the longstanding classical multi-view photometric stereo (MVPS) problem. It is well-known that photometric stereo (PS) is excellent at recovering high-frequency surface details, whereas multi-view stereo (MVS) can help remove the low-frequency d…

Cited by 45PDFScholar
2022

Understanding Bird's-Eye View of Road Semantics Using an Onboard Camera

RA-L 2022

Autonomous navigation requires scene understanding of the action-space to move or anticipate events. For planner agents moving on the ground plane, such as autonomous vehicles, this translates to scene understanding in the bird’s-eye view (BEV). However, the onboard cameras of autonomous cars are cu

Cited by 52SourcecodeScholar
2022

Unsupervised Flow-Aligned Sequence-to-Sequence Learning for Video Restoration

ICML 2022spotlight

How to properly model the inter-frame relation within the video sequence is an important but unsolved challenge for video restoration (VR). In this work, we propose an unsupervised flow-aligned sequence-to-sequence model (S2SVR) to address this problem. On the one hand, the sequence-to-sequence mode…

2022

Zero Pixel Directional Boundary by Vector Transform

ICLR 2022poster

Boundaries or contours are among the primary visual cues used by human and computer vision systems. One of the key problems in boundary detection is the loss formulation, which typically leads to class imbalance and, as a consequence, to thick boundaries which require non-differential post-processin…

2021

3D CNNs With Adaptive Temporal Feature Resolutions

CVPR 2021poster

While state-of-the-art 3D Convolutional Neural Networks (CNN) achieve very good results on action recognition datasets, they are computationally very expensive and require many GFLOPs. While the GFLOPs of a 3D CNN can be decreased by reducing the temporal feature resolution within the network, there…

Cited by 39PDFcodeScholar
2021

ACDC: The Adverse Conditions Dataset With Correspondences for Semantic Driving Scene Understanding

ICCV 2021poster

Level 5 autonomy for self-driving cars requires a robust visual perception system that can parse input images under any visual condition. However, existing semantic segmentation datasets are either dominated by images captured under normal conditions or are small in scale. To address this, we introd…

Cited by 705PDFScholar
2021

Analogical Image Translation for Fog Generation

AAAI 2021technical

Image-to-image translation is to map images from a given style to another given style. While exceptionally successful, current methods assume the availability of training images in both source and target domains, which does not always hold in practice. Inspired by humans' reasoning capability of ana…

Cited by 18SourcePDFScholar
2021

Cluster, Split, Fuse, and Update: Meta-Learning for Open Compound Domain Adaptive Semantic Segmentation

CVPR 2021poster

Open compound domain adaptation (OCDA) is a domain adaptation setting, where target domain is modeled as a compound of multiple unknown homogeneous domains, which brings the advantage of improved generalization to unseen domains. In this work, we propose a principled meta-learning based approach to…

Cited by 45PDFScholar
2021

CompositeTasking: Understanding Images by Spatial Composition of Tasks

CVPR 2021poster

We define the concept of CompositeTasking as the fusion of multiple, spatially distributed tasks, for various aspects of image understanding. Learning to perform spatially distributed tasks is motivated by the frequent availability of only sparse labels across tasks, and the desire for a compact mul…

Cited by 6PDFcodeScholar
2021

DeFlow: Learning Complex Image Degradations From Unpaired Data With Conditional Flows

CVPR 2021poster

The difficulty of obtaining paired data remains a major bottleneck for learning image restoration and enhancement models for real-world applications. Current strategies aim to synthesize realistic training data by modeling noise and degradations that appear in real-world settings. We propose DeFlow,…

Cited by 56PDFcodeScholar
2021

Decoder Fusion RNN: Context and Interaction Aware Decoders for Trajectory Prediction

IROS 2021poster

Forecasting the future behavior of all traffic agents in the vicinity is a key task to achieve safe and reliable autonomous driving systems. It is a challenging problem as agents adjust their behavior depending on their intentions, the others’ actions, and the road layout. In this paper, we propose…

Cited by 17SourceScholar
2021

Deep Reparametrization of Multi-Frame Super-Resolution and Denoising

ICCV 2021poster

We propose a deep reparametrization of the maximum a posteriori formulation commonly employed in multi-frame image restoration tasks. Our approach is derived by introducing a learned error metric and a latent representation of the target image, which transforms the MAP objective to a deep feature sp…

Cited by 74PDFScholar
2021

Designing a Practical Degradation Model for Deep Blind Image Super-Resolution

ICCV 2021poster

It is widely acknowledged that single image super-resolution (SISR) methods would not perform well if the assumed degradation model deviates from those in real images. Although several degradation models take additional factors into consideration, such as blur, they are still not effective enough to…

Cited by 968PDFcodeScholar
2021

Differentiable Multi-Granularity Human Representation Learning for Instance-Aware Human Semantic Parsing

CVPR 2021poster

To address the challenging task of instance-aware human part parsing, a new bottom-up regime is proposed to learn category-level human semantic segmentation as well as multi-person pose estimation in a joint and end-to-end manner. It is a compact, efficient and powerful framework that exploits struc…

Cited by 93PDFcodeScholar
2021

Domain Adaptive Semantic Segmentation With Self-Supervised Depth Estimation

ICCV 2021poster

Domain adaptation for semantic segmentation aims to improve the model performance in the presence of a distribution shift between source and target domain. Leveraging the supervision from auxiliary tasks (such as depth estimation) has the potential to heal this shift because many visual tasks are cl…

Cited by 169PDFcodeScholar
2021

Efficient Conditional GAN Transfer With Knowledge Propagation Across Classes

CVPR 2021poster

Generative adversarial networks (GANs) have shown impressive results in both unconditional and conditional image generation. In recent literature, it is shown that pre-trained GANs, on a different dataset, can be transferred to improve the image generation from a small target data. The same, however…

Cited by 29PDFcodeScholar
2021

End-to-End Urban Driving by Imitating a Reinforcement Learning Coach

ICCV 2021poster

End-to-end approaches to autonomous driving commonly rely on expert demonstrations. Although humans are good drivers, they are not good coaches for end-to-end algorithms that demand dense on-policy supervision. On the contrary, automated experts that leverage privileged information can efficiently g…

Cited by 232PDFcodeScholar
2021

Exploring Cross-Image Pixel Contrast for Semantic Segmentation

ICCV 2021poster

Current semantic segmentation methods focus only on mining "local" context, i.e., dependencies between pixels within individual images, by context-aggregation modules (e.g., dilated convolution, neural attention) or structure-aware optimization criteria (e.g., IoU-like loss). However, they ignore "g…

Cited by 624PDFcodeScholar
2021

Exploring Geometry-Aware Contrast and Clustering Harmonization for Self-Supervised 3D Object Detection

ICCV 2021poster

Current 3D object detection paradigms highly rely on extensive annotation efforts, which makes them not practical in many real-world industrial applications. Inspired by that a human driver can keep accumulating experiences from self-exploring the roads without any tutor's guidance, we first step fo…

Cited by 86PDFcodeScholar
2021

Exploring Relational Context for Multi-Task Dense Prediction

ICCV 2021poster

The timeline of computer vision research is marked with advances in learning and utilizing efficient contextual representations. Most of them, however, are targeted at improving model performance on a single downstream task. We consider a multi-task environment for dense prediction tasks, represente…

Cited by 97PDFcodeScholar
2021

Fast Few-Shot Classification by Few-Iteration Meta-Learning

ICRA 2021poster

Autonomous agents interacting with the real world need to learn new concepts efficiently and reliably. This requires learning in a low-data regime, which is a highly challenging problem. We address this task by introducing a fast optimization-based meta-learning method for few-shot classification. I…

Cited by 9SourcecodeScholar
2021

Flow-Based Kernel Prior With Application to Blind Super-Resolution

CVPR 2021poster

Kernel estimation is generally one of the key problems for blind image super-resolution (SR). Recently, Double-DIP proposes to model the kernel via a network architecture prior, while KernelGAN employs the deep linear network and several regularization losses to constrain the kernel space. However,…

Cited by 176PDFcodeScholar
2021

FoV-Net: Field-of-View Extrapolation Using Self-Attention and Uncertainty

RA-L 2021

The ability to make educated predictions about their surroundings, and associate them with certain confidence, is important for intelligent systems, like autonomous vehicles and robots. It allows them to plan early and decide accordingly. Motivated by this observation, in this letter we utilize info

Cited by 7SourcecodeScholar
2021

Fog Simulation on Real LiDAR Point Clouds for 3D Object Detection in Adverse Weather

ICCV 2021poster

This work addresses the challenging task of LiDAR-based 3D object detection in foggy weather. Collecting and annotating data in such a scenario is very time, labor and cost intensive. In this paper, we tackle this problem by simulating physically accurate fog into clear-weather scenes, so that the a…

Cited by 185PDFcodeScholar
2021

GANmut: Learning Interpretable Conditional Space for Gamut of Emotions

CVPR 2021poster

Humans can communicate emotions through a plethora of facial expressions, each with its own intensity, nuances and ambiguities. The generation of such variety by means of conditional GANs is limited to the expressions encoded in the used label system. These limitations are caused either due to burde…

Cited by 28PDFScholar