← Search

Philip H.S. Torr

58 accepted papers

2025

Vision-Language Models Do Not Understand Negation

CVPR 2025poster

Many practical vision-language applications require models that understand negation, e.g., when using natural language to retrieve images which contain certain objects but not others. Despite advancements in vision-language models (VLMs) through large-scale training, their ability to comprehend nega…

Cited by 7SourcePDFScholar
2024

NeRF-VPT: Learning Novel View Representations with Neural Radiance Fields via View Prompt Tuning

AAAI 2024technical

Neural Radiance Fields (NeRF) have garnered remarkable success in novel view synthesis. Nonetheless, the task of generating high-quality images for novel views persists as a critical challenge. While the existing efforts have exhibited commendable progress, capturing intricate details, enhancing tex…

2023

Computationally Budgeted Continual Learning: What Does Matter?

CVPR 2023poster

Continual Learning (CL) aims to sequentially train models on streams of incoming data that vary in distribution by preserving previous knowledge while adapting to new data. Current CL literature focuses on restricted access to previously seen data, while imposing no constraints on the computational…

2023

Deconstructed Generation-Based Zero-Shot Model

AAAI 2023technical

Recent research on Generalized Zero-Shot Learning (GZSL) has focused primarily on generation-based methods. However, current literature has overlooked the fundamental principles of these methods and has made limited progress in a complex manner. In this paper, we aim to deconstruct the generator-cla…

2023

Deep Deterministic Uncertainty: A New Simple Baseline

CVPR 2023highlight

Reliable uncertainty from deterministic single-forward pass models is sought after because conventional methods of uncertainty quantification are computationally expensive. We take two complex single-forward-pass uncertainty approaches, DUQ and SNGP, and examine whether they mainly rely on a well-re…

Cited by 133SourcePDFScholar
2023

MOSE: A New Dataset for Video Object Segmentation in Complex Scenes

ICCV 2023poster

Video object segmentation (VOS) aims at segmenting a particular object throughout the entire video clip sequence. The state-of-the-art VOS methods have achieved excellent performance (e.g., 90+% J&F) on existing datasets. However, since the target objects in these existing datasets are usually relat…

Cited by 148PDFcodeScholar
2023

MobileBrick: Building LEGO for 3D Reconstruction on Mobile Devices

CVPR 2023poster

High-quality 3D ground-truth shapes are critical for 3D object reconstruction evaluation. However, it is difficult to create a replica of an object in reality, and even 3D reconstructions generated by 3D scanners have artefacts that cause biases in evaluation. To address this issue, we introduce a n…

2023

Open Vocabulary Semantic Segmentation With Patch Aligned Contrastive Learning

CVPR 2023highlight

We introduce Patch Aligned Contrastive Learning (PACL), a modified compatibility function for CLIP's contrastive loss, intending to train an alignment between the patch tokens of the vision encoder and the CLS token of the text encoder. With such an alignment, a model can identify regions of an imag…

2023

Rapid Adaptation in Online Continual Learning: Are We Evaluating It Right?

ICCV 2023poster

We revisit the common practice of evaluating adaptation of Online Continual Learning (OCL) algorithms through the metric of online accuracy, which measures the accuracy of the model on the immediate next few samples. However, we show that this metric is unreliable, as even vacuous blind classifiers,…

Cited by 0PDFcodeScholar
2023

Real-Time Evaluation in Online Continual Learning: A New Hope

CVPR 2023highlight

Current evaluations of Continual Learning (CL) methods typically assume that there is no constraint on training time and computation. This is an unrealistic assumption for any real-world setting, which motivates us to propose: a practical real-time evaluation of continual learning, in which the stre…

2023

Reliability in Semantic Segmentation: Are We on the Right Track?

CVPR 2023poster

Motivated by the increasing popularity of transformers in computer vision, in recent times there has been a rapid development of novel architectures. While in-domain performance follows a constant, upward trend, properties like robustness or uncertainty estimation are less explored -leaving doubts a…

2023

Sample-Dependent Adaptive Temperature Scaling for Improved Calibration

AAAI 2023technical

It is now well known that neural networks can be wrong with high confidence in their predictions, leading to poor calibration. The most common post-hoc approach to compensate for this is to perform temperature scaling, which adjusts the confidences of the predictions on any input by scaling the logi…

2023

Semantics-Aware Dynamic Localization and Refinement for Referring Image Segmentation

AAAI 2023technical

Referring image segmentation segments an image from a language expression. With the aim of producing high-quality masks, existing methods often adopt iterative learning approaches that rely on RNNs or stacked attention layers to refine vision-language features. Despite their complexity, RNN-based me…

Cited by 28SourcePDFScholar
2023

TIPI: Test Time Adaptation With Transformation Invariance

CVPR 2023poster

When deploying a machine learning model to a new environment, we often encounter the distribution shift problem -- meaning the target data distribution is different from the model's training distribution. In this paper, we assume that labels are not provided for this new domain, and that we do not s…

2023

Tem-Adapter: Adapting Image-Text Pretraining for Video Question Answer

ICCV 2023poster

Video-language pre-trained models have shown remarkable success in guiding video question-answering (VideoQA) tasks. However, due to the length of video sequences, training large-scale video-based models incurs considerably higher costs than training image-based ones. This motivates us to leverage t…

Cited by 17PDFcodeScholar
2022

BNV-Fusion: Dense 3D Reconstruction Using Bi-Level Neural Volume Fusion

CVPR 2022poster

Dense 3D reconstruction from a stream of depth images is the key to many mixed reality and robotic applications. Although methods based on Truncated Signed Distance Function (TSDF) Fusion have advanced the field over the years, the TSDF volume representation is confronted with striking a balance bet…

Cited by 44PDFcodeScholar
2022

Combating Adversaries with Anti-adversaries

AAAI 2022technical

Deep neural networks are vulnerable to small input perturbations known as adversarial attacks. Inspired by the fact that these adversaries are constructed by iteratively minimizing the confidence of a network for the true class label, we propose the anti-adversary layer, aimed at countering this eff…

2022

DeformRS: Certifying Input Deformations with Randomized Smoothing

AAAI 2022technical

Deep neural networks are vulnerable to input deformations in the form of vector fields of pixel displacements and to other parameterized geometric deformations e.g. translations, rotations, etc. Current input deformation certification methods either (i) do not scale to deep networks on large input d…

2022

LAVT: Language-Aware Vision Transformer for Referring Image Segmentation

CVPR 2022poster

Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key challenges behind this task is leveraging the referring expression for highlighting relevant positions in the image. A para…

Cited by 385PDFcodeScholar
2022

Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental Learning

CVPR 2022poster

Class Incremental Learning (CIL) aims at learning a classifier in a phase-by-phase manner, in which only data of a subset of the classes are provided at each phase. Previous works mainly focus on mitigating forgetting in phases after the initial one. However, we find that improving CIL at its initia…

Cited by 89PDFcodeScholar
2022

PhysFormer: Facial Video-Based Physiological Measurement With Temporal Difference Transformer

CVPR 2022poster

Remote photoplethysmography (rPPG), which aims at measuring heart activities and physiological signals from facial video without any contact, has great potential in many applications. Recent deep learning approaches focus on mining subtle rPPG clues using convolutional neural networks with limited s…

Cited by 248PDFcodeScholar
2022

Semantic-Aware Auto-Encoders for Self-Supervised Representation Learning

CVPR 2022poster

The resurgence of unsupervised learning can be attributed to the remarkable progress of self-supervised learning, which includes generative (G) and discriminative (D) models. In computer vision, the mainstream self-supervised learning algorithms are D models. However, designing a D model could be ov…

Cited by 11PDFcodeScholar
2022

TransMix: Attend To Mix for Vision Transformers

CVPR 2022poster

Mixup-based augmentation has been found to be effective for generalizing models during training, especially for Vision Transformers (ViTs) since they can easily overfit. However, previous mixup-based methods have an underlying prior knowledge that the linearly interpolated ratio of targets should be…

Cited by 135PDFcodeScholar
2022

YouMVOS: An Actor-Centric Multi-Shot Video Object Segmentation Dataset

CVPR 2022poster

Many video understanding tasks require analyzing multi-shot videos, but existing datasets for video object segmentation (VOS) only consider single-shot videos. To address this challenge, we collected a new dataset---YouMVOS---of 200 popular YouTube videos spanning ten genres, where each video is on…

Cited by 2PDFcodeScholar
2021

Aggregation With Feature Detection

ICCV 2021poster

Aggregating features from different depths of a network is widely adopted to improve the network capability. Lots of modern architectures are equipped with skip connections, which actually makes the feature aggregation happen in all these networks. Since different features tell different semantic m…

Cited by 2PDFScholar
2021

Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With Transformers

CVPR 2021poster

Most recent semantic segmentation methods adopt a fully-convolutional network (FCN) with an encoder-decoder architecture. The encoder progressively reduces the spatial resolution and learns more abstract/semantic visual concepts with larger receptive fields. Since context modeling is critical for se…

Cited by 4009PDFcodeScholar
2021

Separable Flow: Learning Motion Cost Volumes for Optical Flow Estimation

ICCV 2021poster

Full-motion cost volumes play a central role in current state-of-the-art optical flow methods. However, constructed using simple feature correlations, they lack the ability to encapsulate prior, or even non-local, knowledge. This creates artifacts in poorly constrained, ambiguous regions, such as oc…

Cited by 134PDFcodeScholar
2021

Solving Inefficiency of Self-Supervised Representation Learning

ICCV 2021poster

Self-supervised learning (especially contrastive learning) has attracted great interest due to its huge potential in learning discriminative representations in an unsupervised manner. Despite the acknowledged successes, existing contrastive learning methods suffer from very low learning efficiency,…

Cited by 66PDFcodeScholar
2021

Vision Transformer With Progressive Sampling

ICCV 2021poster

Transformers with powerful global relation modeling abilities have been introduced to fundamental computer vision tasks recently. As a typical example, the Vision Transformer (ViT) directly applies a pure transformer architecture on image classification, by simply splitting images into tokens with a…

Cited by 126PDFcodeScholar
2020

AutoSimulate: (Quickly) Learning Synthetic Data Generation

ECCV 2020poster

Simulation is increasingly being used for generating large labelled datasets in many machine learning problems. Recent methods have focused on adjusting simulator parameters with the goal of maximising accuracy on a validation task, usually relying on REINFORCE-like gradient estimators. However thes…

2020

Cross-Modal Deep Face Normals With Deactivable Skip Connections

CVPR 2020oral

We present an approach for estimating surface normals from in-the-wild color images of faces. While data-driven strategies have been proposed for single face images, limited available ground truth data makes this problem difficult. To alleviate this issue, we propose a method that can leverage all a…

Cited by 45PDFScholar
2020

Holistically-Attracted Wireframe Parsing

CVPR 2020poster

This paper presents a fast and parsimonious parsing method to accurately and robustly detect a vectorized wireframe in an input image with a single forward pass. The proposed method is end-to-end trainable, consisting of three components: (i) line segment and junction proposal generation, (ii) line…

Cited by 136PDFcodeScholar
2020

Instance Segmentation of LiDAR Point Clouds

ICRA 2020poster

We propose a robust baseline method for instance segmentation which are specially designed for large-scale outdoor LiDAR point clouds. Our method includes a novel dense feature encoding technique, allowing the localization and segmentation of small, far-away objects, a simple but effective solution…

Cited by 73SourcecodeScholar
2020

Local Class-Specific and Global Image-Level Generative Adversarial Networks for Semantic-Guided Scene Generation

CVPR 2020poster

In this paper, we address the task of semantic-guided scene generation. One open challenge widely observed in global image-level generation methods is the difficulty of generating small objects and detailed local texture. To tackle this issue, in this work we consider learning the scene generation i…

Cited by 192PDFcodeScholar
2020

Meta-Learning Deep Visual Words for Fast Video Object Segmentation

IROS 2020poster

Personal robots and driverless cars need to be able to operate in novel environments and thus quickly and efficiently learn to recognise new object classes. We address this problem by considering the task of video object segmentation. Previous accurate methods for this task finetune a model using th…

Cited by 22SourcecodeScholar
2019

Deep Virtual Networks for Memory Efficient Inference of Multiple Tasks

CVPR 2019poster

Deep networks consume a large amount of memory by their nature. A natural question arises can we reduce that memory requirement whilst maintaining performance. In particular, in this work we address the problem of memory efficient learning for multiple tasks. To this end, we propose a novel network…

Cited by 12PDFScholar
2019

Fast Online Object Tracking and Segmentation: A Unifying Approach

CVPR 2019poster

In this paper we illustrate how to perform both visual object tracking and semi-supervised video object segmentation, in real-time, with a single simple approach. Our method, dubbed SiamMask, improves the offline training procedure of popular fully-convolutional Siamese approaches for object trackin…

Cited by 1768PDFScholar
2019

GA-Net: Guided Aggregation Net for End-To-End Stereo Matching

CVPR 2019oral

In the stereo matching task, matching cost aggregation is crucial in both traditional methods and deep neural network models in order to accurately estimate disparities. We propose two novel neural net layers, aimed at capturing local and the whole-image cost dependencies respectively. The first is…

Cited by 915PDFcodeScholar
2019

Learning to Adapt for Stereo

CVPR 2019poster

Real world applications of stereo depth estimation require models that are robust to dynamic variations in the environment. Even though deep learning based stereo methods are successful, they often fail to generalize to unseen variations in the environment, making them less suitable for practical ap…

Cited by 93PDFcodeScholar
2019

Re-Ranking via Metric Fusion for Object Retrieval and Person Re-Identification

CVPR 2019poster

This work studies the unsupervised re-ranking procedure for object retrieval and person re-identification with a specific concentration on an ensemble of multiple metrics (or similarities). While the re-ranking step is involved by running a diffusion process on the underlying data manifolds, the fus…

Cited by 118PDFScholar
2018

FlipDial: A Generative Model for Two-Way Visual Dialogue

CVPR 2018poster

We present FlipDial, a generative model for Visual Dialogue that simultaneously plays the role of both participants in a visually-grounded dialogue. Given context in the form of an image and an associated caption summarising the contents of the image, FlipDial learns both to answer questions and put…

Cited by 48SourcePDFScholar
2018

Learning to Compare: Relation Network for Few-Shot Learning

CVPR 2018poster

We present a conceptually simple, flexible, and general framework for few-shot learning, where a classifier must learn to recognise new classes given only few examples from each. Our method, called the Relation Network (RN), is trained end-to-end from scratch. During meta-learning, it learns to lear…

Cited by 4722SourcePDFScholar
2018

Long-term Tracking in the Wild: a Benchmark

ECCV 2018poster

We introduce the OxUvA dataset and benchmark for evaluating single-object tracking algorithms. Benchmarks have enabled great strides in the field of object tracking by defining standardized evaluations on large sets of diverse videos. However, these works have focused exclusively on sequences that a…

Cited by 206SourcePDFScholar
2018

Multi-Agent Diverse Generative Adversarial Networks

CVPR 2018poster

We propose MAD-GAN, an intuitive generalization to the Generative Adversarial Networks (GANs) and its conditional variants to address the well known problem of mode collapse. First, MAD-GAN is a multi-agent GAN architecture incorporating multiple generators and one discriminator. Second, to enforce…

Cited by 429SourcePDFScholar
2018

On the Robustness of Semantic Segmentation Models to Adversarial Attacks

CVPR 2018poster

Deep Neural Networks (DNNs) have been demonstrated to perform exceptionally well on most recognition tasks such as image classification and segmentation. However, they have also been shown to be vulnerable to adversarial examples. This phenomenon has recently attracted a lot of attention but it has…

2017

Random forests versus Neural Networks — What's best for camera localization?

ICRA 2017poster

This work addresses the task of camera localization in a known 3D scene given a single input RGB image. State-of-the-art approaches accomplish this in two steps: firstly, regressing for every pixel in the image its 3D scene coordinate and subsequently, using these coordinates to estimate the final 6…

Cited by 90SourceScholar