← Search

Stan Sclaroff

32 accepted papers

2023

Practical Disruption of Image Translation Deepfake Networks

AAAI 2023technical

By harnessing the latest advances in deep learning, image-to-image translation architectures have recently achieved impressive capabilities. Unfortunately, the growing representational power of these architectures has prominent unethical uses. Among these, the threats of (1) face manipulation ("Deep…

Cited by 12SourcePDFScholar
2022

A Broad Study of Pre-training for Domain Generalization and Adaptation

ECCV 2022poster

"Deep models must learn robust and transferable representations in order to perform well on new domains. While domain transfer methods (\eg, domain adaptation, domain generalization) have been proposed to learn transferable representations across domains, they are typically applied to ResNet backbon…

2022

A Unified Framework for Domain Adaptive Pose Estimation

ECCV 2022poster

"While pose estimation is an important computer vision task, it requires expensive annotation and suffers from domain shift. In this paper, we investigate the problem of domain adaptive 2D pose estimation that transfers knowledge learned on a synthetic source domain to a target domain without superv…

2022

Finding Differences Between Transformers and ConvNets Using Counterfactual Simulation Testing

NeurIPS 2022accept

Modern deep neural networks tend to be evaluated on static test sets. One shortcoming of this is the fact that these deep neural networks cannot be easily evaluated for robustness issues with respect to specific scene variations. For example, it is hard to study the robustness of these networks to v…

2022

Simulated Adversarial Testing of Face Recognition Models

CVPR 2022poster

Most machine learning models are validated and tested on fixed datasets. This can give an incomplete picture of the capabilities and weaknesses of the model. Such weaknesses can be revealed at test time in the real world. The risks involved in such failures can be loss of profits, loss of time or ev…

Cited by 17PDFScholar
2021

CDS: Cross-Domain Self-Supervised Pre-Training

ICCV 2021poster

We present a two-stage pre-training approach that improves the generalization ability of standard single-domain pre-training. While standard pre-training on a single large dataset (such as ImageNet) can provide a good initial representation for transfer learning tasks, this approach may result in bi…

Cited by 57PDFScholar
2021

Learning Cross-Modal Contrastive Features for Video Domain Adaptation

ICCV 2021poster

Learning transferable and domain adaptive feature representations from videos is important for video-relevant tasks such as action recognition. Existing video domain adaptation methods mainly rely on adversarial feature alignment, which has been derived from the RGB image space. However, video data…

Cited by 94PDFScholar
2021

Real-Time Semantic Segmentation With Fast Attention

RA-L 2021

In deep CNN based models for semantic segmentation, high accuracy relies on rich spatial context (large receptive fields) and fine spatial details (high resolution), both of which incur high computational costs. In this letter, we propose a novel architecture that addresses both challenges and achie

Cited by 143SourcecodeScholar
2021

Siamese Natural Language Tracker: Tracking by Natural Language Descriptions With Siamese Trackers

CVPR 2021poster

We propose a novel Siamese Natural Language Tracker (SNLT), which brings the advancements in visual tracking to the tracking by natural language (NL) specification task. The proposed SNLT is applicable to a wide range of Siamese trackers, providing a new class of baselines for the tracking by NL tas…

Cited by 93PDFcodeScholar
2021

Tune It the Right Way: Unsupervised Validation of Domain Adaptation via Soft Neighborhood Density

ICCV 2021poster

Unsupervised domain adaptation (UDA) methods can dramatically improve generalization on unlabeled target domains. However, optimal hyper-parameter selection is critical to achieving high accuracy and avoiding negative transfer. Supervised hyper-parameter validation is not possible without labeled ta…

Cited by 75PDFcodeScholar
2020

Temporally Distributed Networks for Fast Video Semantic Segmentation

CVPR 2020poster

We present TDNet, a temporally distributed network designed for fast and accurate video semantic segmentation. We observe that features extracted from a certain high-level layer of a deep CNN can be approximated by composing features extracted from several shallower sub-networks. Leveraging the inhe…

Cited by 250PDFScholar
2020

Universal Domain Adaptation through Self Supervision

NeurIPS 2020poster

Unsupervised domain adaptation methods traditionally assume that all source categories are present in the target domain. In practice, little may be known about the category overlap between the two domains. While some methods address target settings with either partial or open-set categories, they as…

2019

Cost-Aware Fine-Grained Recognition for IoTs Based on Sequential Fixations

ICCV 2019poster

We consider the problem of fine-grained classification on an edge camera device that has limited power. The edge device must sparingly interact with the cloud to minimize communication bits to conserve power, and the cloud upon receiving the edge inputs returns a classification label. To deal with f…

Cited by 3PDFScholar
2019

Generalized Majorization-Minimization

ICML 2019oral

Non-convex optimization is ubiquitous in machine learning. Majorization-Minimization (MM) is a powerful iterative procedure for optimizing non-convex functions that works by optimizing a sequence of bounds on the function. In MM, the bound at each iteration is required to touch the objective functio…

Cited by 23SourcePDFScholar
2019

Language Features Matter: Effective Language Representations for Vision-Language Tasks

ICCV 2019poster

Shouldn't language and vision features be treated equally in vision-language (VL) tasks? Many VL approaches treat the language component as an afterthought, using simple language models that are either built upon fixed word embeddings trained on text-only data or are learned from scratch. We conclud…

Cited by 39PDFScholar
2019

Semi-Supervised Domain Adaptation via Minimax Entropy

ICCV 2019poster

Contemporary domain adaptation methods are very effective at aligning feature distributions of source and target domains without any target supervision. However, we show that these techniques perform poorly when even a few labeled examples are available in the target domain. To address this semi-sup…

Cited by 848PDFScholar
2018

Excitation Backprop for RNNs

CVPR 2018poster

Deep models are state-of-the-art or many vision tasks including video action recognition and video captioning. Models are trained to caption or classify activity in videos, but little is known about the evidence used to make such decisions. Grounding decisions made by deep networks has been studied…

2017

Personalizing Gesture Recognition Using Hierarchical Bayesian Neural Networks

CVPR 2017poster

Building robust classifiers trained on data susceptible to group or subject-specific variations is a challenging pattern recognition problem. We develop hierarchical Bayesian neural networks to capture subject-specific variations and share statistical strength across subjects. Leveraging recent wor…

Cited by 34PDFScholar
2016

Differential Geometric Regularization for Supervised Learning of Classifiers

ICML 2016poster

We study the problem of supervised learning for both binary and multiclass classification from a unified geometric perspective. In particular, we propose a geometric regularization technique to find the submanifold corresponding to an estimator of the class probability P(y|\vec x). The regularizatio…

Cited by 3SourcePDFScholar
2016

Unconstrained Salient Object Detection via Proposal Subset Optimization

CVPR 2016spotlight

We aim at detecting salient objects in unconstrained images. In unconstrained images, the number of salient objects (if any) varies from image to image, and is not given. We present a salient object detection system that directly outputs a compact set of detection windows, if any, for an input image…

Cited by 118PDFScholar
2015

Minimum Barrier Salient Object Detection at 80 FPS

ICCV 2015oral

We propose a highly efficient, yet powerful, salient object detection method based on the Minimum Barrier Distance (MBD) Transform. The MBD transform is robust to pixel-value fluctuation, and thus can be effectively applied on raw pixels without region abstraction. We present an approximate MBD tran…

Cited by 518PDFScholar
2015

Salient Object Subitizing

CVPR 2015poster

People can immediately and precisely identify 1, 2, 3 or 4 items by a simple glance. The phenomenon, known as Subitizing, inspires us to pursue the task of Salient Object Subitizing (SOS), i.e. predicting the existence and the number of salient objects in a scene using holistic cues. To study this p…

Cited by 138SourcePDFScholar