← Search

Ser-nam Lim

83 accepted papers

2026

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

ICLR 2026poster

Existing video-to-audio (V2A) generation methods predominantly rely on text prompts alongside visual information to synthesize audio. However, two critical bottlenecks persist: semantic granularity gaps in training data (e.g., conflating acoustically distinct sounds like different dog barks under co…

Cited by 0SourcecodeScholar
2026

AlignVid: Taming Visual Dominance via Training-Free Attention Modulation in Text-guided Image-to-Video Generation

ICML 2026poster

Text-guided image-to-video generation has made substantial progress, yet it still struggles to execute text-specified edits that require substantial changes to a reference image (e.g., object addition, deletion, or modification). Empirically, our analysis reveals that this stems from **visual domina…

Cited by 0SourceScholar
2026

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA

ICML 2026poster

Document visual question answering requires models not only to answer questions correctly, but also to precisely localize answers within complex document layouts. While large vision-language models (VLMs) achieve strong spatial grounding, their inference cost and latency limit real-world deployment;…

Cited by 0SourceScholar
2026

Next Patch Prediction for AutoRegressive Visual Generation

AAAI 2026technical

Autoregressive models, built based on the Next Token Prediction (NTP) paradigm, show great potential in developing a unified framework that integrates both language and vision tasks. Pioneering works introduce NTP to autoregressive visual generation tasks. In this work, we rethink the NTP for autore

Cited by 0SourcePDFScholar
2026

Think Then Embed: Generative Context Improves Multimodal Embedding

ICLR 2026poster

There is a growing interest in Universal Multimodal Embeddings (UME), where models are required to generate task-specific representations. While recent studies show that Multimodal Large Language Models (MLLMs) perform well on such tasks, they treat MLLMs solely as encoders, overlooking their genera…

Cited by 0SourceScholar
2025

DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

ICCV 2025poster

In this work, we present DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs. Existing approaches struggle with generating coherent, high-quality content in an efficient and user-friendly manner. Concretely, baseline methods relying on only…

2025

ESCA: Contextualizing Embodied Agents via Scene-Graph Generation

NeurIPS 2025spotlight

Multi-modal large language models (MLLMs) are making rapid progress toward general-purpose embodied agents. However, existing MLLMs do not reliably capture fine-grained links between low-level visual features and high-level textual semantics, leading to weak grounding and inaccurate perception. To o…

Cited by 0SourcecodeScholar
2025

Frequency-Guided Masking for Enhanced Vision Self-Supervised Learning

ICLR 2025poster

We present a novel frequency-based Self-Supervised Learning (SSL) approach that significantly enhances its efficacy for pre-training. Prior work in this direction masks out pre-defined frequencies in the input image and employs a reconstruction loss to pre-train the model. While achieving promising…

2025

Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation

NeurIPS 2025poster

Recent advancements in video generation have enabled the creation of high-quality, visually compelling videos. However, generating videos that adhere to the laws of physics remains a critical challenge for applications requiring realism and accuracy. In this work, we propose **PhysHPO**, a novel fra…

Cited by 0SourceScholar
2025

Improving Soft Unification with Knowledge Graph Embedding Methods

ICML 2025poster

Neural Theorem Provers (NTPs) present a promising framework for neuro-symbolic reasoning, combining end-to-end differentiability with the interpretability of symbolic logic programming. However, optimizing NTPs remains a significant challenge due to their complex objective landscape and gradient spa…

Cited by 0SourcePDFScholar
2025

Intervening Anchor Token: Decoding Strategy in Alleviating Hallucinations for MLLMs

ICLR 2025poster

Multimodal large language models (MLLMs) offer a powerful mechanism for interpreting visual information. However, they often suffer from hallucinations, which impede the real-world usage of these models. Existing methods attempt to alleviate this issue by designing special decoding strategies that p…

Cited by 1SourcePDFScholar
2025

LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence

ICML 2025poster

Recent embodied agents are primarily built based on reinforcement learning (RL) or large language models (LLMs). Among them, RL agents are efficient for deployment but only perform very few tasks. By contrast, giant LLM agents (often more than 1000B parameters) present strong generalization while de…

Cited by 1SourcePDFScholar
2025

LASER: A Neuro-Symbolic Framework for Learning Spatio-Temporal Scene Graphs with Weak Supervision

ICLR 2025poster

Supervised approaches for learning spatio-temporal scene graphs (STSG) from video are greatly hindered due to their reliance on STSG-annotated videos, which are labor-intensive to construct at scale. Is it feasible to instead use readily available video captions as weak supervision? To address this…

Cited by 1SourcePDFScholar
2025

Model Reveals What to Cache: Profiling-Based Feature Reuse for Video Diffusion Models

ICCV 2025poster

Video generation using diffusion models has shown remarkable progress, yet it remains computationally expensive due to the repeated processing of redundant features across blocks and steps. To address this, we propose a novel adaptive feature reuse mechanism that dynamically identifies and caches th…

2025

Scaling Up Temporal Domain Generalization via Temporal Experts Averaging

EMNLP 2025

Temporal Domain Generalization (TDG) aims to generalize across temporal distribution shifts, e.g., lexical change over time. Prior work often addresses this by predicting future model weights. However, full model prediction is prohibitively expensive for even reasonably sized models. Thus, recent me

2025

Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models

ICCV 2025poster

Modern retrieval systems often struggle with upgrading to new and more powerful models due to the incompatibility of embeddings between the old and new models. This necessitates a costly process known as backfilling, which involves re-computing the embeddings for a large number of data samples. In v…

Cited by 0SourcePDFScholar
2025

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

NeurIPS 2025poster

Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visual…

Cited by 0SourceScholar
2024

Composing Object Relations and Attributes for Image-Text Matching

CVPR 2024poster

We study the visual semantic embedding problem for image-text matching. Most existing work utilizes a tailored cross-attention mechanism to perform local alignment across the two image and text modalities. This is computationally expensive even though it is more powerful than the unimodal dual-encod…

2024

Fast Encoding and Decoding for Implicit Video Representation

ECCV 2024poster

"Despite the abundant availability and content richness for video data, its high-dimensionality poses challenges for video research. Recent advancements have explored the implicit representation for videos using neural networks, demonstrating strong performance in applications such as video compress…

Cited by 1SourcePDFScholar
2024

Jack of All Tasks Master of Many: Designing General-Purpose Coarse-to-Fine Vision-Language Model

CVPR 2024highlight

The ability of large language models (LLMs) to process visual inputs has given rise to general-purpose vision systems unifying various vision-language (VL) tasks by instruction tuning. However due to the enormous diversity in input-output formats in the vision domain existing general-purpose models…

2024

Label Delay in Online Continual Learning

NeurIPS 2024poster

Online continual learning, the process of training models on streaming data, has gained increasing attention in recent years. However, a critical aspect often overlooked is the label delay, where new data may not be labeled due to slow and costly annotation processes. We introduce a new continual le…

Cited by 2SourcePDFScholar
2024

Language-Free Compositional Action Generation via Decoupling Refinement

ICASSP 2024accepted

Composing simple actions into complex actions is crucial yet challenging. Existing methods largely rely on language annotations to discern composable latent semantics, which is costly and labor-intensive. In this study, we introduce a novel framework to generate compositional actions without languag…

Cited by 0SourceScholar
2024

MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding

CVPR 2024poster

With the success of large language models (LLMs) integrating the vision model into LLMs to build vision-language foundation models has gained much more interest recently. However existing LLM-based large multimodal models (e.g. Video-LLaMA VideoChat) can only take in a limited number of frames for s…

2024

Object Recognition as Next Token Prediction

CVPR 2024highlight

We present an approach to pose object recognition as next token prediction. The idea is to apply a language decoder that auto-regressively predicts the text tokens from image embeddings to form labels. To ground this prediction process in auto-regression we customize a non-causal attention mask for…

2024

On the Robustness of Large Multimodal Models Against Image Adversarial Attacks

CVPR 2024poster

Recent advances in instruction tuning have led to the development of State-of-the-Art Large Multimodal Models (LMMs). Given the novelty of these models the impact of visual adversarial attacks on LMMs has not been thoroughly examined. We conduct a comprehensive study of the robustness of various LMM…

Cited by 39SourcePDFScholar
2024

Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval

CVPR 2024poster

Composed Image Retrieval (CIR) is a task that retrieves images similar to a query based on a provided textual modification. Current techniques rely on supervised learning for CIR models using labeled triplets of the <reference image text target image>. These specific triplets are not as commonly ava…

Cited by 14SourcePDFScholar
2024

uCAP: An Unsupervised Prompting Method for Vision-Language Models

ECCV 2024oral

"This paper addresses a significant limitation that prevents Contrastive Language-Image Pretrained Models (CLIP) from achieving optimal performance on downstream image classification tasks. The key problem with CLIP-style zero-shot classification is that it requires domain-specific context in the fo…

Cited by 0SourcePDFScholar
2023

BT^2: Backward-compatible Training with Basis Transformation

ICCV 2023poster

Modern retrieval system often requires recomputing the representation of every piece of data in the gallery when updating to a better representation model. This process is known as backfilling and can be especially costly in the real world where the gallery often contains billions of samples. Recent…

Cited by 6PDFcodeScholar
2023

Computationally Budgeted Continual Learning: What Does Matter?

CVPR 2023poster

Continual Learning (CL) aims to sequentially train models on streams of incoming data that vary in distribution by preserving previous knowledge while adapting to new data. Current CL literature focuses on restricted access to previously seen data, while imposing no constraints on the computational…

2023

Detecting Everything in the Open World: Towards Universal Object Detection

CVPR 2023poster

In this paper, we formally address universal object detection, which aims to detect every scene and predict every category. The dependence on human annotations, the limited visual information, and the novel categories in the open world severely restrict the universality of traditional detectors. We…

2023

Graph Inductive Biases in Transformers without Message Passing

ICML 2023poster

Transformers for graph data are increasingly widely studied and successful in numerous learning tasks. Graph inductive biases are crucial for Graph Transformers, and previous works incorporate them using message-passing modules and/or positional encodings. However, Graph Transformers that use messag…

2023

HNeRV: A Hybrid Neural Representation for Videos

CVPR 2023poster

Implicit neural representations store videos as neural networks and have performed well for vision tasks such as video compression and denoising. With frame index and/or positional index as input, implicit representations (NeRV, E-NeRV, etc.) reconstruct video frames from fixed and content-agnostic…

2023

Open Vocabulary Semantic Segmentation With Patch Aligned Contrastive Learning

CVPR 2023highlight

We introduce Patch Aligned Contrastive Learning (PACL), a modified compatibility function for CLIP's contrastive loss, intending to train an alignment between the patch tokens of the vision encoder and the CLS token of the text encoder. With such an alignment, a model can identify regions of an imag…

2023

Open-vocabulary Panoptic Segmentation with Embedding Modulation

ICCV 2023poster

Open-vocabulary segmentation is attracting increasing attention due to its critical applications in the real world. Traditional closed-vocabulary segmentation methods are not able to characterize novel objects, whereas several recent open-vocabulary attempts obtain unsatisfactory results, i.e., nota…

Cited by 34PDFScholar
2023

Rapid Adaptation in Online Continual Learning: Are We Evaluating It Right?

ICCV 2023poster

We revisit the common practice of evaluating adaptation of Online Continual Learning (OCL) algorithms through the metric of online accuracy, which measures the accuracy of the model on the immediate next few samples. However, we show that this metric is unreliable, as even vacuous blind classifiers,…

Cited by 0PDFcodeScholar
2023

Riemannian Residual Neural Networks

NeurIPS 2023poster

Recent methods in geometric deep learning have introduced various neural networks to operate over data that lie on Riemannian manifolds. Such networks are often necessary to learn well over graphs with a hierarchical structure or to learn over manifold-valued data encountered in the natural sciences…

Cited by 16SourcePDFScholar
2023

Sample-Dependent Adaptive Temperature Scaling for Improved Calibration

AAAI 2023technical

It is now well known that neural networks can be wrong with high confidence in their predictions, leading to poor calibration. The most common post-hoc approach to compensate for this is to perform temperature scaling, which adjusts the confidences of the predictions on any input by scaling the logi…

2023

TIPI: Test Time Adaptation With Transformation Invariance

CVPR 2023poster

When deploying a machine learning model to a new environment, we often encounter the distribution shift problem -- meaning the target data distribution is different from the model's training distribution. In this paper, we assume that labels are not provided for this new domain, and that we do not s…

2023

Test-Time Distribution Normalization for Contrastively Learned Visual-language Models

NeurIPS 2023poster

Advances in the field of visual-language contrastive learning have made it possible for many downstream applications to be carried out efficiently and accurately by simply taking the dot product between image and text representations. One of the most representative approaches proposed recently know…

2023

Towards Scalable Neural Representation for Diverse Videos

CVPR 2023poster

Implicit neural representations (INR) have gained increasing attention in representing 3D scenes and images, and have been recently applied to encode videos (e.g., NeRV, E-NeRV). While achieving promising results, existing INR-based methods are limited to encoding a handful of short videos (e.g., se…

Cited by 45SourcePDFScholar
2023

Video Dynamics Prior: An Internal Learning Approach for Robust Video Enhancements

NeurIPS 2023poster

In this paper, we present a novel robust framework for low-level vision tasks, including denoising, object removal, frame interpolation, and super-resolution, that does not require any external training data corpus. Our proposed approach directly learns the weights of neural modules by optimizing ov…

Cited by 12SourcePDFScholar
2022

AdaViT: Adaptive Vision Transformers for Efficient Image Recognition

CVPR 2022poster

Built on top of self-attention mechanisms, vision transformers have demonstrated remarkable performance on a variety of vision tasks recently. While achieving excellent performance, they still require relatively intensive computational cost that scales up drastically as the numbers of patches, self-…

Cited by 301PDFcodeScholar
2022

FedSR: A Simple and Effective Domain Generalization Method for Federated Learning

NeurIPS 2022accept

Federated Learning (FL) refers to the decentralized and privacy-preserving machine learning framework in which multiple clients collaborate (with the help of a central server) to train a global model without sharing their data. However, most existing FL methods only focus on maximizing the model's p…

Cited by 119SourcePDFScholar
2022

Few-Shot Fast-Adaptive Anomaly Detection

NeurIPS 2022accept

The ability to detect anomaly has long been recognized as an inherent human ability, yet to date, practical AI solutions to mimic such capability have been lacking. This lack of progress can be attributed to several factors. To begin with, the distribution of ``abnormalities'' is intractable. Anythi…

Cited by 30SourcePDFScholar
2022

GAPX: Generalized Autoregressive Paraphrase-Identification X

NeurIPS 2022accept

Paraphrase Identification is a fundamental task in Natural Language Processing. While much progress has been made in the field, the performance of many state-of- the-art models often suffer from distribution shift during inference time. We verify that a major source of this performance drop comes fr…

2022

HorNet: Efficient High-Order Spatial Interactions with Recursive Gated Convolutions

NeurIPS 2022accept

Recent progress in vision Transformers exhibits great success in various tasks driven by the new spatial modeling mechanism based on dot-product self-attention. In this paper, we show that the key ingredients behind the vision Transformers, namely input-adaptive, long-range and high-order spatial in…

2022

MTFormer: Multi-task Learning via Transformer and Cross-Task Reasoning

ECCV 2022poster

"In this paper, we explore the advantages of utilizing transformer structures for addressing multi-task learning (MTL). Specifically, we demonstrate that models with transformer structures are more appropriate for MTL than convolutional neural networks (CNNs), and we propose a novel transformer-base…

Cited by 67SourcePDFScholar
2022

Object-Centric Unsupervised Image Captioning

ECCV 2022poster

"Image captioning is a longstanding problem in the field of computer vision and natural language processing. To date, researchers have produced impressive state-of-the-art performance in the age of deep learning. Most of these state-of-the-art, however, requires large volume of annotated image-capti…

2022

ObjectFormer for Image Manipulation Detection and Localization

CVPR 2022poster

Recent advances in image editing techniques have posed serious challenges to the trustworthiness of multimedia data, which drives the research of image tampering detection. In this paper, we propose ObjectFormer to detect and localize image manipulations. To capture subtle manipulation traces that a…

Cited by 190PDFScholar
2022

Spartan: Differentiable Sparsity via Regularized Transportation

NeurIPS 2022accept

We present Spartan, a method for training sparse neural network models with a predetermined level of sparsity. Spartan is based on a combination of two techniques: (1) soft top-k masking of low-magnitude parameters via a regularized optimal transportation problem and (2) dual averaging-based paramet…

2022

Teaching with Soft Label Smoothing for Mitigating Noisy Labels in Facial Expressions

ECCV 2022poster

"Recent studies have highlighted the problem of noisy labels in large scale in-the-wild facial expressions datasets due to the uncertainties caused by ambiguous facial expressions, low-quality facial images, and the subjectiveness of annotators. To solve the problem of noisy labels, we propose Soft…

2022

Totems: Physical Objects for Verifying Visual Integrity

ECCV 2022poster

"We introduce a new approach to image forensics: placing physical refractive objects, which we call totems, into a scene so as to protect any photograph taken of that scene. Totems bend and redirect light rays, thus providing multiple, albeit distorted, views of the scene within a single image. A de…

Cited by 3SourcePDFScholar
2022

Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness

NeurIPS 2022accept

We show that the effectiveness of the well celebrated Mixup can be further improved if instead of using it as the sole learning objective, it is utilized as an additional regularizer to the standard cross-entropy loss. This simple change not only improves accuracy but also significantly improves the…

Cited by 103SourcePDFScholar
2022

Visual Prompt Tuning

ECCV 2022poster

"The current modus operandi in adapting pre-trained models involves updating all the backbone parameters, i.e. full fine-tuning. This paper introduces Visual Prompt Tuning (VPT) as an efficient and effective alternative to full fine-tuning for large-scale Transformer models in vision. Taking inspira…

2021

A Continuous Mapping For Augmentation Design

NeurIPS 2021poster

Automated data augmentation (ADA) techniques have played an important role in boosting the performance of deep models. Such techniques mostly aim to optimize a parameterized distribution over a discrete augmentation space. Thus, are restricted by the discretization of the search space which normally…

Cited by 5SourcePDFScholar
2021

Combining Label Propagation and Simple Models out-performs Graph Neural Networks

ICLR 2021poster

Graph Neural Networks (GNNs) are a predominant technique for learning over graphs. However, there is relatively little understanding of why GNNs are successful in practice and whether they are necessary for good performance. Here, we show that for many standard transductive node classification bench…

2021

Cross-Modal Retrieval Augmentation for Multi-Modal Classification

EMNLP 2021finding

Recent advances in using retrieval components over external knowledge sources have shown impressive results for a variety of downstream tasks in natural language processing. Here, we explore the use of unstructured external knowledge sources of images and their corresponding captions for improving v…

Cited by 29SourcePDFScholar
2021

Deep Co-Training With Task Decomposition for Semi-Supervised Domain Adaptation

ICCV 2021poster

Semi-supervised domain adaptation (SSDA) aims to adapt models trained from a labeled source domain to a different but related target domain, from which unlabeled data and a small set of labeled data are provided. Current methods that treat source and target supervision without distinction overlook t…

Cited by 118PDFcodeScholar
2021

Equivariant Manifold Flows

NeurIPS 2021poster

Tractably modelling distributions over manifolds has long been an important goal in the natural sciences. Recent work has focused on developing general machine learning models to learn such distributions. However, for many applications these distributions must respect manifold symmetries—a trait whi…

2021

Exploring Visual Engagement Signals for Representation Learning

ICCV 2021poster

Visual engagement in social media platforms comprises interactions with photo posts including comments, shares, and likes. In this paper, we leverage such visual engagement clues as supervisory signals for representation learning. However, learning from engagement signals is non-trivial as it is not…

Cited by 14PDFcodeScholar
2021

Intentonomy: A Dataset and Study Towards Human Intent Understanding

CVPR 2021poster

An image is worth a thousand words, conveying information that goes beyond the physical visual content therein. In this paper, we study the intent behind social media images with an aim to analyze how visual information can help the recognition of human intent. Towards this goal, we introduce an int…

Cited by 41PDFcodeScholar
2021

Joint Audio-Visual Deepfake Detection

ICCV 2021poster

Deepfakes ("deep learning" + "fake") are synthetically-generated videos from AI algorithms. While they could be entertaining, they could also be misused for falsifying speeches and spreading misinformation. The process to create deepfakes involves both visual and auditory manipulations. Exploration…

Cited by 207PDFScholar
2021

Large Scale Learning on Non-Homophilous Graphs: New Benchmarks and Strong Simple Methods

NeurIPS 2021poster

Many widely used datasets for graph machine learning tasks have generally been homophilous, where nodes with similar labels connect to each other. Recently, new Graph Neural Networks (GNNs) have been developed that move beyond the homophily regime; however, their evaluation has often been conducted…

2021

Learning to Ground Multi-Agent Communication with Autoencoders

NeurIPS 2021poster

Communication requires having a common language, a lingua franca, between agents. This language could emerge via a consensus process, but it may require many generations of trial and error. Alternatively, the lingua franca can be given by the environment, where agents ground their language in repres…

Cited by 74SourcePDFScholar
2021

NeRV: Neural Representations for Videos

NeurIPS 2021poster

We propose a novel neural representation for videos (NeRV) which encodes videos in neural networks. Unlike conventional representations that treat videos as frame sequences, we represent videos as neural networks taking frame index as input. Given a frame index, NeRV outputs the corresponding RGB i…

2021

On Feature Normalization and Data Augmentation

CVPR 2021poster

The moments (a.k.a., mean and standard deviation) of latent features are often removed as noise when training image recognition models, to increase stability and reduce training time. However, in the field of image generation, the moments play a much more central role. Studies have shown that the mo…

Cited by 201PDFcodeScholar
2021

Robustness and Generalization via Generative Adversarial Training

ICCV 2021poster

While deep neural networks have achieved remarkable success in various computer vision tasks, they often fail to generalize to subtle variations of input images. Several defenses have been proposed to improve the robustness against these variations. However, current defenses can only withstand the s…

Cited by 39PDFScholar
2021

When in Doubt: Improving Classification Performance with Alternating Normalization

EMNLP 2021finding

We introduce Classification with Alternating Normalization (CAN), a non-parametric post-processing step for classification. CAN improves classification accuracy for challenging examples by re-adjusting their predicted class probability distribution using the predicted class distributions of high-con…

2020

Curriculum Manager for Source Selection in Multi-Source Domain Adaptation

ECCV 2020poster

The performance of Multi-Source Unsupervised Domain Adaptation (MS-UDA) depends significantly on the effectiveness of transferring from labeled source domain samples. In this paper, we proposed an adversarial agent that learns a dynamic curriculum for source samples, called Curriculum Manager for So…

Cited by 151SourcePDFScholar
2020

Differentiating through the Fréchet Mean

ICML 2020poster

Recent advances in deep representation learning on Riemannian manifolds extend classical deep learning operations to better capture the geometry of the manifold. One possible extension is the Fr{é}chet mean, the generalization of the Euclidean mean; however, it has been difficult to apply because it…

2020

Making an Invisibility Cloak: Real World Adversarial Attacks on Object Detectors

ECCV 2020poster

We present a systematic study of adversarial attacks on state-of-the-art object detection frameworks. Using standard detection datasets, we train patterns that suppress the objectness scores produced by a range of commonly used detectors, and ensembles of detectors. Through extensive experiments, we…

Cited by 339SourcePDFScholar
2020

Quantization Guided JPEG Artifact Correction

ECCV 2020poster

The JPEG image compression algorithm is the most popular method of image compression because of it’s ability for large compression ratios. However, to achieve such high compression, information is lost. For aggressive quantization settings, this leads to a noticeable reduction in image quality. Arti…

2020

What makes fake images detectable? Understanding properties that generalize

ECCV 2020poster

The quality of image generation and manipulation is reaching impressive levels, making it exceedingly difficult for a human to distinguish between what is real and what is fake. However, deep networks can still pick up on the subtle artifacts in these doctored images. We seek to understand what prop…

2019

Cross-X Learning for Fine-Grained Visual Categorization

ICCV 2019poster

Recognizing objects from subcategories with very subtle differences remains a challenging task due to the large intra-class and small inter-class variation. Recent work tackles this problem in a weakly-supervised manner: object parts are first detected and the corresponding part-specific features ar…

Cited by 241PDFcodeScholar
2019

Enhancing Adversarial Example Transferability With an Intermediate Level Attack

ICCV 2019poster

Neural networks are vulnerable to adversarial examples, malicious inputs crafted to fool trained models. Adversarial examples often exhibit black-box transfer, meaning that adversarial examples for one model can fool another model. However, adversarial examples are typically overfit to exploit the p…

Cited by 306PDFcodeScholar