← Search

Yu-Xiong Wang

98 accepted papers

2026

BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning

ICLR 2026poster

Recent advances in Vision-Language Models (VLMs) have propelled embodied agents by enabling direct perception, reasoning, and planning task-oriented actions from visual inputs. However, such vision-driven embodied agents open a new attack surface: visual backdoor attacks, where the agent behaves no…

Cited by 0SourcecodeScholar
2026

Capturing Visual Environment Structure Correlates with Control Performance

ICLR 2026poster

The choice of visual representation is key to scaling generalist robot policies. However, direct evaluation via policy rollouts is expensive, even in simulation. Existing proxy metrics focus on the representation's capacity to capture narrow aspects of the visual world, like object shape, limiting g…

Cited by 0SourceScholar
2026

HandX: Scaling Bimanual Motion and Interaction Generation

CVPR 2026

Synthesizing human motion has advanced rapidly, yet realistic hand motion and bimanual interaction remain underexplored. Whole-body models often miss the fine-grained cues that drive dexterous behavior, finger articulation, contact timing, and inter-hand coordination, and existing resources lack hig

Cited by 0SourcecodeScholar
2026

InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions

CVPR 2026

Humans rarely plan whole-body interactions with objects at the level of explicit whole-body movements. High-level intentions, such as affordance, define the goal, while coordinated balance, contact, and manipulation can emerge naturally from underlying physical and motor priors. Scaling such priors

Cited by 0SourceScholar
2026

LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight

CVPR 2026

To act in the world, a model must name what it sees and know where it is in 3D. Today's vision-language models excel at open-ended 2D description and grounding, yet multi-object 3D detection remains largely missing from the VLM toolbox. We present LocateAnything3D, a VLM-native recipe that casts 3D

Cited by 0SourcecodeScholar
2026

Think in Latent, Explain in Language: Self-Explainable Latent Reasoning

ICML 2026poster

Latent reasoning has emerged as a powerful alternative to text-based Chain-of-Thought (CoT), offering significant gains in computational efficiency by compressing verbose reasoning into compact embeddings. However, compressing reasoning into the latent space renders the thinking opaque, hindering it…

Cited by 0SourceScholar
2026

Unleashing Guidance Without Classifiers for Human-Object Interaction Animation

ICLR 2026poster

Generating realistic human-object interaction (HOI) animations remains challenging because it requires jointly modeling dynamic human actions and diverse object geometries. Prior diffusion-based approaches often rely on handcrafted contact priors or human-imposed kinematic constraints to improve con…

Cited by 0SourceScholar
2025

AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark

NeurIPS 2025poster

We present **AgMMU**, a challenging real‑world benchmark for evaluating and advancing vision-language models (VLMs) in the knowledge‑intensive domain of agriculture. Unlike prior datasets that rely on crowdsourced prompts, AgMMU is distilled from 116,231 authentic dialogues between everyday growers…

Cited by 0SourceScholar
2025

Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception

ICLR 2025poster

With success in image generation, generative diffusion models are increasingly adopted for discriminative scenarios because generating pixels is a unified and natural perception interface. Although directly re-purposing their generative denoising process has established promising progress in special…

2025

Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought

CVPR 2025poster

Recent advances in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks, yet they often struggle with vision-centric scenarios where precise visual focus is needed for accurate reasoning. In this paper, we introduce Argus to address these limita…

Cited by 0SourcePDFScholar
2025

Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference Scoped Exploration

CoRL 2025poster

Hand–object motion-capture (MoCap) repositories provide abundant, contact-rich human demonstrations for scaling dexterous manipulation on robots. Yet demonstration inaccuracy and embodiment gaps between human and robot hands challenge direct policy learning. Existing pipelines adapt a three-stage wo…

Cited by 0SourceScholar
2025

Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models

ICLR 2025poster

Beyond high-fidelity image synthesis, diffusion models have recently exhibited promising results in dense visual perception tasks. However, most existing work treats diffusion models as a standalone component for perception tasks, employing them either solely for off-the-shelf data augmentation or a…

2025

Floating No More: Object-Ground Reconstruction from a Single Image

CVPR 2025poster

Recent advancements in 3D object reconstruction from single images have primarily focused on improving the accuracy of object shapes. Yet, these techniques often fail to accurately capture the inter-relation between the object, ground, and camera. As a result, the reconstructed objects often appear…

Cited by 3SourcePDFScholar
2025

InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation

CVPR 2025poster

While large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to dataset limitations. Existing datasets often lack extensive, high-quality motion and annotation and exhibit artifacts s…

Cited by 2SourcePDFScholar
2025

InterMimic: Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions

CVPR 2025highlight

Achieving realistic simulations of humans interacting with a wide range of objects has long been a fundamental goal. Extending physics-based motion imitation to complex human-object interactions (HOIs) is challenging due to intricate human-object coupling, variability in object geometries, and artif…

2025

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding

NeurIPS 2025poster

Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames. With each video frame typically expanding into tens or hundreds of tokens, the limited context length of large language models (LLMs) forces the VLMs to perceive the frames…

Cited by 0SourceScholar
2025

Proposer-Agent-Evaluator (PAE): Autonomous Skill Discovery For Foundation Model Internet Agents

ICML 2025poster

A generalist foundation model agent needs to have a large and diverse skill repertoire, such as finding directions between two travel locations and buying specific items from the Internet. If each skill needs to be specified manually through a fixed set of human-annotated instructions, the agent’s s…

Cited by 11SourcePDFScholar
2025

RandAR: Decoder-only Autoregressive Visual Generation in Random Orders

CVPR 2025poster

We introduce RandAR, a decoder-only visual autoregressive (AR) model capable of generatng images in arbitrary token orders. Unlike previous decoder-only AR models that rely on a predefined generation order, RandAR removes this inductive bias, unlocking new capabilities in decoder-only generation. Ou…

2025

Refer to Any Segmentation Mask Group With Vision-Language Prompts

ICCV 2025poster

Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language and vision. This limitation reduces their effectiveness in applications that requi…

2025

ReferEverything: Towards Segmenting Everything We Can Speak of in Videos

ICCV 2025poster

We present REM, a framework for segmenting a wide range of concepts in video that can be described through natural language. Our method leverages the universal visual-language mapping learned by video diffusion models on Internet-scale data by fine-tuning them on small-scale Referring Object Segment…

Cited by 0SourcePDFScholar
2025

Self-Guided Hierarchical Exploration for Generalist Foundation Model Web Agents

NeurIPS 2025poster

Foundation models have recently shown strong potential as web agents, capable of interpreting high-level instructions and interacting with complex web interfaces. However, existing training paradigms for these agents often rely on predefined task datasets and curated demonstrations, limiting their s…

Cited by 0SourceScholar
2025

Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning

ICLR 2025poster

Vision Foundation Models (VFMs) have demonstrated outstanding performance on numerous downstream tasks. However, due to their inherent representation biases originating from different training paradigms, VFMs exhibit advantages and disadvantages across distinct vision tasks. Although amalgamating th…

Cited by 1SourcePDFScholar
2025

Virtual Fitting Room: Generating Arbitrarily Long Videos of Virtual Try-On from a Single Image

NeurIPS 2025poster

This paper proposes Virtual Fitting Room (VFR), a novel video generative model that produces arbitrarily long virtual try-on videos. Our VFR models long video generation tasks as an auto-regressive, segment-by-segment generation process, eliminating the need for resource-intensive generation and len…

Cited by 0SourcecodeScholar
2025

Visual Program Distillation with Template-Based Augmentation

EMNLP 2025

Adapting visual programming or prompting large language models (LLMs) to generate executable code for visual tasks like visual question answering (VQA) for specialized tasks or domains remains challenging due to high annotation and inference costs. We propose a low-cost visual program distillation m

2024

ATraDiff: Accelerating Online Reinforcement Learning with Imaginary Trajectories

ICML 2024poster

Training autonomous agents with sparse rewards is a long-standing problem in online reinforcement learning (RL), due to low data efficiency. Prior work overcomes this challenge by extracting useful knowledge from offline data, often accomplished through the learning of action distribution from offli…

2024

AlignDiff: Aligning Diffusion Models for General Few-Shot Segmentation

ECCV 2024oral

"Text-to-image diffusion models have shown remarkable success in synthesizing photo-realistic images. Apart from creative applications, can we use such models to synthesize samples that aid the few-shot training of discriminative models? In this work, we propose AlignDiff, a general framework for sy…

2024

Aligning Large Multimodal Models with Factually Augmented RLHF

ACL 2024findings

Large Multimodal Models (LMM) are built across modalities and the misalignment between two modalities can result in “hallucination”, generating textual outputs that are not grounded by the multimodal information in context. To address the multimodal misalignment issue, we adapt the Reinforcement Lea…

2024

ConsistDreamer: 3D-Consistent 2D Diffusion for High-Fidelity Scene Editing

CVPR 2024poster

This paper proposes ConsistDreamer - a novel framework that lifts 2D diffusion models with 3D awareness and 3D consistency thus enabling high-fidelity instruction-guided scene editing. To overcome the fundamental limitation of missing 3D consistency in 2D diffusion models our key insight is to intro…

Cited by 9SourcePDFScholar
2024

Frozen Transformers in Language Models Are Effective Visual Encoder Layers

ICLR 2024spotlight

This paper reveals that large language models (LLMs), despite being trained solely on text data, are surprisingly}strong encoders for purely visual tasks in the absence of language. Even more intriguingly, this can be achieved by a simple yet previously overlooked strategy -- employing a frozen tran…

2024

Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion

CVPR 2024poster

This paper proposes Instruct 4D-to-4D that achieves 4D awareness and spatial-temporal consistency for 2D diffusion models to generate high-quality instruction-guided dynamic scene editing results. Traditional applications of 2D diffusion models in dynamic scene editing often result in inconsistency…

Cited by 9SourcePDFScholar
2024

InstructG2I: Synthesizing Images from Multimodal Attributed Graphs

NeurIPS 2024poster

In this paper, we approach an overlooked yet critical task Graph2Image: generating images from multimodal attributed graphs (MMAGs). This task poses significant challenges due to the explosion in graph size, dependencies among graph entities, and the need for controllability in graph conditions. To…

2024

InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object Interaction

NeurIPS 2024poster

Text-conditioned human motion generation has experienced significant advancements with diffusion models trained on extensive motion capture data and corresponding textual annotations. However, extending such success to 3D dynamic human-object interaction (HOI) generation faces notable challenges, pr…

Cited by 23SourcePDFScholar
2024

Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models

ICML 2024poster

While language models (LMs) have shown potential across a range of decision-making tasks, their reliance on simple acting processes limits their broad deployment as autonomous agents. In this paper, we introduce Language Agent Tree Search (LATS) -- the first general framework that synergizes the cap…

2024

Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding

NeurIPS 2024poster

Complex 3D scene understanding has gained increasing attention, with scene encoding strategies built on top of visual foundation models playing a crucial role in this success. However, the optimal scene encoding strategies for various scenarios remain unclear, particularly compared to their image-ba…

Cited by 14SourcePDFScholar
2024

Offline Imitation from Observation via Primal Wasserstein State Occupancy Matching

ICML 2024poster

In real-world scenarios, arbitrary interactions with the environment can often be costly, and actions of expert demonstrations are not always available. To reduce the need for both, offline Learning from Observations (LfO) is extensively studied: the agent learns to solve a task given only expert st…

2024

Region-Based Representations Revisited

CVPR 2024poster

We investigate whether region-based representations are effective for recognition. Regions were once a mainstay in recognition approaches but pixel and patch-based features are now used almost exclusively. We show that recent class-agnostic segmenters like SAM can be effectively combined with strong…

2024

Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision Transformers

NeurIPS 2024spotlight

Decision Transformers have recently emerged as a new and compelling paradigm for offline Reinforcement Learning (RL), completing a trajectory in an autoregressive way. While improvements have been made to overcome initial shortcomings, online finetuning of decision transformers has been surprisingl…

2024

Robust Model-Based Optimization for Challenging Fitness Landscapes

ICLR 2024poster

Protein design, a grand challenge of the day, involves optimization on a fitness landscape, and leading methods adopt a model-based approach where a model is trained on a training set (protein sequences and fitness) and proposes candidates to explore next. These methods are challenged by sparsity of…

2024

SOHES: Self-supervised Open-world Hierarchical Entity Segmentation

ICLR 2024poster

Open-world entity segmentation, as an emerging computer vision task, aims at segmenting entities in images without being restricted by pre-defined classes, offering impressive generalization capabilities on unseen images and concepts. Despite its promise, existing entity segmentation methods like Se…

2024

TAMM: TriAdapter Multi-Modal Learning for 3D Shape Understanding

CVPR 2024poster

The limited scale of current 3D shape datasets hinders the advancements in 3D shape understanding and motivates multi-modal learning approaches which transfer learned knowledge from data-abundant 2D image and language modalities to 3D shapes. However even though the image and language representation…

2023

A Simple Solution for Offline Imitation from Observations and Examples with Possibly Incomplete Trajectories

NeurIPS 2023poster

Offline imitation from observations aims to solve MDPs where only task-specific expert states and task-agnostic non-expert state-action pairs are available. Offline imitation is useful in real-world scenarios where arbitrary interactions are costly and expert actions are unavailable. The state-of-th…

2023

Contrastive Learning Relies More on Spatial Inductive Bias Than Supervised Learning: An Empirical Study

ICCV 2023poster

Though self-supervised contrastive learning (CL) has shown its potential to achieve state-of-the-art accuracy without any supervision, its behavior still remains under investigated by academia. Different from most previous work that understands CL from learning objectives, we focus on an unexplored…

Cited by 2PDFScholar
2023

Contrastive Mean Teacher for Domain Adaptive Object Detectors

CVPR 2023poster

Object detectors often suffer from the domain gap between training (source domain) and real-world applications (target domain). Mean-teacher self-training is a powerful paradigm in unsupervised domain adaptation for object detection, but it struggles with low-quality pseudo-labels. In this work, we…

2023

Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models

NeurIPS 2023poster

We propose a conceptually simple and lightweight framework for improving the robustness of vision models through the combination of knowledge distillation and data augmentation. We address the conjecture that larger models do not make for better teachers by showing strong gains in out-of-distributio…

2023

DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception

IROS 2023poster

Closing the domain gap between training and deployment and incorporating multiple sensor modalities are two challenging yet critical topics for self-driving. Existing work only focuses on single one of the above topics, overlooking the simultaneous domain and modality shift which pervasively exists…

Cited by 5SourcecodeScholar
2023

HASSOD: Hierarchical Adaptive Self-Supervised Object Detection

NeurIPS 2023poster

The human visual perception system demonstrates exceptional capabilities in learning without explicit supervision and understanding the part-to-whole composition of objects. Drawing inspiration from these two abilities, we propose Hierarchical Adaptive Self-Supervised Object Detection (HASSOD), a no…

2023

Improving Equivariance in State-of-the-Art Supervised Depth and Normal Predictors

ICCV 2023poster

Dense depth and surface normal predictors should possess the equivariant property to cropping-and-resizing -- cropping the input image should result in cropping the same output image. However, we find that state-of-the-art depth and normal predictors, despite having strong performances, surprisingly…

Cited by 1PDFcodeScholar
2023

InterDiff: Generating 3D Human-Object Interactions with Physics-Informed Diffusion

ICCV 2023poster

This paper addresses a novel task of anticipating 3D human-object interactions (HOIs). Most existing research on HOI synthesis lacks comprehensive whole-body interactions with dynamic objects, e.g., often limited to manipulating small or static objects. Our task is significantly more challenging, as…

Cited by 116PDFcodeScholar
2023

Learning Lightweight Object Detectors via Multi-Teacher Progressive Distillation

ICML 2023poster

Resource-constrained perception systems such as edge computing and vision-for-robotics require vision models to be both accurate and lightweight in computation and memory usage. While knowledge distillation is a proven strategy to enhance the performance of lightweight classification models, its app…

2023

NeuralEditor: Editing Neural Radiance Fields via Manipulating Point Clouds

CVPR 2023poster

This paper proposes NeuralEditor that enables neural radiance fields (NeRFs) natively editable for general shape editing tasks. Despite their impressive results on novel-view synthesis, it remains a fundamental challenge for NeRFs to edit the shape of the scene. Our key insight is to exploit the exp…

2023

Object Discovery From Motion-Guided Tokens

CVPR 2023poster

Object discovery -- separating objects from the background without manual labels -- is a fundamental open challenge in computer vision. Previous methods struggle to go beyond clustering of low-level cues, whether handcrafted (e.g., color, texture) or learned (e.g., from auto-encoders). In this work,…

2023

Standing Between Past and Future: Spatio-Temporal Modeling for Multi-Camera 3D Multi-Object Tracking

CVPR 2023poster

This work proposes an end-to-end multi-camera 3D multi-object tracking (MOT) framework. It emphasizes spatio-temporal continuity and integrates both past and future reasoning for tracked objects. Thus, we name it "Past-and-Future reasoning for Tracking" (PF-Track). Specifically, our method adapts th…

2023

YouTubePD: A Multimodal Benchmark for Parkinson’s Disease Analysis

NeurIPS 2023poster

The healthcare and AI communities have witnessed a growing interest in the development of AI-assisted systems for automated diagnosis of Parkinson's Disease (PD), one of the most prevalent neurodegenerative disorders. However, the progress in this area has been significantly impeded by the absence o…

Cited by 3SourcePDFScholar
2022

CEIP: Combining Explicit and Implicit Priors for Reinforcement Learning with Demonstrations

NeurIPS 2022accept

Although reinforcement learning has found widespread use in dense reward settings, training autonomous agents with sparse rewards remains challenging. To address this difficulty, prior work has shown promising results when using not only task-specific demonstrations but also task-agnostic albeit som…

2022

Continual Learning with Evolving Class Ontologies

NeurIPS 2022accept

Lifelong learners must recognize concept vocabularies that evolve over time. A common yet underexplored scenario is learning with class labels that continually refine/expand old classes. For example, humans learn to recognize ${\tt dog}$ before dog breeds. In practical settings, dataset ${\it versio…

Cited by 12SourcePDFScholar
2022

DIVeR: Real-Time and Accurate Neural Radiance Fields With Deterministic Integration for Volume Rendering

CVPR 2022oral

DIVeR builds on the key ideas of NeRF and its variants -- density models and volume rendering -- to learn 3D object models that can be rendered realistically from small numbers of images. In contrast to all previous NeRF methods, DIVeR uses deterministic rather than stochastic estimates of the volum…

Cited by 85PDFcodeScholar
2022

Discovering Objects That Can Move

CVPR 2022poster

This paper studies the problem of object discovery -- separating objects from the background without manual labels. Existing approaches utilize appearance cues, such as color, texture, and location, to group pixels into object-like regions. However, by relying on appearance alone, these methods fail…

Cited by 52PDFcodeScholar
2022

Diverse Human Motion Prediction Guided by Multi-level Spatial-Temporal Anchors

ECCV 2022poster

"Predicting diverse human motions given a sequence of historical poses has received increasing attention. Despite rapid progress, existing work captures the multi-modal nature of human motions primarily through likelihood-based sampling, where the mode collapse has been widely observed. In this pape…

2022

Embracing Single Stride 3D Object Detector With Sparse Transformer

CVPR 2022poster

In LiDAR-based 3D object detection for autonomous driving, the ratio of the object size to input scene size is significantly smaller compared to 2D detection cases. Overlooking this difference, many 3D detectors directly follow the common practice of 2D detectors, which downsample the feature maps e…

Cited by 305PDFcodeScholar
2022

On the Importance of Firth Bias Reduction in Few-Shot Classification

ICLR 2022spotlight

Learning accurate classifiers for novel categories from very few examples, known as few-shot image classification, is a challenging task in statistical machine learning and computer vision. The performance in few-shot classification suffers from the bias in the estimation of classifier parameters; h…

2021

Bowtie Networks: Generative Modeling for Joint Few-Shot Recognition and Novel-View Synthesis

ICLR 2021poster

We propose a novel task of joint few-shot recognition and novel-view synthesis: given only one or few images of a novel object from arbitrary views with only category annotation, we aim to simultaneously learn an object classifier and generate images of that type of object from new viewpoints. While…

2021

DAP: Detection-Aware Pre-Training With Weak Supervision

CVPR 2021poster

This paper presents a detection-aware pre-training (DAP) approach, which leverages only weakly-labeled classification-style datasets (e.g., ImageNet) for pre-training, but is specifically tailored to benefit object detection tasks. In contrast to the widely used image classification-based pre-traini…

Cited by 21PDFcodeScholar
2021

Image-Level or Object-Level? A Tale of Two Resampling Strategies for Long-Tailed Detection

ICML 2021spotlight

Training on datasets with long-tailed distributions has been challenging for major recognition tasks such as classification and detection. To deal with this challenge, image resampling is typically introduced as a simple but effective approach. However, we observe that long-tailed detection differs…

2021

Learning To Hallucinate Examples From Extrinsic and Intrinsic Supervision

ICCV 2021poster

Learning to hallucinate additional examples has recently been shown as a promising direction to address few-shot learning tasks. This work investigates two important yet overlooked natural supervision signals for guiding the hallucination process -- (i) extrinsic: classifiers trained on hallucinated…

Cited by 8PDFScholar
2021

Pixel Contrastive-Consistent Semi-Supervised Semantic Segmentation

ICCV 2021poster

We present a novel semi-supervised semantic segmentation method which jointly achieves two desiderata of segmentation model regularities: the label-space consistency property between image augmentations and the feature-space contrastive property among different pixels. We leverage the pixel-level L2…

Cited by 226PDFScholar
2019

Image Deformation Meta-Networks for One-Shot Learning

CVPR 2019oral

Humans can robustly learn novel visual concepts even when images undergo various deformations and loose certain information. Mimicking the same behavior and synthesizing deformed instances of new concepts may help visual recognition systems perform better one-shot learning, i.e., learning concepts f…

Cited by 303PDFcodeScholar
2018

Adversarial Geometry-Aware Human Motion Prediction

ECCV 2018poster

We explore an approach to forecasting human motion in a few milliseconds given an input 3D skeleton sequence based on a recurrent encoder-decoder framework. Current approaches suffer from the problem of prediction discontinuities and may fail to predict human-like motion in longer time horizons due…

Cited by 317SourcePDFScholar
2018

Few-Shot Human Motion Prediction via Meta-Learning

ECCV 2018poster

Human motion prediction, forecasting human motion in a few milliseconds conditioning on a historical 3D skeleton sequence, is a long-standing problem in computer vision and robotic vision. Existing forecasting algorithms rely on extensive annotated motion capture data and are brittle to novel action…

Cited by 155SourcePDFScholar
2018

Teaching Robots to Predict Human Motion

IROS 2018poster

Teaching a robot to predict and mimic how a human moves or acts in the near future by observing a series of historical human movements is a crucial first step in human-robot interaction and collaboration. In this paper, we instrument a robot with such a prediction ability by leveraging recent deep l…

Cited by 137SourceScholar
2016

Learning from Small Sample Sets by Combining Unsupervised Meta-Training with CNNs

NeurIPS 2016accepted

This work explores CNNs for the recognition of novel categories from few examples. Inspired by the transferability properties of CNNs, we introduce an additional unsupervised meta-training stage that exposes multiple top layer units to a large amount of unlabeled real-world images. By encouraging th…

Cited by 94SourcePDFScholar