← Search

Stella X. Yu

45 accepted papers

2026

GeoSANE: Learning Geospatial Representations from Models, Not Data

CVPR 2026

Recent advances in remote sensing have led to an increase in the number of available foundation models; each trained on different modalities, datasets, and objectives, yet capturing only part of the vast geospatial knowledge landscape. While these models show strong results within their respective d

Cited by 0SourcecodeScholar
2025

Novel View Synthesis from A Few Glimpses via Test-Time Natural Video Completion

NeurIPS 2025poster

Given just a few glimpses of a scene, can you imagine the movie playing out as the camera glides through it? That’s the lens we take on sparse-input novel view synthesis, not only as filling spatial gaps between widely spaced views, but also as completing a natural video unfolding through space. We…

Cited by 0SourcecodeScholar
2025

Open Ad-hoc Categorization with Contextualized Feature Learning

CVPR 2025poster

Adaptive categorization of visual scenes is essential for AI agents to handle changing tasks. Unlike fixed common categories for plants or animals, ad-hoc categories, such as things to sell at a garage sale, are created dynamically to achieve specific tasks. We study open ad-hoc categorization, wh…

2025

Test-Time Canonicalization by Foundation Models for Robust Perception

ICML 2025poster

Real-world visual perception requires invariance to diverse transformations, yet current methods rely heavily on specialized architectures or training on predefined augmentations, limiting generalization. We propose FoCal, a test-time, data-driven framework that achieves robust perception by leverag…

2025

Visually Consistent Hierarchical Image Classification

ICLR 2025poster

Hierarchical classification predicts labels across multiple levels of a taxonomy, e.g., from coarse-level \textit{Bird} to mid-level \textit{Hummingbird} to fine-level \textit{Green hermit}, allowing flexible recognition under varying visual conditions. It is commonly framed as multiple single-leve…

Cited by 0SourcePDFScholar
2024

Learning Hierarchical Image Segmentation For Recognition and By Recognition

ICLR 2024spotlight

Large vision and language models learned directly through image-text associations often lack detailed visual substantiation, whereas image segmentation tasks are treated separately from recognition, supervisedly learned without interconnections. Our key observation is that, while an image can be r…

2024

Pose-Aware Self-Supervised Learning with Viewpoint Trajectory Regularization

ECCV 2024oral

"Learning visual features from unlabeled images has proven successful for semantic categorization, often by mapping different views of the same object to the same feature to achieve recognition invariance. However, visual recognition involves not only identifying what an object is but also understan…

2024

Unsupervised Feature Learning with Emergent Data-Driven Prototypicality

CVPR 2024poster

Given a set of images our goal is to map each image to a point in a feature space such that not only point proximity indicates visual similarity but where it is located directly encodes how prototypical the image is according to the dataset. Our key insight is to perform unsupervised feature learnin…

Cited by 4SourcePDFScholar
2023

Bootstrapping Objectness From Videos by Relaxed Common Fate and Visual Grouping

CVPR 2023poster

We study learning object segmentation from unlabeled videos. Humans can easily segment moving objects without knowing what they are. The Gestalt law of common fate, i.e., what move at the same speed belong together, has inspired unsupervised object discovery based on motion segmentation. However, co…

2023

Cut and Learn for Unsupervised Object Detection and Instance Segmentation

CVPR 2023poster

We propose Cut-and-LEaRn (CutLER), a simple approach for training unsupervised object detection and segmentation models. We leverage the property of self-supervised models to 'discover' objects without supervision and amplify it to train a state-of-the-art localization model without any human labels…

2023

Learning to Transform for Generalizable Instance-wise Invariance

ICCV 2023poster

Computer vision research has long aimed to build systems that are robust to transformations found in natural data. Traditionally, this is done using data augmentation or hard-coding invariances into the architecture. However, too much or too little invariance can hurt, and the correct amount is un…

Cited by 1PDFcodeScholar
2023

ResoNet: Noise-Trained Physics-Informed MRI Off-Resonance Correction

NeurIPS 2023poster

Magnetic Resonance Imaging (MRI) is a powerful medical imaging modality that offers diagnostic information without harmful ionizing radiation. Unlike optical imaging, MRI sequentially samples the spatial Fourier domain (k-space) of the image. Measurements are collected in multiple shots, or readout…

2023

The Audio-Visual BatVision Dataset for Research on Sight and Sound

IROS 2023poster

Vision research showed remarkable success in understanding our world, propelled by datasets of images and videos. Sensor data from radar, LiDAR and cameras supports research in robotics and autonomous driving for at least a decade. However, while visual sensors may fail in some conditions, sound has…

Cited by 7SourcecodeScholar
2022

Unsupervised Hierarchical Semantic Segmentation With Multiview Cosegmentation and Clustering Transformers

CVPR 2022oral

Unsupervised semantic segmentation aims to discover groupings within and across images that capture object- and view-invariance of a category without external supervision. Grouping naturally has levels of granularity, creating ambiguity in unsupervised segmentation. Existing methods avoid this ambig…

Cited by 59PDFcodeScholar
2022

Unsupervised Selective Labeling for More Effective Semi-Supervised Learning

ECCV 2022poster

"Given an unlabeled dataset and an annotation budget, we study how to selectively label a fixed number of instances so that semi-supervised learning (SSL) on such a partially labeled dataset is most effective. We focus on selecting the right data to label, in addition to usual SSL’s propagating labe…

2021

Unsupervised Discriminative Learning of Sounds for Audio Event Classification

ICASSP 2021accepted

Recent progress in network-based audio event classification has shown the benefit of pre-training models on visual data such as ImageNet. While this process allows knowledge transfer across different domains, training a model on large-scale visual datasets is time consuming. On several audio event c…

Cited by 5SourceScholar
2020

BatVision: Learning to See 3D Spatial Layout with Two Ears

ICRA 2020poster

Many species have evolved advanced non-visual perception while artificial systems fall behind. Radar and ultrasound complement camera-based vision but they are often too costly and complex to set up for very limited information gain. In nature, sound is used effectively by bats, dolphins, whales, an…

Cited by 0SourcecodeScholar
2019

Adversarial Structure Matching for Structured Prediction Tasks

CVPR 2019poster

Pixel-wise losses, i.e., cross-entropy or L2, have been widely used in structured prediction tasks as a spatial extension of generic image classification or regression. However, its i.i.d. assumption neglects the structural regularity present in natural images. Various attempts have been made to inc…

Cited by 18PDFcodeScholar
2019

Large-Scale Long-Tailed Recognition in an Open World

CVPR 2019oral

Real world data often have a long-tailed and open-ended distribution. A practical recognition system must classify among majority and minority classes, generalize from a few known instances, and acknowledge novelty upon a never seen instance. We define Open Long-Tailed Recognition (OLTR) as learning…

Cited by 1471PDFcodeScholar
2019

SegSort: Segmentation by Discriminative Sorting of Segments

ICCV 2019poster

Almost all existing deep learning approaches for semantic segmentation tackle this task as a pixel-wise classification problem. Yet humans understand a scene not in terms of pixels, but by decomposing it into perceptual groups and structures that are the basic building blocks of recognition. This mo…

Cited by 165PDFScholar
2018

Improving Generalization via Scalable Neighborhood Component Analysis

ECCV 2018poster

Current visual recognition is dominated by the end-to-end formulation of classification problems implemented by the parametric softmax classifiers. Such formulation makes a closed world assumption with a fixed set of categories. This becomes problematic for open-set scenarios where new categories ar…

2018

Unsupervised Feature Learning via Non-Parametric Instance Discrimination

CVPR 2018poster

Neural net classifiers trained on data with annotated class labels can also capture apparent visual similarity among categories without being directed to do so. We study whether this observation can be extended beyond the conventional domain of supervised learning: Can we learn a good feature repres…

2017

Am I a Baller? Basketball Performance Assessment From First-Person Videos

ICCV 2017poster

This paper presents a method to assess a basketball player's performance from his/her first-person video. A key challenge lies in the fact that the evaluation metric is highly subjective and specific to a particular evaluator. We leverage the first-person camera to address this challenge. The spatio…

Cited by 106PDFScholar
2017

Convolutional Random Walk Networks for Semantic Image Segmentation

CVPR 2017poster

Most current semantic segmentation methods rely on fully convolutional networks (FCNs). However, their use of large receptive fields and many pooling layers cause low spatial resolution inside the deep layers. This leads to predictions with poor localization around the boundaries. Prior work has att…

Cited by 175PDFScholar
2017

Unsupervised Learning of Important Objects From First-Person Videos

ICCV 2017poster

A first-person camera, placed at a person's head, captures, which objects are important to the camera wearer. Most prior methods for this task learn to detect such important objects from the manually labeled first-person data in a supervised fashion. However, important objects are strongly related t…

Cited by 33PDFScholar
2016

Affinity CNN: Learning Pixel-Centric Pairwise Relations for Figure/Ground Embedding

CVPR 2016spotlight

Spectral embedding provides a framework for solving perceptual organization problems, including image segmentation and figure/ground organization. From an affinity matrix describing pairwise relationships between pixels, it clusters pixels into regions, and, using a complex-valued extension, orders…

Cited by 76PDFScholar
2015

Direct Intrinsics: Learning Albedo-Shading Decomposition by Convolutional Regression

ICCV 2015poster

We introduce a new approach to intrinsic image decomposition, the task of decomposing a single image into albedo and shading components. Our strategy, which we term direct intrinsics, is to learn a convolutional neural network (CNN) that directly predicts output albedo and shading channels from an…

Cited by 216PDFScholar
2015

FlowWeb: Joint Image Set Alignment by Weaving Consistent, Pixel-Wise Correspondences

CVPR 2015poster

Given a set of poorly aligned images of the same visual concept without any annotations, we propose an algorithm to jointly bring them into pixel-wise correspondence by estimating a FlowWeb representation of the image set. FlowWeb is a fully-connected correspondence flow graph with each node represe…

Cited by 186SourcePDFScholar