← Search

Jianbo Shi

34 accepted papers

2026

Vibe Spaces for Creatively Connecting and Expressing Visual Concepts

CVPR 2026

Creating new visual concepts often requires connecting distinct ideas through their most relevant shared attributes--their vibe. We introduce Vibe Blending, a novel task for generating coherent and meaningful hybrids that reveals these shared attributes between images. Achieving such blends is chall

Cited by 0SourcecodeScholar
2025

CoralSRT: Revisiting Coral Reef Semantic Segmentation by Feature Rectification via Self-supervised Guidance

ICCV 2025poster

We investigate coral reef semantic segmentation, in which coral reefs are governed by multifaceted factors, like genes, environmental changes, and internal interactions. Unlike segmenting structural units/instances, which are predictable and follow a set pattern, also referred to as commonsense or p…

Cited by 0SourcePDFScholar
2024

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

CVPR 2024poster

We present Ego-Exo4D a diverse large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g. sports music dance bike repair). 740 participants from 13 cities worldwide perform…

2023

Beyond mAP: Towards Better Evaluation of Instance Segmentation

CVPR 2023highlight

Correctness of instance segmentation constitutes counting the number of objects, correctly localizing all predictions and classifying each localized prediction. Average Precision is the de-facto metric used to measure all these constituents of segmentation. However, this metric does not penalize dup…

Cited by 10SourcePDFScholar
2023

Perceptual Artifacts Localization for Image Synthesis Tasks

ICCV 2023poster

Recent advancements in deep generative models have facilitated the creation of photo-realistic images across various tasks. However, these generated images often exhibit perceptual artifacts in specific regions, necessitating manual correction. In this study, we present a comprehensive empirical exa…

Cited by 23PDFcodeScholar
2023

Starting From Non-Parametric Networks for 3D Point Cloud Analysis

CVPR 2023poster

We present a Non-parametric Network for 3D point cloud analysis, Point-NN, which consists of purely non-learnable components: farthest point sampling (FPS), k-nearest neighbors (k-NN), and pooling operations, with trigonometric functions. Surprisingly, it performs well on various 3D tasks, requiring…

2023

iQuery: Instruments As Queries for Audio-Visual Sound Separation

CVPR 2023poster

Current audio-visual separation methods share a standard architecture design where an audio encoder-decoder network is fused with visual encoding features at the encoder bottleneck. This design confounds the learning of multi-modal feature encoding with robust sound decoding for audio separation. To…

2022

"Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications"

ECCV 2022poster

"Egocentric videos offer fine-grain information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding viewer’s behaviors and intentions. We provide a labeled dataset consisting of 11,235 egocentric images with per-pixel segmentation labe…

2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Inpainting at Modern Camera Resolution by Guided PatchMatch with Auto-Curation

ECCV 2022poster

"Recently, deep models have established SOTA performance for low-resolution image inpainting, but they lack fidelity at resolutions associated with modern cameras such as 4K or more, and for large holes. We contribute an inpainting benchmark dataset of photos at 4K and above representative of modern…

Cited by 9SourcePDFScholar
2022

Perceptual Artifacts Localization for Inpainting

ECCV 2022poster

"Image inpainting is an essential task for multiple practical applications like object removal and image editing. Deep GAN-based models greatly improve the inpainting performance in structures and textures within the hole, but might also generate unexpected artifacts like broken structures or color…

2020

Learning Object Placement by Inpainting for Compositional Data Augmentation

ECCV 2020poster

We study the problem of common sense placement of visual objects in an image.  This involves multiple aspects of visual recognition: the instance segmentation of the scene, 3D layout, and common knowledge of how objects are placed and where objects are moving in the 3D scene. This seemingly simple t…

Cited by 78SourcePDFScholar
2020

Nested Scale-Editing for Conditional Image Synthesis

CVPR 2020poster

We propose an image synthesis approach that provides stratified navigation in the latent code space. With a tiny amount of partial or very low-resolution image, our approach can consistently out-perform state-of-the-art counterparts in terms of generating the closest sampled image to the ground trut…

Cited by 7PDFScholar
2019

Adversarial Structure Matching for Structured Prediction Tasks

CVPR 2019poster

Pixel-wise losses, i.e., cross-entropy or L2, have been widely used in structured prediction tasks as a spatial extension of generic image classification or regression. However, its i.i.d. assumption neglects the structural regularity present in natural images. Various attempts have been made to inc…

Cited by 18PDFcodeScholar
2019

Learning Temporal Pose Estimation from Sparsely-Labeled Videos

NeurIPS 2019poster

Modern approaches for multi-person pose estimation in video require large amounts of dense annotations. However, labeling every frame in a video is costly and labor intensive. To reduce the need for dense annotations, we propose a PoseWarper network that leverages training videos with sparse annotat…

2019

SegSort: Segmentation by Discriminative Sorting of Segments

ICCV 2019poster

Almost all existing deep learning approaches for semantic segmentation tackle this task as a pixel-wise classification problem. Yet humans understand a scene not in terms of pixels, but by decomposing it into perceptual groups and structures that are the basic building blocks of recognition. This mo…

Cited by 165PDFScholar
2019

Zoom-In-To-Check: Boosting Video Interpolation via Instance-Level Discrimination

CVPR 2019poster

We propose a light-weight video frame interpolation algorithm. Our key innovation is an instance-level supervision that allows information to be learned from the high-resolution version of similar objects. Our experiment shows that the proposed method can generate state-of-the-art results across di…

Cited by 31PDFScholar
2017

Am I a Baller? Basketball Performance Assessment From First-Person Videos

ICCV 2017poster

This paper presents a method to assess a basketball player's performance from his/her first-person video. A key challenge lies in the fact that the evaluation metric is highly subjective and specific to a particular evaluator. We leverage the first-person camera to address this challenge. The spatio…

Cited by 106PDFScholar
2017

Convolutional Random Walk Networks for Semantic Image Segmentation

CVPR 2017poster

Most current semantic segmentation methods rely on fully convolutional networks (FCNs). However, their use of large receptive fields and many pooling layers cause low spatial resolution inside the deep layers. This leads to predictions with poor localization around the boundaries. Prior work has att…

Cited by 175PDFScholar
2017

Unsupervised Learning of Important Objects From First-Person Videos

ICCV 2017poster

A first-person camera, placed at a person's head, captures, which objects are important to the camera wearer. Most prior methods for this task learn to detect such important objects from the manually labeled first-person data in a supervised fashion. However, important objects are strongly related t…

Cited by 33PDFScholar
2015

DeepEdge: A Multi-Scale Bifurcated Deep Network for Top-Down Contour Detection

CVPR 2015poster

Contour detection has been a fundamental component in many image segmentation and object detection systems. Most previous work utilizes low-level features such as texture or saliency to detect contours and then use them as cues for a higher-level task such as object detection. However, we claim that…

Cited by 648SourcePDFScholar
2015

High-for-Low and Low-for-High: Efficient Boundary Detection From Deep Object Features and its Applications to High-Level Vision

ICCV 2015poster

Most of the current boundary detection systems rely exclusively on low-level features, such as color and texture. However, perception studies suggest that humans employ object-level reasoning when judging if a particular pixel is a boundary. Inspired by this observation, in this work we show how to…

Cited by 227PDFScholar