← Search

Jin Sun

12 accepted papers

2025

Concept-Centric Token Interpretation for Vector-Quantized Generative Models

ICML 2025poster

Vector-Quantized Generative Models (VQGMs) have emerged as powerful tools for image generation. However, the key component of VQGMs---the codebook of discrete tokens---is still not well understood, e.g., which tokens are critical to generate an image of a certain concept? This paper introduces Conce…

2025

Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data

ICLR 2025poster

Large Multimodal Models (LMMs), or Vision-Language Models (VLMs), have shown impressive capabilities in a wide range of visual tasks. However, they often struggle with fine-grained visual reasoning, failing to identify domain-specific objectives and provide justifiable explanations for their predict…

2024

Neural Gaffer: Relighting Any Object via Diffusion

NeurIPS 2024poster

Single-image relighting is a challenging task that involves reasoning about the complex interplay between geometry, materials, and lighting. Many prior methods either support only specific categories of images, such as portraits, or require special capture conditions, like using a flashlight. Altern…

Cited by 14SourcePDFScholar
2023

Black-box Backdoor Defense via Zero-shot Image Purification

NeurIPS 2023poster

Backdoor attacks inject poisoned samples into the training data, resulting in the misclassification of the poisoned input during a model's deployment. Defending against such attacks is challenging, especially for real-world black-box models where only query access is permitted. In this paper, we pro…

2021

Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal Vision

ICCV 2021poster

The abundance and richness of Internet photos of landmarks and cities has led to significant progress in 3D vision over the past two decades, including automated 3D reconstructions of the world's landmarks from tourist photos. However, a major source of information available for these 3D-augmented c…

Cited by 20PDFcodeScholar
2020

Hidden Footprints: Learning Contextual Walkability from 3D Human Trails

ECCV 2020poster

Predicting where people can walk in a scene is important for many tasks, including autonomous driving systems and human behavior analysis. Yet learning a computational model for this purpose is challenging due to semantic ambiguity and a lack of labeled data: current datasets only have labels on whe…

2018

Label Denoising Adversarial Network (LDAN) for Inverse Lighting of Faces

CVPR 2018poster

Lighting estimation from faces is an important task and has applications in many areas such as image editing, intrinsic image decomposition, and image forgery detection. We propose to train a deep Convolutional Neural Network (CNN) to regress lighting parameters from a single face image. Lacking mas…

Cited by 23SourcePDFScholar
2017

Generating Holistic 3D Scene Abstractions for Text-Based Image Retrieval

CVPR 2017poster

Spatial relationships between objects provide important information for text-based image retrieval. As users are more likely to describe a scene from a real world perspective, using 3D spatial relationships rather than 2D relationships that assume a particular viewing direction, one of the main chal…

Cited by 35PDFScholar