← Search

Yao Zhao

111 accepted papers

2026

CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style Transfer

ICLR 2026poster

Transferring visual style between images while preserving semantic correspondence between similar objects remains a central challenge in computer vision. While existing methods have made great strides, most of them operate at global level but overlook region-wise and even pixel-wise semantic corresp…

Cited by 0SourceScholar
2026

Fixed Budget is No Harder Than Fixed Confidence in Best-Arm Identification up to Logarithmic Factors

ICML 2026poster

The best-arm identification (BAI) problem is one of the most fundamental problems in interactive machine learning, which has two flavors: the fixed-budget setting (FB) and the fixed-confidence setting (FC). For $K$-armed bandits with the unique best arm, the optimal sample complexities for both sett…

Cited by 0SourceScholar
2026

HFR and HDR Video from Multi-Attenuated Spikes Using a Rapidly Rotating SpokeND Filter

CVPR 2026

Capturing scenes with both high dynamic range (HDR) and high-speed motion remains challenging for conventional cameras. Existing alternating-exposure approaches exacerbate temporal resolution loss, making them unsuitable for high-speed scenes. Consequently, current solutions typically compromise eit

Cited by 0SourceScholar
2026

Hunting Normality from Query Sample via Residual Learning for Generalist Anomaly Detection

CVPR 2026

Generalist Anomaly Detection (GAD) seeks to overcome the domain-specific limitations of traditional anomaly detection by training a unified model that can generalize to unseen classes. A promising GAD strategy involves using residual features to create a class-invariant space. However, existing meth

Cited by 0SourceScholar
2026

PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion

ICML 2026spotlight

Achieving a complete and explorable 360-degree visual world is a cornerstone of immersive content creation. While recent advances in video generation have achieved impressive results, they follow a 2D paradigm that treats content generation as transitions of 2D pixels, lacking an intrinsic understan…

Cited by 0SourceScholar
2026

RAIN: Redundancy-Aware Latent Injection for Quality-Preserving Image Watermarking

AAAI 2026technical

Diffusion models have gained widespread adoption due to their ability to generate highly realistic images, yet their rapid proliferation also raises security and traceability concerns. To address issues of ownership verification and accountability, current watermarking technique

Cited by 0SourcePDFScholar
2026

Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images

ICLR 2026poster

The rapid advancement of AI-generated content (AIGC) has enabled the synthesis of visually convincing images; however, many such outputs exhibit subtle \textbf{semantic anomalies}, including unrealistic object configurations, violations of physical laws, or commonsense inconsistencies, which comprom…

Cited by 0SourceScholar
2026

StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation

CVPR 2026

The growing adoption of XR devices has fueled strong demand for high-quality stereo video, yet its production remains costly and artifact-prone.To address this challenge, we present **StereoWorld**, an **end-to-end framework** that repurposes a pretrained video generator for high-fidelity monocular-

Cited by 0SourceScholar
2026

Subspace-Aware Feature Reshaping for Open-Set Graph Class-Incremental Learning

ICML 2026poster

Graph class-incremental learning (GCIL) has emerged to address the challenge of learning from dynamically evolving graphs, which continuously learns new classes over a sequence of tasks while retaining performance on previously seen classes. However, existing GCIL methods assume a closed-set test di…

Cited by 0SourceScholar
2026

ThinkGen: Generalized Thinking for Visual Generation

CVPR 2026

Recent progress in Multimodal Large Language Models (MLLMs) demonstrates that Chain-of-Thought (CoT) reasoning enables systematic solutions to complex understanding tasks. However, its extension to generation tasks remains nascent and limited by scenario-specific mechanisms that hinder generalizatio

Cited by 0SourcecodeScholar
2026

VideoWorld 2: Learning Transferable Knowledge from Real-world Videos

CVPR 2026

Learning transferable knowledge from unlabeled video data and applying it in new environments is a fundamental capability of intelligent agents. This work presents VideoWorld 2, which extends VideoWorld and provides the first investigation of learning transferable knowledge for complex, long-horizon

Cited by 0SourceScholar
2025

ALIC: Adaptive Fusion Entropy Model for Learned Image Compression

ICASSP 2025accepted

Recently, learned image compression algorithms have achieved significant performance. The entropy model is crucial for improving the rate-distortion performance by estimating the probability distribution of latent representation. In this paper, we propose an adaptive fusion entropy model for learned…

Cited by 0SourceScholar
2025

Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning

AAAI 2025technical

Zero-shot learning (ZSL) endeavors to transfer knowledge from the seen categories to recognize unseen categories, which mostly relies on the semantic-visual interactions between image and attribute tokens. Recently, the prompt learning has emerged in ZSL and demonstrated significant potential as it…

2025

C2P-CLIP: Injecting Category Common Prompt in CLIP to Enhance Generalization in Deepfake Detection

AAAI 2025technical

This work focuses on AIGC detection to develop universal detectors capable of identifying various types of forgery images. Recent studies have found large pre-trained models, such as CLIP, are effective for generalizable deepfake detection along with linear classifiers. However, two critical issues…

2025

CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting

ICCV 2025poster

Recent works in 3D representation learning and multimodal pre-training have made remarkable progress. However, typically multimodal 3D models are only capable of handling point clouds. Compared to the emerging 3D representation technique, 3D Gaussian Splatting (3DGS), the spatially sparse point clou…

Cited by 0SourcePDFScholar
2025

CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation

AAAI 2025technical

The emergence of long-context text applications utilizing large language models (LLMs) has presented significant scalability challenges, particularly in memory footprint. The linear growth of the Key-Value (KV) cache, which stores attention keys and values to reduce redundant computations, can signi…

Cited by 1SourcePDFScholar
2025

CharaConsist: Fine-Grained Consistent Character Generation

ICCV 2025poster

In text-to-image generation, producing a series of consistent contents that preserve the same identity is highly valuable for real-world applications. Although a few works have explored training-free methods to enhance the consistency of generated subjects, we observe that they suffer from the follo…

2025

ClassDiffusion: More Aligned Personalization Tuning with Explicit Class Guidance

ICLR 2025poster

Recent text-to-image customization works have proven successful in generating images of given concepts by fine-tuning diffusion models on a few examples. However, tuning-based methods inherently tend to overfit the concepts, resulting in failure to create the concept under multiple conditions (*e.g.…

2025

DATA-VSR: Dynamic Trajectory Attention and Texture Adaptive Rooter for Video Super-Resolution

ICASSP 2025accepted

Video Super-Resolution (VSR) is essential for reconstructing high-definition sequences from correlated video frames. While Transformer-based VSR methods have improved reconstruction quality, they require substantial computational resources, limiting deployment on resource-constrained devices. To tac…

Cited by 0SourceScholar
2025

DCI: Dual-Conditional Inversion for Boosting Diffusion-Based Image Editing

NeurIPS 2025poster

Diffusion models have achieved remarkable success in image generation and editing tasks. Inversion within these models aims to recover the latent noise representation for a real or generated image, enabling reconstruction, editing, and other downstream tasks. However, to date, most inversion approac…

Cited by 0SourcecodeScholar
2025

DIDS: Domain Impact-aware Data Sampling for Large Language Model Training

EMNLP 2025

Large language models (LLMs) are commonly trained on multi-domain datasets, where domain sampling strategies significantly impact model performance due to varying domain importance across downstream tasks. Existing approaches for optimizing domain-level sampling strategies struggle with maintaining

2025

Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans?

CVPR 2025poster

To detect prohibited items in challenging categories, human inspectors typically rely on images from two distinct views (vertical and side). Can AI detect prohibited items from dual-view X-ray images in the same way humans do? Existing X-ray datasets often suffer from limitations, such as single-vie…

2025

EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with Events

CVPR 2025highlight

Continuous space-time video super-resolution (C-STVSR) endeavors to upscale videos simultaneously at arbitrary spatial and temporal scales, which has recently garnered increasing interest. However, prevailing methods struggle to yield satisfactory videos at out-of-distribution spatial and temporal s…

2025

Fixing the Loose Brake: Exponential-Tailed Stopping Time in Best Arm Identification

ICML 2025poster

The best arm identification problem requires identifying the best alternative (i.e., arm) in active experimentation using the smallest number of experiments (i.e., arm pulls), which is crucial for cost-efficient and timely decision-making processes. In the fixed confidence setting, an algorithm must…

Cited by 0SourcePDFScholar
2025

FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction

NeurIPS 2025poster

This work challenges the residual prediction paradigm in visual autoregressive modeling and presents FlexVAR, a new Flexible Visual AutoRegressive image generation paradigm. FlexVAR facilitates autoregressive learning with ground-truth prediction, enabling each step to independently produce plausibl…

Cited by 0SourcecodeScholar
2025

Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation

NeurIPS 2025poster

In this paper, we propose \textbf{Jasmine}, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD’s visual priors to enhance the sharpness and generalization of unsupervised prediction. Previous SD-based methods are all supervi…

Cited by 0SourceScholar
2025

Learning from Disjoint Views: A Contrastive Prototype Matching Network for Fully Incomplete Multi-View Clustering

NeurIPS 2025poster

Multi-view clustering aims to enhance clustering performance by leveraging information from diverse sources. However, its practical application is often hindered by a barrier: the lack of correspondences across views. This paper focuses on the understudied problem of fully incomplete multi-view clus…

Cited by 0SourceScholar
2025

LiPO: Listwise Preference Optimization through Learning-to-Rank

NAACL 2025long

Aligning language models (LMs) with curated human feedback is critical to control their behaviors in real-world applications. Several recent policy optimization methods, such as DPO and SLiC, serve as promising alternatives to the traditional Reinforcement Learning from Human Feedback (RLHF) approac…

Cited by 44SourcePDFScholar
2025

Making RALM Robust to Irrelevant Contexts via Layer Knowledge Guided Attention

ACL 2025finding

Retrieval-augmented language models (RALMs) aim to incorporate external knowledge to address the issues of factual hallucination and knowledge obsolescence faced by large language models (LLMs). Inevitably, the retrieved passages based on similarity search may be irrelevant to the given question, an…

2025

Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions

NeurIPS 2025poster

The synthesis of realistic Martian landscape videos, essential for mission rehearsal and robotic simulation, presents unique challenges. These primarily stem from the scarcity of high-quality Martian data and the significant domain gap relative to terrestrial imagery. To address these challenges, we…

Cited by 0SourceScholar
2025

Memory Efficient Matting with Adaptive Token Routing

AAAI 2025technical

Transformer-based models have recently achieved outstanding performance in image matting. However, their application to high-resolution images remains challenging due to the quadratic complexity of global self-attention. To address this issue, we propose MEMatte, a memory-efficient matting framework…

2025

NTClick: Achieving Precise Interactive Segmentation With Noise-tolerant Clicks

CVPR 2025highlight

Interactive segmentation is a pivotal task in computer vision, focused on predicting precise masks with minimal user input. Although the click has recently become the most prevalent form of interaction due to its flexibility and efficiency, its advantages diminish as the complexity and details of ta…

Cited by 0SourcePDFScholar
2025

ODDN: Addressing Unpaired Data Challenges in Open-World Deepfake Detection on Online Social Networks

AAAI 2025technical

Despite significant advances in deepfake detection, handling varying image quality, especially due to different compressions on online social networks (OSNs), remains challenging. Current methods succeed by leveraging correlations between paired images, whether raw or compressed. However, in open-wo…

2025

Once-for-All: Controllable Generative Image Compression with Dynamic Granularity Adaptation

ICLR 2025poster

Although recent generative image compression methods have demonstrated impressive potential in optimizing the rate-distortion-perception trade-off, they still face the critical challenge of flexible rate adaptation to diverse compression necessities and scenarios. To overcome this challenge, this pa…

Cited by 1SourcePDFScholar
2025

PixelStitch: Structure-Preserving Pixel-Wise Bidirectional Warps for Unsupervised Image Stitching

ICCV 2025poster

We propose PixelStitch, a pixel-wise bidirectional warp that learns to stitch images as well as preserve structure in an unsupervised paradigm. To produce natural stitched images, we first determine the middle plane through homography decomposition and globally project the original images toward the…

2025

PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling

NeurIPS 2025poster

Audio-visual event parsing plays a crucial role in understanding multimodal video content, but existing methods typically rely on offline processing of entire videos with huge model sizes, limiting their real-time applicability. We introduce Online Audio-Visual Event Parsing (On-AVEP), a novel parad…

Cited by 0SourcecodeScholar
2025

ReCoT: Reflective Self-Correction Training for Mitigating Confirmation Bias in Large Vision-Language Models

ICCV 2025poster

Recent advancements in Large Vision-Language Models (LVLMs) have greatly improved their ability to understand both visual and text information. However, a common problem in LVLMs is confirmation bias, where models tend to repeat previous assumptions and follow earlier viewpoints instead of reflectin…

Cited by 0SourcePDFScholar
2025

TokMan:Tokenize Manhattan Mask Optimization for Inverse Lithography

NeurIPS 2025poster

Manhattan representations, defined by axis-aligned, orthogonal structures, are widely used in vision, robotics, and semiconductor design for their geometric regularity and algorithmic simplicity. In integrated circuit (IC) design, Manhattan geometry is key for routing, design rule checking, and lith…

Cited by 0SourceScholar
2025

Towards Pre-trained Graph Condensation via Optimal Transport

NeurIPS 2025poster

Graph condensation (GC) aims to distill the original graph into a small-scale graph, mitigating redundancy and accelerating GNN training. However, conventional GC approaches heavily rely on rigid GNNs and task-specific supervision. Such a dependency severely restricts their reusability and generaliz…

Cited by 0SourceScholar
2025

Unifying Reconstruction and Density Estimation via Invertible Contraction Mapping in One-Class Classification

NeurIPS 2025poster

Due to the difficulty in collecting all unexpected abnormal patterns, One-Class Classification (OCC) has become the most popular approach to anomaly detection (AD). Reconstruction-based AD method relies on the discrepancy between inputs and reconstructed results to identify unobserved anomalies. How…

Cited by 0SourceScholar
2025

Unsupervised Region-Based Image Editing of Denoising Diffusion Models

AAAI 2025technical

Although diffusion models have achieved remarkable success in the field of image generation, their latent space remains under-explored. Current methods for identifying semantics within latent space often rely on external supervision, such as textual information and segmentation masks. In this paper,…

Cited by 0SourcePDFScholar
2025

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

CVPR 2025poster

This work explores whether a deep generative model can learn complex knowledge solely from visual input, in contrast to the prevalent focus on text-based models like large language models (LLMs). We develop VideoWorld, an auto-regressive video generation model trained on unlabeled video data, and te…

Cited by 8SourcePDFScholar
2025

Visual Relation Diffusion for Human-Object Interaction Detection

ICCV 2025poster

Human-object interaction (HOI) detection relies on fine-grained visual understanding to distinguish complex relationships between humans and objects. While recent generative diffusion models have demonstrated remarkable capability in learning detailed visual concepts through pixel-level generation,…

Cited by 0SourcePDFScholar
2024

Diffusion for Natural Image Matting

ECCV 2024poster

"Existing natural image matting algorithms inevitably have flaws in their predictions on difficult cases, and their one-step prediction manner cannot further correct these errors. In this paper, we investigate a multi-step iterative approach for the first time to tackle the challenging natural image…

2024

Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion Models

NeurIPS 2024poster

The availability of large-scale multimodal datasets and advancements in diffusion models have significantly accelerated progress in 4D content generation. Most prior approaches rely on multiple images or video diffusion models, utilizing score distillation sampling for optimization or generating pse…

Cited by 32SourcePDFScholar
2024

Eliminating Warping Shakes for Unsupervised Online Video Stitching

ECCV 2024poster

"In this paper, we retarget video stitching to an emerging issue, named warping shake, when extending image stitching to video stitching. It unveils the temporal instability of warped content in non-overlapping regions, despite image stitching having endeavored to preserve the natural structures. Th…

2024

Endow SAM with Keen Eyes: Temporal-spatial Prompt Learning for Video Camouflaged Object Detection

CVPR 2024poster

The Segment Anything Model (SAM) a prompt-driven foundational model has demonstrated remarkable performance in natural image segmentation. However its application in video camouflaged object detection (VCOD) encounters challenges chiefly stemming from the overlooked temporal-spatial associations and…

Cited by 11SourcePDFScholar
2024

Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection

CVPR 2024poster

In this paper we study the problem of generalizable synthetic image detection aiming to detect forgery images from diverse generative methods e.g. GANs and diffusion models. Cutting-edge solutions start to explore the benefits of pre-trained models and mainly follow the fixed paradigm of solely trai…

2024

Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning

AAAI 2024technical

This research addresses the challenge of developing a universal deepfake detector that can effectively identify unseen deepfake images despite limited training data. Existing frequency-based paradigms have relied on frequency-level artifacts introduced during the up-sampling in GAN pipelines to det…

2024

Frozen CLIP: A Strong Backbone for Weakly Supervised Semantic Segmentation

CVPR 2024highlight

Weakly supervised semantic segmentation has witnessed great achievements with image-level labels. Several recent approaches use the CLIP model to generate pseudo labels for training an individual segmentation model while there is no attempt to apply the CLIP model as the backbone to directly segment…

2024

PixelLM: Pixel Reasoning with Large Multimodal Model

CVPR 2024poster

While large multimodal models (LMMs) have achieved remarkable progress generating pixel-level masks for image reasoning tasks involving multiple open-world targets remains a challenge. To bridge this gap we introduce PixelLM an effective and efficient LMM for pixel-level reasoning and understanding.…

Cited by 84SourcePDFScholar
2024

Region-Adaptive Transform with Segmentation Prior for Image Compression

ECCV 2024poster

"Learned Image Compression (LIC) has shown remarkable progress in recent years. Existing works commonly employ CNN-based or Transformer-based modules as transform methods for compression. However, there is no prior research on neural transform that focuses on specific regions. In response, we introd…

2024

Region-Native Visual Tokenization

ECCV 2024poster

"We explore an innovative region-based visual token representation and present the REgion-native AutoencoDER (Reader). In contrast to the majority of previous methods, which represent each image as a grid-shaped tokens map, Reader perceives each image into sequential region-based tokens, with each t…

2024

SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-Resolution

NeurIPS 2024poster

Diffusion-based Video Super-Resolution (VSR) is renowned for generating perceptually realistic videos, yet it grapples with maintaining detail consistency across frames due to stochastic fluctuations. The traditional approach of pixel-level alignment is ineffective for diffusion-processed frames bec…

2024

Semantic Lens: Instance-Centric Semantic Alignment for Video Super-resolution

AAAI 2024technical

As a critical clue of video super-resolution (VSR), inter-frame alignment significantly impacts overall performance. However, accurate pixel-level alignment is a challenging task due to the intricate motion interweaving in the video. In response to this issue, we introduce a novel paradigm for VSR n…

2024

Statistical Rejection Sampling Improves Preference Optimization

ICLR 2024poster

Improving the alignment of language models with human preferences remains an active research challenge. Previous approaches have primarily utilized online Reinforcement Learning from Human Feedback (RLHF). Recently, offline methods such as Sequence Likelihood Calibration (SLiC) and Direct Preference…

Cited by 199SourcePDFScholar
2024

Towards the Uncharted: Density-Descending Feature Perturbation for Semi-supervised Semantic Segmentation

CVPR 2024poster

Semi-supervised semantic segmentation allows model to mine effective supervision from unlabeled data to complement label-guided training. Recent research has primarily focused on consistency regularization techniques exploring perturbation-invariant training at both the image and feature levels. In…

2024

Transferable and Principled Efficiency for Open-Vocabulary Segmentation

CVPR 2024poster

Recent success of pre-trained foundation vision-language models makes Open-Vocabulary Segmentation (OVS) possible. Despite the promising performance this approach introduces heavy computational overheads for two challenges: 1) large model sizes of the backbone; 2) expensive costs during the fine-tun…

2024

WeatherDepth: Curriculum Contrastive Learning for Self-Supervised Depth Estimation under Adverse Weather Conditions

ICRA 2024poster

Depth estimation models have shown promising performance on clear scenes but fail to generalize to adverse weather conditions due to illumination variations, weather particles, etc. In this paper, we propose WeatherDepth, a self-supervised robust depth estimation model with curriculum contrastive le…

Cited by 15SourcecodeScholar
2023

An In-Depth Exploration of Person Re-Identification and Gait Recognition in Cloth-Changing Conditions

CVPR 2023poster

The target of person re-identification (ReID) and gait recognition is consistent, that is to match the target pedestrian under surveillance cameras. For the cloth-changing problem, video-based ReID is rarely studied due to the lack of a suitable cloth-changing benchmark, and gait recognition is ofte…

2023

CTP:Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation

ICCV 2023poster

Vision-Language Pretraining (VLP) has shown impressive results on diverse downstream tasks by offline training on large-scale datasets. Regarding the growing nature of real-world data, such an offline training paradigm on ever-expanding data is unsustainable, because models lack the continual learni…

Cited by 33PDFcodeScholar
2023

Calibrating Sequence likelihood Improves Conditional Language Generation

ICLR 2023poster

Conditional language models are predominantly trained with maximum likelihood estimation (MLE), giving probability mass to sparsely observed target sequences. While MLE trained models assign high probability to plausible sequences given the context, the model probabilities often do not accurately ra…

Cited by 140SourcePDFScholar
2023

Disentangling Orthogonal Planes for Indoor Panoramic Room Layout Estimation With Cross-Scale Distortion Awareness

CVPR 2023poster

Based on the Manhattan World assumption, most existing indoor layout estimation schemes focus on recovering layouts from vertically compressed 1D sequences. However, the compression procedure confuses the semantics of different planes, yielding inferior performance with ambiguous interpretability. T…

2023

Global Knowledge Calibration for Fast Open-Vocabulary Segmentation

ICCV 2023poster

Recent advancements in pre-trained vision-language models, such as CLIP, have enabled the segmentation of arbitrary concepts solely from textual inputs, a process commonly referred to as open-vocabulary semantic segmentation (OVS). However, existing OVS techniques confront a fundamental challenge: t…

Cited by 55PDFScholar
2023

Group Pose: A Simple Baseline for End-to-End Multi-Person Pose Estimation

ICCV 2023poster

In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.g., regarding pose estimation as keypoint box detection and combining with human detection in ED-Pose, hierarchically pr…

Cited by 41PDFcodeScholar
2023

Improving the Robustness of Summarization Models by Detecting and Removing Input Noise

EMNLP 2023long findings

The evaluation of abstractive summarization models typically uses test data that is identically distributed as training data. In real-world practice, documents to be summarized may contain input noise caused by text extraction artifacts or data pipeline bugs. The robustness of model performance unde…

Cited by 0SourceScholar
2023

Investigating Efficiently Extending Transformers for Long Input Summarization

EMNLP 2023long main

While large pretrained Transformer models have proven highly capable at tackling natural language tasks, handling long sequence inputs still poses a significant challenge. One such task is long input summarization, where inputs are longer than the maximum input context of most models. Through an ext…

Cited by 0SourcecodeScholar
2023

Learning Mask-aware CLIP Representations for Zero-Shot Segmentation

NeurIPS 2023poster

Recently, pre-trained vision-language models have been increasingly used to tackle the challenging zero-shot segmentation task. Typical solutions follow the paradigm of first generating mask proposals and then adopting CLIP to classify them. To maintain the CLIP's zero-shot transferability, previous…

2023

Learning To Segment Every Referring Object Point by Point

CVPR 2023poster

Referring Expression Segmentation (RES) can facilitate pixel-level semantic alignment between vision and language. Most of the existing RES approaches require massive pixel-level annotations, which are expensive and exhaustive. In this paper, we propose a new partially supervised training paradigm f…

2023

Learning on Gradients: Generalized Artifacts Representation for GAN-Generated Images Detection

CVPR 2023poster

Recently, there has been a significant advancement in image generation technology, known as GAN. It can easily generate realistic fake images, leading to an increased risk of abuse. However, most image detectors suffer from sharp performance drops in unseen domains. The key of fake image detection i…

2023

Locating Noise is Halfway Denoising for Semi-Supervised Segmentation

ICCV 2023poster

We investigate semi-supervised semantic segmentation with self-training, where a teacher model generates pseudo masks to exploit the benefits of a large amount of unlabeled images. We notice that the noisy label from the generated pseudo masks is the major obstacle to achieving good performance. Pre…

Cited by 13PDFScholar
2023

Out-of-Distribution Detection and Selective Generation for Conditional Language Models

ICLR 2023top-25%

Machine learning algorithms typically assume independent and identically distributed samples in training and at test time (IID). Much work has shown that high-performing ML classifiers can degrade significantly and provide overly-confident, wrong classification predictions, particularly for out-of-…

Cited by 106SourcePDFScholar
2023

Progressive Semantic-Visual Mutual Adaption for Generalized Zero-Shot Learning

CVPR 2023highlight

Generalized Zero-Shot Learning (GZSL) identifies unseen categories by knowledge transferred from the seen domain, relying on the intrinsic interactions between visual and semantic information. Prior works mainly localize regions corresponding to the sharing attributes. When various visual appearance…

2023

RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments

NeurIPS 2023poster

Intention-oriented object detection aims to detect desired objects based on specific intentions or requirements. For instance, when we desire to "lie down and rest", we instinctively seek out a suitable option such as a "bed" or a "sofa" that can fulfill our needs. Previous work in this area is limi…

Cited by 14SourcePDFScholar
2023

RecRecNet: Rectangling Rectified Wide-Angle Images by Thin-Plate Spline Model and DoF-based Curriculum Learning

ICCV 2023poster

The wide-angle lens shows appealing applications in VR technologies, but it introduces severe radial distortion into its captured image. To recover the realistic scene, previous works devote to rectifying the content of the wide-angle image. However, such a rectification solution inevitably distorts…

Cited by 16PDFcodeScholar
2023

Revisiting Simple Regret: Fast Rates for Returning a Good Arm

ICML 2023poster

Simple regret is a natural and parameter-free performance criterion for pure exploration in multi-armed bandits yet is less popular than the probability of missing the best arm or an $\epsilon$-good arm, perhaps due to lack of easy ways to characterize it. In this paper, we make a significant progre…

Cited by 19SourcePDFScholar
2023

SIGVIC: Spatial Importance Guided Variable-Rate Image Compression

ICASSP 2023accepted

Variable-rate mechanism has improved the flexibility and efficiency of learning-based image compression that trains multiple models for different rate-distortion tradeoffs. One of the most common approaches for variable-rate is to channel- wisely or spatial-uniformly scale the internal features. How…

Cited by 0SourceScholar
2023

SegRefiner: Towards Model-Agnostic Segmentation Refinement with Discrete Diffusion Process

NeurIPS 2023poster

In this paper, we explore a principal way to enhance the quality of object masks produced by different segmentation models. We propose a model-agnostic solution called SegRefiner, which offers a novel perspective on this problem by interpreting segmentation refinement as a data generation process. A…

2023

Spatiotemporal Deformation Perception for Fisheye Video Rectification

AAAI 2023technical

Although the distortion correction of fisheye images has been extensively studied, the correction of fisheye videos is still an elusive challenge. For different frames of the fisheye video, the existing image correction methods ignore the correlation of sequences, resulting in temporal jitter in the…

2023

Towards Reliable Image Outpainting: Learning Structure-Aware Multimodal Fusion with Depth Guidance

ICASSP 2023accepted

Image outpainting technology generates visually plausible content regardless of authenticity, making it unreliable to be applied in practice. Thus, we propose a reliable image outpainting task, introducing the sparse depth from LiDARs (Light Detection And Ranging devices) to extrapolate authentic RG…

Cited by 0SourceScholar
2023

Unsupervised OmniMVS: Efficient Omnidirectional Depth Inference via Establishing Pseudo-Stereo Supervision

IROS 2023poster

Omnidirectional multi-view stereo (MVS) vision is attractive for its ultra-wide field-of-view (FoV), enabling machines to perceive 360°3D surroundings. However, the existing solutions require expensive dense depth labels for supervision, making them impractical in real-world applications. In this pa…

Cited by 8SourcecodeScholar
2022

A Well-Composed Text is Half Done! Composition Sampling for Diverse Conditional Generation

ACL 2022long

We propose Composition Sampling, a simple but effective method to generate diverse outputs for conditional generation of higher quality compared to previous stochastic decoding strategies. It builds on recently proposed plan-based neural generation models (FROST, Narayan et al, 2021) that are traine…

2022

Exploring Complementarity of Global and Local Spatiotemporal Information for Fake Face Video Detection

ICASSP 2022accepted

The spread of fake face videos leads to severe social concerns, which promotes the development of detection methods for these videos. Existing patch-based methods focus on local regions to find forgery common clues, while ignoring the important role of the global information. In this paper, a novel…

Cited by 0SourceScholar
2022

Implicit Relation Linking for Question Answering over Knowledge Graph

ACL 2022findings

Relation linking (RL) is a vital module in knowledge-based question answering (KBQA) systems. It aims to link the relations expressed in natural language (NL) to the corresponding ones in knowledge graph (KG). Existing methods mainly rely on the textual similarities between NL and KG to build relati…

Cited by 6SourcePDFScholar
2022

Mask Matching Transformer for Few-Shot Segmentation

NeurIPS 2022accept

In this paper, we aim to tackle the challenging few-shot segmentation task from a new perspective. Typical methods follow the paradigm to firstly learn prototypical features from support images and then match query features in pixel-level to obtain segmentation results. However, to obtain satisfacto…

2022

PanoFormer: Panorama Transformer for Indoor 360° Depth Estimation

ECCV 2022poster

"Existing panoramic depth estimation methods based on convolutional neural networks (CNNs) focus on removing panoramic distortions, failing to perceive panoramic structures efficiently due to the fixed receptive field in CNNs. This paper proposes the panorama Transformer (named PanoFormer) to estima…

2022

SiRi: A Simple Selective Retraining Mechanism for Transformer-Based Visual Grounding

ECCV 2022poster

"In this paper, we investigate how to achieve better referring visual grounding with modern vision-language transformers, and propose a simple yet powerful Selective Retraining (SiRi) mechanism. Particularly, SiRi conveys a significant principle to the research of visual grounding, i.e, a better ini…

2022

Slim Scissors: Segmenting Thin Object from Synthetic Background

ECCV 2022poster

"Existing interactive segmentation algorithms typically fail when segmenting objects with elongated thin structures (bicycle spokes). Though some recent efforts attempt to address this challenge by introducing a new synthetic dataset and a three-stream network design, they suffer two limitations: 1)…

Cited by 7SourcePDFScholar
2021

Double Low-Rank Representation With Projection Distance Penalty for Clustering

CVPR 2021poster

This paper presents a novel, simple yet robust self-representation method, i.e., Double Low-Rank Representation with Projection Distance penalty (DLRRPD) for clustering. With the learned optimal projected representations, DLRRPD is capable of obtaining an effective similarity graph to capture the mu…

Cited by 35PDFScholar
2021

GradingNet: Towards Providing Reliable Supervisions for Weakly Supervised Object Detection by Grading the Box Candidates

AAAI 2021technical

Weakly-Supervised Object Detection (WSOD) aims at training a model with limited and coarse annotations for precisely locating the regions of objects. Existing works solve the WSOD problem by using a two-stage framework, i.e., generating candidate bounding boxes with weak supervision information and…

Cited by 17SourcePDFScholar
2021

Multi-Level Curriculum for Training a Distortion-Aware Barrel Distortion Rectification Model

ICCV 2021poster

Barrel distortion rectification aims at removing the radial distortion in a distorted image captured by a wide-angle lens. Previous deep learning methods mainly solve this problem by learning the implicit distortion parameters or the nonlinear rectified mapping function in a direct manner. However,…

Cited by 16PDFScholar
2021

Progressively Complementary Network for Fisheye Image Rectification Using Appearance Flow

CVPR 2021poster

Distortion rectification is often required for fisheye images. The generation-based method is one mainstream solution due to its label-free property, but its naive skip-connection and overburdened decoder will cause blur and incomplete correction. First, the skip-connection directly transfers the im…

Cited by 57PDFcodeScholar
2021

Towards Complete Scene and Regular Shape for Distortion Rectification by Curve-Aware Extrapolation

ICCV 2021poster

The wide-angle lens gains increasing attention since it can capture a wide field-of-view scene (FoV). However, the obtained image is contaminated with radial distortion, making the scene not realistic. Previous distortion rectification methods rectify the image in a rectangle or invagination, failin…

Cited by 8PDFScholar
2021

Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline

CVPR 2021poster

Depth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR) is a practical and valuable task, which upscales the depth map into high-resolution (HR) space. However, limited by the…

Cited by 100PDFScholar
2020

CoADNet: Collaborative Aggregation-and-Distribution Networks for Co-Salient Object Detection

NeurIPS 2020poster

Co-Salient Object Detection (CoSOD) aims at discovering salient objects that repeatedly appear in a given query group containing two or more relevant images. One challenging issue is how to effectively capture co-saliency cues by modeling and exploiting inter-image relationships. In this paper, we p…

2020

Distribution-Induced Bidirectional Generative Adversarial Network for Graph Representation Learning

CVPR 2020poster

Graph representation learning aims to encode all nodes of a graph into low-dimensional vectors that will serve as input of many computer vision tasks. However, most existing algorithms ignore the existence of inherent data distribution and even noises. This may significantly increase the phenomenon…

Cited by 48PDFcodeScholar
2020

Fast Template Matching and Update for Video Object Tracking and Segmentation

CVPR 2020poster

In this paper, the main task we aim to tackle is the multi-instance semi-supervised video object segmentation across a sequence of frames where only the first-frame box-level ground-truth is provided. Detection-based algorithms are widely adopted to handle this task, and the challenges lie in the se…

Cited by 82PDFcodeScholar
2020

Interactive Object Segmentation With Inside-Outside Guidance

CVPR 2020oral

This paper explores how to harvest precise object segmentation masks while minimizing the human interaction cost. To achieve this, we propose an Inside-Outside Guidance (IOG) approach in this work. Concretely, we leverage an inside point that is clicked near the object center and two outside points…

Cited by 158PDFcodeScholar
2020

PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization

ICML 2020poster

Recent work pre-training Transformers with self-supervised objectives on large text corpora has shown great success when fine-tuned on downstream NLP tasks including text summarization. However, pre-training objectives tailored for abstractive text summarization have not been explored. Furthermore t…

2017

Object Region Mining With Adversarial Erasing: A Simple Classification to Semantic Segmentation Approach

CVPR 2017oral

We investigate a principle way to progressively mine discriminative object regions using classification networks to address the weakly-supervised semantic segmentation problems. Classification networks are only responsive to small and sparse discriminative regions from the object of interest, which…

Cited by 1023PDFScholar