← Search

Xiaobo Zhang

17 accepted papers

2026

Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection

AAAI 2026technical

Existing core-set selection methods predominantly rely on heuristic scoring signals such as training dynamics or model uncertainty, lacking explicit modeling of data likelihood. This omission may hinder the constructed subset from capturing subtle yet critical distributional structures that underpin

Cited by 0SourcePDFScholar
2026

Zero-Shot Image Denoising via Hybrid Prior-Guided Pseudo Sample Generation

CVPR 2026

Zero-shot image denoising has gained prominence in recent years, as it inherently relies on the intrinsic priors of images rather than learning from external data. Nevertheless, most existing methods either fail to fully exploit global priors, or do not properly preserve the fine-grained details gov

Cited by 0SourceScholar
2025

AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation

NeurIPS 2025poster

Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Medical Large Multimodal Models (MLMMs) have enabled automated CXR interpretation, improving diagnostic accuracy and efficiency. However, despite their strong visual understanding,…

Cited by 0SourcecodeScholar
2025

GLiM: Integrating Graph Transformer and LLM for Document-Level Biomedical Relation Extraction with Incomplete Labeling

ACL 2025finding

Document-level relation extraction (DocRE) identifies relations between entities across an entire document. However, as the number and complexity of entities and entity-pair relations grow, the problem space expands quadratically, causing incomplete annotations and frequent false negatives, especial…

2025

Influence-Guided Diffusion for Dataset Distillation

ICLR 2025poster

Dataset distillation aims to streamline the training process by creating a compact yet effective dataset for a much larger original dataset. However, existing methods often struggle with distilling large, high-resolution datasets due to prohibitive resource costs and limited performance, primarily s…

2025

Long-Range Multi-Scale Fusion for Efficient Single Image Super-Resolution

ICASSP 2025accepted

Improving the performance of single image super-resolution (SISR) via extending the effective receptive field (ERF) of the model has become an admired paradigm in the field due to the universal self-similarity prior of natural images. However, it cannot fully explore model capability by solely incre…

Cited by 0SourceScholar
2025

RoBGuard: Enhancing LLMs to Assess Risk of Bias in Clinical Trial Documents

COLING 2025main

Randomized Controlled Trials (RCTs) are rigorous clinical studies crucial for reliable decision-making, but their credibility can be compromised by bias. The Cochrane Risk of Bias tool (RoB 2) assesses this risk, yet manual assessments are time-consuming and labor-intensive. Previous approaches have…

Cited by 0SourcePDFScholar
2025

Tracing Copied Pixels and Regularizing Patch Affinity in Copy Detection

ICCV 2025poster

Image Copy Detection (ICD) aims to identify manipulated content between image pairs through robust feature representation learning. While self-supervised learning (SSL) has advanced ICD systems, existing view-level contrastive methods struggle with sophisticated edits due to insufficient fine-graine…

Cited by 0SourcePDFScholar
2025

Uncertainty-Aware Dynamic Fusion for Multimodal Clinical Prediction Tasks

ICASSP 2025accepted

Multimodal fusion offers significant potential for enhancing medical diagnosis, particularly in the Intensive Care Unit (ICU), where integrating diverse data sources is crucial. Traditional static fusion models often fail to account for sample-wise variations in modality importance, which can impact…

Cited by 0SourceScholar
2024

Cross-Image Distillation for Semi-Supervised Semantic Segmentation

ICASSP 2024accepted

Semi-supervised semantic segmentation approaches have drawn much more attention in recent years, which aim to exploit a large amount of unlabeled data together with a small number of labeled data. However, existing models usually regarded segmentation as pixel-wise classification, neglecting global…

Cited by 0SourceScholar
2024

Let All Be Whitened: Multi-Teacher Distillation for Efficient Visual Retrieval

AAAI 2024technical

Visual retrieval aims to search for the most relevant visual items, e.g., images and videos, from a candidate gallery with a given query item. Accuracy and efficiency are two competing objectives in retrieval tasks. Instead of crafting a new method pursuing further improvement on accuracy, in this p…

2024

Self-Supervised Video Copy Localization with Regional Token Representation

ECCV 2024poster

"The task of video copy localization aims at finding the start and end timestamps of all copied segments within a pair of untrimmed videos. Recent approaches usually extract frame-level features and generate a frame-to-frame similarity map for the video pair. Learned detectors are used to identify d…

2024

Towards Evidential and Class Separable Open Set Object Detection

AAAI 2024technical

Detecting in open-world scenarios poses a formidable challenge for models intended for real-world deployment. The advanced closed set object detectors achieve impressive performance under the closed set setting, but often produce overconfident misprediction on unknown objects due to the lack of supe…

2023

Boundary-Aware Backward-Compatible Representation via Adversarial Learning in Image Retrieval

CVPR 2023poster

Image retrieval plays an important role in the Internet world. Usually, the core parts of mainstream visual retrieval systems include an online service of the embedding model and a large-scale vector database. For traditional model upgrades, the old model will not be replaced by the new one until th…

2023

Large Language Models are Complex Table Parsers

EMNLP 2023long main

With the Generative Pre-trained Transformer 3.5 (GPT-3.5) exhibiting remarkable reasoning and comprehension abilities in Natural Language Processing (NLP), most Question Answering (QA) research has primarily centered around general QA tasks based on GPT, neglecting the specific challenges posed by C…

Cited by 0SourceScholar
2023

TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible Supervision

AAAI 2023technical

Video copy localization aims to precisely localize all the copied segments within a pair of untrimmed videos in video retrieval applications. Previous methods typically start from frame-to-frame similarity matrix generated by cosine similarity between frame-level features of the input video pair, an…

2022

A Large-Scale Comprehensive Dataset and Copy-Overlap Aware Evaluation Protocol for Segment-Level Video Copy Detection

CVPR 2022poster

In this paper, we introduce VCSL (Video Copy Segment Localization), a new comprehensive segment-level annotated video copy dataset. Compared with existing copy detection datasets restricted by either video-level annotation or small-scale, VCSL not only has two orders of magnitude more segment-level…

Cited by 18PDFcodeScholar