← Search

Xiaobin Zhu

14 accepted papers

2026

Dual-Geometry Graph Network: Unifying Local and Global Priors for Few-Shot Learning

AAAI 2026technical

In few-shot learning, utilizing local and global geometric priors to capture both subtle local class metrics and coarse global structures within the meta-task are important to obtain discriminative embeddings. However, existing graph-based and curvature-based few-shot approaches only focus on either

Cited by 0SourcePDFScholar
2026

PlantRSR: A New Plant Dataset and Method for Reference-based Super-Resolution

ICLR 2026poster

Single image super-resolution (SISR) often struggles to reconstruct high-resolution (HR) details from heavily degraded low-resolution (LR) inputs. Instead, reference-based super-resolution (RefSR) methods offer an alternative solution to generate promising results using high-quality reference (Ref)…

Cited by 0SourcecodeScholar
2026

Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection

CVPR 2026

Multimodal misinformation poses an escalating challenge that often evades traditional detectors, which are opaque black boxes and fragile against new manipulation tactics. We present Probabilistic Concept Graph Reasoning (PCGR), an interpretable, modular, and evolvable framework that reframes multim

Cited by 0SourcecodeScholar
2026

Semantic-Enhanced Time-Series Forecasting via Large Language Models

ICLR 2026poster

Time series forecasting plays a significant role in finance, energy, meteorology, and IoT applications. Recent studies have leveraged the generalization capabilities of large language models (LLMs) to adapt to time series forecasting, achieving promising performance. However, existing studies focus…

Cited by 0SourcecodeScholar
2026

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

AAAI 2026technical

Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) have demonstrated significant potential in caption generation, existing benchmar

Cited by 0SourcePDFScholar
2025

DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid Framework

CVPR 2025poster

Optical flow estimation is essential for video processing tasks, such as restoration and action recognition. The quality of videos is constantly increasing, with current standards reaching 8K resolution. However, optical flow methods are usually designed for low resolution and do not generalize to l…

2024

Arbitrary Time Information Modeling via Polynomial Approximation for Temporal Knowledge Graph Embedding

COLING 2024main

Distinguished from traditional knowledge graphs (KGs), temporal knowledge graphs (TKGs) must explore and reason over temporally evolving facts adequately. However, existing TKG approaches still face two main challenges, i.e., the limited capability to model arbitrary timestamps continuously and the…

2024

LayoutFormer: Hierarchical Text Detection Towards Scene Text Understanding

CVPR 2024poster

Existing scene text detectors generally focus on accurately detecting single-level (i.e. word-level line-level or paragraph-level) text entities without exploring the relationships among different levels of text entities. To comprehensively understand scene texts detecting multi-level texts while ex…

Cited by 2SourcePDFScholar
2024

RAPIDFlow: Recurrent Adaptable Pyramids with Iterative Decoding for Efficient Optical Flow Estimation

ICRA 2024poster

Extracting motion information from videos with optical flow estimation is vital in multiple practical robot applications. Current optical flow approaches show remarkable accuracy, but top-performing methods have high computational costs and are unsuitable for embedded devices. Although some previous…

Cited by 8SourcecodeScholar
2024

Recurrent Partial Kernel Network for Efficient Optical Flow Estimation

AAAI 2024technical

Optical flow estimation is a challenging task consisting of predicting per-pixel motion vectors between images. Recent methods have employed larger and more complex models to improve the estimation accuracy. However, this impacts the widespread adoption of optical flow methods and makes it harder to…

2023

Learning Correction Filter via Degradation-Adaptive Regression for Blind Single Image Super-Resolution

ICCV 2023poster

Although existing image deep learning super-resolution (SR) methods achieve promising performance on benchmark datasets, they still suffer from severe performance drops when the degradation of the low-resolution (LR) input is not covered in training. To address the problem, we propose an innovative…

Cited by 32PDFcodeScholar
2022

Learning Aligned Cross-Modal Representation for Generalized Zero-Shot Classification

AAAI 2022technical

Learning a common latent embedding by aligning the latent spaces of cross-modal autoencoders is an effective strategy for Generalized Zero-Shot Classification (GZSC). However, due to the lack of fine-grained instance-wise annotations, it still easily suffer from the domain shift problem for the disc…

Cited by 22SourcePDFScholar
2021

Adaptive Boundary Proposal Network for Arbitrary Shape Text Detection

ICCV 2021poster

Arbitrary shape text detection is a challenging task due to the high complexity and variety of scene texts. In this work, we propose a novel adaptive boundary proposal network for arbitrary shape text detection, which can learn to directly produce accurate boundary for arbitrary shape text without a…

Cited by 121PDFcodeScholar
2020

Deep Relational Reasoning Graph Network for Arbitrary Shape Text Detection

CVPR 2020oral

Arbitrary shape text detection is a challenging task due to the high variety and complexity of scenes texts. In this paper, we propose a novel unified relational reasoning graph network for arbitrary shape text detection. In our method, an innovative local graph bridges a text proposal model via Con…

Cited by 281PDFcodeScholar