← Search

Xu-Cheng Yin

25 accepted papers

2026

Dual-Geometry Graph Network: Unifying Local and Global Priors for Few-Shot Learning

AAAI 2026technical

In few-shot learning, utilizing local and global geometric priors to capture both subtle local class metrics and coarse global structures within the meta-task are important to obtain discriminative embeddings. However, existing graph-based and curvature-based few-shot approaches only focus on either

Cited by 0SourcePDFScholar
2026

PlantRSR: A New Plant Dataset and Method for Reference-based Super-Resolution

ICLR 2026poster

Single image super-resolution (SISR) often struggles to reconstruct high-resolution (HR) details from heavily degraded low-resolution (LR) inputs. Instead, reference-based super-resolution (RefSR) methods offer an alternative solution to generate promising results using high-quality reference (Ref)…

Cited by 0SourcecodeScholar
2026

Predicting Emergent Tool Use in LLMs Before It Emerges: A Proxy Perspective

AAAI 2026technical

Tool-use capabilities fundamentally transform large language models (LLMs) from passive language generators into active agents with real-world utility, drawing intense research focus. Yet, their emergent nature renders traditional scaling laws ineffective for early-stage prediction, obstructing prin

Cited by 0SourcePDFScholar
2026

Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection

CVPR 2026

Multimodal misinformation poses an escalating challenge that often evades traditional detectors, which are opaque black boxes and fragile against new manipulation tactics. We present Probabilistic Concept Graph Reasoning (PCGR), an interpretable, modular, and evolvable framework that reframes multim

Cited by 0SourcecodeScholar
2026

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

AAAI 2026technical

Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) have demonstrated significant potential in caption generation, existing benchmar

Cited by 0SourcePDFScholar
2025

AtomNet: Designing Tiny Models from Operators Under Extreme MCU Constraints

AAAI 2025technical

Tiny machine learning (TinyML) has attracted heightened attention for its ability to provide low-cost and instantaneous performance on edge devices. Particularly, the commonly used microcontroller unit (MCU) imposes extreme constraints on peak memory (SRAM) and storage (Flash). Existing TinyML metho…

Cited by 0SourcePDFScholar
2025

Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition

ICASSP 2025accepted

Recently, end-to-end automatic speech recognition has become the mainstream approach in both industry and academia. To optimize system performance in specific scenarios, the Weighted Finite-State Transducer (WFST) is extensively used to integrate acoustic and language models, leveraging its capacity…

Cited by 0SourceScholar
2025

DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid Framework

CVPR 2025poster

Optical flow estimation is essential for video processing tasks, such as restoration and action recognition. The quality of videos is constantly increasing, with current standards reaching 8K resolution. However, optical flow methods are usually designed for low resolution and do not generalize to l…

2025

FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles

AAAI 2025technical

Humans can perceive speakers’ characteristics (e.g., identity, gender, personality and emotion) by their appearance, which are generally aligned to their voice style. Recently, vision-driven Text-to-speech ( TTS ) scholars grounded their investigations on real-person faces, thereby restricting effec…

2025

Joint Feature and Kernel Fusion for Improved Depth-Aware Panoptic Segmentation

ICASSP 2025accepted

Depth-aware Panoptic Segmentation, which combines panoptic segmentation and monocular depth estimation, is a challenging task that requires a comprehensive understanding of both scene geometry and object semantics. Recent multi-task learning approaches have leveraged dynamic kernel methods to tackle…

Cited by 0SourceScholar
2025

Tool Playgrounds: A Comprehensive and Analyzable Benchmark for LLM Tool Invocation

ICASSP 2025accepted

The rapid advancement of large language models (LLMs) has paved the way for their use in solving real-world problems, which in turn has significantly driven the development of tool-assisted LLMs. This progress necessitates thorough evaluation methods. However, existing benchmarks typically only prov…

Cited by 0SourceScholar
2024

Arbitrary Time Information Modeling via Polynomial Approximation for Temporal Knowledge Graph Embedding

COLING 2024main

Distinguished from traditional knowledge graphs (KGs), temporal knowledge graphs (TKGs) must explore and reason over temporally evolving facts adequately. However, existing TKG approaches still face two main challenges, i.e., the limited capability to model arbitrary timestamps continuously and the…

2024

Attention Decoupling for Query-Based Object Detection

ICASSP 2024accepted

Benefiting from attention mechanisms, query-based detectors have a strong model capacity. They predict classification and regression by utilizing their shared queries and features in the decoder. Inter-task biases cause multi-directional gradients that disturb each other to limit model optimization.…

Cited by 0SourceScholar
2024

LayoutFormer: Hierarchical Text Detection Towards Scene Text Understanding

CVPR 2024poster

Existing scene text detectors generally focus on accurately detecting single-level (i.e. word-level line-level or paragraph-level) text entities without exploring the relationships among different levels of text entities. To comprehensively understand scene texts detecting multi-level texts while ex…

Cited by 2SourcePDFScholar
2024

RAPIDFlow: Recurrent Adaptable Pyramids with Iterative Decoding for Efficient Optical Flow Estimation

ICRA 2024poster

Extracting motion information from videos with optical flow estimation is vital in multiple practical robot applications. Current optical flow approaches show remarkable accuracy, but top-performing methods have high computational costs and are unsuitable for embedded devices. Although some previous…

Cited by 8SourcecodeScholar
2024

Recurrent Partial Kernel Network for Efficient Optical Flow Estimation

AAAI 2024technical

Optical flow estimation is a challenging task consisting of predicting per-pixel motion vectors between images. Recent methods have employed larger and more complex models to improve the estimation accuracy. However, this impacts the widespread adoption of optical flow methods and makes it harder to…

2023

Learning Correction Filter via Degradation-Adaptive Regression for Blind Single Image Super-Resolution

ICCV 2023poster

Although existing image deep learning super-resolution (SR) methods achieve promising performance on benchmark datasets, they still suffer from severe performance drops when the degradation of the low-resolution (LR) input is not covered in training. To address the problem, we propose an innovative…

Cited by 32PDFcodeScholar
2023

Self-Convolution for Automatic Speech Recognition

ICASSP 2023accepted

Self-attention plays a significant role in recent automatic speech recognition (ASR) models with promising results. However, it suffers from high computational complexity and weak capability in modeling local information. In contrast, the convolutional neural network (CNN) is computationally effecti…

Cited by 0SourceScholar
2023

VLPD: Context-Aware Pedestrian Detection via Vision-Language Semantic Self-Supervision

CVPR 2023poster

Detecting pedestrians accurately in urban scenes is significant for realistic applications like autonomous driving or video surveillance. However, confusing human-like objects often lead to wrong detections, and small scale or heavily occluded pedestrians are easily missed due to their unusual appea…

2022

Learning Aligned Cross-Modal Representation for Generalized Zero-Shot Classification

AAAI 2022technical

Learning a common latent embedding by aligning the latent spaces of cross-modal autoencoders is an effective strategy for Generalized Zero-Shot Classification (GZSC). However, due to the lack of fine-grained instance-wise annotations, it still easily suffer from the domain shift problem for the disc…

Cited by 22SourcePDFScholar
2022

Non-Autoregressive Transformer with Unified Bidirectional Decoder for Automatic Speech Recognition

ICASSP 2022accepted

Non-autoregressive (NAR) transformer models have been studied intensively in automatic speech recognition (ASR), and many NAR transformer models is to use the causal mask to limit token dependencies. However, the causal mask is designed for the left-to-right decoding process of the non-parallel auto…

Cited by 0SourceScholar
2021

Adaptive Boundary Proposal Network for Arbitrary Shape Text Detection

ICCV 2021poster

Arbitrary shape text detection is a challenging task due to the high complexity and variety of scene texts. In this work, we propose a novel adaptive boundary proposal network for arbitrary shape text detection, which can learn to directly produce accurate boundary for arbitrary shape text without a…

Cited by 121PDFcodeScholar
2021

Adaptive Pattern-Parameter Matching for Robust Pedestrian Detection

AAAI 2021technical

Pedestrians with challenging patterns, e.g. small scale or heavy occlusion, appear frequently in practical applications like autonomous driving, which remains tremendous obstacle to higher robustness of detectors. Although plenty of previous works have been dedicated to these problems, properly matc…

Cited by 19SourcePDFScholar
2020

Deep Relational Reasoning Graph Network for Arbitrary Shape Text Detection

CVPR 2020oral

Arbitrary shape text detection is a challenging task due to the high variety and complexity of scenes texts. In this paper, we propose a novel unified relational reasoning graph network for arbitrary shape text detection. In our method, an innovative local graph bridges a text proposal model via Con…

Cited by 281PDFcodeScholar