← Search

Xiaoyu Yue

11 accepted papers

2026

EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language Models

ICLR 2026poster

Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or cover isolated subdomains, lacking holistic evaluation…

Cited by 0SourceScholar
2026

Transition Models: Rethinking the Generative Learning Objective

CVPR 2026

A fundamental dilemma in generative modeling persists: iterative diffusion models achieve outstanding fidelity, but at a significant computational cost, while efficient few-step alternatives are constrained by a hard quality ceiling. This conflict between generation steps and output quality arises f

Cited by 0SourcecodeScholar
2025

InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis

ICCV 2025poster

Arbitrary resolution image generation provides a consistent visual experience across devices, having extensive applications for producers and consumers. Current diffusion models increase computational demand quadratically with resolution, causing 4K image generation delays over 100 seconds. To solve…

2025

RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts

NeurIPS 2025poster

Quality analysis of weather forecasts is an essential topic in meteorology. Although traditional score-based evaluation metrics can quantify certain forecast errors, they are still far from meteorological experts in terms of descriptive capability, interpretability, and understanding of dynamic evol…

Cited by 0SourceScholar
2025

Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation

NeurIPS 2025poster

Recent studies have demonstrated the importance of high-quality visual representations in image generation and have highlighted the limitations of generative models in image understanding. As a generative paradigm originally designed for natural language, autoregressive models face similar challenge…

Cited by 0SourceScholar
2023

OV-PARTS: Towards Open-Vocabulary Part Segmentation

NeurIPS 2023poster

Segmenting and recognizing diverse object parts is a crucial ability in applications spanning various computer vision and robotic tasks. While significant progress has been made in object-level Open-Vocabulary Semantic Segmentation (OVSS), i.e., segmenting objects with arbitrary text, the correspond…

2022

Rethinking the Two-Stage Framework for Grounded Situation Recognition

AAAI 2022technical

Grounded Situation Recognition (GSR), i.e., recognizing the salient activity (or verb) category in an image (e.g.,buying) and detecting all corresponding semantic roles (e.g.,agent and goods), is an essential step towards “human-like” event understanding. Since each verb is associated with a specifi…

2021

Aggregation With Feature Detection

ICCV 2021poster

Aggregating features from different depths of a network is widely adopted to improve the network capability. Lots of modern architectures are equipped with skip connections, which actually makes the feature aggregation happen in all these networks. Since different features tell different semantic m…

Cited by 2PDFScholar
2021

Vision Transformer With Progressive Sampling

ICCV 2021poster

Transformers with powerful global relation modeling abilities have been introduced to fundamental computer vision tasks recently. As a typical example, the Vision Transformer (ViT) directly applies a pure transformer architecture on image classification, by simply splitting images into tokens with a…

Cited by 126PDFcodeScholar
2020

RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition

ECCV 2020poster

The attention-based encoder-decoder framework has recently achieved impressive results for scene text recognition, and many variants have emerged with improvements in recognition quality. However, it performs poorly on contextless texts (e.g., random character sequences) which is unacceptable in mos…

2019

Geometry Normalization Networks for Accurate Scene Text Detection

ICCV 2019poster

Large geometry (e.g., orientation) variances are the key challenges in the scene text detection. In this work, we first conduct experiments to investigate the capacity of networks for learning geometry variances on detecting scene texts, and find that networks can handle only limited text geometry v…

Cited by 40PDFcodeScholar