← Search

Xi Yang

69 accepted papers

2026

Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT

CVPR 2026

Multiple-choice question answering (MCQA) has been a popular format for evaluating and reinforcement fine-tuning (RFT) of modern multimodal language models. Its constrained output format allows for simplified, deterministic automatic verification.However, we find that the options may leak exploitabl

Cited by 0SourceScholar
2026

Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench

CVPR 2026

Reading measurement instruments is effortless for humans and requires relatively little domain expertise, yet it remains surprisingly challenging for current vision-language models (VLMs) as we find in preliminary evaluation. In this work, we introduce MeasureBench, a benchmark on visual measurement

Cited by 0SourcecodeScholar
2026

GSFixer: Improving 3D Gaussian Splatting with Reference-Guided Video Diffusion Priors

ICML 2026poster

Reconstructing 3D scenes using 3D Gaussian Splatting (3DGS) from sparse views is an ill-posed problem due to insufficient information, often resulting in noticeable artifacts. While recent approaches have sought to leverage generative priors to complete information for under-constrained regions, the…

Cited by 0SourceScholar
2026

Learning Compact Latent Space for Representing Neural Signed Distance Functions with High-fidelity Geometry Details

AAAI 2026technical

Neural signed distance functions (SDFs) have been a vital representation to represent 3D shapes or scenes with neural networks. An SDF is an implicit function that can query signed distances at specific coordinates for recovering a 3D surface. Although implicit functions work well on a single shape

Cited by 0SourcePDFScholar
2026

Out-of-Context Misinformation Detection via Variational Domain-Invariant Learning with Test-Time Training

AAAI 2026technical

Out-of-context misinformation (OOC) is a low-cost form of misinformation in news reports, which refers to place authentic images into out-of-context or fabricated image-text pairings. This problem has attracted significant attention from researchers in recent years. Current methods focus on assessin

Cited by 0SourcePDFScholar
2026

Prompt Tuning for CLIP on the Pretrained Manifold

ICML 2026poster

Prompt tuning introduces learnable prompt vectors that adapt pretrained vision-language models to downstream tasks in a parameter-efficient manner. However, under limited supervision, prompt tuning alters pretrained representations and drives downstream features away from the pretrained manifold tow…

Cited by 0SourceScholar
2026

SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks

ICLR 2026poster

With the rapid advancement of Large Language Models (LLMs), the safety of LLMs has been a critical concern requiring precise assessment. Current benchmarks primarily concentrate on single-turn dialogues or a single jailbreak attack method to assess the safety. Additionally, these benchmarks have not…

Cited by 0SourcecodeScholar
2026

Sequential Information Bottleneck Fusion: Towards Robust and Generalizable Multi-Modal Brain Tumor Segmentation

ICLR 2026poster

Brain tumor segmentation in multi-modal MRIs poses significant challenges when one or more modalities are missing. Recent approaches commonly employ parallel fusion strategies; however, these methods often risk losing crucial shared information across modalities, which can degrade segmentation perfo…

Cited by 0SourceScholar
2026

TacpAgent: Enhancing Student Engagement in Classroom Exercises Through LLM-Generated Feedback

AAAI 2026technical

Classroom exercises are imperative for reinforcing learning. However, in conventional instruction, students frequently lack timely and personalized feedback. To address this, we present TacpAgent(Teaching Agent for Classroom Practice),a generative LLM-based agent that delivers detailed, individualiz

Cited by 0SourcePDFScholar
2026

Task-Driven Subspace Decomposition for Knowledge Sharing and Isolation in LoRA-based Continual Learning

ICML 2026poster

Continual Learning (CL) requires models to sequentially adapt to new tasks without forgetting old knowledge. Recently, Low-Rank Adaptation (LoRA), a representative Parameter-Efficient Fine-Tuning (PEFT) method, has gained increasing attention in CL. Several LoRA-based CL methods reduce interference …

Cited by 0SourceScholar
2025

A Self-Evolving Framework for Multi-Agent Medical Consultation Based on Large Language Models

ICASSP 2025accepted

We propose a multi-agent approach (SeM-Agents) based on large language models for medical consultations. This framework incorporates various doctor roles and auxiliary roles, with agents communicating through natural language. Using a residual structure, the system conducts multi-round medical consu…

Cited by 0SourceScholar
2025

ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5

ACL 2025long

Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children’s speech remains challenging due to differences in pronunciation, tone,…

2025

Clustering Properties of Self-Supervised Learning

ICML 2025poster

Self-supervised learning (SSL) methods via joint embedding architectures have proven remarkably effective at capturing semantically rich representations with strong clustering properties, magically in the absence of label supervision. Despite this, few of them have explored leveraging these untapped…

Cited by 0SourcePDFScholar
2025

DIH-CLIP: Unleashing the Diversity of Multi-Head Self-Attention for Training-Free Open-Vocabulary Semantic Segmentation

ICCV 2025poster

Recent Training-Free Open-Vocabulary Semantic Segmentation (TF-OVSS) leverages a pre-training vision-language model to segment images from open-set visual concepts without training and fine-tuning. The key of TF-OVSS is to improve the local spatial representation of CLIP by leveraging self-correlati…

2025

Disentangling Tabular Data Towards Better One-Class Anomaly Detection

AAAI 2025technical

Tabular anomaly detection under the one-class classification setting poses a significant challenge, as it involves accurately conceptualizing "normal" derived exclusively from a single category to discern anomalies from normal data variations. Capturing the intrinsic correlation among attributes wit…

2025

Dual Domain Control via Active Learning for Remote Sensing Domain Incremental Object Detection

ICCV 2025poster

Domain incremental object detection in remote sensing addresses the challenge of adapting to continuously emerging domains with distinct characteristics. Unlike natural images, remote sensing data vary significantly due to differences in sensors, altitudes, and geographic locations, leading to data…

Cited by 0SourcePDFScholar
2025

Dual Information Purification for Lightweight SAR Object Detection

AAAI 2025technical

Synthetic aperture radar (SAR) object detection requires accurate identification and localization of targets at various scales within SAR images. However, background clutter and speckle noise can obscure key features and mislead the knowledge distillation process. To address these challenges, we int…

Cited by 1SourcePDFScholar
2025

IdentityLock: An Identity-aware Backdoor strategy for Face Swapping Defense

ICASSP 2025accepted

DeepFakes have become capable of producing highly realistic fabricated faces, posing significant threats to personal privacy and social security. The uncontrolled spread of such forged content, especially when influential figures are targeted, could lead to catastrophic consequences for society. Exi…

Cited by 0SourceScholar
2025

LineArt: A Knowledge-guided Training-free High-quality Appearance Transfer for Design Drawing with Diffusion Model

CVPR 2025poster

Image rendering from line drawings is vital in design and image generation technologies reduce costs, yet professional line drawings demand preserving complex details. Text prompts struggle with accuracy, and image translation struggles with consistency and fine-grained control. We present LineArt,…

Cited by 1SourcePDFScholar
2025

MLVU: Benchmarking Multi-task Long Video Understanding

CVPR 2025poster

The evaluation of Long Video Understanding (LVU) performance poses an important but challenging research problem. Despite previous efforts, the existing video understanding benchmarks are severely constrained by several issues, especially the insufficient lengths of videos, a lack of diversity in vi…

2025

Multi-Label Prototype Visual Spatial Search for Weakly Supervised Semantic Segmentation

CVPR 2025highlight

Existing Weakly Supervised Semantic Segmentation (WSSS) relies on the CNN-based Class Activation Map (CAM) and Transformer-based self-attention map to generate class-specific masks for semantic segmentation. However, CAM and self-attention maps usually cause incomplete segmentation due to classifica…

Cited by 0SourcePDFScholar
2025

PO3AD: Predicting Point Offsets toward Better 3D Point Cloud Anomaly Detection

CVPR 2025poster

Point cloud anomaly detection under the anomaly-free setting poses significant challenges as it requires accurately capturing the features of 3D normal data to identify deviations indicative of anomalies. Current efforts focus on devising reconstruction tasks, such as acquiring normal data represent…

2025

RoPaSS: Robust Watermarking for Partial Screen-Shooting Scenarios

AAAI 2025technical

Screen-shooting robust watermarking is an effective means of preventing screen content leakage from unauthorized camera shooting, as it can trace the leaked source through the watermark extraction thereby providing an effective deterrent. However, current screen-shooting resilient watermarking schem…

Cited by 0SourcePDFScholar
2025

SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors

NeurIPS 2025poster

While voice technologies increasingly serve aging populations, current systems exhibit significant performance gaps due to inadequate training data capturing elderly-specific vocal characteristics like presbyphonia and dialectal variations. The limited data available on super-aged individuals in exi…

Cited by 0SourcecodeScholar
2025

Surrogate Prompt Learning: Towards Efficient and Diverse Prompt Learning for Vision-Language Models

ICML 2025poster

Prompt learning is a cutting-edge parameter-efficient fine-tuning technique for pre-trained vision-language models (VLMs). Instead of learning a single text prompt, recent works have revealed that learning diverse text prompts can effectively boost the performances on downstream tasks, as the divers…

Cited by 0SourcePDFScholar
2025

Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant Representation

ICCV 2025poster

Variations in medical imaging modalities and individual anatomical differences pose challenges to cross-modality generalization in multi-modal tasks. Existing methods often concentrate exclusively on common anatomical patterns, thereby neglecting individual differences and consequently limiting thei…

2025

Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs

NeurIPS 2025poster

The increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily focus on model vulnerabilities exposed by static image inputs, ignoring the temporal dynamics of video that may induce…

Cited by 0SourcecodeScholar
2024

CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning

IJCAI 2024poster

Multi-modal large language models(MLLMs) have achieved remarkable progress and demonstrated powerful knowledge comprehension and reasoning abilities. However, the mastery of domain-specific knowledge, which is essential for evaluating the intelligence of MLLMs, continues to be a challenge. Current m…

2024

DPIC: Decoupling Prompt and Intrinsic Characteristics for LLM Generated Text Detection

NeurIPS 2024poster

Large language models (LLMs) have the potential to generate texts that pose risks of misuse, such as plagiarism, planting fake reviews on e-commerce platforms, or creating inflammatory false tweets. Consequently, detecting whether a text is generated by LLMs has become increasingly important. Existi…

Cited by 3SourcePDFScholar
2024

Feature-Level Adversarial Attacks and Ranking Disruption for Visible-Infrared Person Re-identification

NeurIPS 2024poster

Visible-infrared person re-identification (VIReID) is widely used in fields such as video surveillance and intelligent transportation, imposing higher demands on model security. In practice, the adversarial attacks based on VIReID aim to disrupt output ranking and quantify the security risks of mode…

Cited by 1SourcePDFScholar
2024

Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification

NeurIPS 2024spotlight

Vision models excel in image classification but struggle to generalize to unseen data, such as classifying images from unseen domains or discovering novel categories. In this paper, we explore the relationship between logical reasoning and deep learning generalization in visual classification. A log…

2024

MuST: Robust Image Watermarking for Multi-Source Tracing

AAAI 2024technical

In recent years, with the popularity of social media applications, massive digital images are available online, which brings great convenience to image recreation. However, the use of unauthorized image materials in multi-source composite images is still inadequately regulated, which may cause signi…

2024

Off-Policy Selection for Initiating Human-Centric Experimental Design

NeurIPS 2024poster

In human-centric applications like healthcare and education, the \textit{heterogeneity} among patients and students necessitates personalized treatments and instructional interventions. While reinforcement learning (RL) has been utilized in those tasks, off-policy selection (OPS) is pivotal to close…

Cited by 0SourcePDFScholar
2024

On Trajectory Augmentations for Off-Policy Evaluation

ICLR 2024poster

In the realm of reinforcement learning (RL), off-policy evaluation (OPE) holds a pivotal position, especially in high-stake human-involved scenarios such as e-learning and healthcare. Applying OPE to these domains is often challenging with scarce and underrepresentative offline training trajectories…

Cited by 4SourcePDFScholar
2024

Optimizing IT FinOps and Sustainability through Unsupervised Workload Characterization

AAAI 2024technical

The widespread adoption of public and hybrid clouds, along with elastic resources and various automation tools for dynamic deployment, has accelerated the rapid provisioning of compute resources as needed. Despite these advancements, numerous resources persist unnecessarily due to factors such as po…

Cited by 2SourcePDFScholar
2024

PairingNet: A Learning-based Pair-searching and -matching Network for Image Fragments

ECCV 2024poster

"In this paper, we propose a learning-based image fragment pair-searching and -matching approach to solve the challenging restoration problem. Existing works use rule-based methods to match similar contour shapes or textures, which are always difficult to tune hyperparameters for extensive data and…

2024

Point Deformable Network with Enhanced Normal Embedding for Point Cloud Analysis

AAAI 2024technical

Recently MLP-based methods have shown strong performance in point cloud analysis. Simple MLP architectures are able to learn geometric features in local point groups yet fail to model long-range dependencies directly. In this paper, we propose Point Deformable Network (PDNet), a concise MLP-based ne…

Cited by 3SourcePDFScholar
2024

Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization

ECCV 2024poster

"Weakly Supervised Object Localization (WSOL), which aims to localize objects by only using image-level labels, has attracted much attention because of its low annotation cost in real applications. Current studies focus on the Class Activation Map (CAM) of CNN and the self-attention map of transform…

Cited by 1SourcePDFScholar
2024

Rethinking Multi-domain Generalization with A General Learning Objective

CVPR 2024poster

Multi-domain generalization (mDG) is universally aimed to minimize the discrepancy between training and testing distributions to enhance marginal-to-label distribution mapping. However existing mDG literature lacks a general learning objective paradigm and often imposes constraints on static target…

2024

Unraveling Batch Normalization for Realistic Test-Time Adaptation

AAAI 2024technical

While recent test-time adaptations exhibit efficacy by adjusting batch normalization to narrow domain disparities, their effectiveness diminishes with realistic mini-batches due to inaccurate target estimation. As previous attempts merely introduce source statistics to mitigate this issue, the funda…

2023

A Benchmark for Chinese-English Scene Text Image Super-Resolution

ICCV 2023poster

Scene Text Image Super-resolution (STISR) aims to recover high-resolution (HR) scene text images with visually pleasant and readable text content from the given low-resolution (LR) input. Most existing works focus on recovering English texts, which have simple structures in the characters, while lit…

Cited by 12PDFcodeScholar
2023

AutoStegaFont: Synthesizing Vector Fonts for Hiding Information in Documents

AAAI 2023technical

Hiding information in text documents has been a hot topic recently, with the most typical schemes of utilizing fonts. By constructing several fonts with similar appearances, information can be effectively represented and embedded in documents. However, due to the unstructured characteristic, font ve…

Cited by 3SourcePDFScholar
2023

Divide and Conquer: 3D Point Cloud Instance Segmentation With Point-Wise Binarization

ICCV 2023poster

Instance segmentation on point clouds is crucially important for 3D scene understanding. Most SOTAs adopt distance clustering, which is typically effective but does not perform well in segmenting adjacent objects with the same semantic label (especially when they share neighboring points). Due to th…

Cited by 32PDFcodeScholar
2023

Fault Injection Based Interventional Causal Learning for Distributed Applications

AAAI 2023technical

We apply the machinery of interventional causal learning with programmable interventions to the domain of applications management. Modern applications are modularized into interdependent components or services (e.g. microservices) for ease of development and management. The communication graph among…

2023

IPMix: Label-Preserving Data Augmentation Method for Training Robust Classifiers

NeurIPS 2023poster

Data augmentation has been proven effective for training high-accuracy convolutional neural network classifiers by preventing overfitting. However, building deep neural networks in real-world scenarios requires not only high accuracy on clean data but also robustness when data distributions shift. W…

2023

MaskBooster: End-to-End Self-Training for Sparsely Supervised Instance Segmentation

AAAI 2023technical

The present paper introduces sparsely supervised instance segmentation, with the datasets being fully annotated bounding boxes and sparsely annotated masks. A direct solution to this task is self-training, which is not fully explored for instance segmentation yet. In this paper, we propose MaskBoost…

Cited by 0SourcePDFScholar
2023

Multi-Granularity Archaeological Dating of Chinese Bronze Dings Based on a Knowledge-Guided Relation Graph

CVPR 2023poster

The archaeological dating of bronze dings has played a critical role in the study of ancient Chinese history. Current archaeology depends on trained experts to carry out bronze dating, which is time-consuming and labor-intensive. For such dating, in this study, we propose a learning-based approach t…

2023

Parts2Words: Learning Joint Embedding of Point Clouds and Texts by Bidirectional Matching Between Parts and Words

CVPR 2023poster

Shape-Text matching is an important task of high-level shape understanding. Current methods mainly represent a 3D shape as multiple 2D rendered views, which obviously can not be understood well due to the structural ambiguity caused by self-occlusion in the limited number of views. To resolve this i…

2023

Rethinking Data Augmentation for Single-Source Domain Generalization in Medical Image Segmentation

AAAI 2023technical

Single-source domain generalization (SDG) in medical image segmentation is a challenging yet essential task as domain shifts are quite common among clinical image datasets. Previous attempts most conduct global-only/random augmentation. Their augmented samples are usually insufficient in diversity…

2022

3D Random Occlusion and Multi-layer Projection for Deep Multi-Camera Pedestrian Localization

ECCV 2022poster

"Although deep-learning based methods for monocular pedestrian detection have made a great progress, they are still vulnerable to heavy occlusions. Using multi-view information fusion is a potential solution but has limited applications, due to the lack of annotated training samples in existing mult…

2022

A Differentiable Two-Stage Alignment Scheme for Burst Image Reconstruction With Large Shift

CVPR 2022poster

Denoising and demosaicking are two essential steps to reconstruct a clean full-color image from the raw data. Recently, joint denoising and demosaicking (JDD) for burst images, namely JDD-B, has attracted much attention by using multiple raw images captured in a short time to reconstruct a single hi…

Cited by 14PDFcodeScholar
2022

A Reinforcement Learning-Informed Pattern Mining Framework for Multivariate Time Series Classification

IJCAI 2022poster

Multivariate time series (MTS) classification is a challenging and important task in various domains and real-world applications. Much of prior work on MTS can be roughly divided into neural network (NN)- and pattern-based methods. The former can lead to robust classification performance, but many o…

2022

Explore Unsupervised Structures in Pretrained Models for Relation Extraction

EMNLP 2022finding

Syntactic trees have been widely applied in relation extraction (RE). However, since parsing qualities are not stable on different text domains and a pre-defined grammar may not well fit the target relation schema, the introduction of syntactic structures sometimes fails to improve RE performances c…

2022

Fine-tuning Deep Neural Networks by Interactively Refining the 2D Latent Space of Ambiguous Images

IJCAI 2022poster

Deep neural networks (DNNs) have achieved excellent results currently in classification, while they may still suffer from ambiguous images which are similar across classes. By contrast, humans have a relatively good ability to distinguish these categories of images. Therefore, we propose a human-in-…

Cited by 5SourcePDFScholar
2022

Towards Semi-Supervised Deep Facial Expression Recognition With an Adaptive Confidence Margin

CVPR 2022poster

Only parts of unlabeled data are selected to train models for most semi-supervised learning methods, whose confidence scores are usually higher than the pre-defined threshold (i.e., the confidence margin). We argue that the recognition performance should be further improved by making full use of all…

Cited by 116PDFcodeScholar
2022

Tracing Text Provenance via Context-Aware Lexical Substitution

AAAI 2022technical

Text content created by humans or language models is often stolen or misused by adversaries. Tracing text provenance can help claim the ownership of text content or identify the malicious users who distribute misleading content like machine-generated fake news. There have been some attempts to achie…

Cited by 70SourcePDFScholar
2021

Real-World Video Super-Resolution: A Benchmark Dataset and a Decomposition Based Learning Scheme

ICCV 2021poster

Video super-resolution (VSR) aims to improve the spatial resolution of low-resolution (LR) videos. Existing VSR methods are mostly trained and evaluated on synthetic datasets, where the LR videos are uniformly downsampled from their high-resolution (HR) counterparts by some simple operators (e.g., b…

Cited by 62PDFcodeScholar
2021

Syncretic Modality Collaborative Learning for Visible Infrared Person Re-Identification

ICCV 2021poster

Visible infrared person re-identification (VI-REID) aims to match pedestrian images between the daytime visible and nighttime infrared camera views. The large cross-modality discrepancies have become the bottleneck which limits the performance of VI-REID. Existing methods mainly focus on capturing c…

Cited by 178PDFScholar
2021

Training Binary Neural Network without Batch Normalization for Image Super-Resolution

AAAI 2021technical

Recently, binary neural network (BNN) based super-resolution (SR) methods have enjoyed initial success in the SR field. However, there is a noticeable performance gap between the binarized model and the full-precision one. Furthermore, the batch normalization (BN) in binary SR networks introduces…

Cited by 46SourcePDFScholar