← Search

Xi Zhang

71 accepted papers

2026

Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and Benchmark

AAAI 2026technical

Accurate point tracking in surgical environments remains challenging due to complex visual conditions, including smoke occlusion, specular reflections, and tissue deformation. While existing surgical tracking datasets provide coordinate information, they lack the semantic context necessary to unders

Cited by 0SourcePDFScholar
2026

Constant Degree Matrix-Driven Incomplete Multi-View Clustering via Connectivity-Structure and Embedding Tensor Learning

ICLR 2026poster

Tensor-based incomplete multi-view clustering has attracted significant research attention due to its capability to exploit high-order correlations across different views for revealing underlying cluster structures from partially observed multi-view data. However, most existing approaches construct…

Cited by 0SourceScholar
2026

Diffusion with a Linguistic Compass: Steering the Generation of Clinically Plausible Future sMRI Representations for Early MCI Conversion Prediction

CVPR 2026

Early prediction of Mild Cognitive Impairment (MCI) conversion is hampered by a trade-off between immediacy--making fast predictions from a single baseline sMRI--and accuracy--leveraging longitudinal scans to capture disease progression. We propose MCI-Diff, a diffusion-based framework that synthesi

Cited by 0SourceScholar
2026

Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products

ICLR 2026poster

Recent theoretical advances reveal that the Hadamard product induces nonlinear representations and implicit high-dimensional mappings for the field of deep learning, yet their practical deployment in efficient vision models remains underdeveloped. To address this gap, we introduce the Adaptive Cross…

Cited by 0SourceScholar
2026

METP: Multi-Granularity Integration of External Covariates for Temporal Point Processes

AAAI 2026technical

Accurate modeling of temporal point processes is critical for reliable event forecasting and informed decision-making. While historical event sequences provide a foundation for intensity estimation, existing approaches often neglect external covariates whose lagged effects impact future intensities

Cited by 0SourcePDFScholar
2026

MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror Crafting

AAAI 2026technical

Defending large language models (LLMs) against jailbreak attacks is crucial for ensuring their safe deployment. Existing defense strategies typically rely on predefined static criteria to differentiate between harmful and benign prompts. However, such rigid rules fail to accommodate the inherent com

Cited by 0SourcePDFScholar
2026

OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents

ICLR 2026poster

With advances in decision-making and reasoning capabilities, multimodal agents show strong potential in computer application scenarios. Past evaluations have mainly assessed GUI interaction skills, while tool invocation abilities, such as those enabled by the Model Context Protocol (MCP), have been…

Cited by 0SourcecodeScholar
2026

SecMoE: Communication-Efficient Secure MoE Inference via Select-Then-Compute

AAAI 2026technical

Privacy-preserving Transformer inference has gained attention due to the potential leakage of private information. Despite recent progress, existing frameworks still fall short of practical model scales, with gaps up to a hundredfold. A possible way to close this gap is the Mixture of Experts (MoE)

Cited by 0SourcePDFScholar
2026

Structured Labeling Enables Faster Vision-Language Models for End-To-End Autonomous Driving

ICRA 2026poster

Vision-Language Models (VLMs) offer a promising approach to end-to-end autonomous driving due to their human-like reasoning capabilities. However, troublesome gaps remains between current VLMs and real-world autonomous driving applications. One major limitation is that existing datasets with loosely…

2026

Towards the Explainability of Temporal Graph Networks via Memory Backtracking and Topological Attribution

ICML 2026spotlight

Temporal graphs are ubiquitous in real-world applications such as social networks and finance, where Temporal Graph Networks (TGNs) capture both structural and temporal dependencies, achieving in superior predictive accuracy. Understanding which historical events drive specific model predictions can…

Cited by 0SourceScholar
2026

TruthfulRAG: Resolving Factual-level Conflicts in Retrieval-Augmented Generation with Knowledge Graphs

AAAI 2026technical

Retrieval-Augmented Generation (RAG) has emerged as a powerful framework for enhancing the capabilities of Large Language Models (LLMs) by integrating retrieval-based methods with generative models. As external knowledge repositories continue to expand and the parametric knowledge within models beco

Cited by 0SourcePDFScholar
2026

X-EviProbe: Post-hoc Parameter-free Evidential Uncertainty Quantification for Frozen Graph Neural Networks

ICML 2026poster

Reliable uncertainty quantification (UQ) is crucial for deploying graph neural networks (GNNs) in safety-critical settings, yet dominant solutions either rely on costly multi-pass sampling or require retraining—often using *black-box auxiliary* models—to obtain evidential semantics. We propose **X-E…

Cited by 0SourceScholar
2026

XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

ICML 2026poster

Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hindered by the limited availability of sensor data and resource constraints of edge systems. While transferring pre-trained models to different sensing ap…

Cited by 0SourceScholar
2025

Automated Detection of Pre-training Text in Black-box LLMs

IJCAI 2025

Detecting whether a given text is a member in the pre-training data of Large Language Models (LLMs) is crucial for ensuring data privacy and copyright protection. Most existing methods rely on the LLM's hidden information (e.g., model parameters or token probabilities), making them ineffective in th

2025

Beyond Text: Fine-Grained Multi-Modal Fact Verification with Hypergraph Transformers

AAAI 2025technical

Fact verification has become increasingly vital in the internet age, driven by the proliferation of false claims and political misinformation. While traditional methods rely predominantly on text-based evidence, multi-modal evidence introduces richer sources of information, offering valuable insigh…

Cited by 0SourcePDFScholar
2025

CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models

ACL 2025long

Faithfulness hallucinations are claims generated by a Large Language Model (LLM) not supported by contexts provided to the LLM. Lacking assessment standards, existing benchmarks focus on “factual statements” that rephrase source materials while overlooking “cognitive statements” that involve making…

2025

Curriculum Hierarchical Knowledge Distillation for Bias-Free Survival Prediction

IJCAI 2025

Survival prediction is a pivotal task for estimating mortality risk within a given timeframe based on whole slide images (WSIs). Conventional models typically assume that WSIs across patients are independent and identically distributed, an assumption that may not hold due to inherent variability in

Cited by 0SourcePDFScholar
2025

DSG-MCTS: A Dynamic Strategy-Guided Monte Carlo Tree Search for Diversified Reasoning in Large Language Models

EMNLP 2025

Large language models (LLMs) have shown strong potential in complex reasoning tasks. However, as task complexity increases, their performance often degrades, resulting in hallucinations, errors, and logical inconsistencies. To enhance reasoning capabilities, Monte Carlo Tree Search (MCTS) has been i

Cited by 0SourcePDFScholar
2025

Decision-Making for Autonomous Driving via a Coupled Reinforcement Learning Network Combined With Risk Assessment

RA-L 2025

The realization of autonomous driving(AV) is closely linked to the development of intelligent decision-making modules that can operate safely in dynamic, uncertain environments. To address issues such as delayed response and poor coupling in highway scenarios, this paper proposes a hierarchical Coup

Cited by 1SourceScholar
2025

Deep Rank-One Tensor Functional Factorization for Multi-Dimensional Data Recovery

AAAI 2025technical

Many real-world data are inherently multi-dimensional, e.g., color images, videos, and hyperspectral images. How to effectively and compactly represent these multi-dimensional data within a unified framework is an important pursuit. Previous methods focus on tensor factorizations, convolutional netw…

Cited by 0SourcePDFScholar
2025

From Representation Space to Prognostic Insights: Whole Slide Image Generation with Hierarchical Diffusion Model for Survival Prediction

AAAI 2025technical

Deep learning has significantly enhanced survival prediction using whole slide images (WSIs) by adopting a two-stage learning paradigm: WSI preparation and patient-level prediction. While existing research generally concentrates on developing advanced patient-level prediction modules, the critical i…

Cited by 0SourcePDFScholar
2025

Ghidorah: Towards Robust Multi-Scale Information Diffusion Prediction via Test-Time Training

AAAI 2025technical

Information diffusion prediction (IDP) is a pivotal task for understanding the dynamics of information propagation within social networks. Conventional models typically adhere to a fixed learning-based paradigm, where the trained prediction model remains static during the inference phase. This parad…

Cited by 0SourcePDFScholar
2025

Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression

NeurIPS 2025poster

Large Language Models (LLMs) have demonstrated remarkable capabilities but typically require extensive computational resources and memory for inference. Post-training quantization (PTQ) can effectively reduce these demands by storing weights in lower bit-width formats. However, standard uniform quan…

Cited by 0SourcecodeScholar
2025

Libra: Leveraging Temporal Images for Biomedical Radiology Analysis

ACL 2025finding

Radiology report generation (RRG) requires advanced medical image analysis, effective temporal reasoning, and accurate text generation. While multimodal large language models (MLLMs) align with pre-trained vision encoders to enhance visual-language understanding, most existing methods rely on single…

2025

Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation

NeurIPS 2025poster

In recent years, Multimodal Large Language Models (MLLMs) have been extensively utilized for multimodal reasoning tasks, including Graphical User Interface (GUI) automation. Unlike general offline multimodal tasks, GUI automation is executed in online interactive environments, necessitating step-by-…

Cited by 0SourcecodeScholar
2025

MRR-FV: Unlocking Complex Fact Verification with Multi-Hop Retrieval and Reasoning

AAAI 2025technical

The pervasive spread of misinformation on social networks highlights the critical necessity for effective fact verification systems. Traditional approaches primarily focus on pairwise correlations between claims and evidence, often neglecting comprehensive multi-hop retrieval and reasoning, which re…

Cited by 0SourcePDFScholar
2025

Meta Flow Matching: Integrating Vector Fields on the Wasserstein Manifold

ICLR 2025poster

Numerous biological and physical processes can be modeled as systems of interacting entities evolving continuously over time, e.g. the dynamics of communicating cells or physical particles. Learning the dynamics of such systems is essential for predicting the temporal evolution of populations across…

Cited by 6SourcePDFScholar
2025

Multirate Neural Image Compression with Adaptive Lattice Vector Quantization

CVPR 2025highlight

Recent research has explored integrating lattice vector quantization (LVQ) into learned image compression models. Due to its more efficient Voronoi covering of vector space than scalar quantization (SQ), LVQ achieves better rate-distortion (R-D) performance than SQ, while still retaining the low com…

2025

One SPACE to Rule Them All: Jointly Mitigating Factuality and Faithfulness Hallucinations in LLMs

NeurIPS 2025poster

LLMs have demonstrated unprecedented capabilities in natural language processing, yet their practical deployment remains hindered by persistent factuality and faithfulness hallucinations. While existing methods address these hallucination types independently, they inadvertently induce performance tr…

Cited by 0SourceScholar
2025

Online Functional Tensor Decomposition via Continual Learning for Streaming Data Completion

NeurIPS 2025spotlight

Online tensor decompositions are powerful and proven techniques that address the challenges in processing high-velocity streaming tensor data, such as traffic flow and weather system. The main aim of this work is to propose a novel online functional tensor decomposition (OFTD) framework, which repre…

Cited by 0SourceScholar
2025

Overcoming Dual Drift for Continual Long-Tailed Visual Question Answering

ICCV 2025poster

Visual Question Answering (VQA) is a widely explored multimodal task aimed at answering questions based on images. Recently, a few studies have started to investigate continual learning in VQA to cope with evolving multimodal data streams. However, these studies fall short of tackling another critic…

Cited by 0SourcePDFScholar
2025

Robust Explanations of Graph Neural Networks via Graph Curvatures

NeurIPS 2025poster

Explaining graph neural networks (GNNs) is a key approach to improve the trustworthiness of GNN in high-stakes applications, such as finance and healthcare. However, existing methods are vulnerable to perturbations, raising concerns about explanation reliability. Prior methods enhance explanation ro…

Cited by 0SourcecodeScholar
2025

SCCD: A Session-based Dataset for Chinese Cyberbullying Detection

COLING 2025main

The rampant spread of cyberbullying content poses a growing threat to societal well-being. However, research on cyberbullying detection in Chinese remains underdeveloped, primarily due to the lack of comprehensive and reliable datasets. Notably, no existing Chinese dataset is specifically tailored f…

2025

Semantic-Guided Illumination-Aware Deformable Transformer for RGB-T Object Detection

RA-L 2025

RGB-T object detection in autonomous driving has been researched increasingly in recent years. Nevertheless, several problems limit the performance of RGB-T fusion perception. Initially, although illumination awareness is a mature technology to guide fusion process, the outputs of previous methods l

Cited by 1SourceScholar
2025

When Open-Vocabulary Visual Question Answering Meets Causal Adapter: Benchmark and Approach

AAAI 2025technical

Visual Question Answering (VQA) is a multifaceted task that integrates computer vision and natural language processing to produce textual answers from images and questions. Existing VQA benchmarks predominantly adhere to a closed-set paradigm, limiting their ability to address arbitrary, unseen answ…

Cited by 0SourcePDFScholar
2025

Zero-shot Stance Detection with Logically Consistent Data Augmentation

ICASSP 2025accepted

Zero-shot stance detection (ZSSD) is a challenging task that requires classifying stances towards unseen targets without large, well-curated training datasets. Existing data augmentation methods for ZSSD often suffer from semantic inconsistencies, hindering their effectiveness. To address these limi…

Cited by 0SourceScholar
2024

A New Pre-Training Paradigm for Offline Multi-Agent Reinforcement Learning with Suboptimal Data

ICASSP 2024accepted

Offline multi-agent reinforcement learning (MARL) with pre-training paradigm, which uses a large quantity of trajectories for offline pre-training and online deployment, has become fashionable lately. While performing well on various tasks, conventional pre-trained decision-making models based on im…

Cited by 0SourceScholar
2024

BaitAttack: Alleviating Intention Shift in Jailbreak Attacks via Adaptive Bait Crafting

EMNLP 2024main

Jailbreak attacks enable malicious queries to evade detection by LLMs. Existing attacks focus on meticulously constructing prompts to disguise harmful intentions. However, the incorporation of sophisticated disguising prompts may incur the challenge of “intention shift”. Intention shift occurs when…

Cited by 3SourcePDFScholar
2024

Enhancing Robustness of Graph Neural Networks on Social Media with Explainable Inverse Reinforcement Learning

NeurIPS 2024spotlight

Adversarial attacks against graph neural networks (GNNs) through perturbations of the graph structure are increasingly common in social network tasks like rumor detection. Social media platforms capture diverse attack sequence samples through both machine and manual screening processes. Investigatin…

Cited by 2SourcePDFScholar
2024

Evidence Retrieval is almost All You Need for Fact Verification

ACL 2024findings

Current fact verification methods generally follow the two-stage training paradigm: evidence retrieval and claim verification. While existing works focus on developing sophisticated claim verification modules, the fundamental importance of evidence retrieval is largely ignored. Existing approaches u…

Cited by 4SourcePDFScholar
2024

Fast Point Cloud Geometry Compression with Context-based Residual Coding and INR-based Refinement

ECCV 2024poster

"Compressing a set of unordered points is far more challenging than compressing images/videos of regular sample grids, because of the difficulties in characterizing neighboring relations in an irregular layout of points. Many researchers resort to voxelization to introduce regularity, but this appro…

2024

Linear Uncertainty Quantification of Graphical Model Inference

NeurIPS 2024poster

Uncertainty Quantification (UQ) is vital for decision makers as it offers insights into the potential reliability of data and model, enabling more informed and risk-aware decision-making. Graphical models, capable of representing data with complex dependencies, are widely used across domains. Exist…

Cited by 0SourcePDFScholar
2024

MIBench: Evaluating Multimodal Large Language Models over Multiple Images

EMNLP 2024main

Built on the power of LLMs, numerous multimodal large language models (MLLMs) have recently achieved remarkable performance on various vision-language tasks. However, most existing MLLMs and benchmarks primarily focus on single-image input scenarios, leaving the performance of MLLMs when handling re…

Cited by 10SourcePDFScholar
2024

Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration

NeurIPS 2024poster

Mobile device operation tasks are increasingly becoming a popular multi-modal AI application scenario. Current Multi-modal Large Language Models (MLLMs), constrained by their training data, lack the capability to function effectively as operation assistants. Instead, MLLM-based agents, which enhance…

2024

Training for Stable Explanation for Free

NeurIPS 2024poster

To foster trust in machine learning models, explanations must be faithful and stable for consistent insights. Existing relevant works rely on the $\ell_p$ distance for stability assessment, which diverges from human perception. Besides, existing adversarial training (AT) associated with intensive co…

2024

Trajectory Flow Matching with Applications to Clinical Time Series Modelling

NeurIPS 2024spotlight

Modeling stochastic and irregularly sampled time series is a challenging problem found in a wide range of applications, especially in medicine. Neural stochastic differential equations (Neural SDEs) are an attractive modeling technique for this problem, which parameterize the drift and diffusion ter…

Cited by 7SourcePDFScholar
2023

Anytime User Engagement Prediction in Information Cascades for Arbitrary Observation Periods

AAAI 2023technical

Predicting user engagement -- whether a user will engage in a given information cascade -- is an important problem in the context of social media, as it is useful to online marketing and misinformation mitigation just to name a couple major applications. Based on split population multi-variate survi…

Cited by 3SourcePDFScholar
2023

LVQAC: Lattice Vector Quantization Coupled With Spatially Adaptive Companding for Efficient Learned Image Compression

CVPR 2023poster

Recently, numerous end-to-end optimized image compression neural networks have been developed and proved themselves as leaders in rate-distortion performance. The main strength of these learnt compression methods is in powerful nonlinear analysis and synthesis transforms that can be facilitated by d…

Cited by 30SourcePDFScholar
2023

VQACL: A Novel Visual Question Answering Continual Learning Setting

CVPR 2023poster

Research on continual learning has recently led to a variety of work in unimodal community, however little attention has been paid to multimodal tasks like visual question answering (VQA). In this paper, we establish a novel VQA Continual Learning setting named VQACL, which contains two key componen…

2022

ABO: Dataset and Benchmarks for Real-World 3D Object Understanding

CVPR 2022poster

We introduce Amazon Berkeley Objects (ABO), a new large-scale dataset designed to help bridge the gap between real and virtual 3D worlds. ABO contains product catalog images, metadata, and artist-created 3D models with complex geometries and physically-based materials that correspond to real, househ…

Cited by 225PDFcodeScholar
2022

Anytime Information Cascade Popularity Prediction via Self-Exciting Processes

ICML 2022spotlight

One important aspect of understanding behaviors of information cascades is to be able to accurately predict their popularity, that is, their message counts at any future time. Self-exciting Hawkes processes have been widely adopted for such tasks due to their success in describing cascading behavior…

2022

DDGCN: Dual Dynamic Graph Convolutional Networks for Rumor Detection on Social Media

AAAI 2022technical

Detecting rumors on social media has become particular important due to the rapid dissemination and adverse impacts on our lives. Though a set of rumor detection models have exploited the message propagation structural or temporal information, they seldom model them altogether to enjoy the best of b…

Cited by 104SourcePDFScholar
2022

Explainable Survival Analysis with Convolution-Involved Vision Transformer

AAAI 2022technical

Image-based survival prediction models can facilitate doctors in diagnosing and treating cancer patients. With the advance of digital pathology technologies, the big whole slide images (WSIs) provide increasing resolution and more details for diagnosis. However, the gigabyte-size WSIs would make mos…

2022

MFAN: Multi-modal Feature-enhanced Attention Networks for Rumor Detection

IJCAI 2022poster

Rumor spreaders are increasingly taking advantage of multimedia content to attract and mislead news consumers on social media. Although recent multimedia rumor detection models have exploited both textual and visual features for classification, they do not integrate the social structure features sim…

Cited by 76SourcePDFScholar
2021

Inconsistency Matters: A Knowledge-guided Dual-inconsistency Network for Multi-modal Rumor Detection

EMNLP 2021finding

Rumor spreaders are increasingly utilizing multimedia content to attract the attention and trust of news consumers. Though a set of rumor detection models have exploited the multi-modal data, they seldom consider the inconsistent relationships among images and texts. Moreover, they also fail to find…

2020

DAVD-Net: Deep Audio-Aided Video Decompression of Talking Heads

CVPR 2020oral

Close-up talking heads are among the most common and salient object in video contents, such as face-to-face conversations in social media, teleconferences, news broadcasting, talk shows, etc. Due to the high sensitivity of human visual system to faces, compression distortions in talking heads videos…

Cited by 38PDFScholar
2020

Rumor Detection on Social Media with Graph Structured Adversarial Learning

IJCAI 2020poster

The wide spread of rumors on social media has caused tremendous effects in both the online and offline world. In addition to text information, recent detection methods began to exploit the graph structure in the propagation network. However, without a rigorous design, rumors may evade such graph mod…

Cited by 0SourcePDFScholar
2019

Nonlinear Prediction of Multidimensional Signals via Deep Regression with Applications to Image Coding

ICASSP 2019accepted

Deep convolutional neural networks (DCNN) have enjoyed great successes in many signal processing applications because they can learn complex, non-linear causal relationships from input to output. In this light, DCNNs are well suited for the task of sequential prediction of multidimensional signals,…

Cited by 0SourceScholar
2018

Intervention Aided Reinforcement Learning for Safe and Practical Policy Optimization in Navigation

CoRL 2018

Combining deep neural networks with reinforcement learning has shown great potential in the next-generation intelligent control. However, there are challenges in terms of safety and cost in practical applications. In this pa- per, we propose the Intervention Aided Reinforcement Learning (IARL) frame

2017

An Automatic Condition Detection Approach for Quality Assurance in Solar Cell Manufacturing Processes

RA-L 2017

Solar conversion efficiency is one of the most important quality metrics in solar cell production. During the production process, many operating factors might potentially influence the process conditions, thereby leading to unstable solar conversion efficiency of solar cell products. However, most s

Cited by 3SourceScholar
2015

Tracking benchmark and evaluation for manipulation tasks

ICRA 2015poster

In this paper we present a public dataset to evaluate trackers used for human and robot manipulation tasks. For these tasks both high DOF motion and high accuracy is needed. We describe in detail, both the process of recording the sequences and how ground truth data was generated for the videos. The…

Cited by 43SourceScholar