← Search

Xi Zhou

21 accepted papers

2026

M3UCD: A Multi-task Multimodal Metaphor Understanding Challenge Dataset for LLMs

AAAI 2026technical

Understanding multimodal metaphors represents a crucial pathway for machines to comprehend human cognition. However, current research remains constrained by superficial dataset annotations, insufficient systematic evaluation of large language models, and fragmented task frameworks. To bridge these g

Cited by 0SourcePDFScholar
2025

Beyond Inherent Cognition Biases in LLM-Based Event Forecasting: A Multi-Cognition Agentic Framework

EMNLP 2025

Large Language Models (LLMs) exhibit strong reasoning capabilities and are widely applied in event forecasting. However, studies have demonstrated that LLMs exhibit human-like cognitive biases, systematic patterns of deviation from rationality in decision-making. To explore the cognitive biases in e

Cited by 0SourcePDFScholar
2025

FedSe: Group-Based Sequential Training Strategies for Mitigating Label Skew in Federated Learning

ICASSP 2025accepted

Federated Learning (FL) has emerged as a promising approach for distributed machine learning, enabling clients to collaboratively train models without sharing their data. However, existing FL methods continue to face challenges when dealing with non-IID data, particularly under conditions of extreme…

Cited by 0SourceScholar
2025

Low-Resource Language Expansion and Translation Capacity Enhancement for LLM: A Study on the Uyghur

COLING 2025main

Although large language models have significantly advanced natural language generation, their potential in low-resource machine translation has not yet been fully explored, especially for languages that translation models have not been trained on. In this study, we provide a detailed demonstration o…

2025

MambaTrack: Exploiting Dual-Enhancement for Night UAV Tracking

ICASSP 2025accepted

Night unmanned aerial vehicle (UAV) tracking is impeded by the challenges of poor illumination, with previous daylight-optimized methods demonstrating suboptimal performance in low-light conditions, limiting the utility of UAV applications. To this end, we propose an efficient mamba-based tracker, l…

Cited by 0SourceScholar
2025

Open-Set Cross-Network Node Classification via Unknown-Excluded Adversarial Graph Domain Alignment

AAAI 2025technical

Existing cross-network node classification methods are mainly proposed for closed-set setting, where the source network and the target network share exactly the same label space. Such a setting is restricted in real-world applications, since the target network might contain additional classes that a…

2025

OpenForecast: A Large-Scale Open-Ended Event Forecasting Dataset

COLING 2025main

Complex events generally exhibit unforeseen, multifaceted, and multi-step developments, and cannot be well handled by existing closed-ended event forecasting methods, which are constrained by a limited answer space. In order to accelerate the research on complex event forecasting, we introduce OpenF…

2024

On the Federated Learning Framework for Cooperative Perception

RA-L 2024

Cooperative perception (CP) is essential to enhance the efficiency and safety of future transportation systems, requiring extensive data sharing among vehicles on the road, which raises significant privacy concerns. Federated learning offers a promising solution by enabling data privacy-preserving c

Cited by 10SourceScholar
2024

WebUOT-1M: Advancing Deep Underwater Object Tracking with A Million-Scale Benchmark

NeurIPS 2024poster

Underwater Object Tracking (UOT) is essential for identifying and tracking submerged objects in underwater videos, but existing datasets are limited in scale, diversity of target categories and scenarios covered, impeding the development of advanced tracking algorithms. To bridge this gap, we take t…

2023

A Domain-Transfer Meta Task Design Paradigm for Few-Shot Slot Tagging

AAAI 2023technical

Few-shot slot tagging is an important task in dialogue systems and attracts much attention of researchers. Most previous few-shot slot tagging methods utilize meta-learning procedure for training and strive to construct a large number of different meta tasks to simulate the testing situation of insu…

Cited by 0SourcePDFScholar
2023

A Slot-Shared Span Prediction-Based Neural Network for Multi-Domain Dialogue State Tracking

ICASSP 2023accepted

There are a large number of candidate values shared among slots in multi-domain dialogue state tracking (DST). The existing span prediction-based DST methods generally adopt slot-independent value extraction architecture, which ignore the value sharing. Besides, the slot-independent design leads to…

Cited by 0SourceScholar
2023

Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos

ICCV 2023poster

Recently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, annotating point cloud videos is usually notoriously expensive. Moreover, training via one or only a few traditional task…

Cited by 18PDFcodeScholar
2023

Neighbor Contrastive Learning on Learnable Graph Augmentation

AAAI 2023technical

Recent years, graph contrastive learning (GCL), which aims to learn representations from unlabeled graphs, has made great progress. However, the existing GCL methods mostly adopt human-designed graph augmentations, which are sensitive to various graph datasets. In addition, the contrastive losses or…

2023

PointCMP: Contrastive Mask Prediction for Self-Supervised Learning on Point Cloud Videos

CVPR 2023poster

Self-supervised learning can extract representations of good quality from solely unlabeled data, which is appealing for point cloud videos due to their high labelling cost. In this paper, we propose a contrastive mask prediction (PointCMP) framework for self-supervised learning on point cloud videos…

2022

Diversity Features Enhanced Prototypical Network for Few-shot Intent Detection

IJCAI 2022poster

Few-shot Intent Detection (FSID) is a challenging task in dialogue systems due to the scarcity of available annotated utterances. Although existing few-shot learning approaches have made remarkable progress, they fall short in adapting to the Generalized Few-shot Intent Detection (GFSID) task where…

Cited by 11SourcePDFScholar
2021

Filling the Gap of Utterance-aware and Speaker-aware Representation for Multi-turn Dialogue

AAAI 2021technical

A multi-turn dialogue is composed of multiple utterances from two or more different speaker roles. Thus utterance- and speaker-aware clues are supposed to be well captured in models. However, in the existing retrieval-based multi-turn dialogue modeling, the pre-trained language models (PrLMs) as enc…

2021

Relation-aware Video Reading Comprehension for Temporal Language Grounding

EMNLP 2021main

Temporal language grounding in videos aims to localize the temporal span relevant to the given query sentence. Previous methods treat it either as a boundary regression task or a span extraction task. This paper will formulate temporal language grounding into video reading comprehension and propose…

2021

Semantics-Aware Inferential Network for Natural Language Understanding

AAAI 2021technical

For natural language understanding tasks, either machine reading comprehension or natural language inference, both semantics-aware and inference are favorable features of the concerned modeling for better understanding performance. Thus we propose a Semantics-Aware Inferential Network (SAIN) to meet…

2018

Joint 3D Face Reconstruction and Dense Alignment with Position Map Regression Network

ECCV 2018poster

We propose a straightforward method that simultaneously reconstructs the 3D facial structure and provides dense alignment. To achieve this, we design a 2D representation called UV position map which records the 3D shape of a complete face in UV space, then train a simple Convolutional Neural Network…

2017

A Deep Regression Architecture With Two-Stage Re-Initialization for High Performance Facial Landmark Detection

CVPR 2017poster

Regression based facial landmark detection methods usually learns a series of regression functions to update the landmark positions from an initial estimation. Most of existing approaches focus on learning effective mapping functions with robust image features to improve performance. The approach to…

Cited by 308PDFScholar