← Search

Xinxin Zhang

9 accepted papers

2026

Boomda: Balanced Multi-objective Optimization for Multimodal Domain Adaptation

AAAI 2026technical

Multimodal learning, while contributing to numerous success stories across various fields, faces the challenge of prohibitively expensive manual annotation. To address the scarcity of annotated data, a popular solution is unsupervised domain adaptation, which has been extensively studied in unimodal

Cited by 0SourcePDFScholar
2025

Adversarial Alignment with Anchor Dragging Drift (A3D2): Multimodal Domain Adaptation with Partially Shifted Modalities

ACL 2025long

Multimodal learning has celebrated remarkable success across diverse areas, yet faces the challenge of prohibitively expensive data collection and annotation when adapting models to new environments. In this context, domain adaptation has gained growing popularity as a technique for knowledge transf…

2025

Disparity-Guided Cross-View Transformer For Stereo Image Super-Resolution

ICASSP 2025accepted

Although transformer-based methods excel in stereo image super-resolution, the full potential of the distinctive, complementary information inherent in stereo images has not been fully utilized. We propose a Disparity-Guided Cross-View Transformer (DCT) to extract features across dimensions and view…

Cited by 0SourceScholar
2025

Improving Generalization in Meta-Learning via Meta-Gradient Augmentation

IJCAI 2025

Meta-learning methods typically follow a two-loop framework, where each loop potentially suffers from notorious overfitting, hindering rapid adaptation and generalization to new tasks. Existing methods address this by enhancing the mutual-exclusivity or diversity of training samples, but these data

2025

LOFI: Harnessing Attention Dynamics for Facial Expression Recognition with Noisy Labels

ICASSP 2025accepted

Facial expression recognition (FER) faces unique challenges from expression ambiguity and noisy labels, degrading performance in real-world applications. While leveraging attention, existing methods frequently neglect attention dynamic mechanism of dispersion followed by focus and the spatially stru…

Cited by 0SourceScholar
2025

Towards Reliable LLM-based Robots Planning via Combined Uncertainty Estimation

NeurIPS 2025poster

Large language models (LLMs) demonstrate advanced reasoning abilities, enabling robots to understand natural language instructions and generate high-level plans with appropriate grounding. However, LLM hallucinations present a significant challenge, often leading to overconfident yet potentially mis…

Cited by 0SourcecodeScholar
2024

Amanda: Adaptively Modality-Balanced Domain Adaptation for Multimodal Emotion Recognition

ACL 2024findings

This paper investigates unsupervised multimodal domain adaptation for multimodal emotion recognition, which is a solution for data scarcity yet remains under studied. Due to the varying distribution discrepancies of different modalities between source and target domains, the primary challenge lies i…

2024

RedCore: Relative Advantage Aware Cross-Modal Representation Learning for Missing Modalities with Imbalanced Missing Rates

AAAI 2024technical

Multimodal learning is susceptible to modality missing, which poses a major obstacle for its practical applications and, thus, invigorates increasing research interest. In this paper, we investigate two challenging problems: 1) when modality missing exists in the training data, how to exploit the in…

2024

Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation

ECCV 2024poster

"Open-Vocabulary Video Instance Segmentation (VIS) is attracting increasing attention due to its ability to segment and track arbitrary objects. However, the recent Open-Vocabulary VIS attempts obtained unsatisfactory results, especially in terms of generalization ability of novel categories. We dis…