← Search

Xin Shen

16 accepted papers

2026

IVQA-LD: Inclusive Multimodal Understanding for Population with Limb-Deficiency

ICML 2026poster

People with limb differences often face significant challenges in accessing inclusive AI services, largely due to the lack of structured, high-quality resources centered on disability contexts. In this work, we introduce a limb-deficiency aware body-centric learning and evaluation paradigm that invo…

Cited by 0SourceScholar
2026

LEARNING DOMAIN-ROBUST BIOACOUSTIC REPRESENTATIONS FOR MOSQUITO SPECIES CLASSIFICATION WITH CONTRASTIVE LEARNING AND DISTRIBUTION ALIGNMENT

ICASSP 2026poster

Mosquito Species Classification (MSC) is crucial for vector surveillance and disease control. The collection of mosquito bioacoustic data is often limited by mosquito activity seasons and fieldwork. Mosquito recordings across regions, habitats, and laboratories often show non-biological variations f…

Cited by 0SourcePDFScholar
2026

MAKP: Multi-Mode Accurate Kicking Policy for Humanoid Robots

ICRA 2026poster

Humanoid robot soccer players face fundamental challenges in achieving stable motion execution and ball trajectory control, particularly under balance constraints during single-leg support phases. In this paper, we introduce MAKP (Multi-mode Accurate Kicking Policy), a novel motion generation-based …

Cited by 0Scholar
2026

PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling

CVPR 2026

Consistent image generation requires faithfully preserving identities, styles, and logical coherence across multiple images,which is essential for applications such as storytelling and character design.Supervised training approaches struggle with this task due to the lack of large-scale datasets cap

Cited by 0SourcecodeScholar
2026

Stay in Character, Stay Safe: Dual-Cycle Adversarial Self-Evolution for Role-Playing Agents

IJCAI 2026

LLM-based role-playing has rapidly improved in fidelity, yet stronger adherence to persona constraints commonly increases vulnerability to jailbreak attacks, especially for risky or negative personas. Most prior work mitigates this issue with training-time solutions (e.g., data curation or alignment

Cited by 0Scholar
2026

Zero-shot Implicit Neural Manifold Representation (INMR) for Ultra-high Temporal Resolution Dynamic MRI

AAAI 2026technical

Capturing accurate dynamic information of moving organs is essential for functional assessment using non-invasive imaging modalities. Achieving high temporal resolution visualization of physiological processes remains a critical challenge in dynamic magnetic resonance imaging (MRI) when reconstructi

Cited by 0SourcePDFScholar
2025

Blind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion Model

CVPR 2025poster

Bitstream-corrupted video recovery aims to fill in realistic video content due to bitstream corruption during video storage or transmission. Most existing methods typically assume that the predefined masks of the corrupted regions are known in advance. However, manually annotating these masks is lab…

Cited by 0SourcePDFScholar
2025

Cross-View Isolated Sign Language Recognition via View Synthesis and Feature Disentanglement

ICCV 2025poster

Cross-view isolated sign language recognition (CV-ISLR) addresses the challenge of identifying isolated signs from viewpoints unseen during training, a problem aggravated by the scarcity of multi-view data in existing benchmarks. To bridge this gap, we introduce a novel two-stage framework comprisin…

Cited by 0SourcePDFScholar
2025

M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world Settings

CVPR 2025poster

Human pose estimation is a critical task in computer vision for applications in sports analysis, healthcare monitoring, and human-computer interaction. However, existing human pose datasets are collected either from custom-configured laboratories with complex devices or they only include data on sin…

Cited by 0SourcePDFScholar
2025

Multimodal Retina Image Analysis Survey: Datasets, Tasks and Methods

IJCAI 2025

Retina images provide a noninvasive view of the central nervous system and microvasculature, making it essential for clinical applications. Changes in the retina often indicate both ophthalmic and systemic diseases, aiding in diagnosis and early intervention. While deep learning algorithms have adva

Cited by 0SourcePDFScholar
2024

MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition Dataset

NeurIPS 2024poster

Isolated Sign Language Recognition (ISLR) focuses on identifying individual sign language glosses. Considering the diversity of sign languages across geographical regions, developing region-specific ISLR datasets is crucial for supporting communication and research. Auslan, as a sign language specif…

Cited by 0SourcePDFScholar
2023

Auslan-Daily: Australian Sign Language Translation for Daily Communication and News

NeurIPS 2023poster

Sign language translation (SLT) aims to convert a continuous sign language video clip into a spoken language. Considering different geographic regions generally have their own native sign languages, it is valuable to establish corresponding SLT datasets to support related communication and research.…

Cited by 18SourcePDFScholar
2023

MNER-QG: An End-to-End MRC Framework for Multimodal Named Entity Recognition with Query Grounding

AAAI 2023technical

Multimodal named entity recognition (MNER) is a critical step in information extraction, which aims to detect entity spans and classify them to corresponding entity types given a sentence-image pair. Existing methods either (1) obtain named entities with coarse-grained visual clues from attention me…

Cited by 53SourcePDFScholar
2021

Learning to Select Context in a Hierarchical and Global Perspective for Open-Domain Dialogue Generation

ICASSP 2021accepted

Open-domain multi-turn conversations mainly have three features, which are hierarchical semantic structure, redundant information, and long-term dependency. Grounded on these, selecting relevant context becomes a challenge step for multiturn dialogue generation. However, existing methods cannot diff…

Cited by 0SourceScholar