← Search

Xinyue Zhang

19 accepted papers

2026

G-VTM: A Multimodal Vision-Trajectory Model for Generalized Vehicle Trajectory Prediction

IJCAI 2026

Generalized vehicle trajectory prediction across diverse junctions, including urban intersections and roundabouts, remains a fundamental task in Cooperative Vehicle–Infrastructure Systems (CVIS). This study faces two key challenges: (1) Generalize across junctions with heterogeneous map semantics an

Cited by 0Scholar
2026

Learnability-Driven Knowledge Assimilation for Class-Incremental Semantic Segmentation

ICML 2026poster

Class-incremental semantic segmentation learns new classes while retaining old ones without access to past data. Although existing methods alleviate catastrophic forgetting on old classes, new-class performance remains limited. We identify the key bottleneck arises from low-margin regions, where the…

Cited by 0SourceScholar
2026

Learning from Disagreement: A Group Decision Simulation Framework for Robust Medical Image Segmentation

ICASSP 2026poster

Medical image segmentation annotation suffers from inter-rater variability (IRV) due to differences in annotators' expertise and the inherent blurriness of medical images. Standard approaches that simply average expert labels are flawed, as they discard the valuable clinical uncertainty revealed in…

Cited by 0SourcePDFScholar
2026

TrajAR: Long-Term Trajectory Prediction at Urban Intersections via Multi-scale Interaction Perception

IJCAI 2026

Accurate trajectory prediction of multiple road users at urban intersections--including motorized and nonmotorized vehicles and pedestrians--is critical for cooperative vehicle-infrastructure systems and intelligent transportation systems. This study focuses on multiple road users' trajectory predic

Cited by 0Scholar
2026

VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning

ICML 2026poster

Multi-model learning has attracted great attention in visual-text tasks. However, visual-tabular data, which plays a pivotal role in high-stakes domains like healthcare and industry, remains underexplored. In this paper, we introduce \textit{VT-Bench}, the first unified benchmark for standardizing v…

Cited by 0SourceScholar
2026

YuE: Scaling Open Foundation Models for Long-Form Music Generation

ICLR 2026poster

We tackle the task of long-form music generation, particularly the challenging \textbf{lyrics-to-song} problem, by introducing \textbf{YuE (乐)}, a family of open-source music generation foundation models. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while…

Cited by 0SourcecodeScholar
2025

Constructing Your Model’s Value Distinction: Towards LLM Alignment with Anchor Words Tuning

EMNLP 2025

With the widespread applications of large language models (LLMs), aligning LLMs with human values has emerged as a critical challenge. For alignment, we always expect LLMs to be honest, positive, harmless, etc. And LLMs appear to be capable of generating the desired outputs after the alignment tunin

2025

Counterfactual Thinking Driven Emotion Regulation for Image Sentiment Recognition

IJCAI 2025

Image sentiment recognition (ISR) facilitates the practical application of affective computing on rapidly growing social platforms. Nowadays, region-based ISR methods that use affective regions to guide emotion prediction have gained significant attention. However, existing methods lack a causality-

Cited by 0SourcePDFScholar
2025

Entire-Space Variational Information Exploitation for Post-Click Conversion Rate Prediction

AAAI 2025technical

In recommender systems, post-click conversion rate (CVR) estimation is an essential task to model user preferences for items and estimate the value of recommendations. Sample selection bias (SSB) and data sparsity (DS) are two persistent challenges for post-click conversion rate (CVR) estimation. Cu…

2025

Neural-Link: Non-Overlapping MPC Fusion and Passive Inertial Sensing on Soft Platforms

IROS 2025

Soft, elastic platforms may pose an intricate challenge towards sensor fusion as forces acting on the structure render extrinsic transformations variable over time. The present paper tackles this problem by introducing an elastic deformation model and embedding it into a sensor fusion scheme. The co

Cited by 0SourceScholar
2025

OmniBench: Towards The Future of Universal Omni-Language Models

NeurIPS 2025poster

Recent advancements in multimodal large language models (MLLMs) have focused on integrating multiple modalities, yet their ability to simultaneously process and reason across different inputs remains underexplored. We introduce OmniBench, a novel benchmark designed to evaluate models’ ability to rec…

Cited by 0SourcecodeScholar
2025

Simulator HC: Regression-based Online Simulation of Starting Problem-Solution Pairs for Homotopy Continuation in Geometric Vision

CVPR 2025highlight

While automatically generated polynomial elimination templates have sparked great progress in the field of 3D computer vision, there remain many problems for which the degree of the constraints or the number of unknowns leads to intractability. In recent years, homotopy continuation has been introdu…

Cited by 0SourcePDFScholar
2024

CausVSR: Causality Inspired Visual Sentiment Recognition

IJCAI 2024poster

Visual Sentiment Recognition (VSR) is an evolving field that aims to detect emotional tendencies within visual content. Despite its growing significance, detecting emotions depicted in visual content, such as images, faces challenges, notably the emergence of misleading or spurious correlations of t…

Cited by 2SourcePDFScholar
2024

Driving Style Alignment for LLM-powered Driver Agent

IROS 2024poster

Recently, LLM-powered driver agents have demonstrated considerable potential in the field of autonomous driving, showcasing human-like reasoning and decision-making abilities. However, current research on aligning driver agent behaviors with human driving styles remains limited, partly due to the sc…

Cited by 12SourceScholar
2024

InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

NeurIPS 2024poster

The Large Vision-Language Model (LVLM) field has seen significant advancements, yet its progression has been hindered by challenges in comprehending fine-grained visual content due to limited resolution. Recent efforts have aimed to enhance the high-resolution understanding capabilities of LVLMs, ye…

2024

Mobility-LLM: Learning Visiting Intentions and Travel Preference from Human Mobility Data with Large Language Models

NeurIPS 2024poster

Location-based services (LBS) have accumulated extensive human mobility data on diverse behaviors through check-in sequences. These sequences offer valuable insights into users’ intentions and preferences. Yet, existing models analyzing check-in sequences fail to consider the semantics contained in…

Cited by 5SourcePDFScholar
2023

Lightweight Portrait Segmentation Via Edge-Optimized Attention

ICASSP 2023accepted

With the outbreak of COVID-19 around the world, the frequency of video conferencing at home is increasing. Therefore, a segmentation architecture that can quickly carry out close-range portrait segmentation has become a current need. However, the current portrait segmentation architectures cannot me…

Cited by 0SourceScholar