← Search

Xinyao Wang

9 accepted papers

2026

Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking

AAAI 2026technical

3D multi-object tracking is a critical and challenging task in the field of autonomous driving. A common paradigm relies on modeling individual object motion, e.g., Kalman filters, to predict trajectories. While effective in simple scenarios, this approach often struggles in crowded environments or

Cited by 0SourcePDFScholar
2026

InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization

AAAI 2026technical

The emergence of Multimodal Large Language Models (MLLMs) has propelled the development of autonomous agents that operate on Graphical User Interfaces (GUIs) using pure visual input. A fundamental challenge is robustly grounding natural language instructions. This requires a precise spatial alignmen

Cited by 0SourcePDFScholar
2026

MODEL MERGING SCALING LAWS IN LARGE LANGUAGE MODELS

ICML 2026poster

We study empirical scaling laws for language model merging measured by cross-entropy. Despite its wide practical use, merging lacks a quantitative rule that predicts returns as we add experts or scale the model size. We identify a compact power law that links model size and expert number: the size-d…

Cited by 0SourceScholar
2025

Attribute Association Driven Multi-Task Learning for Session-based Recommendation

IJCAI 2025

Session-based Recommendation (SBR) aims to predict users’ next interaction based on their current session without relying on long-term profiles. Despite its effectiveness in privacy-preserving and real-time scenarios, SBR remains challenging due to limited behavioral signals. Prior methods often ove

Cited by 0SourcePDFScholar
2025

DiffLM: Controllable Synthetic Data Generation via Diffusion Language Models

ACL 2025finding

Recent advancements in large language models (LLMs) have significantly enhanced their knowledge and generative capabilities, leading to a surge of interest in leveraging LLMs for high-quality data synthesis. However, synthetic data generation via prompting LLMs remains challenging due to LLMs’ limit…

2024

CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts

NeurIPS 2024poster

Recent advancements in Multimodal Large Language Models (LLMs) have focused primarily on scaling by increasing text-image pair data and enhancing LLMs to improve performance on multimodal tasks. However, these scaling approaches are computationally expensive and overlook the significance of efficien…

2022

End-to-End Compressed Video Representation Learning for Generic Event Boundary Detection

CVPR 2022poster

Generic event boundary detection aims to localize the generic, taxonomy-free event boundaries that segment videos into chunks. Existing methods typically require video frames to be decoded before feeding into the network, which demands considerable computational power and storage space. To that end,…

Cited by 20PDFScholar
2021

Stable and Effective One-Step Method for Person Search

ICASSP 2021accepted

Person search, which requires both pedestrian detection and person re-identification, is a challenging computer vision task applied to real-world scenarios. The challenges faced by detection and re-identification, such as occlusion, poor illumination, confusing background, are still urgent for perso…

Cited by 0SourceScholar