← Search

Xiangdong Zhang

9 accepted papers

2026

Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditions

CVPR 2026

The rapid movements and agile maneuvers of unmanned aerial vehicles (UAVs) induce significant observational challenges for multi-object tracking (MOT). However, existing UAV-perspective MOT benchmarks often lack these complexities, featuring predominantly predictable camera dynamics and linear motio

Cited by 0SourcecodeScholar
2026

NITP: Next Implicit Token Prediction for LLM Pre-training

ICML 2026poster

Standard Next-Token Prediction (NTP) supervises language models solely through discrete labels in the output logit space. We argue that this sparse, one-hot supervision leaves the latent representation space under-constrained, allowing hidden states to drift into degenerate and anisotropic configura…

Cited by 0SourceScholar
2026

ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning

ICLR 2026poster

Data quality plays a critical role in enhancing supervised fine-tuning (SFT) for large language models (LLMs), and token-level data selection has emerged as a promising direction for its fine-grained nature. Despite their strong empirical performance, existing token-level selection methods share two…

Cited by 0SourcecodeScholar
2025

A Simple yet Effective Hypergraph Clustering Network

IJCAI 2025

Hypergraph Clustering has gained significant attention due to its capability of capturing high order structural information. Among different approaches, contrastive learning-based methods leverage self-supervised learning and data augmentation, exhibiting impressive performance. However, most of the

Cited by 0SourcePDFScholar
2025

Fair Incomplete Multi-View Clustering via Distribution Alignment

IJCAI 2025

Incomplete multi-view clustering (IMVC) extracts consistent and complementary information from multi-source/modality data with missing views, aiming to partition the data into different clusters. It can effectively address the problem of unsupervised multi-source data analysis in complex environment

Cited by 0SourcePDFScholar
2025

Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views

ICCV 2025poster

Point cloud learning, especially in a self-supervised way without manual labels, has gained growing attention in both vision and learning communities due to its potential utility in a wide range of applications. Most existing generative approaches for point cloud self-supervised learning focus on re…

2025

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

NeurIPS 2025poster

Recent advancements in text-to-video (T2V) diffusion models have enabled high-fidelity and realistic video synthesis. However, current T2V models often struggle to generate physically plausible content due to their limited inherent ability to accurately understand physics. We found that while the re…

Cited by 0SourcecodeScholar
2024

PCP-MAE: Learning to Predict Centers for Point Masked Autoencoders

NeurIPS 2024spotlight

Masked autoencoder has been widely explored in point cloud self-supervised learning, whereby the point cloud is generally divided into visible and masked parts. These methods typically include an encoder accepting visible patches (normalized) and corresponding patch centers (position) as input, with…

2021

Free-Form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloud

ICCV 2021poster

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of…

Cited by 100PDFcodeScholar