← Search

Xiao Sun

46 accepted papers

2026

From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking

CVPR 2026

End-to-end multi-object tracking (MOT) methods have recently achieved remarkable progress by unifying detection and association within a single framework. Despite their strong detection performance, these methods suffer from relatively low association accuracy. Through detailed analysis, we observe

Cited by 0SourcecodeScholar
2026

Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence

ICML 2026oral

The pursuit of spatial intelligence fundamentally relies on access to large-scale, fine-grained 3D data. However, existing approaches predominantly construct spatial understanding benchmarks by generating question–answer (QA) pairs from a limited number of manually annotated datasets, rather than sy…

Cited by 0SourceScholar
2026

Motion-Aware Animatable Gaussian Avatars Deblurring

CVPR 2026

The creation of 3D human avatars from multi-view videos is a significant yet challenging task in computer vision. However, existing techniques rely on high-quality, sharp images as input, which are often impractical to obtain in real-world scenarios due to variations in human motion speed and intens

Cited by 0SourcecodeScholar
2026

Proxy-GS: Unified Occlusion Priors for Training and Inference in Structured 3D Gaussian Splatting

CVPR 2026

3D Gaussian Splatting (3DGS) has emerged as an efficient approach for photorealistic rendering. Recent MLP-based variants further improve visual fidelity but introduce substantial decoding overhead during rendering. To reduce the computational cost, several pruning strategies and level-of-detail (LO

Cited by 0SourcecodeScholar
2026

RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis

AAAI 2026technical

We introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research

Cited by 0SourcePDFScholar
2025

CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction

ICCV 2025poster

Despite its significant achievements in large-scale scene reconstruction, 3D Gaussian Splatting still faces substantial challenges, including slow processing, high computational costs, and limited geometric accuracy. These core issues arise from its inherently unstructured design and the absence of…

Cited by 0SourcePDFScholar
2025

DEFormer: DCT-driven Enhancement Transformer for Low-light Image and Dark Vision

ICASSP 2025accepted

Low-light image enhancement restores the colors and details of a single image and improves high-level visual tasks. However, restoring the lost details in the dark area is still a challenge relying only on the RGB domain. In this paper, we delve into frequency as a new clue into the model and propos…

Cited by 0SourceScholar
2025

GRPose: Learning Graph Relations for Human Image Generation with Pose Priors

AAAI 2025technical

Recent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. However, existing efforts are still struggling to generate high-quality images with consistent pose alignment, resulting in unsatisfactory output. In this…

2025

MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks

CVPR 2025poster

While 3D Gaussian Splatting (3DGS) has demonstrated remarkable performance in novel view synthesis and real-time rendering, the high memory consumption due to the use of millions of Gaussians limits its practicality. To mitigate this issue, improvements have been made by pruning unnecessary Gaussian…

2025

MultiAgentESC: A LLM-based Multi-Agent Collaboration Framework for Emotional Support Conversation

EMNLP 2025

The development of Emotional Support Conversation (ESC) systems is critical for delivering mental health support tailored to the needs of help-seekers. Recent advances in large language models (LLMs) have contributed to progress in this domain, while most existing studies focus on generating respons

Cited by 0SourcePDFScholar
2025

OccMamba: Semantic Occupancy Prediction with State Space Models

CVPR 2025poster

Training deep learning models for semantic occupancy prediction is challenging due to factors such as a large number of occupancy cells, severe occlusion, limited visual cues, complicated driving scenarios, etc. Recent methods often adopt transformer-based architectures given their strong capability…

2025

PEARL: Parallel Speculative Decoding with Adaptive Draft Length

ICLR 2025poster

Speculative decoding (SD), where an extra draft model is employed to provide multiple **draft** tokens first and then the original target model verifies these tokens in parallel, has shown great power for LLM inference acceleration. However, existing SD methods suffer from the mutual waiting problem…

Cited by 0SourcePDFScholar
2025

Safety-Critical Traffic Simulation with Adversarial Transfer of Driving Intentions

ICRA 2025

Traffic simulation, complementing real-world data with a long-tail distribution, allows for effective evaluation and enhancement of the ability of autonomous vehicles to handle accident-prone scenarios. Simulating such safety-critical scenarios is nontrivial, however, from log data that are typicall

Cited by 2SourceScholar
2025

Sequential Gaussian Avatars with Hierarchical Motion Context

ICCV 2025poster

The emergence of neural rendering has significantly advanced the rendering quality of 3D human avatars, with the recently popular 3DGS technique enabling real-time performance. However, SMPL-driven 3DGS human avatars still struggle to capture fine appearance details due to the complex mapping from p…

2025

Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition

ICASSP 2025accepted

Speech Emotion Recognition (SER) plays a critical role in enhancing user experience within human-computer interaction. However, existing methods are overwhelmed by temporal domain analysis, overlooking the valuable envelope structures of the frequency domain that are equally important for robust emo…

Cited by 0SourceScholar
2025

Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human Avatars

ICCV 2025poster

In this paper, we highlight a critical yet often overlooked factor in most 3D human tasks, namely modeling complicated 3D human with with hand-held objects or loose-fitting clothing. It is known that the parameterized formulation of SMPL is able to fit human skin; while hand-held objects and loose-f…

2024

"Clearer Frames, Anytime: Resolving Velocity Ambiguity in Video Frame Interpolation"

ECCV 2024oral

"Existing video frame interpolation (VFI) methods blindly predict where each object is at a specific timestep t (“time indexing”), which struggles to predict precise object movements. Given two images of a baseball, there are infinitely many possible trajectories: accelerating or decelerating, strai…

2024

Aleth-NeRF: Illumination Adaptive NeRF with Concealing Field Assumption

AAAI 2024technical

The standard Neural Radiance Fields (NeRF) paradigm employs a viewer-centered methodology, entangling the aspects of illumination and material reflectance into emission solely from 3D points. This simplified rendering approach presents challenges in accurately modeling images captured under adverse…

2024

Bi-ViT: Pushing the Limit of Vision Transformer Quantization

AAAI 2024technical

Vision transformers (ViTs) quantization offers a promising prospect to facilitate deploying large pre-trained networks on resource-limited devices. Fully-binarized ViTs (Bi-ViT) that pushes the quantization of ViTs to its limit remain largely unexplored and a very challenging task yet, due to their…

2024

DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement

CVPR 2024poster

We present Dive Into the Boundaries (DIBS) a novel pretraining framework for dense video captioning (DVC) that elaborates on improving the quality of the generated event captions and their associated pseudo event boundaries from unlabeled videos. By leveraging the capabilities of diverse large langu…

Cited by 10SourcePDFScholar
2024

LDIP: Real-time on-road object detection with depth estimation from a single image

IROS 2024poster

Detecting on-road objects with absolute depth information is one of the most crucial tasks in autonomous driving to ensure safety. Traditional 2D object detection aims to classify and locate objects in image space, but it cannot acquire in-depth information. While 3D object detection and pixel-level…

Cited by 0SourcecodeScholar
2024

Learning 1-Bit Tiny Object Detector with Discriminative Feature Refinement

ICML 2024poster

1-bit detectors show impressive performance comparable to their real-valued counterparts when detecting commonly sized objects while exhibiting significant performance degradation on tiny objects. The challenge stems from the fact that high-level features extracted by 1-bit convolutions seem less co…

Cited by 1SourcePDFScholar
2024

LucidAction: A Hierarchical and Multi-model Dataset for Comprehensive Action Quality Assessment

NeurIPS 2024poster

Action Quality Assessment (AQA) research confronts formidable obstacles due to limited, mono-modal datasets sourced from one-shot competitions, which hinder the generalizability and comprehensiveness of AQA models. To address these limitations, we present LucidAction, the first systematically collec…

Cited by 2SourcePDFScholar
2024

PREFER: Prompt Ensemble Learning via Feedback-Reflect-Refine

AAAI 2024technical

As an effective tool for eliciting the power of Large Language Models (LLMs), prompting has recently demonstrated unprecedented abilities across a variety of complex tasks. To further improve the performance, prompt ensemble has attracted substantial interest for tackling the hallucination and insta…

2023

Q-DM: An Efficient Low-bit Quantized Diffusion Model

NeurIPS 2023poster

Denoising diffusion generative models are capable of generating high-quality data, but suffers from the computation-costly generation process, due to a iterative noise estimation using full-precision networks. As an intuitive solution, quantization can significantly reduce the computational and mem…

Cited by 39SourcePDFScholar
2023

Randomized Quantization: A Generic Augmentation for Data Agnostic Self-supervised Learning

ICCV 2023poster

Self-supervised representation learning follows a paradigm of withholding some part of the data and tasking the network to predict it from the remaining part. Among many techniques, data augmentation lies at the core for creating the information gap. Towards this end, masking has emerged as a generi…

Cited by 11PDFcodeScholar
2022

A Simple Multi-Modality Transfer Learning Baseline for Sign Language Translation

CVPR 2022poster

This paper proposes a simple transfer learning baseline for sign language translation. Existing sign language datasets (e.g. PHOENIX-2014T, CSL-Daily) contain only about 10K-20K pairs of sign videos, gloss annotations and texts, which are an order of magnitude smaller than typical parallel data for…

Cited by 182PDFcodeScholar
2022

Animation from Blur: Multi-modal Blur Decomposition with Motion Guidance

ECCV 2022poster

"We study the challenging problem of recovering detailed motion from a single motion-blurred image. Existing solutions to this problem estimate a single image sequence without considering the motion ambiguity for each region. Therefore, the results tend to converge to the mean of the multi-modal pos…

2022

Bringing Rolling Shutter Images Alive with Dual Reversed Distortion

ECCV 2022poster

"Rolling shutter (RS) distortion can be interpreted as the result of picking a row of pixels from instant global shutter (GS) frames over time during the exposure of the RS camera. This means that the information of each instant GS frame is partially, yet sequentially, embedded into the row-dependen…

2022

Cross-Model Pseudo-Labeling for Semi-Supervised Action Recognition

CVPR 2022oral

Semi-supervised action recognition is a challenging but important task due to the high cost of data annotation. A common approach to this problem is to assign unlabeled data with pseudo-labels, which are then used as additional supervision in training. Typically in recent work, the pseudo-labels are…

Cited by 75PDFScholar
2022

Unsupervised Learning of Efficient Geometry-Aware Neural Articulated Representations

ECCV 2022poster

"We propose an unsupervised method for 3D geometry-aware representation learning of articulated objects, in which no image-pose pairs or foreground masks are used for training. Though photorealistic images of articulated objects can be rendered with explicit pose control through existing 3D neural r…

2021

Learning Skeletal Graph Neural Networks for Hard 3D Pose Estimation

ICCV 2021poster

Various deep learning techniques have been proposed to solve the single-view 2D-to-3D pose estimation problem. While the average prediction accuracy has been improved significantly over the years, the performance on hard poses with depth ambiguity, self-occlusion, and complex or rare poses is still…

Cited by 163PDFScholar
2020

Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation

ECCV 2020poster

Instance Segmentation and Pose Estimation","A recent approach for object detection and human pose estimation is to regress bounding boxes or human keypoints from a central point on the object or person. While this center-point regression is simple and efficient, we argue that the image features extr…

2020

SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach

ECCV 2020poster

Human poses that are rare or unseen in a training set are challenging for a network to predict. Similar to the long-tailed distribution problem in visual recognition, the small number of examples for such poses limits the ability of networks to model them. Interestingly, local pose distributions suf…

2020

ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed Training

NeurIPS 2020poster

Large-scale distributed training of Deep Neural Networks (DNNs) on state-of-the-art platforms are expected to be severely communication constrained. To overcome this limitation, numerous gradient compression techniques have been proposed and have demonstrated high compression ratios. However, most e…

Cited by 79SourcePDFScholar
2020

Ultra-Low Precision 4-bit Training of Deep Neural Networks

NeurIPS 2020oral

In this paper, we propose a number of novel techniques and numerical representation formats that enable, for the very first time, the precision of training systems to be aggressively scaled from 8-bits to 4-bits. To enable this advance, we explore a novel adaptive Gradient Scaling technique (Gradsca…

2019

Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks

NeurIPS 2019poster

Reducing the numerical precision of data and computation is extremely effective in accelerating deep learning training workloads. Towards this end, 8-bit floating point representations (FP8) were recently proposed for DNN training. However, its applicability was demonstrated on a few selected models…

2018

End-effector with a Hook and Two Fingers for the Locomotion and Simple Work of a Four-limbed Robot

IROS 2018poster

In this paper, we propose an end-effector for realizing various locomotion modes and simple work of a legged robot. The locomotion modes include climbing a vertical ladder, crawling, and walking. The simple work includes grasping and switching motions required at a disaster site. The developed end-e…

Cited by 3SourceScholar
2017

A four-limbed disaster-response robot having high mobility capabilities in extreme environments

IROS 2017poster

This paper describes a novel four-limbed robot having high mobility capability in extreme environments. At disaster sites, there are various types of environments where a robot must move such as rough terrain with possibility of collapse, narrow places, stairs, vertical ladders and so forth. In this…

Cited by 18SourceScholar
2017

Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach

ICCV 2017poster

In this paper, we study the task of 3D human pose estimation in the wild. This task is challenging due to lack of training data, as existing datasets are either in the wild images with 2D pose or in the lab images with 3D pose. We propose a weakly-supervised transfer learning method that uses mixed…

Cited by 749PDFcodeScholar