← Search

Xing Zhang

16 accepted papers

2025

AniMo: Species-Aware Model for Text-Driven Animal Motion Generation

CVPR 2025poster

Text-driven motion generation has made significant strides in recent years. However, most existing works focus on human motion, largely overlooking the rich and diverse behaviors of animals. Understanding and synthesizing animal motion have important applications in wildlife conservation, animal eco…

2025

Circuit Transformer: A Transformer That Preserves Logical Equivalence

ICLR 2025poster

Implementing Boolean functions with circuits consisting of logic gates is fundamental in digital computer design. However, the implemented circuit must be exactly equivalent, which hinders generative neural approaches on this task due to their occasionally wrong predictions. In this study, we introd…

2025

Skeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation

EMNLP 2025

Code translation benchmarks are essential for evaluating the accuracy and efficiency of LLM-based systems. Existing benchmarks mainly target individual functions, overlooking repository-level challenges like intermodule coherence and dependency management. Recent repository-level efforts exist, but

Cited by 0SourcePDFScholar
2024

Boosting 3D Visual Grounding by Object-Centric Referring Network

IROS 2024poster

3D visual grounding is tasked with locating a specific object within a 3D scene, as described by a given textual reference. This task is challenging because it requires (1) the accurate recognition of various objects in a 3D scene and (2) the understanding of spatial relations in the description. Ho…

Cited by 0SourceScholar
2024

Efficient Adaptation of Pre-trained Vision Transformer via Householder Transformation

NeurIPS 2024poster

A common strategy for Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViTs) involves adapting the model to downstream tasks by learning a low-rank adaptation matrix. This matrix is decomposed into a product of down-projection and up-projection matrices, with the bottleneck…

Cited by 1SourcePDFScholar
2024

J-MAE: Jigsaw Meets Masked Autoencoders in X-Ray Security Inspection

ICASSP 2024accepted

The X-ray security inspection aims to identify any restricted items to protect public safety. Due to the lack of focus on unsupervised learning in this field, using pre-trained models on natural images leads to suboptimal results in downstream tasks. Previous works would lose the relative positional…

Cited by 0SourceScholar
2024

Low-Rank Rescaled Vision Transformer Fine-Tuning: A Residual Design Approach

CVPR 2024poster

Parameter-efficient fine-tuning for pre-trained Vision Transformers aims to adeptly tailor a model to downstream tasks by learning a minimal set of new adaptation parameters while preserving the frozen majority of pre-trained parameters. Striking a balance between retaining the generalizable represe…

2024

MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing

ECCV 2024poster

"The diffusion model is widely leveraged for either video generation or video editing. As each field has its task-specific problems, it is difficult to merely develop a single diffusion for completing both tasks simultaneously. Video diffusion sorely relying on the text prompt can be adapted to unif…

2024

SweepMM: A High-Quality Multimodal Dataset for Sweeping Robots in Home Scenarios for Vision-Language Model

ICASSP 2024accepted

Embodied intelligence based on vision-language models aims to learn from interactions and derive general intelligence. However, existing generalized vision-language models cannot understand domain knowledge in home scenarios due to the lack of sweeping robot multimodal datasets. In this paper, we pr…

Cited by 0SourceScholar
2023

Toward Robust Diagnosis: A Contour Attention Preserving Adversarial Defense for COVID-19 Detection

AAAI 2023technical

As the COVID-19 pandemic puts pressure on healthcare systems worldwide, the computed tomography image based AI diagnostic system has become a sustainable solution for early diagnosis. However, the model-wise vulnerability under adversarial perturbation hinders its deployment in practical situation.…

2022

Nonlinear ICA Using Volume-Preserving Transformations

ICLR 2022poster

Nonlinear ICA is a fundamental problem in machine learning, aiming to identify the underlying independent components (sources) from data which is assumed to be a nonlinear function (mixing function) of these sources. Recent works prove that if the sources have some particular structures (e.g. tempor…

Cited by 23SourcePDFScholar
2021

On Effective Scheduling of Model-based Reinforcement Learning

NeurIPS 2021poster

Model-based reinforcement learning has attracted wide attention due to its superior sample efficiency. Despite its impressive success so far, it is still unclear how to appropriately schedule the important hyperparameters to achieve adequate performance, such as the real data ratio for policy optimi…

2021

VideoLT: Large-Scale Long-Tailed Video Recognition

ICCV 2021poster

Label distributions in real-world are oftentimes long-tailed and imbalanced, resulting in biased models towards dominant labels. While long-tailed recognition has been extensively studied for image classification tasks, limited effort has been made for video domain. In this paper, we introduce Video…

Cited by 53PDFcodeScholar
2019

Simultaneous DFT and IDFT through Widely Linear CLMS

ICASSP 2019accepted

Complex least mean square (CLMS) based adaptive computation of discrete orthogonal transforms has been extensively investigated in the literature. However, all of these results provide only a means for the calculation of either forward orthogonal transforms or their inverse orthogonal transforms, se…

Cited by 0SourceScholar
2016

Multimodal Spontaneous Emotion Corpus for Human Behavior Analysis

CVPR 2016poster

Emotion is expressed in multiple modalities, yet most research has considered at most one or two. This stems in part from the lack of large, diverse, well-annotated, multimodal databases with which to develop and test algorithms. We present a well-annotated, multimodal, multidimensional spontaneous…

Cited by 521PDFScholar