← Search

Chong Zhang

53 accepted papers

2026

Beyond Pixels: Mining Compressed Domain Artifacts for Efficient AI-Generated Video Detection

ICML 2026poster

With the rapid advancement of high-fidelity video generation models, robust AI-generated video (AIGV) detection has become increasingly needed. While most AIGV detection methods operate in the decoded pixel domain, we observe that detection in the pixel domain inevitably entangles task-irrelevant se…

Cited by 0SourceScholar
2026

Diffusion Implicit Policy for Unpaired Scene-aware Motion Synthesis

AAAI 2026technical

Scene-aware motion synthesis has been widely researched recently due to its numerous applications. Prevailing methods rely heavily on paired motion-scene data, while it is difficult to generalize to diverse scenes when trained only on a few specific ones. Thus, we propose a unified framework, termed

Cited by 0SourcePDFScholar
2026

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

ICLR 2026poster

Recent studies on end-to-end (E2E) speech generation with large language models (LLMs) have attracted significant community attention, with multiple works extending text-based LLMs to generate discrete speech tokens. Existing E2E approaches primarily fall into two categories: (1) Methods that genera…

Cited by 0SourceScholar
2026

Privacy on the Fly: A Predictive Adversarial Transformation Network for Mobile Sensor Data

AAAI 2026technical

Mobile motion sensors such as accelerometers and gyroscopes are now ubiquitously accessible by third-party apps via standard APIs. While enabling rich functionalities like activity recognition and step counting, this openness has also enabled unregulated inference of sensitive user traits, such as g

Cited by 0SourcePDFScholar
2025

A Data-Efficient Progressive Learning Framework for Robot Scooping Task

ICRA 2025

Robot scooping is a challenging and important task in robotic tool manipulation research due to the complex relationship between the robot, the tool, and target objects/environment. Taking into account different tools, different target objects and varying environments, the required scooping manipula

Cited by 0SourceScholar
2025

A Hierarchical Multi Robot Coverage Strategy for Large Maps With Reinforcement Learning and Dense Segmented Siamese Network

RA-L 2025

Complete coverage of multiple robots for a large map is an important collaborative planning task, which is widely used in disaster search and rescue, forest fire prevention, resource exploration, and other fields. It generally focuses on coverage completion with less robot (mostly drone) occupation

Cited by 3SourceScholar
2025

Conditional Latent Diffusion-Based Speech Enhancement via Dual Context Learning

ICASSP 2025accepted

Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or spectral domains, leading to increased generation complexity and…

Cited by 0SourceScholar
2025

DocFusion: A Unified Framework for Document Parsing Tasks

ACL 2025finding

Document parsing involves layout element detection and recognition, essential for extracting information. However, existing methods often employ multiple models for these tasks, leading to increased system complexity and maintenance overhead. While some models attempt to unify detection and recognit…

2025

Gamma Distribution PCA-Enhanced Feature Learning for Angle-Robust SAR Target Recognition

ICML 2025poster

Scattering characteristics of synthetic aperture radar (SAR) targets are typically related to observed azimuth and depression angles. However, in practice, it is difficult to obtain adequate training samples at all observation angles, which probably leads to poor robustness of deep networks. In thi…

2025

HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution

ICASSP 2025accepted

The application of generative adversarial networks (GANs) has recently advanced speech super-resolution (SR) based on intermediate representations like mel-spectrograms. However, existing SR methods that typically rely on independently trained and concatenated networks may lead to inconsistent repre…

Cited by 7SourceScholar
2025

Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models

EMNLP 2025

Large Language Models (LLMs) are increasingly equipped with capabilities of real-time web search and integrated with protocols like the Model Context Protocol (MCP). This extension could introduce new security vulnerabilities. We present a systematic investigation of LLM vulnerabilities to hidden ad

2025

MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation

ACL 2025long

Recent multimodal large language models (MLLMs) have demonstrated significant potential in open-ended conversation, generating more accurate and personalized responses. However, their abilities to memorize, recall, and reason in sustained interactions within real-world scenarios remain underexplored…

2025

ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation

ICCV 2025poster

End-to-end (E2E) autonomous driving methods still struggle to make correct decisions in interactive closed-loop evaluation due to limited causal reasoning capability. Current methods attempt to leverage the powerful understanding and reasoning abilities of Vision-Language Models (VLMs) to resolve th…

Cited by 0SourcePDFScholar
2025

Revisiting Adversarial Patch Defenses on Object Detectors: Unified Evaluation, Large-Scale Dataset, and New Insights

ICCV 2025poster

Developing reliable defenses against patch attacks on object detectors has attracted increasing interest. However, we identify that existing defense evaluations lack a unified and comprehensive framework, resulting in inconsistent and incomplete assessments of current methods. To address this issue,…

2025

UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook

ACL 2025long

The emergence of audio language models is empowered by neural audio codecs, which establish critical mappings between continuous waveforms and discrete tokens compatible with language model paradigms. The evolutionary trends from multi-layer residual vector quantizer to single-layer quantizer are be…

2024

Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion

RSS 2024poster

Legged robots navigating cluttered environments must be jointly agile for efficient task execution and safe to avoid collisions with obstacles or humans. Existing studies either develop conservative controllers (< 1.0 m/s) to ensure safety, or focus on agility without considering potentially fatal c…

2024

Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?

ICASSP 2024accepted

Large self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternative by utilizing minimal soft prompt guidance, enhancing portability while also maintaining competitive performance. Howe…

Cited by 0SourceScholar
2024

Learning Highly Dynamic Behaviors for Quadrupedal Robots

ICRA 2024poster

Learning highly dynamic behaviors for robots has been a longstanding challenge. Traditional approaches have demonstrated robust locomotion, but the exhibited behaviors lack diversity and agility. They employ approximate models, which lead to compromises in performance. Data-driven approaches have be…

Cited by 5SourceScholar
2024

Learning Human-to-Humanoid Real-Time Whole-Body Teleoperation

IROS 2024poster

We present Human to Humanoid (H2O), a reinforcement learning (RL) based framework that enables real-time whole-body teleoperation of a full-sized humanoid robot with only an RGB camera. To create a large-scale retargeted motion dataset of human movements for humanoid robots, we propose a scalable "s…

Cited by 83SourceScholar
2024

Loss Masking Is Not Needed In Decoder-Only Transformer For Discrete-Token-Based ASR

ICASSP 2024accepted

Recently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and use a shared vocabulary for both text and speech tokens. Then they train a singl…

Cited by 0SourceScholar
2024

Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding

EMNLP 2024main

Modeling and leveraging layout reading order in visually-rich documents (VrDs) is critical in document intelligence as it captures the rich structure semantics within documents.Previous works typically formulated layout reading order as a permutation of layout elements, i.e. a sequence containing al…

2024

MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation

ICASSP 2024accepted

Our previously proposed MossFormer has achieved promising performance in monaural speech separation. However, it predominantly adopts a self-attention-based MossFormer module, which tends to emphasize longer-range, coarser-scale dependencies, with a deficiency in effectively modelling finer-scale re…

Cited by 0SourceScholar
2024

Multi-Model Wireless Federated Learning with Downlink Beamforming

ICASSP 2024accepted

This paper studies the design of wireless federated learning (FL) for simultaneously training multiple machine learning models. We consider round robin device-model assignment and downlink beamforming for concurrent multiple model updates. After formulating the joint downlink-uplink transmission pro…

Cited by 0SourceScholar
2024

OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning

CoRL 2024poster

We present OmniH2O (Omni Human-to-Humanoid), a learning-based system for whole-body humanoid teleoperation and autonomy. Using kinematic pose as a universal control interface, OmniH2O enables various ways for a human to control a full-sized humanoid with dexterous hands, including using real-time te…

Cited by 69SourcecodeScholar
2024

PDF-to-Tree: Parsing PDF Text Blocks into a Tree

EMNLP 2024finding

In many PDF documents, the reading order of text blocks is missing, which can hinder machine understanding of the document’s content.Existing works try to extract one universal reading order for a PDF file.However, applications, like Retrieval Augmented Generation (RAG), require breaking long articl…

2024

Resilient Legged Local Navigation: Learning to Traverse with Compromised Perception End-to-End

ICRA 2024poster

Autonomous robots must navigate reliably in unknown environments even under compromised exteroceptive perception, or perception failures. Such failures often occur when harsh environments lead to degraded sensing, or when the perception algorithm misinterprets the scene due to limited generalization…

Cited by 16SourceScholar
2024

Rethinking Robustness Assessment: Adversarial Attacks on Learning-based Quadrupedal Locomotion Controllers

RSS 2024poster

Legged locomotion has recently achieved remarkable success with the progress of machine learning techniques, especially deep reinforcement learning (RL). Controllers employing neural networks have demonstrated empirical and qualitative robustness against real-world uncertainties, including sensor no…

Cited by 21SourcePDFScholar
2024

SPGM: Prioritizing Local Features for Enhanced Speech Separation Performance

ICASSP 2024accepted

Dual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, w…

Cited by 0SourceScholar
2024

WoCoCo: Learning Whole-Body Humanoid Control with Sequential Contacts

CoRL 2024poster

Humanoid activities involving sequential contacts are crucial for complex robotic interactions and operations in the real world and are traditionally solved by model-based motion planning, which is time-consuming and often relies on simplified dynamics models. Although model-free reinforcement lear…

Cited by 45SourcecodeScholar
2023

A Holistic Approach to Undesired Content Detection in the Real World

AAAI 2023technical

We present a holistic approach to building a robust and useful natural language classification system for real-world content moderation. The success of such a system relies on a chain of carefully designed and executed steps, including the design of content taxonomies and labeling instructions, data…

2023

Adaptive Knowledge Distillation Between Text and Speech Pre-Trained Models

ICASSP 2023accepted

Learning on a massive amount of speech corpus leads to the recent success of many self-supervised speech models. With knowledge distillation, these models may also benefit from the knowledge encoded by language models that are pre-trained on rich sources of texts. The distillation process, however,…

Cited by 0SourceScholar
2023

Auxiliary Pooling Layer For Spoken Language Understanding

ICASSP 2023accepted

End-to-end spoken language understanding requires speech data annotated with semantic information and may suffer from the shortage of annotated data. Recent progresses leverage unlabelled speech data to pre-train a speech encoder. However, it remains a challenge for the pre-trained speech encoder to…

Cited by 0SourceScholar
2023

Contrastive Speech Mixup for Low-Resource Keyword Spotting

ICASSP 2023accepted

Most of the existing neural-based models for keyword spotting (KWS) in smart devices require thousands of training samples to learn a decent audio representation. However, with the rising demand for smart devices to become more person-alized, KWS models need to adapt quickly to smaller user samples.…

Cited by 15SourceScholar
2023

De'hubert: Disentangling Noise in a Self-Supervised Model for Robust Speech Recognition

ICASSP 2023accepted

Existing self-supervised pre-trained speech models have offered an effective way to leverage massive unannotated corpora to build good automatic speech recognition (ASR). However, many current models are trained on a clean corpus from a single source, which tends to do poorly when noise is present d…

Cited by 0SourceScholar
2023

Ditto: A Simple and Efficient Approach to Improve Sentence Embeddings

EMNLP 2023short main

Prior studies diagnose the anisotropy problem in sentence representations from pre-trained language models, e.g., BERT, without fine-tuning. Our analysis reveals that the sentence embeddings from BERT suffer from a bias towards uninformative words, limiting the performance in semantic textual simila…

Cited by 0SourcecodeScholar
2023

Generating a Terrain-Robustness Benchmark for Legged Locomotion: A Prototype via Terrain Authoring and Active Learning

ICRA 2023poster

Terrain-aware locomotion has become an emerging topic in legged robotics. However, it is hard to generate diverse, challenging, and realistic unstructured terrains in simulation, which limits the way researchers evaluate their locomotion policies. In this paper, we prototype the generation of a terr…

Cited by 3SourceScholar
2023

HiTIN: Hierarchy-aware Tree Isomorphism Network for Hierarchical Text Classification

ACL 2023long

Hierarchical text classification (HTC) is a challenging subtask of multi-label classification as the labels form a complex hierarchical structure. Existing dual-encoder methods in HTC achieve weak performance gains with huge memory overheads and their structure encoders heavily rely on domain knowle…

2023

Learning Terrain-Adaptive Locomotion with Agile Behaviors by Imitating Animals

IROS 2023poster

In this paper, we present a general learning framework for controlling a quadruped robot that can mimic the behavior of real animals and traverse challenging terrains. Our method consists of two steps: an imitation learning step to learn from motions of real animals, and a terrain adaptation step to…

Cited by 13SourceScholar
2023

Reading Order Matters: Information Extraction from Visually-rich Documents by Token Path Prediction

EMNLP 2023long main

Recent advances in multimodal pre-trained models have significantly improved information extraction from visually-rich documents (VrDs), in which named entity recognition (NER) is treated as a sequence-labeling task of predicting the BIO entity tags for tokens, following the typical setting of NLP.…

Cited by 0SourcecodeScholar
2022

Hierarchical Information Matters: Text Classification via Tree Based Graph Neural Network

COLING 2022main

Text classification is a primary task in natural language processing (NLP). Recently, graph neural networks (GNNs) have developed rapidly and been applied to text classification tasks. As a special kind of graph data, the tree has a simpler data structure and can provide rich hierarchical informatio…

Cited by 11SourcePDFScholar
2022

Training language models to follow instructions with human feedback

NeurIPS 2022accept

Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users. In this paper, we sho…

2021

A Partition Filter Network for Joint Entity and Relation Extraction

EMNLP 2021main

In joint entity and relation extraction, existing work either sequentially encode task-specific features, leading to an imbalance in inter-task feature interaction where features extracted later have no direct contact with those that come first. Or they encode entity features and relation features i…

2021

DROID: Minimizing the Reality Gap Using Single-Shot Human Demonstration

RA-L 2021

Reinforcement learning (RL) has demonstrated great success in the past several years. However, most of the scenarios focus on simulated environments. One of the main challenges of transferring the policy learned in a simulated environment to real world, is the discrepancy between the dynamics of the

Cited by 36SourceScholar
2021

Double Perturbation: On the Robustness of Robustness and Counterfactual Bias Evaluation

NAACL 2021long

Robustness and counterfactual bias are usually evaluated on a test dataset. However, are these evaluations robust? If the test dataset is perturbed slightly, will the evaluation results keep the same? In this paper, we propose a “double perturbation” framework to uncover model weaknesses beyond the…

2021

First-Order Fast Algorithm for Structurally Optimal Multi-Group Multicast Beamforming in Large-Scale Systems

ICASSP 2021accepted

We consider multi-group multicast beamforming in large-scale systems to minimize the transmit power subject to the signal-to-interference-plus-noise ratio (SINR) requirements. Based on the optimal multicast beamforming structure, we propose a fast first-order algorithm to obtain the beamforming solu…

Cited by 0SourceScholar
2021

Learning Implicit Sentiment in Aspect-based Sentiment Analysis with Supervised Contrastive Pre-Training

EMNLP 2021main

Aspect-based sentiment analysis aims to identify the sentiment polarity of a specific aspect in product reviews. We notice that about 30% of reviews do not contain obvious opinion words, but still convey clear human-aware sentiment orientation, which is known as implicit sentiment. However, recent n…

2021

Simultaneous Actuation and Localization of Magnetic Robots Using Mobile Coils and Eye-In-Hand Hall-Effect Sensors

IROS 2021poster

Large workspace localization of magnetic robots is important for medical applications. This paper presents a novel localization strategy to achieve simultaneous localization and actuation of magnetic robots using hall-effect sensors. We integrate 25 sensors into a sensing probe and mount it on to th…

Cited by 7SourceScholar
2020

Exploring Parameter Space with Structured Noise for Meta-Reinforcement Learning

IJCAI 2020poster

Efficient exploration is a major challenge in Reinforcement Learning (RL) and has been studied extensively. However, for a new task existing methods explore either by taking actions that maximize task agnostic objectives (such as information gain) or applying a simple dithering strategy (such as noi…

Cited by 0SourcePDFScholar