← Search

He Huang

33 accepted papers

2026

Budget-Feasible Mechanisms for Submodular Welfare Maximization in Procurement Auctions

ICML 2026poster

Budget-feasible procurement auctions play a pivotal role in various AI-driven marketplaces, such as data acquisition and crowdsourcing, where a buyer with a limited budget seeks to procure services from strategic sellers with private costs. While numerous budget-feasible mechanisms have been propose…

Cited by 0SourceScholar
2026

Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics

ICRA 2026poster

近年来,视觉-语言-行动(VLA)模型通过无缝整合视觉感知、语言理解和动作生成,在端到端的学习框架中彻底革新了机器人作。然而,由于这些模型设计为直接与物理世界和人类交互,其安全性至关重要,即使是小漏洞也可能导致灾难性故障。在本研究中,我们提出了通用对抗对象,这是一种表面纹理优化的球体,当置于机器人视野内时,任务成功率会显著降低。具体来说,我们的方法引入了一个多层次攻击框架,能够共同干扰轨迹规划、任务执行和动作控制。我们在模拟和现实机器人环境中验证了我们的方法。实验结果表明,对抗对象在两种代表性VLA模型(Pi0和RDT&#

Cited by 0Scholar
2025

A Hierarchical Compression Technique for 3D Gaussian Splatting Compression

ICASSP 2025accepted

3D Gaussian Splatting (GS) demonstrates excellent rendering quality and generation speed in novel view synthesis. However, substantial data size poses challenges for storage and transmission, making 3D GS compression an essential technology. Current 3D GS compression research primarily focuses on de…

Cited by 0SourceScholar
2025

ADC-GS: Anchor-Driven Deformable and Compressed Gaussian Splatting for Dynamic Scene Reconstruction

IJCAI 2025

Existing 4D Gaussian Splatting methods rely on per-Gaussian deformation from a canonical space to target frames, which overlooks redundancy among adjacent Gaussian primitives and result in suboptimal performance. To address this limitation, we propose Anchor-Driven Deformable and Compressed Gaussian

2025

Egocentric Speaker Diarization with Vision-Guided Clustering and Adaptive Speech Re-detection

ICASSP 2025accepted

Speaker diarization aims to identify "who spoke when" in multi-person conversational scenarios. State-of-the-art audio-only diarization methods divide the task into multi-stages of speech segmentation, neural speaker embedding and unsupervised clustering. Egocentric speaker diarization (i.e., diariz…

Cited by 0SourceScholar
2025

LINR-PCGC: Lossless Implicit Neural Representations for Point Cloud Geometry Compression

ICCV 2025poster

Existing AI-based point cloud compression methods struggle with dependence on specific training data distributions, which limits their real-world deployment. Implicit Neural Representation (INR) methods solve the above problem by encoding overfitted network parameters to the bitstream, resulting in…

Cited by 0SourcePDFScholar
2025

META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR

ICASSP 2025accepted

We propose a novel end-to-end multi-talker automatic speech recognition (ASR) framework that enables both multi-speaker (MS) ASR and target-speaker (TS) ASR. Our proposed model is trained in a fully end-to-end manner, incorporating speaker supervision from a pre-trained speaker diarization module. W…

Cited by 0SourceScholar
2025

NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks

ICASSP 2025accepted

Self-supervised learning (SSL) has been proved to benefit a wide range of speech processing tasks, such as speech recognition/translation, speaker verification and diarization, etc. However, most of current speech SSL approaches are computationally expensive. In this paper, we introduce a simplified…

Cited by 0SourceScholar
2025

Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

ICML 2025poster

Sortformer is an encoder-based speaker diarization model designed for supervising speaker tagging in speech-to-text models. Instead of relying solely on permutation invariant loss (PIL), Sortformer introduces Sort Loss to resolve the permutation problem, either independently or in tandem with PIL. I…

Cited by 0SourcePDFScholar
2025

VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning

NAACL 2025long

Recent studies have augmented large language models (LLMs) with speech capabilities, leading to the development of speech language models (SpeechLMs). Earlier SpeechLMs focused on single-turn speech-based question answering (QA), where user input comprised a speech context and a text question. More…

2024

SALM: Speech-Augmented Language Model with in-Context Learning for Speech Recognition and Translation

ICASSP 2024accepted

We present a novel Speech Augmented Language Model (SALM) with multitask and in-context learning capabilities. SALM comprises a frozen text LLM, a audio encoder, a modality adapter module, and LoRA layers to accommodate speech input and associated task instructions. The unified SALM not only achieve…

Cited by 0SourceScholar
2023

A Robotic Assistance Personalization Control Approach of Hip Exoskeletons for Gait Symmetry Improvement

IROS 2023poster

Healthy human locomotion functions with good gait symmetry depend on rhythmic coordination of the left and right legs, which can be deteriorated by neurological disorders like stroke and spinal cord injury. Powered exoskeletons are promising devices to improve impaired people's locomotion functions,…

Cited by 6SourceScholar
2023

DSR: Dynamical Surface Representation as Implicit Neural Networks for Protein

NeurIPS 2023poster

We propose a novel neural network-based approach to modeling protein dynamics using an implicit representation of a protein’s surface in 3D and time. Our method utilizes the zero-level set of signed distance functions (SDFs) to represent protein surfaces, enabling temporally and spatially continuous…

2023

Efficient Sequence Transduction by Jointly Predicting Tokens and Durations

ICML 2023poster

This paper introduces a novel Token-and-Duration Transducer (TDT) architecture for sequence-to-sequence tasks. TDT extends conventional RNN-Transducer architectures by jointly predicting both a token and its duration, i.e. the number of input frames covered by the emitted token. This is achieved by…

2023

MM-Fi: Multi-Modal Non-Intrusive 4D Human Dataset for Versatile Wireless Sensing

NeurIPS 2023poster

4D human perception plays an essential role in a myriad of applications, such as home automation and metaverse avatar simulation. However, existing solutions which mainly rely on cameras and wearable devices are either privacy intrusive or inconvenient to use. To address these issues, wireless sensi…

2023

Practical Parallel Algorithms for Submodular Maximization Subject to a Knapsack Constraint with Nearly Optimal Adaptivity

AAAI 2023technical

Submodular maximization has wide applications in machine learning and data mining, where massive datasets have brought the great need for designing efficient and parallelizable algorithms. One measure of the parallelizability of a submodular maximization algorithm is its adaptivity complexity, which…

Cited by 8SourcePDFScholar
2022

A New Robotic Knee Impedance Control Parameter Optimization Method Facilitated by Inverse Reinforcement Learning

RA-L 2022

Recent efforts in the design of intelligent controllers for configuring robotic prostheses have demonstrated new possibilities in improving mobility and restoring locomotion for individuals with lower-limb disabilities. In these efforts, personalizing the controller of the robotic device is a crucia

Cited by 18SourceScholar
2022

Characterizing Prosthesis Control Fault During Human-Prosthesis Interactive Walking Using Intrinsic Sensors

RA-L 2022

The physical interactions between wearable lower limb robots and humans have been investigated to inform effective robot design for walking augmentation. However, human-robot interactions when internal faults occur within robots have not been systematically reported, but it is essential to improve t

Cited by 8SourceScholar
2022

Human-Robotic Prosthesis as Collaborating Agents for Symmetrical Walking

NeurIPS 2022accept

This is the first attempt at considering human influence in the reinforcement learning control of a robotic lower limb prosthesis toward symmetrical walking in real world situations. We propose a collaborative multi-agent reinforcement learning (cMARL) solution framework for this highly complex and…

Cited by 12SourcePDFScholar
2022

Reinforcement Learning Impedance Control of a Robotic Prosthesis to Coordinate With Human Intact Knee Motion

RA-L 2022

This study aims to demonstrate reinforcement learning tracking control for automatically configuring the impedance parameters of a robotic knee prosthesis. While our previous studies involving human subjects have focused on tuning the impedance control parameters to meet a fixed, subjectively prescr

Cited by 28SourceScholar
2022

Resembled Tactile Feedback for Object Recognition Using a Prosthetic Hand

RA-L 2022

Tactile feedback in the hand is essential for interaction with objects. Here, we evaluated how artificial tactile sensation affected the recognition of object properties using a myoelectrically controlled prosthetic hand. Electromyogram signals from the flexor and extensor finger muscles were used t

Cited by 13SourceScholar
2021

Randomized Algorithms for Submodular Function Maximization with a $k$-System Constraint

ICML 2021spotlight

Submodular optimization has numerous applications such as crowdsourcing and viral marketing. In this paper, we study the problem of non-negative submodular function maximization subject to a $k$-system constraint, which generalizes many other important constraints in submodular optimization such as…

Cited by 16SourcePDFScholar
2021

SUTD-TrafficQA: A Question Answering Benchmark and an Efficient Network for Video Reasoning Over Traffic Events

CVPR 2021poster

Traffic event cognition and reasoning in videos is an important task that has a wide range of applications in intelligent transportation, assisted driving, and autonomous vehicles. In this paper, we create a novel dataset, SUTD-TrafficQA (Traffic Question Answering), which takes the form of video QA…

Cited by 106PDFcodeScholar
2019

Generative Dual Adversarial Network for Generalized Zero-Shot Learning

CVPR 2019poster

This paper studies the problem of generalized zero-shot learning which requires the model to train on image-label pairs from some seen classes and test on the task of classifying new images from both seen and unseen classes. In this paper, we propose a novel model that provides a unified framework…

Cited by 283PDFcodeScholar