← Search

Yichen Huang

6 accepted papers

2025

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

NAACL 2025system demonstrations

In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface with flexible configuration and dependency control, making it user-friendly and efficient. With full installation, VERSA of…

2023

Learning Graph Dynamics With External Contact for Deformable Linear Objects Shape Control

RA-L 2023

This letter focuses on the shape control manipulation of deformable linear objects (DLO) with a dual-arm robotic system. One significant challenge of DLO shape control is the underactuated control system, which means that finite robotic manipulators can not fully control DLO's shape due to the lack

Cited by 19SourceScholar
2023

Learning Interpretable Low-dimensional Representation via Physical Symmetry

NeurIPS 2023poster

We have recently seen great progress in learning interpretable music representations, ranging from basic factors, such as pitch and timbre, to high-level concepts, such as chord and texture. However, most methods rely heavily on music domain knowledge. It remains an open question what general comput…

2023

Robustness Tests for Automatic Machine Translation Metrics with Adversarial Attacks

EMNLP 2023short findings

We investigate MT evaluation metric performance on adversarially-synthesized texts, to shed light on metric robustness. We experiment with word- and character-level attacks on three popular machine translation metrics: BERTScore, BLEURT, and COMET. Our human experiments validate that automatic metri…

Cited by 0SourcecodeScholar
2020

Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning"

ICML 2020poster

Visual reasoning tasks such as visual question answering (VQA) require an interplay of visual perception with reasoning about the question semantics grounded in perception. However, recent advances in this area are still primarily driven by perception improvements (e.g. scene graph generation) rathe…