← Search

Yong Cao

13 accepted papers

2026

Design of Bio-manta Based on Continuously Programmable Soft Pneumatic Actuators (CPSPAs)

RA-L 2026

Bionic autonomous underwater vehicles (AUVs) have been prevalent in underwater equipment for their excel features. Manta rays have been widely studied due to their highly efficient and maneuverable pectoral fin propulsion. Traditional single-hinge actuation methods have impeded bionic pectoral fin a

Cited by 0SourceScholar
2026

FrankenMotion: Part-level Human Motion Generation and Composition

CVPR 2026

Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptions due to the absence of fine-grained, part-level motion annotations. This limits their controllability over individual

Cited by 0SourcecodeScholar
2025

Application of soft constraints on mirror position to improve robustness of optical target positioning in shallow water

IROS 2025

The unique optical characteristics of the underwater environment, such as light refraction and loss of salient features, pose a significant challenge to traditional vision sensors, especially in the swarm operation scenario where multiple autonomous underwater vehicles (AUVs) cooperate with the moth

Cited by 0SourceScholar
2025

Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive Reasoning

EMNLP 2025

Introducing **MARK**, the **M**ulti-st**A**ge **R**easoning framewor**K** for cultural value survey response simulation, designed to enhance the accuracy, steerability, and interpretability of large language models in this task. The system is inspired by the type dynamics theory in the MBTI psycholo

Cited by 0SourcePDFScholar
2025

ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

ACL 2025long

Dialogue agents powered by Large Language Models (LLMs) show superior performance in various tasks. Despite the better user understanding and human-like responses, their **lack of controllability** remains a key challenge, often leading to unfocused conversations or task failure. To address this, we…

2025

Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural Knowledge

NAACL 2025long

Recent studies have highlighted the presence of cultural biases in Large Language Models (LLMs), yet often lack a robust methodology to dissect these phenomena comprehensively. Our work aims to bridge this gap by delving into the Food domain—a universally relevant yet culturally diverse aspect of hu…

2025

GLO: General LiDAR-Only Odometry With High Efficiency and Low Drift

RA-L 2025

This study proposes GLO, a general LiDAR-only odometry method with high efficiency and low drift. First, we propose a map data structure using multilevel voxels to improve map update efficiency. Each voxel node actively maintains plane features, minimizing redundant fitting and enhancing matching ef

Cited by 2SourceScholar
2025

Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations

NAACL 2025long

Large-scale surveys are essential tools for informing social science research and policy, but running surveys is costly and time-intensive. If we could accurately simulate group-level survey results, this would therefore be very valuable to social science research. Prior work has explored the use of…

2024

MLPs Compass: What is Learned When MLPs are Combined with PLMs?

ICASSP 2024accepted

While Transformer-based pre-trained language models and their variants exhibit strong semantic representation capabilities, the question of comprehending the information gain derived from the additional components of PLMs remains an open question in this field. Motivated by recent efforts that prove…

Cited by 0SourceScholar
2024

Rethinking the Effectiveness of Graph Classification Datasets in Benchmarks for Assessing GNNs

IJCAI 2024poster

Graph classification benchmarks, vital for assessing and developing graph neural network (GNN) models, have recently been scrutinized, as simple methods like MLPs have demonstrated comparable performance. This leads to an important question: Do these benchmarks effectively distinguish the advancemen…

2023

Pay More Attention to Relation Exploration for Knowledge Base Question Answering

ACL 2023findings

Knowledge base question answering (KBQA) is a challenging task that aims to retrieve correct answers from large-scale knowledge bases. Existing attempts primarily focus on entity representation and final answer reasoning, which results in limited supervision for this task. Moreover, the relations, w…

2022

Explore More Guidance: A Task-aware Instruction Network for Sign Language Translation Enhanced with Data Augmentation

NAACL 2022findings

Sign language recognition and translation first uses a recognition module to generate glosses from sign language videos and then employs a translation module to translate glosses into spoken sentences. Most existing works focus on the recognition step, while paying less attention to sign language tr…

2020

Large-Scale Unsupervised Pre-Training for End-to-End Spoken Language Understanding

ICASSP 2020accepted

End-to-end Spoken Language Understanding (SLU) is proposed to infer the semantic meaning directly from audio features without intermediate text representation. In this paper, we explore unsupervised pre-training for End-to-end SLU models by learning representations from large-scale raw audios. The p…

Cited by 0SourceScholar