← Search

Yaoxian Song

5 accepted papers

2025

Evaluating Semantic Variation in Text-to-Image Synthesis: A Causal Perspective

ICLR 2025poster

Accurate interpretation and visualization of human instructions are crucial for text-to-image (T2I) synthesis. However, current models struggle to capture semantic variations from word order changes, and existing evaluations, relying on indirect metrics like text-image similarity, fail to reliably…

2024

Multi-Task Domain Adaptation for Language Grounding with 3D Objects

ECCV 2024poster

"The existing works on object-level language grounding with 3D objects mostly focus on improving performance by utilizing the off-the-shelf pre-trained models to capture features, such as viewpoint selection or geometric priors. However, they have failed to consider exploring the cross-modal represe…

Cited by 1SourcePDFScholar
2023

Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and Future

EMNLP 2023long main

Machine learning (ML) systems in natural language processing (NLP) face significant challenges in generalizing to out-of-distribution (OOD) data, where the test distribution differs from the training data distribution. This poses important questions about the robustness of NLP models and their high…

Cited by 0SourceScholar
2022

Human-in-the-loop Robotic Grasping Using BERT Scene Representation

COLING 2022main

Current NLP techniques have been greatly applied in different domains. In this paper, we propose a human-in-the-loop framework for robotic grasping in cluttered scenes, investigating a language interface to the grasping process, which allows the user to intervene by natural language commands. This f…

2020

Multimodal Aggregation Approach for Memory Vision-Voice Indoor Navigation with Meta-Learning

IROS 2020poster

Vision and voice are two vital keys for agents’ interaction and learning. In this paper, we present a novel indoor navigation model called Memory Vision-Voice Indoor Navigation (MVV-IN), which receives voice commands and analyzes multimodal information of visual observation in order to enhance robot…

Cited by 24SourceScholar