← Search

Zhihao Xu

11 accepted papers

2026

AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration

ICML 2026poster

Language agents have shown strong promise for task automation. Realizing this promise for increasingly complex, long-horizon tasks has driven the rise of a subagent-as-tools paradigm for multi-turn task solving. However, existing designs still lack a dynamic abstraction view of sub-agents, thereby h…

Cited by 0SourceScholar
2026

Information-Theoretic Decomposition for Multimodal Interaction Learning

CVPR 2026

Multimodal learning hinges on capturing redundant, unique, and synergistic information across modalities, which collectively constitute multimodal interactions. A critical yet underexplored challenge is that these implicit interactions vary dynamically across samples. In this work, we present the fi

Cited by 0SourcecodeScholar
2026

OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment

ICLR 2026poster

Large language model (LLM) alignment faces a critical dilemma when addressing multiple human preferences: improvements in one dimension frequently come at the expense of others, creating unavoidable trade-offs between competing objectives like helpfulness and harmlessness. While prior work mainly fo…

Cited by 0SourcecodeScholar
2026

SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval

CVPR 2026

For video-text retrieval, the use of CLIP has been a de facto standard. However, as CLIP provides only image and text encoders, this consensus has led to a biased paradigm that entirely ignores the sound track of videos. While several attempts have been made to reintroduce audio -- typically by inco

Cited by 0SourcecodeScholar
2025

Internal Value Alignment in Large Language Models through Controlled Value Vector Activation

ACL 2025long

Aligning Large Language Models (LLMs) with human values has attracted increasing attention since it provides clarity, transparency, and the ability to adapt to evolving scenarios. In this paper, we introduce a Controlled Value Vector Activation (ConVA) method that directly aligns the internal values…

2025

U-ViLAR: Uncertainty-Aware Visual Localization for Autonomous Driving via Differentiable Association and Registration

ICCV 2025poster

Accurate localization using visual information is a critical yet challenging task, especially in urban environments where nearby buildings and construction sites significantly degrade GNSS (Global Navigation Satellite System) signal quality. This issue underscores the importance of visual localizati…

Cited by 0SourcePDFScholar
2024

An Optimal Transport View for Subspace Clustering and Spectral Clustering

AAAI 2024technical

Clustering is one of the most fundamental problems in machine learning and data mining, and many algorithms have been proposed in the past decades. Among them, subspace clustering and spectral clustering are the most famous approaches. In this paper, we provide an explanation for subspace clustering…

Cited by 4SourcePDFScholar
2024

Evaluating Readability and Faithfulness of Concept-based Explanations

EMNLP 2024main

With the growing popularity of general-purpose Large Language Models (LLMs), comes a need for more global explanations of model behaviors. Concept-based explanations arise as a promising avenue for explaining high-level patterns learned by LLMs. Yet their evaluation poses unique challenges, especial…

2024

KMTalk: Speech-Driven 3D Facial Animation with Key Motion Embedding

ECCV 2024poster

"We present a novel approach for synthesizing 3D facial motions from audio sequences using key motion embeddings. Despite recent advancements in data-driven techniques, accurately mapping between audio signals and 3D facial meshes remains challenging. Direct regression of the entire sequence often l…

2024

Uncovering Safety Risks of Large Language Models through Concept Activation Vector

NeurIPS 2024poster

Despite careful safety alignment, current large language models (LLMs) remain vulnerable to various attacks. To further unveil the safety risks of LLMs, we introduce a Safety Concept Activation Vector (SCAV) framework, which effectively guides the attacks by accurately interpreting LLMs' safety mech…

2022

A Framework of Rehabilitation-assisted Robot Skill Representation, Learning, and Modulation via Manifold-Mappings and Gaussian Processes

IROS 2022poster

Stroke survivors usually have dyskinesia, who have an urgent need for rehabilitation-assist training. To reduce the labor of rehabilitation therapists, this paper attempts to investigate an effective rehabilitation-assisted robot skill acquisition framework, which is inspired by the scheme of robot…

Cited by 5SourceScholar