← Search

Xiang Gao

57 accepted papers

2026

Boomda: Balanced Multi-objective Optimization for Multimodal Domain Adaptation

AAAI 2026technical

Multimodal learning, while contributing to numerous success stories across various fields, faces the challenge of prohibitively expensive manual annotation. To address the scarcity of annotated data, a popular solution is unsupervised domain adaptation, which has been extensively studied in unimodal

Cited by 0SourcePDFScholar
2026

BuildingGPT: Auto-Regressive Building Wireframe Reconstruction Model with Reinforcement Learning

CVPR 2026

In this paper, we propose BuildingGPT, a novel auto-regressive model for building wireframe reconstruction from point clouds with reinforcement learning.Unlike prior works based on detection or diffusion models, BuildingGPT reformulates the building wireframe reconstruction task into a sequence pred

Cited by 0SourcecodeScholar
2026

Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction

ICLR 2026poster

The evaluation of code-generating Large Language Models (LLMs) is fundamentally constrained by two intertwined challenges: a reliance on static, easily contaminated problem sources and the use of superficial, low-rigor testing. This paper introduces a new benchmark construction philosophy, Dual Scal…

Cited by 0SourcecodeScholar
2026

DiscoX: Benchmarking Discourse-Level Translation in Expert Domains

ICLR 2026poster

The evaluation of discourse-level translation in expert domains remains inadequate, despite its centrality to knowledge dissemination and cross-lingual scholarly communication. While these translations demand discourse-level coherence and strict terminological precision, current evaluation methods p…

Cited by 0SourcecodeScholar
2026

FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning

ICLR 2026poster

Search has emerged as core infrastructure for LLM-based agents and is widely viewed as critical on the path toward more general intelligence. Finance is a particularly demanding proving ground: analysts routinely conduct complex, multi-step searches over time-sensitive, domain-specific data, making…

Cited by 0SourcecodeScholar
2026

FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction

ICLR 2026poster

Future prediction is a complex task for LLM agents, requiring a high level of analytical thinking, information gathering, contextual understanding, and decision-making under uncertainty. Agents must not only gather and interpret vast amounts of dynamic information but also integrate diverse data sou…

Cited by 0SourceScholar
2026

Inverse Rendering for High-Genus 3D Surface Meshes from Multi-view Images with Persistent Homology Priors

ICASSP 2026poster

Reconstructing 3D objects from images is inherently an ill-posed problem due to ambiguities in geometry, appearance, and topology. This paper introduces collaborative inverse rendering with persistent homology priors, a novel strategy that leverages topological constraints to resolve these ambiguiti…

Cited by 0SourcePDFScholar
2026

NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents

ICML 2026poster

Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks primarily evaluate short-horizon behaviors such as localized code generation, scaffolded completion, or repository repair, leaving it unclear whether agents can sustain coherent re…

Cited by 0SourceScholar
2026

PGS: Effective LLM Code Refinement via Property-Oriented and Structurally Minimal Feedback

ICML 2026poster

LLMs excel at code generation, yet ensuring the functional correctness of their outputs remains a persistent challenge. While recent studies have applied Test-Driven Development (TDD) to refine code, these methods are often undermined by poor feedback quality, stemming from the scarcity of high-qual…

Cited by 0SourceScholar
2026

REMem: Reasoning with Episodic Memory in Language Agent

ICLR 2026poster

Humans excel at remembering concrete experiences along spatiotemporal contexts and performing reasoning across those events, i.e., the capacity for episodic memory. In contrast, memory in language agents remains mainly semantic, and current agents are not yet capable of effectively recollecting and…

Cited by 0SourcecodeScholar
2026

WIET: Harmonizing Group-aware Model Weighting and Worker Allocation for Ensemble Temporal Prediction MaaS

AAAI 2026technical

Ensemble Temporal Prediction Model-as-a-Service (ETP-MaaS) has become crucial in fields like financial modeling and cloud monitoring. Existing solutions fail to co-optimally address a two-fold challenge of dynamic collaboration and heterogeneity, treating models as independent entities and employing

Cited by 0SourcePDFScholar
2025

BWFormer: Building Wireframe Reconstruction from Airborne LiDAR Point Cloud with Transformer

CVPR 2025highlight

In this paper, we present BWFormer, a novel Transformer-based model for building wireframe reconstruction from airborne LiDAR point cloud. The problem is solved in a ground-up manner here by detecting the building corners in 2D, lifting and connecting them in 3D space afterwards with additional dat…

2025

CABS: Conflict-Aware and Balanced Sparsification for Enhancing Model Merging

ICML 2025poster

Model merging based on task vectors, i.e., the parameter differences between fine-tuned models and a shared base model, provides an efficient way to integrate multiple task-specific models into a multitask model without retraining. Recent works have endeavored to address the conflicts between task v…

Cited by 0SourcePDFScholar
2025

GIFStream: 4D Gaussian-based Immersive Video with Feature Stream

CVPR 2025poster

Immersive video offers a 6-Dof-free viewing experience, potentially playing a key role in future video technology. Recently, 4D Gaussian Splatting has gained attention as an effective approach for immersive video due to its high rendering efficiency and quality, though maintaining quality with manag…

Cited by 0SourcePDFScholar
2025

Gradient-guided Attention Map Editing: Towards Efficient Contextual Hallucination Mitigation

NAACL 2025findings

In tasks such as summarization and open-book question answering (QA), Large Language Models (LLMs) frequently experience “contextual hallucination”, where they generate irrelevant or incorrect responses despite having access to accurate information in the input. This issue often stems from the model…

2025

Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval

NeurIPS 2025poster

While an image is worth more than a thousand words, only a few provide crucial information for a given task and thus should be focused on. In light of this, ideal text-to-image (T2I) retrievers should prioritize specific visual attributes relevant to queries. To evaluate current retrievers on handli…

Cited by 0SourceScholar
2025

Learning to Search Effective Example Sequences for In-Context Learning

NAACL 2025findings

Large language models (LLMs) demonstrate impressive few-shot learning capabilities, but their performance varies widely based on the sequence of in-context examples. Key factors influencing this include the sequence’s length, composition, and arrangement, as well as its relation to the specific quer…

Cited by 1SourcePDFScholar
2025

PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion Model

CVPR 2025poster

Optical illusion hidden picture is an interesting visual perceptual phenomenon where an image is cleverly integrated into another picture. Established on the off-the-shelf text-to-image (T2I) diffusion model, we propose a novel text-guided image-to-image (I2I) translation framework dubbed as Phase-T…

2025

SSFSL: Self-Supervised and Few-Shot Learning for Cross-Domain Hyperspectral Image Classification

ICASSP 2025accepted

Few-shot learning (FSL) has gained increasing attention in hyperspectral image (HSI) classification due to its ability to perform cross-domain classification with minimal labeled samples. However, existing FSL methods overlook the continuity of HSI spectral sequences and fail to utilize the large am…

Cited by 0SourceScholar
2025

The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs

ACL 2025finding

Large Language Model (LLM)-based agents have significantly impacted Task-Oriented Dialog Systems (TODS) but continue to face notable performance challenges, especially in zero-shot scenarios. While prior work has noted this performance gap, the behavioral factors driving the performance gap remain u…

2025

UGM2N: An Unsupervised and Generalizable Mesh Movement Network via M-Uniform Loss

NeurIPS 2025poster

Partial differential equations (PDEs) form the mathematical foundation for modeling physical systems in science and engineering, where numerical solutions demand rigorous accuracy-efficiency tradeoffs. Mesh movement techniques address this challenge by dynamically relocating mesh nodes to rapidly-va…

Cited by 0SourceScholar
2025

UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action Generation

EMNLP 2025

Mobile task automation is an emerging technology that leverages AI to automatically execute routine tasks by users’ commands on mobile devices like Android, thus enhancing efficiency and productivity. While large language models (LLMs) excel at general mobile tasks through training on massive datase

2024

Bi2Lane: Bi-Directional Temporal Refinement with Bi-Level Feature Aggregation for 3D Lane Detection

ICRA 2024poster

Monocular 3D lane detection has recently received increasing research attention in autonomous driving due to its application effectiveness and simplicity. However, depending solely on the limited semantic information from a single image makes current monocular detection methods unable to deal with c…

Cited by 0SourceScholar
2024

Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation

AAAI 2024technical

Recently, text-to-image diffusion models have emerged as a powerful tool for image-to-image translation (I2I), allowing flexible image translation via user-provided text prompts. This paper proposes frequency-controlled diffusion model (FCDiffusion), an end-to-end diffusion-based framework contribut…

2024

Generative 3D Part Assembly via Part-Whole-Hierarchy Message Passing

CVPR 2024poster

Generative 3D part assembly involves understanding part relationships and predicting their 6-DoF poses for assembling a realistic 3D shape. Prior work often focus on the geometry of individual parts neglecting part-whole hierarchies of objects. Leveraging two key observations: 1) super-part poses pr…

2024

InvariantOODG: Learning Invariant Features of Point Clouds for Out-of-Distribution Generalization

ICASSP 2024accepted

The convenience of 3D sensors has led to an increase in the use of 3D point clouds in various applications. However, the differences in acquisition devices or scenarios lead to divergence in the data distribution of point clouds, which requires good generalization of point cloud representation learn…

Cited by 0SourceScholar
2024

Mitigating Hallucination in Fictional Character Role-Play

EMNLP 2024finding

Role-playing has wide-ranging applications in customer support, embodied agents, and computational social science. The influence of parametric world knowledge of large language models (LLMs) often causes role-playing characters to act out of character and to hallucinate about things outside the scop…

2024

PVALane: Prior-Guided 3D Lane Detection with View-Agnostic Feature Alignment

AAAI 2024technical

Monocular 3D lane detection is essential for a reliable autonomous driving system and has recently been rapidly developing. Existing popular methods mainly employ a predefined 3D anchor for lane detection based on front-viewed (FV) space, aiming to mitigate the effects of view transformations. Howev…

Cited by 6SourcePDFScholar
2024

PolyRoom: Room-aware Transformer for Floorplan Reconstruction

ECCV 2024poster

"Reconstructing geometry and topology structures from raw unstructured data has always been an important research topic in indoor mapping research. In this paper, we aim to reconstruct the floorplan with a vectorized representation from point clouds. Despite significant advancements achieved in rece…

2023

Coarse-to-fine Few-shot Learning for Named Entity Recognition

ACL 2023findings

Recently, Few-shot Named Entity Recognition has received wide attention with the growing need for NER models to learn new classes with minimized annotation costs. However, one common yet understudied situation is to transfer a model trained with coarse-grained classes to recognize fine-grained class…

2023

Divide Rows and Conquer Cells: Towards Structure Recognition for Large Tables

IJCAI 2023poster

Recent advanced Table Structure Recognition (TSR) models adopt image-to-text solutions to parse table structure. These methods can be formulated as image caption problem, i.e., input a single-table image and output table structure description in a specific text format, e.g., HTML. With the impressiv…

Cited by 20SourcePDFScholar
2023

Farewell to Aimless Large-scale Pretraining: Influential Subset Selection for Language Model

ACL 2023findings

Pretrained language models have achieved remarkable success in various natural language processing tasks. However, pretraining has recently shifted toward larger models and larger data, which has resulted in significant computational and energy costs. In this paper, we propose Influence Subset Selec…

2023

Learning “O” Helps for Learning More: Handling the Unlabeled Entity Problem for Class-incremental NER

ACL 2023long

As the categories of named entities rapidly increase, the deployed NER models are required to keep updating toward recognizing more entity types, creating a demand for class-incremental learning for NER. Considering the privacy concerns and storage constraints, the standard paradigm for class-increm…

2023

Look and Think: Intrinsic Unification of Self-Attention and Convolution for Spatial-Channel Specificity

ICASSP 2023accepted

Convolution and self-attention are popular paradigms and many works take them as two separate components to explore their potential combination. In this work, we consider their intrinsic properties in spatial and channel domains for vision representation. Convolution has the great property of channe…

Cited by 0SourceScholar
2023

Open Set Relation Extraction via Unknown-Aware Training

ACL 2023long

The existing supervised relation extraction methods have achieved impressive performance in a closed-set setting, in which the relations remain the same during both training and testing. In a more realistic open-set setting, unknown relations may appear in the test set. Due to the lack of supervisio…

2023

Structured Neural Networks for Density Estimation and Causal Inference

NeurIPS 2023poster

Injecting structure into neural networks enables learning functions that satisfy invariances with respect to subsets of inputs. For instance, when learning generative models using neural networks, it is advantageous to encode the conditional independence structure of observed variables, often in the…

Cited by 7SourcePDFScholar
2022

Anderson Acceleration for on-Manifold Iterated Error State Kalman Filters

RA-L 2022

Iterated Extended Kalman Filter is a promising and widely-used estimator for real-time localization applications. It iterates the observation equation to find a better linearization point and, simultaneously, only maintains the state estimation in a single time to save the computation resources. Ins

Cited by 10SourceScholar
2022

Faster-LIO: Lightweight Tightly Coupled Lidar-Inertial Odometry Using Parallel Sparse Incremental Voxels

RA-L 2022

This letter presents an incremental voxel-based lidar-inertial odometry (LIO) method for fast-tracking spinning and solid-state lidar scans. To achieve the high tracking speed, we neither use complicated tree-based structures to divide the spatial point cloud nor the strict k nearest neighbor (k-NN)

Cited by 338SourceScholar
2022

Human-Robotic Prosthesis as Collaborating Agents for Symmetrical Walking

NeurIPS 2022accept

This is the first attempt at considering human influence in the reinforcement learning control of a robotic lower limb prosthesis toward symmetrical walking in real world situations. We propose a collaborative multi-agent reinforcement learning (cMARL) solution framework for this highly complex and…

Cited by 12SourcePDFScholar
2022

Learning From Demonstrations Via Multi-Level and Multi-Attention Domain-Adaptive Meta-Learning

RA-L 2022

Despite significant advances in few-shot classification, object detection, or speech recognition in recent years, training an effective robot to adapt to previously unseen environments in a small data regime is still a long-lasting problem for learning from demonstrations (LfD). A promising solution

Cited by 3SourceScholar
2022

Learning With Dual Demonstration Domains: Random Domain-Adaptive Meta-Learning

RA-L 2022

Although robots have been widely applied in various fields, allowing a robot to perform a wide range of tasks like humans is a significant challenge. One promising method is meta-learning, which enables robots to learn from demonstrations with the concept of “learning to learn.” Howeve

Cited by 7SourceScholar
2022

Learning to Incorporate Texture Saliency Adaptive Attention to Image Cartoonization

ICML 2022spotlight

Image cartoonization is recently dominated by generative adversarial networks (GANs) from the perspective of unsupervised image-to-image translation, in which an inherent challenge is to precisely capture and sufficiently transfer characteristic cartoon styles (e.g., clear edges, smooth color shadin…

2022

RetGen: A Joint Framework for Retrieval and Grounded Text Generation Modeling

AAAI 2022technical

Recent advances in large-scale pre-training such as GPT-3 allow seemingly high quality text to be generated from a given prompt. However, such generation systems often suffer from problems of hallucinated facts, and are not inherently designed to incorporate useful external information. Grounded gen…

2021

A Controllable Model of Grounded Response Generation

AAAI 2021technical

Current end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, as attested by pretrained language…

2021

NICE: Neural Image Commenting with Empathy

EMNLP 2021finding

Emotion and empathy are examples of human qualities lacking in many human-machine interactions. The goal of our work is to generate engaging dialogue grounded in a user-shared image with increased emotion and empathy while minimizing socially inappropriate or offensive outputs. We release the Neural…

Cited by 7SourcePDFScholar
2021

RGLN: Robust Residual Graph Learning Networks via Similarity-Preserving Mapping on Graphs

ICASSP 2021accepted

Graph Convolutional Neural Networks (GCNNs) extend CNNs to irregular graph data domain, such as brain networks, citation networks and 3D point clouds. It is critical to identify an appropriate graph for basic operations in GCNNs. Existing methods often manually construct or learn one fixed graph bas…

Cited by 0SourceScholar
2020

GraphTER: Unsupervised Learning of Graph Transformation Equivariant Representations via Auto-Encoding Node-Wise Transformations

CVPR 2020poster

Recent advances in Graph Convolutional Neural Networks (GCNNs) have shown their efficiency for nonEuclidean data on graphs, which often require a large amount of labeled data with high cost. It it thus critical to learn graph feature representations in an unsupervised manner in practice. To this end…

Cited by 56PDFcodeScholar
2020

Knowledge-Guided Reinforcement Learning Control for Robotic Lower Limb Prosthesis

ICRA 2020poster

Robotic prostheses provide new opportunities to better restore lost functions than passive prostheses for trans-femoral amputees. But controlling a prosthesis device automatically for individual users in different task environments is an unsolved problem. Reinforcement learning (RL) is a naturally p…

Cited by 24SourceScholar
2019

Offline Policy Iteration Based Reinforcement Learning Controller for Online Robotic Knee Prosthesis Parameter Tuning

ICRA 2019poster

This paper aims to develop an optimal controller that can automatically provide personalized control of robotic knee prosthesis in order to best support gait of individual prosthesis wearers. We introduced a new reinforcement learning (RL) controller for this purpose based on the promising ability o…

Cited by 27SourceScholar
2018

Challenges in Monocular Visual Odometry: Photometric Calibration, Motion Bias, and Rolling Shutter Effect

RA-L 2018

Monocular visual odometry (VO) and simultaneous localization and mapping (SLAM) have seen tremendous improvements in accuracy, robustness, and efficiency, and have gained increasing popularity over recent years. Nevertheless, not so many discussions have been carried out to reveal the influences of

Cited by 117SourceScholar