← Search

Ge Gao

36 accepted papers

2026

ArcDAE: Asymmetric Rectified Contrastive Diffusion Autoencoder for Unified Representation Learning

ICML 2026poster

The unification of generative details and discriminative semantics presents a structural paradox in \textit{diffusion-based representation learning}. Early approaches decouple semantics from generation, inevitably compromising representational completeness (i.e., \textit{information split}). While r…

Cited by 0SourceScholar
2026

Human-in-the-Loop Capacitive Microphone Sensors-Based Muscle Sensing System for Predictive and Adaptive Exoskeleton Assistance

RA-L 2026

Mobility impairments among older adults and individuals with neuromuscular weakness motivate the need for timely and adaptive exoskeleton assistance. This paper presents a human-in-the-loop muscle sensing and control system based on capacitive microphone sensors (CMS) that capture subtle mechanical

Cited by 0SourceScholar
2025

An Interdisciplinary Approach to Human-Centered Machine Translation

EMNLP 2025

Machine Translation (MT) tools are widely used today, often in contexts where professional translators are not present. Despite progress in MT technology, a gap persists between system development and real-world usage, particularly for non-expert users who may struggle to assess translation reliabil

Cited by 0SourcePDFScholar
2025

FatesGS: Fast and Accurate Sparse-View Surface Reconstruction Using Gaussian Splatting with Depth-Feature Consistency

AAAI 2025technical

Recently, Gaussian Splatting has sparked a new trend in the field of computer vision. Apart from novel view synthesis, it has also been extended to the area of multi-view reconstruction. The latest methods facilitate complete, detailed surface reconstruction while ensuring fast training speed. Howev…

Cited by 2SourcePDFScholar
2025

From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora

EMNLP 2025

Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limits its ability to effectively capture cross-lingual semantics. In contrast, mult

2025

HIIF: Hierarchical Encoding based Implicit Image Function for Continuous Super-resolution

CVPR 2025poster

Recent advances in implicit neural representations (INRs) have shown significant promise in modeling visual signals for various low-vision tasks including image super-resolution (ISR). INR-based ISR methods typically learn continuous representations, providing flexibility for generating high-resolut…

2025

Principled Fine-tuning of LLMs from User-Edits: A Medley of Preference, Supervision, and Reward

NeurIPS 2025poster

We study how to fine-tune LLMs using user-edit deployment data consisting of a set of context, an agent's response, and user edits. This deployment data is naturally generated by users in applications such as LLMs-based writing assistants and coding agents. The _natural_ origin of user edits makes i…

Cited by 0SourceScholar
2025

Sparis: Neural Implicit Surface Reconstruction of Indoor Scenes from Sparse Views

AAAI 2025technical

In recent years, reconstructing indoor scene geometry from multi-view images has achieved encouraging accomplishments. Current methods incorporate monocular priors into neural implicit surface models to achieve high-quality reconstructions. However, these methods require hundreds of images for scene…

Cited by 2SourcePDFScholar
2025

Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect Translations

EMNLP 2025

As Machine Translation (MT) becomes increasingly commonplace, understanding how the general public perceives and relies on imperfect MT is crucial for contextualizing MT research in real-world applications. We present a human study conducted in a public museum (n=452), investigating how fluency and

Cited by 0SourcePDFScholar
2024

Aligning LLM Agents by Learning Latent Preference from User Edits

NeurIPS 2024poster

We study interactive learning of language agents based on user edits made to the agent's output. In a typical setting such as writing assistants, the user interacts with a language agent to generate a response given a context, and may optionally edit the agent response to personalize it based on the…

2024

Enhancing Emotion Prediction in News Headlines: Insights from ChatGPT and Seq2Seq Models for Free-Text Generation

COLING 2024main

Predicting emotions elicited by news headlines can be challenging as the task is largely influenced by the varying nature of people’s interpretations and backgrounds. Previous works have explored classifying discrete emotions directly from news headlines. We provide a different approach to tackling…

Cited by 1SourcePDFScholar
2024

GridFormer: Point-Grid Transformer for Surface Reconstruction

AAAI 2024technical

Implicit neural networks have emerged as a crucial technology in 3D surface reconstruction. To reconstruct continuous surfaces from discrete point clouds, encoding the input points into regular grid features (plane or volume) has been commonly employed in existing approaches. However, these methods…

2024

I Could’ve Asked That: Reformulating Unanswerable Questions

EMNLP 2024main

When seeking information from unfamiliar documents, users frequently pose questions that cannot be answered by the documents. While existing large language models (LLMs) identify these unanswerable questions, they do not assist users in reformulating their questions, thereby reducing their overall u…

2024

Implicit Filtering for Learning Neural Signed Distance Functions from 3D Point Clouds

ECCV 2024poster

"Neural signed distance functions (SDFs) have shown powerful ability in fitting the shape geometry. However, inferring continuous signed distance fields from discrete unoriented point clouds still remains a challenge. The neural network typically fits the shape with a rough surface and omits fine-gr…

2024

NeuSurf: On-Surface Priors for Neural Surface Reconstruction from Sparse Input Views

AAAI 2024technical

Recently, neural implicit functions have demonstrated remarkable results in the field of multi-view reconstruction. However, most existing methods are tailored for dense views and exhibit unsatisfactory performance when dealing with sparse views. Several latest methods have been proposed for general…

Cited by 22SourcePDFScholar
2024

Off-Policy Selection for Initiating Human-Centric Experimental Design

NeurIPS 2024poster

In human-centric applications like healthcare and education, the \textit{heterogeneity} among patients and students necessitates personalized treatments and instructional interventions. While reinforcement learning (RL) has been utilized in those tasks, off-policy selection (OPS) is pivotal to close…

Cited by 0SourcePDFScholar
2024

On Trajectory Augmentations for Off-Policy Evaluation

ICLR 2024poster

In the realm of reinforcement learning (RL), off-policy evaluation (OPE) holds a pivotal position, especially in high-stake human-involved scenarios such as e-learning and healthcare. Applying OPE to these domains is often challenging with scarce and underrepresentative offline training trajectories…

Cited by 4SourcePDFScholar
2023

HiNeRV: Video Compression with Hierarchical Encoding-based Neural Representation

NeurIPS 2023poster

Learning-based video compression is currently a popular research topic, offering the potential to compete with conventional standard video codecs. In this context, Implicit Neural Representations (INRs) have previously been used to represent and compress image and video content, demonstrating relati…

Cited by 55SourcePDFScholar
2023

Off-Policy Evaluation for Human Feedback

NeurIPS 2023poster

Off-policy evaluation (OPE) is important for closing the gap between offline training and evaluation of reinforcement learning (RL), by estimating performance and/or rank of target (evaluation) policies using offline trajectories only. It can improve the safety and efficiency of data collection and…

Cited by 8SourcePDFScholar
2023

Physician Detection of Clinical Harm in Machine Translation: Quality Estimation Aids in Reliance and Backtranslation Identifies Critical Errors

EMNLP 2023long main

A major challenge in the practical use of Machine Translation (MT) is that users lack information on translation quality to make informed decisions about how to rely on outputs. Progress in quality estimation research provides techniques to automatically assess MT quality, but these techniques have…

Cited by 0SourcecodeScholar
2022

A Reinforcement Learning-Informed Pattern Mining Framework for Multivariate Time Series Classification

IJCAI 2022poster

Multivariate time series (MTS) classification is a challenging and important task in various domains and real-world applications. Much of prior work on MTS can be roughly divided into neural network (NN)- and pattern-based methods. The former can lead to robust classification performance, but many o…

2022

Cross-Linked Unified Embedding for cross-modality representation learning

NeurIPS 2022accept

Multi-modal learning is essential for understanding information in the real world. Jointly learning from multi-modal data enables global integration of both shared and modality-specific information, but current strategies often fail when observa- tions from certain modalities are incomplete or missi…

Cited by 31SourcePDFScholar
2022

Simulating Bandit Learning from User Feedback for Extractive Question Answering

ACL 2022long

We study learning from user feedback for extractive question answering by simulating feedback using supervised data. We cast the problem as contextual bandit learning, and analyze the characteristics of several learning scenarios with focus on reducing data annotation. We show that systems initially…

2021

CloudAAE: Learning 6D Object Pose Regression with On-line Data Synthesis on Point Clouds

ICRA 2021poster

It is often desired to train 6D pose estimation systems on synthetic data because manual annotation is expensive. However, due to the large domain gap between the synthetic and real images, synthesizing color images is expensive. In contrast, this domain gap is considerably smaller and easier to fil…

Cited by 58SourcecodeScholar
2021

Neural Image Compression via Attentional Multi-Scale Back Projection and Frequency Decomposition

ICCV 2021poster

In recent years, neural image compression emerges as a rapidly developing topic in computer vision, where the state-of-the-art approaches now exhibit superior compression performance than their conventional counterparts. Despite the great progress, current methods still have limitations in preservin…

Cited by 89PDFScholar
2020

6D Object Pose Regression via Supervised Learning on Point Clouds

ICRA 2020poster

This paper addresses the task of estimating the 6 degrees of freedom pose of a known 3D object from depth information represented by a point cloud. Deep features learned by convolutional neural networks from color information have been the dominant features to be used for inferring object poses, whi…

Cited by 111SourcecodeScholar