← Search

Yuxuan Liu

53 accepted papers

2026

Causality-inspired Federated Learning for Dynamic Spatio-Temporal Graphs

AAAI 2026technical

Federated Graph Learning (FGL) has emerged as a powerful paradigm for decentralized training of graph neural networks while preserving data privacy. However, existing FGL methods are predominantly designed for static graphs and rely on parameter averaging or distribution alignment, which implicitly

Cited by 0SourcePDFScholar
2026

CoME: Empowering Channel-of-Mobile-Experts with Informative Hybrid-Capabilities Reasoning

ICML 2026poster

Mobile Agents can autonomously execute user instructions, which requires hybrid-capabilities reasoning, including screen summary, subtask planning, action decision and action function. However, existing agents struggle to achieve both decoupled enhancement and balanced integration of these capabilit…

Cited by 0SourceScholar
2026

GenAlign: Towards Unified Alignment Framework of MLLMs via Generative Reward Model

ICML 2026poster

Aligning Multimodal Large Language Models (MLLMs) with human preferences remains a fundamental challenge. While Generative Reward Models (GRMs) offer a promising reasoning-based alternative to scalar models, they are often hindered by severe position bias and prohibitively high computational overhea…

Cited by 0SourceScholar
2026

Learning Task-Invariant Properties Via Dreamer: Enabling Efficient Policy Transfer for Quadruped Robots

ICRA 2026poster

Achieving quadruped robot locomotion across diverse and dynamic terrains presents significant challenges, primarily due to the discrepancies between simulation environments and real-world conditions. Traditional sim-to-real transfer methods often rely on manual feature design or costly real-world fi…

2026

MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning

ICLR 2026poster

The Chain of Action-Planning Thoughts (CoaT) paradigm has been shown to improve the reasoning performance of VLM-based mobile agents in GUI tasks. However, the scarcity of diverse CoaT trajectories limits the expressiveness and generalization ability of such agents. While self-training is commonly e…

Cited by 0SourceScholar
2026

SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy Tasks

ICLR 2026poster

VLM-based mobile agents are increasingly popular due to their capabilities to interact with smartphone GUIs and XML-structured texts and to complete daily tasks. However, existing online benchmarks fail to obtain stable critical reward signals under dynamic environmental changes, and neglect the inf…

Cited by 0SourcecodeScholar
2026

Scalable Spatio-Temporal SE(3) Diffusion for Long-Horizon Protein Dynamics

ICLR 2026poster

Molecular dynamics (MD) simulations remain the gold standard for studying protein dynamics, but their computational cost limits access to biologically relevant timescales. Recent generative models have shown promise in accelerating simulations, yet they struggle with long-horizon generation due to a…

Cited by 0SourceScholar
2025

BiDeV: Bilateral Defusing Verification for Complex Claim Fact-Checking

AAAI 2025technical

Complex claim fact-checking performs a crucial role in disinformation detection. However, existing fact-checking methods struggle with claim vagueness, specifically in effectively handling latent information and complex relations within claims. Moreover, evidence redundancy, where non-essential info…

2025

Cross-modal Ship Re-Identification via Optical and SAR Imagery: A Novel Dataset and Method

ICCV 2025poster

Detecting and tracking ground objects using earth observation imagery remains a significant challenge in the field of remote sensing. Continuous maritime ship tracking is crucial for applications such as maritime search and rescue, law enforcement, and shipping analysis. However, most current ship t…

2025

Deep Coarse-to-Fine Networks for Robust Segmentation and Pose Estimation of Surgical Suturing Threads

IROS 2025

Autonomous suturing is a critical challenge in robot-assisted surgery, where accurate segmentation and pose estimation of suturing threads are essential prerequisites. However, suturing threads are easily occluded by moving instruments and embedded in deformable tissues which make the task much more

Cited by 0SourceScholar
2025

FisheyeDepth: A Real Scale Self-Supervised Depth Estimation Model for Fisheye Camera

ICRA 2025

Accurate depth estimation is crucial for 3D scene comprehension in robotics and autonomous vehicles. Fisheye cameras, known for their wide field of view, have inherent geometric benefits. However, their use in depth estimation is restricted by a scarcity of ground truth data and image distortions. W

Cited by 7SourcecodeScholar
2025

Self-Deformable Magnetic Miniature Robot for Traction Assistance in Endoscopic Submucosal Dissection

ICRA 2025

Between 1999 and 2020, gastrointestinal cancers were responsible for over three million deaths, emphasizing the critical role of minimally invasive surgical techniques like Endoscopic Submucosal Dissection (ESD) in managing such life-threatening conditions. ESD, which dissects the connective tissue

Cited by 1SourceScholar
2025

TSCLIP: Robust CLIP Fine-Tuning for Worldwide Cross-Regional Traffic Sign Recognition

ICRA 2025

Traffic sign is a critical map feature for navigation and traffic control. Nevertheless, current methods for traffic sign recognition rely on traditional deep learning models, which typically suffer from significant performance degradation considering the variations in data distribution across diffe

Cited by 6SourcecodeScholar
2025

Towards Accurate Brain Electrode Implantation via Cross-modality Fusion of White-light and Photoacoustic Microscopy

IROS 2025

Invasive flexible neural electrodes are becoming increasingly prevalent in monitoring and modulating brain neural activity, necessitating the precise and minimally invasive implantation of these electrodes to a depth of a few millimeters beneath the cerebral surface. Although Neuralink has pioneered

Cited by 0SourceScholar
2024

Advancing Video Synchronization with Fractional Frame Analysis: Introducing a Novel Dataset and Model

AAAI 2024technical

Multiple views play a vital role in 3D pose estimation tasks. Ideally, multi-view 3D pose estimation tasks should directly utilize naturally collected videos for pose estimation. However, due to the constraints of video synchronization, existing methods often use expensive hardware devices to synchr…

2024

Calibrating LLM-Based Evaluator

COLING 2024main

Recent advancements in large language models (LLMs) and their emergent capabilities make LLM a promising reference-free evaluator on the quality of natural language generation, and a competent alternative to human evaluation. However, hindered by the closed-source or high computational demand to hos…

Cited by 73SourcePDFScholar
2024

Disentangle Estimation of Causal Effects from Cross-Silo Data

ICASSP 2024accepted

Estimating causal effects among different events is of great importance to critical fields such as drug development. Nevertheless, the data features associated with events may be distributed across various silos and remain private within respective parties, impeding direct information exchange betwe…

Cited by 0SourceScholar
2024

Every Dataset Counts: Scaling up Monocular 3D Object Detection with Joint Datasets Training

IROS 2024poster

Monocular 3D object detection is essential for autonomous driving. However, current monocular 3D detection algorithms rely on expensive 3D labels from LiDAR scans, making it difficult to use in new datasets and unfamiliar environments. This study explores training a monocular 3D object detection mod…

Cited by 4SourceScholar
2024

Fast Photoacoustic Microscopy with Robot Controlled Microtrajectory Optimization

ICRA 2024poster

Photoacoustic Microscopy (PAM) is a relatively new imaging modality in biomedicine. However, point-by-point raster scanning in PAM suffers from low imaging speed. Sparse sampling has been studied in recent years and with the development of deep learning algorithms, extensive efforts have been devote…

Cited by 0SourceScholar
2024

HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition

ACL 2024long

Large language models (LLMs) have emerged as a promising alternative to expensive human evaluations. However, the alignment and coverage of LLM-based evaluations are often limited by the scope and potential bias of the evaluation prompts and criteria. To address this challenge, we propose HD-Eval, a…

2024

Modeling Adaptive Inter-Task Feature Interactions via Sentiment-Aware Contrastive Learning for Joint Aspect-Sentiment Prediction

AAAI 2024technical

Aspect prediction (AP) and sentiment prediction (SP) are representative applications in fine-grained sentiment anal- ysis. They can be considered as sequential tasks, where AP identifies mentioned aspects in a sentence, and SP infers fine-grained sentiments for these aspects. Recent models perform t…

Cited by 8SourcePDFScholar
2024

ODGEN: Domain-specific Object Detection Data Generation with Diffusion Models

NeurIPS 2024poster

Modern diffusion-based image generative models have made significant progress and become promising to enrich training data for the object detection task. However, the generation quality and the controllability for complex scenes containing multi-class objects and dense objects with occlusions remain…

Cited by 5SourcePDFScholar
2024

Text Diffusion with Reinforced Conditioning

AAAI 2024technical

Diffusion models have demonstrated exceptional capability in generating high-quality images, videos, and audio. Due to their adaptiveness in iterative refinement, they provide a strong potential for achieving better non-autoregressive sequence generation. However, existing text diffusion models stil…

Cited by 1SourcePDFScholar
2023

CenterLineDet: CenterLine Graph Detection for Road Lanes with Vehicle-mounted Sensors by Transformer for HD Map Generation

ICRA 2023poster

With the fast development of autonomous driving technologies, there is an increasing demand for high-definition (HD) maps, which provide reliable and robust prior information about the static part of the traffic environments. As one of the important elements in HD maps, road lane centerline is criti…

Cited by 18SourcecodeScholar
2023

Democratizing Reasoning Ability: Tailored Learning from Large Language Model

EMNLP 2023long main

Large language models (LLMs) exhibit impressive emergent abilities in natural language processing, but their democratization is hindered due to huge computation requirements and closed-source nature. Recent research on advancing open-source smaller LMs by distilling knowledge from black-box LLMs has…

Cited by 0SourcecodeScholar
2023

Distributional Instance Segmentation: Modeling Uncertainty and High Confidence Predictions with Latent-MaskRCNN

ICRA 2023poster

Object recognition and instance segmentation are fundamental skills in any robotic or autonomous system. Existing state-of-the-art methods are often unable to capture meaningful uncertainty in challenging or ambiguous scenes, and as such can cause critical errors in high-performance applications. In…

Cited by 4SourceScholar
2023

EasyGaze3D: Towards Effective and Flexible 3D Gaze Estimation from a Single RGB Camera

IROS 2023poster

Eye gaze can convey rich information of human intentions, which enables the social robots to comprehend the cognition and behavior of human targets. However, the existing 3D gaze estimation methods generally have high requirements either on the dedicated hardware or the quantity and quality of train…

Cited by 3SourceScholar
2023

EgoHMR: Egocentric Human Mesh Recovery via Hierarchical Latent Diffusion Model

ICRA 2023poster

Egocentric vision has gained increasing popularity in social robotics, demonstrating great potentials for personal assistance and human-centric behavior analysis. Holistic per-ception of human body itself is a prerequisite for downstream applications, including action recognition and anticipation. E…

Cited by 12SourceScholar
2023

Fairness-aware Contrastive Learning with Partially Annotated Sensitive Attributes

ICLR 2023poster

Learning high-quality representation is important and essential for visual recognition. Unfortunately, traditional representation learning suffers from fairness issues since the model may learn information of sensitive attributes. Recently, a series of studies have been proposed to improve fairness…

Cited by 35SourcePDFScholar
2023

Learning Instrumental Variable from Data Fusion for Treatment Effect Estimation

AAAI 2023technical

The advent of the big data era brought new opportunities and challenges to draw treatment effect in data fusion, that is, a mixed dataset collected from multiple sources (each source with an independent treatment assignment mechanism). Due to possibly omitted source labels and unmeasured confounders…

2023

RNGDet++: Road Network Graph Detection by Transformer With Instance Segmentation and Multi-Scale Features Enhancement

RA-L 2023

The road network graph is a critical component for downstream tasks in autonomous driving, such as global route planning and navigation. In the past years, road network graphs are usually annotated by human experts manually, which is time-consuming and labor-intensive. To annotate road network graph

Cited by 50SourceScholar
2022

Autoregressive Uncertainty Modeling for 3D Bounding Box Prediction

ECCV 2022poster

"3D bounding boxes are a widespread intermediate representation in many computer vision applications. However, predicting them is a challenging task, largely due to partial observability, which motivates the need for a strong sense of uncertainty. While many recent methods have explored better archi…

Cited by 7SourcePDFScholar
2022

Ego+X: An Egocentric Vision System for Global 3D Human Pose Estimation and Social Interaction Characterization

IROS 2022poster

Egocentric vision is an emerging topic, which has demonstrated great potential in assistive healthcare scenarios, ranging from human-centric behavior analysis to personal social assistance. Within this field, due to the heterogeneity of visual perception from first-person views, egocentric pose esti…

Cited by 9SourceScholar
2022

Integrating Dependency Tree into Self-Attention for Sentence Representation

ICASSP 2022accepted

Recent progress on parse tree encoder for sentence representation learning is notable. However, these works mainly en-code tree structures recursively, which is not conducive to parallelization. On the other hand, these works rarely take into account the labels of arcs in dependency trees. To addres…

Cited by 0SourceScholar
2022

PoseSDF: Simultaneous 3D Human Shape Reconstruction and Gait Pose Estimation Using Signed Distance Functions

ICRA 2022poster

Vision-based 3D human pose estimation and shape reconstruction play important roles in robot-assisted healthcare monitoring and personal assistance. However, 3D data captured from a single viewpoint always encounter occlusions and exhibit substantial heterogeneity across different views, resulting i…

Cited by 7SourceScholar
2022

Tackling Long-Tailed Category Distribution under Domain Shifts

ECCV 2022poster

"Machine learning models fail to perform well on real-world applications when 1) the category distribution P(Y) of the training dataset suffers from long-tailed distribution and 2) the test data is drawn from different conditional distributions P(X|Y). Existing approaches cannot handle the scenario…

2022

csBoundary: City-Scale Road-Boundary Detection in Aerial Images for High-Definition Maps

RA-L 2022

High-Definition (HD) maps can provide precise geometric and semantic information of static traffic environments for autonomous driving. Road-boundary is one important information presented in HD maps since it distinguishes between road areas and off-road areas, which can guide vehicles to drive with

Cited by 37SourceScholar
2021

In Defense of Knowledge Distillation for Task Incremental Learning and Its Application in 3D Object Detection

RA-L 2021

Making robots learn skills incrementally is an efficient way to design real intelligent agents. To achieve this, researchers adopt knowledge distillation to transfer old-task knowledge from old models to new ones. However, when the length of the task sequence increases, the effectiveness of knowledg

Cited by 23SourceScholar
2021

Vision-Based Autonomous Car Racing Using Deep Imitative Reinforcement Learning

RA-L 2021

Autonomous car racing is a challenging task in the robotic control area. Traditional modular methods require accurate mapping, localization and planning, which makes them computationally inefficient and sensitive to environmental changes. Recently, deep-learning-based end-to-end systems have shown p

Cited by 78SourcecodeScholar
2019

Learning to Drive from Simulation without Real World Labels

ICRA 2019poster

Simulation can be a powerful tool for under-standing machine learning systems and designing methods to solve real-world problems. Training and evaluating methods purely in simulation is often “doomed to succeed” at the desired task in a simulated environment, but the resulting models are incapable o…

Cited by 146SourceScholar
2018

Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation

ICRA 2018poster

Imitation learning is an effective approach for autonomous systems to acquire control policies when an explicit reward function is unavailable, using supervision provided as demonstrations from an expert, typically a human operator. However, standard imitation learning methods assume that the agent…

Cited by 455SourceScholar
2018

Meta-Reinforcement Learning of Structured Exploration Strategies

NeurIPS 2018spotlight

Exploration is a fundamental challenge in reinforcement learning (RL). Many current exploration methods for deep RL use task-agnostic objectives, such as information gain or bonuses based on state visitation. However, many practical applications of RL involve learning more than a single task, and pr…

2018

Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings

ICML 2018oral

In this work, we take a representation learning perspective on hierarchical reinforcement learning, where the problem of learning lower layers in a hierarchy is transformed into the problem of learning trajectory-level generative models. We show that we can learn continuous latent representations of…

Cited by 193SourcePDFScholar
2017

Learning Invariant Feature Spaces to Transfer Skills with Reinforcement Learning

ICLR 2017poster

People can learn a wide range of tasks from their own experience, but can also learn from observing other creatures. This can accelerate acquisition of new skills even when the observed agent differs substantially from the learning agent in terms of morphology. In this paper, we examine how reinforc…

Cited by 369SourceScholar