← Search

Chen Liu

81 accepted papers

2026

A Computationally Efficient Nonparametric Approach for Robot Imitation Learning

ICRA 2026poster

Transferring human skills to robots through learning from demonstrations has been an important topic in the robotics community, and many models have been developed for learning and adapting such skills. Among them, nonparametric representations are an appealing choice, since nonparametric solutions …

Cited by 0Scholar
2026

AlphaBench: Benchmarking Large Language Models in Formulaic Alpha Factor Mining

ICLR 2026poster

Formulaic alpha factor mining (FAFM) is a central problem in quantitative investment, where interpretable formulas are designed to extract predictive signals from historical financial series. With the emergence of large language models (LLMs), recent studies have begun to explore their roles in FAFM…

Cited by 0SourceScholar
2026

CTR-LORA: CURVATURE-AWARE AND TRUST-REGION GUIDED LOW-RANK ADAPTATION FOR LARGE LANGUAGE MODELS

ICASSP 2026oral

Parameter-efficient fine-tuning (PEFT) has become the standard approach for adapting large language models under limited compute and memory budgets. Although previous methods improve efficiency through low-rank updates, quantization, or heuristic budget reallocation, they often decouple the allocati…

Cited by 0SourcePDFScholar
2026

Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models

ICML 2026poster

Large language models (LLMs) achieve remarkable performance through ever-increasing parameter counts, but scaling incurs steep computational costs. To better understand LLM scaling, we study representational differences between LLMs and their smaller counterparts, with the goal of replicating the re…

Cited by 0SourceScholar
2026

DualOptim+: Bridging Shared and Decoupled Optimizer States for Better Machine Unlearning in Large Language Models

ICML 2026poster

We propose **DualOptim+**, a novel optimization framework for improving machine unlearning in large language models. It introduces a base state to capture common representations shared by forgetting and retaining objectives and delta states to preserve objective-specific residuals. This architecture…

Cited by 0SourceScholar
2026

RCP-LO: A Relative Coordinate Prediction Framework for Generalizable Deep LiDAR Odometry

AAAI 2026technical

LiDAR odometry is a critical component of SLAM in autonomous driving and robotics. Learning-based methods have shown remarkable performance by regressing relative poses in an end-to-end manner. However, when applying these trained models, originally developed on the widely used KITTI dataset, to oth

Cited by 0SourcePDFScholar
2026

Tool-Grasp: A 6-DoF Functional Grasping Framework for General-Purpose Hand Tools

ICRA 2026poster

Detecting functional grasp poses for tool operation is critical for robots in complex real-world tasks, yet existing methods lack this capability. Key challenges are: 1) Scarce realworld datasets with fine-grained functional labels and task-valid grasp annotations, as their construction requires dom…

Cited by 0Scholar
2026

UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction as Reasoning

ICLR 2026poster

GUI grounding, which maps natural-language instructions to actionable UI elements, is a core capability of GUI agents. Prior work largely treats instructions as a static proxy for user intent, overlooking the impact of instruction diversity on grounding performance. Through a careful investigation o…

Cited by 0SourcecodeScholar
2026

Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images

CVPR 2026

Creating realistic and simulation-ready 3D assets is crucial for autonomous driving research and virtual environment construction. However, existing 3D vehicle generation methods are often trained on synthetic data with significant domain gaps from real-world distributions. The generated models ofte

Cited by 0SourcecodeScholar
2025

CourtReasoner: Can LLM Agents Reason Like Judges?

EMNLP 2025

LLMs are increasingly applied in the legal domain in tasks such as summarizing legal texts and providing basic legal advice. Yet, their capacity to draft full judicial analyses in U.S. court opinions is still largely uncharted, such as generating entire judicial reasoning sections in U.S. court deci

2025

Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models

NeurIPS 2025poster

Accuracy remains a standard metric for evaluating AI systems, but it offers limited insight into how models arrive at their solutions. In this work, we introduce a benchmark based on brainteasers written in long narrative form to probe more deeply into the types of reasoning strategies that models…

Cited by 0SourceScholar
2025

Data Selection Matters: Towards Robust Instruction Tuning of Large Multimodal Models

NeurIPS 2025poster

Selecting a compact subset of visual instruction–following data has emerged as an effective way to align large multimodal models with human intentions while avoiding the high cost of full-dataset training. Yet we observe that both full-data training and existing state-of-the-art data selection metho…

Cited by 1SourcecodeScholar
2025

Design2GarmentCode: Turning Design Concepts to Tangible Garments Through Program Synthesis

CVPR 2025poster

Sewing patterns, the essential blueprints for fabric cutting and tailoring, act as a crucial bridge between design concepts and producible garments. However, existing uni-modal sewing pattern generation models struggle to effectively encode complex design concepts with a multi-modal nature and corre…

2025

DiffKillR: Killing and Recreating Diffeomorphisms for Cell Annotation in Dense Microscopy Images

ICASSP 2025accepted

The proliferation of digital microscopy images, driven by advances in automated whole slide scanning, presents significant opportunities for biomedical research and clinical diagnostics. However, accurately annotating densely packed information in these images remains a major challenge. To address t…

Cited by 0SourceScholar
2025

DiffLO: Semantic-Aware LiDAR Odometry with Diffusion-Based Refinement

CVPR 2025poster

LiDAR odometry is a critical module in autonomous driving systems, responsible for accurate localization by estimating the relative pose transformation between consecutive point cloud frames. However, existing studies frequently encounter challenges with unreliable pose estimation, due to the lack o…

2025

DualOptim: Enhancing Efficacy and Stability in Machine Unlearning with Dual Optimizers

NeurIPS 2025poster

Existing machine unlearning (MU) approaches exhibit significant sensitivity to hyperparameters, requiring meticulous tuning that limits practical deployment. In this work, we first empirically demonstrate the instability and suboptimal performance of existing popular MU methods when deployed in diff…

Cited by 0SourceScholar
2025

Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics

CVPR 2025poster

Sound-guided object segmentation has drawn considerable attention for its potential to enhance multimodal perception. Previous methods primarily focus on developing advanced architectures to facilitate effective audio-visual interactions, without fully addressing the inherent challenges posed by aud…

2025

Generative Video Bi-flow

ICCV 2025poster

We propose a novel generative video model to robustly learn temporal change as a neural Ordinary Differential Equation (ODE) flow with a bilinear objective which combines two aspects: The first is to map from the past into future video frames directly. Previous work has mapped the noise to new frame…

Cited by 0SourcePDFScholar
2025

Hierarchy UGP: Hierarchy Unified Gaussian Primitive for Large-Scale Dynamic Scene Reconstruction

ICCV 2025poster

Recent advances in differentiable rendering have significantly improved dynamic street scene reconstruction. However, the complexity of large-scale scenarios and dynamic elements, such as vehicles and pedestrians, remains a substantial challenge. Existing methods often struggle to scale to large sce…

Cited by 0SourcePDFScholar
2025

Hyperedge Representations with Hypergraph Wavelets: Applications to Spatial Transcriptomics

ICASSP 2025accepted

In many data-driven applications, higher-order relationships among multiple objects are essential in capturing complex interactions. Hypergraphs, which generalize graphs by allowing edges to connect any number of nodes, provide a flexible and powerful framework for modeling such higher-order relatio…

Cited by 0SourceScholar
2025

ImageFlowNet: Forecasting Multiscale Image-Level Trajectories of Disease Progression with Irregularly-Sampled Longitudinal Medical Images

ICASSP 2025accepted

Advances in medical imaging technologies have enabled the collection of longitudinal images, which involve repeated scanning of the same patients over time, to monitor disease progression. However, predictive modeling of such data remains challenging due to high dimensionality, irregular sampling, a…

Cited by 0SourceScholar
2025

LightLoc: Learning Outdoor LiDAR Localization at Light Speed

CVPR 2025poster

Scene coordinate regression achieves impressive results in outdoor LiDAR localization but requires days of training. Since training needs to be repeated for each new scene, long training times make these impractical for applications requiring time-sensitive system upgrades, such as autonomous drivin…

2025

Multimodal Disease Progression Modeling via Spatiotemporal Disentanglement and Multiscale Alignment

NeurIPS 2025spotlight

Longitudinal multimodal data, including electronic health records (EHR) and sequential chest X-rays (CXRs), is critical for modeling disease progression, yet remains underutilized due to two key challenges: (1) redundancy in consecutive CXR sequences, where static anatomical regions dominate over cl…

Cited by 0SourceScholar
2025

Not All Frame Features Are Equal: Video-to-4D Generation via Decoupling Dynamic-Static Features

ICCV 2025poster

Recently, the generation of dynamic 3D objects from a video has shown impressive results. Existing methods directly optimize Gaussians using whole information in frames. However, when dynamic regions are interwoven with static regions within frames, particularly if the static regions account for a l…

2025

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration

CVPR 2025poster

Closed-loop simulation is crucial for end-to-end autonomous driving. Existing sensor simulation methods (e.g., NeRF and 3DGS) reconstruct driving scenes based on conditions that closely mirror training data distributions. However, these methods struggle with rendering novel trajectories, such as lan…

Cited by 11SourcePDFScholar
2025

Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment

CVPR 2025poster

Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from ambiguous audio-visual correspondences--such as nearby visually similar but acou…

Cited by 0SourcePDFScholar
2025

Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data

ICCV 2025poster

AI for tumor segmentation is limited by the lack of large, voxel-wise annotated datasets, which are hard to create and require medical experts. In our proprietary JHH dataset of 3,000 annotated pancreatic tumor scans, we found that AI performance stopped improving after 1,500 scans. With synthetic d…

2025

TS-Net: Assembling Task-specific Features from Multiple Feature Levels for Multi-task Learning

ICASSP 2025accepted

Multi-task learning (MTL) has become an attractive topic that leverages shared knowledge to improve performance and enhance generalization. However, most existing works neglect the varying contribution of multi-level features to sub-task representations. In this paper, we explore the impact of multi…

Cited by 0SourceScholar
2025

Towards Reliable and Holistic Visual In-Context Learning Prompt Selection

NeurIPS 2025poster

Visual In-Context Learning (VICL) has emerged as a prominent approach for adapting visual foundation models to novel tasks, by effectively exploiting contextual information embedded in in-context examples, which can be formulated as a global ranking problem of potential candidates. Current VICL meth…

Cited by 0SourceScholar
2025

Understanding and Improving Fast Adversarial Training against $l_0$ Bounded Perturbations

NeurIPS 2025poster

This work studies fast adversarial training against sparse adversarial perturbations bounded by $l_0$ norm. We first demonstrate the unique challenges of employing $1$-step attacks on $l_0$ bounded perturbations, especially catastrophic overfitting (CO) that cannnot be properly addressed by existing…

Cited by 0SourceScholar
2024

Addressing Asynchronicity in Clinical Multimodal Fusion via Individualized Chest X-ray Generation

NeurIPS 2024poster

Integrating multi-modal clinical data, such as electronic health records (EHR) and chest X-ray images (CXR), is particularly beneficial for clinical prediction tasks. However, in a temporal setting, multi-modal data are often inherently asynchronous. EHR can be continuously collected but CXR is gene…

2024

Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings

NAACL 2024long

Large language models (LLMs) are highly adept at question answering and reasoning tasks, but when reasoning in a situational context, human expectations vary depending on the relevant cultural common ground. As languages are associated with diverse cultures, LLMs should also be culturally-diverse re…

2024

Benchmarking Audio Visual Segmentation for Long-Untrimmed Videos

CVPR 2024poster

Existing audio-visual segmentation datasets typically focus on short-trimmed videos with only one pixel-map annotation for a per-second video clip. In contrast for untrimmed videos the sound duration start- and end-sounding time positions and visual deformation of audible objects vary significantly.…

2024

CPT-VR: Improving Surface Rendering via Closest Point Transform with View-Reflection Appearance

ECCV 2024poster

"Differentiable surface rendering has significantly advanced 3D reconstruction. Existing surface rendering methods assume that the local surface is planar, and thus employ linear approximation based on the Singed Distance Field (SDF) values to predict the point on the surface. However, this assumpti…

Cited by 0SourcePDFScholar
2024

Concentrate Attention: Towards Domain-Generalizable Prompt Optimization for Language Models

NeurIPS 2024poster

Recent advances in prompt optimization have notably enhanced the performance of pre-trained language models (PLMs) on downstream tasks. However, the potential of optimized prompts on domain generalization has been under-explored. To explore the nature of prompt generalization on unknown domains, we…

2024

FUN with Fisher: Improving Generalization of Adapter-Based Cross-lingual Transfer with Scheduled Unfreezing

NAACL 2024long

Standard fine-tuning of language models typically performs well on in-distribution data, but suffers with generalization to distribution shifts. In this work, we aim to improve the generalization of adapter-based cross-lingual task transfer where such cross-language distribution shifts are imminent.…

2024

Large Language Model Guided Knowledge Distillation for Time Series Anomaly Detection

IJCAI 2024poster

Self-supervised methods have gained prominence in time series anomaly detection due to the scarcity of available annotations. Nevertheless, they typically demand extensive training data to acquire a generalizable representation map, which conflicts with scenarios of a few available samples, thereby…

Cited by 21SourcePDFScholar
2024

Mixture of Adversarial LoRAs: Boosting Robust Generalization in Meta-Tuning

NeurIPS 2024poster

This paper introduces AMT, an \textbf{A}dversarial \textbf{M}eta-\textbf{T}uning methodology, to boost the robust generalization of pre-trained models in the out-of-domain (OOD) few-shot learning. To address the challenge of transferring knowledge from source domains to unseen target domains, we con…

2024

Refinement Bird's Eye View Feature for 3D Lane Detection with Dual-Branch View Transformation Module

ICASSP 2024accepted

Detecting 3D lane lines from images is a fundamental challenge and an ill-posed problem in autonomous driving. Existing methods are limited by scene robustness and computational efficiency. This paper introduces an innovative 3D lane detection method that addresses the challenges of lane detection i…

Cited by 0SourceScholar
2024

StablePT : Towards Stable Prompting for Few-shot Learning via Input Separation

EMNLP 2024finding

Large language models have shown their ability to become effective few-shot learners with prompting, revoluting the paradigm of learning with data scarcity. However, this approach largely depends on the quality of prompt initialization and always exhibits large variability among different runs. Such…

2024

Towards Efficient Training and Evaluation of Robust Models against $l_0$ Bounded Adversarial Perturbations

ICML 2024poster

This work studies sparse adversarial perturbations bounded by $l_0$ norm. We propose a white-box PGD-like attack method named sparse-PGD to effectively and efficiently generate such perturbations. Furthermore, we combine sparse-PGD with a black-box attack to comprehensively and more reliably evaluat…

2024

Towards Global Optimal Visual In-Context Learning Prompt Selection

NeurIPS 2024poster

Visual In-Context Learning (VICL) is a prevailing way to transfer visual foundation models to new tasks by leveraging contextual information contained in in-context examples to enhance learning and prediction of query sample. The fundamental problem in VICL is how to select the best prompt to activa…

Cited by 4SourcePDFScholar
2024

Treemil: A Multi-Instance Learning Framework for Time Series Anomaly Detection with Inexact Supervision

ICASSP 2024accepted

Time series anomaly detection (TSAD) plays a vital role in various domains such as healthcare, networks and industry. Considering labels are crucial for detection but difficult to obtain, we turn to TSAD with inexact supervision: only series-level labels are provided during the training phase, while…

Cited by 0SourceScholar
2024

USCILab3D: A Large-scale, Long-term, Semantically Annotated Outdoor Dataset

NeurIPS 2024poster

In this paper, we introduce the \textbf{USCILab3D dataset}, a large-scale, annotated outdoor dataset designed for versatile applications across multiple domains, including computer vision, robotics, and machine learning. The dataset was acquired using a mobile robot equipped with 5 cameras and a 32-…

Cited by 0SourcePDFScholar
2023

Adaptive Patchwork: Real-Time Ground Segmentation for 3D Point Cloud With Adaptive Partitioning and Spatial-Temporal Context

RA-L 2023

Ground segmentation is a fundamental task in the field of 3D perception using 3D LiDAR sensors. Several ground segmentation methods have been proposed, but they often suffer from mis-segmentation, especially missed detection, due to poor noise removal, unreasonable ground pre-blocking and lack of re

Cited by 8SourceScholar
2023

Behavior Prior Representation learning for Offline Reinforcement Learning

ICLR 2023poster

Offline reinforcement learning (RL) struggles in environments with rich and noisy inputs, where the agent only has access to a fixed dataset without environment interactions. Past works have proposed common workarounds based on the pre-training of state representations, followed by policy training.…

2023

Diverse 3D Hand Gesture Prediction From Body Dynamics by Bilateral Hand Disentanglement

CVPR 2023poster

Predicting natural and diverse 3D hand gestures from the upper body dynamics is a practical yet challenging task in virtual avatar creation. Previous works usually overlook the asymmetric motions between two hands and generate two hands in a holistic manner, leading to unnatural results. In this wor…

2023

Towards Stable and Efficient Adversarial Training against $l_1$ Bounded Adversarial Attacks

ICML 2023poster

We address the problem of stably and efficiently training a deep neural network robust to adversarial perturbations bounded by an $l_1$ norm. We demonstrate that achieving robustness against $l_1$-bounded perturbations is more challenging than in the $l_2$ or $l_\infty$ cases, because adversarial tr…

2022

Exploring Label Hierarchy in a Generative Way for Hierarchical Text Classification

COLING 2022main

Hierarchical Text Classification (HTC), which aims to predict text labels organized in hierarchical space, is a significant task lacking in investigation in natural language processing. Existing methods usually encode the entire hierarchical structure and fail to construct a robust label-dependent m…

2022

FigMemes: A Dataset for Figurative Language Identification in Politically-Opinionated Memes

EMNLP 2022main

Real-world politically-opinionated memes often rely on figurative language to cloak propaganda and radical ideas to help them spread. It is not only a scientific challenge to develop machine learning models to recognize them in memes, but also sociologically beneficial to understand hidden meanings…

Cited by 22SourcePDFScholar
2022

Robust Binary Models by Pruning Randomly-initialized Networks

NeurIPS 2022accept

Robustness to adversarial attacks was shown to require a larger model capacity, and thus a larger memory footprint. In this paper, we introduce an approach to obtain robust yet compact models by pruning randomly-initialized binary networks. Unlike adversarial training, which learns the model paramet…

2021

A Probabilistic Model for Segmentation of Ambiguous 3D Lung Nodule

ICASSP 2021accepted

Many medical images domains suffer from inherent ambiguities. A feasible approach to resolve the ambiguity of lung nodule in the segmentation task is to learn a distribution over segmentations based on a given 2D lung nodule image. Whereas lung nodule with 3D structure contains dense 3D spatial info…

Cited by 0SourceScholar
2021

An Architecture for Accelerated Large-Scale Inference of Transformer-Based Language Models

NAACL 2021industry

This work demonstrates the development process of a machine learning architecture for inference that can scale to a large volume of requests. We used a BERT model that was fine-tuned for emotion analysis, returning a probability distribution of emotions given a paragraph. The model was deployed as a…

2021

DFDM: A Deep Feature Decoupling Module for Lung Nodule Segmentation

ICASSP 2021accepted

In this paper, we propose a novel feature decoupling method to tackle two critical problems in the lung nodule segmentation task: (i) ambiguity of nodule boundary leads to the imprecise segmentation boundary and (ii) the high false positive rate of segmentation result. Our motivation is that an accu…

Cited by 0SourceScholar
2021

Deepnodule: Multi-Task Learning of Segmentation Bootstrap for Pulmonary Nodule Detection

ICASSP 2021accepted

Pulmonary nodule detection and segmentation are the necessary successively steps in lung cancer screening with low-dose computed tomography (CT) scans. However, the state-of-the-art models focus on solving tasks separately, thereby ignore the correlation between each task. Besides, most nodule detec…

Cited by 0SourceScholar
2021

FLiText: A Faster and Lighter Semi-Supervised Text Classification with Convolution Networks

EMNLP 2021main

In natural language processing (NLP), state-of-the-art (SOTA) semi-supervised learning (SSL) frameworks have shown great performance on deep pre-trained language models such as BERT, and are expected to significantly reduce the demand for manual labeling. However, our empirical studies indicate that…

2021

Learning Dynamic Alignment via Meta-Filter for Few-Shot Learning

CVPR 2021poster

Few-shot learning (FSL), which aims to recognise new classes by adapting the learned knowledge with extremely limited few-shot (support) examples, remains an important open problem in computer vision. Most of the existing methods for feature alignment in few-shot learning only consider image-level o…

Cited by 150PDFScholar
2021

Learning a Few-shot Embedding Model with Contrastive Learning

AAAI 2021technical

Few-shot learning (FSL) aims to recognize target classes by adapting the prior knowledge learned from source classes. Such knowledge usually resides in a deep embedding model for a general matching purpose of the support and query image pairs. The objective of this paper is to repurpose the contrast…

Cited by 217SourcePDFScholar
2021

MG-DVD: A Real-time Framework for Malware Variant Detection Based on Dynamic Heterogeneous Graph Learning

IJCAI 2021poster

Detecting the newly emerging malware variants in real time is crucial for mitigating cyber risks and proactively blocking intrusions. In this paper, we propose MG-DVD, a novel detection framework based on dynamic heterogeneous graph learning, to detect malware variants in real time. Particularly…

2020

DessiLBI: Exploring Structural Sparsity of Deep Networks via Differential Inclusion Paths

ICML 2020poster

Over-parameterization is ubiquitous nowadays in training neural networks to benefit both optimization in seeking global optima and generalization in reducing prediction error. However, compressive networks are desired in many real world applications and direct training of small networks may be trapp…

2020

On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them

NeurIPS 2020poster

We analyze the influence of adversarial training on the loss landscape of machine learning models. To this end, we first provide analytical studies of the properties of adversarial loss functions under different adversarial budgets. We then demonstrate that the adversarial loss landscape is less fav…

2020

π-Map: A Decision-Based Sensor Fusion with Global Optimization for Indoor Mapping

IROS 2020poster

In this paper, we propose π-map, a tightly coupled fusion mechanism that dynamically consumes LiDAR and sonar data to generate reliable and scalable indoor maps for autonomous robot navigation. The key novelty of π-map over previous attempts is the utilization of a fusion mechanism that works in thr…

Cited by 3SourceScholar
2019

Floor-SP: Inverse CAD for Floorplans by Sequential Room-Wise Shortest Path

ICCV 2019poster

This paper proposes a new approach for automated floorplan reconstruction from RGBD scans, a major milestone in indoor mapping research. The approach, dubbed Floor-SP, formulates a novel optimization problem, where room-wise coordinate descent sequentially solves shortest path problems to optimize t…

Cited by 115PDFcodeScholar
2019

PlaneRCNN: 3D Plane Detection and Reconstruction From a Single Image

CVPR 2019oral

This paper proposes a deep neural architecture, PlaneRCNN, that detects and reconstructs piecewise planar regions from a single RGB image. PlaneRCNN employs a variant of Mask R-CNN to detect planes with their plane parameters and segmentation masks. PlaneRCNN then refines an arbitrary number of segm…

Cited by 275PDFScholar
2018

FloorNet: A Unified Framework for Floorplan Reconstruction from 3D Scans

ECCV 2018poster

The ultimate goal of this indoor mapping research is to automatically reconstruct a floorplan simply by walking through a house with a smartphone in a pocket. This paper tackles this problem by proposing FloorNet, a novel deep neural architecture. The challenge lies in the processing of RGBD streams…

2018

PlaneNet: Piece-Wise Planar Reconstruction From a Single RGB Image

CVPR 2018poster

This paper proposes a deep neural network (DNN) for piece-wise planar depthmap reconstruction from a single RGB image. While DNNs have brought remarkable progress to single-image pixel-wise depth prediction, piece-wise planar depthmap reconstruction requires a structured geometry representation, an…