← Search

Ying Liu

39 accepted papers

2026

AnyBand-Diff: A Unified Remote Sensing Image Generation and Band Repair Framework with Spectral Priors

ICML 2026poster

Existing diffusion models have made significant progress in generating realistic images. However, their direct adaptation to remote sensing imagery often disregards intrinsic physical laws. This oversight frequently leads to spectral distortion and radiometric inconsistency, severely limiting the sc…

Cited by 0SourceScholar
2026

PointSLAM++: Robust Dense Neural Gaussian Point Cloud-based SLAM

AAAI 2026technical

Real-time 3D reconstruction is crucial for robotics and augmented reality, yet current simultaneous localization and mapping(SLAM) approaches often struggle to maintain structural consistency and robust pose estimation in the presence of depth noise. This work introduces PointSLAM++, a novel RGB-D S

Cited by 0SourcePDFScholar
2026

Scheduling Adaptive Imitation Learning for Long-Horizon Dexterous Robot Micromanipulation of Deformable Cell

RA-L 2026

Robots performing collaborative, long-horizon dexterity cell micromanipulation tasks are challenging and practically significant, such as stripping intact cell membranes, which is considered as one of the most technically demanding procedures. Imitation learning approach is expected to address the c

Cited by 2SourceScholar
2026

Sparse-Scale Transformer with Bidirectional Awareness for Time Series Forecasting

AAAI 2026technical

Time series forecasting (TSF) plays a crucial role in many real-world applications, such as weather prediction and economic planning. While Transformer-based models have shown strong capabilities in modeling long-range dependencies, effectively capturing the multi-scale temporal dynamics inherent in

Cited by 0SourcePDFScholar
2026

VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs

ICLR 2026poster

Large Multimodal Models have achieved remarkable progress in integrating vision and language, enabling strong performance across perception, reasoning, and domain-specific tasks. However, their capacity to reason over multiple, visually similar inputs remains insufficiently explored. Such fine-grain…

Cited by 0SourcecodeScholar
2025

A Top-down Graph-based Tool for Modeling Classical Semantic Maps: A Case Study of Supplementary Adverbs

NAACL 2025long

Semantic map models (SMMs) construct a network-like conceptual space from cross-linguistic instances or forms, based on the connectivity hypothesis. This approach has been widely used to represent similarity and entailment relationships in cross-linguistic concept comparisons. However, most SMMs are…

2025

Complete Coverage Path Planning Algorithm Based on Improved Biologically Inspired Neural Networks in Spray Painting

RA-L 2025

Intelligent putty coating technology is the main way to improve the degree of automation of railroad vehicle painting workshops. The two-component putty, which is currently used in the vehicle coating system, has extremely low fluidity, which requires improving the full coverage of the spray path wh

Cited by 3SourceScholar
2025

Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark

ICRA 2025

Open-vocabulary detectors are proposed to locate and recognize objects in novel classes. However, variations in vision-aware language vocabulary data used for open-vocabulary learning can lead to unfair and unreliable evaluations. Recent evaluation methods have attempted to address this issue by inc

Cited by 2SourcecodeScholar
2025

Iterative Self-Training with Class-Aware Text-to-Image Synthesis for Visual Task Learning

AAAI 2025technical

Generative models are widely used to produce synthetic images with annotations, alleviating the burden of image collection and annotation for training deep visual models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often…

Cited by 0SourcePDFScholar
2025

TD-GS: Few-shot Object View Synthesis via Task-Disentangled 3D Gaussian Splatting

ICASSP 2025accepted

3D Gaussian Splatting (3D-GS) has exhibited impressive progress in novel view synthesis. When given the sparse views, its performance degrades severely, causing many problems like novel views collapse and excessive floaters. Many recent methods take into account fitting input views, inferring missin…

Cited by 0SourceScholar
2024

A Multiscale Objective Function for Camera Color Correction

ICASSP 2024accepted

Color correction (CC) plays a pivotal role in camera imaging. Existing approaches usually conduct CC tuning by minimizing ∆E (e.g. ∆E2000), a standard metric proposed by CIE for representing color differences in LAB space. However, we observe that not all the colors with identical ∆E error to the ta…

Cited by 0SourceScholar
2024

Approaches and Challenges for Resolving Different Representations of Fictional Characters for Chinese Novels

COLING 2024main

Due to the huge scale of literary works, automatic text analysis technologies are urgently needed for literary studies such as Digital Humanities. However, the domain-generality of existing NLP technologies limits their effectiveness on in-depth literary studies. It is valuable to explore how to ada…

2024

CDUMA: An Adaptive Approach for Mitigating Confounder for MCQA

ICASSP 2024accepted

Multiple-choice question answering (MCQA) requires the model to select the correct answer from a set of candidate options when given a passage and a question. Previous research has achieved promising results with the assistance of Pre-trained Language Models(PrLMs). However, it has been observed tha…

Cited by 0SourceScholar
2024

CausalME: Balancing bi-modalities in Visual Question Answering

ICASSP 2024accepted

Mitigating linguistic bias and attaining modal equilibrium in Visual Question Answering (VQA) tasks constitute a pivotal concern. Previous work has mainly focused on data augmentation or a uni-modal approach, which is insufficient to fully utilize bi-modal information. In this work, we propose a new…

Cited by 0SourceScholar
2024

Clear Up Confusion: Advancing Cross-Domain Few-Shot Relation Extraction through Relation-Aware Prompt Learning

NAACL 2024short

Cross-domain few-shot Relation Extraction (RE) aims to transfer knowledge from a source domain to a different target domain to address low-resource problems.Previous work utilized label descriptions and entity information to leverage the knowledge of the source domain.However, these models are prone…

Cited by 0SourcePDFScholar
2024

Evaluating Moral Beliefs across LLMs through a Pluralistic Framework

EMNLP 2024finding

Proper moral beliefs are fundamental for language models, yet assessing these beliefs poses a significant challenge. This study introduces a novel three-module framework to evaluate the moral beliefs of four prominent large language models. Initially, we constructed a dataset containing 472 moral ch…

2024

Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics

ACL 2024findings

Large language models have achieved remarkable success in general language understanding tasks. However, as a family of generative methods with the objective of next token prediction, the semantic evolution with the depth of these models are not fully explored, unlike their predecessors, such as BER…

2024

Fusion Makes Perfection: An Efficient Multi-Grained Matching Approach for Zero-Shot Relation Extraction

NAACL 2024short

Predicting unseen relations that cannot be observed during the training phase is a challenging task in relation extraction. Previous works have made progress by matching the semantics between input instances and label descriptions. However, fine-grained matching often requires laborious manual annot…

2024

Quite Good, but Not Enough: Nationality Bias in Large Language Models - a Case Study of ChatGPT

COLING 2024main

While nationality is a pivotal demographic element that enhances the performance of language models, it has received far less scrutiny regarding inherent biases. This study investigates nationality bias in ChatGPT (GPT-3.5), a large language model (LLM) designed for text generation. The research cov…

2024

Small Object Detection on the Water Surface Based on Radar and Camera Fusion

ICASSP 2024accepted

With the growing applications of water operations, water surface object detection tasks are facing new challenges. In this paper, we focus on improving the performance of water surface small object detection. Due to the limitations of single-sensor in water environments, we propose RCFNet, a novel s…

Cited by 0SourceScholar
2023

Always the Best Fit: Adaptive Domain Gap Filling from Causal Perspective for Few-Shot Relation Extraction

EMNLP 2023short findings

Cross-domain Relation Extraction aims to transfer knowledge from a source domain to a different target domain to address low-resource challenges. However, the semantic gap caused by data bias between domains is a major challenge, especially in few-shot scenarios. Previous work has mainly focused on…

Cited by 0SourceScholar
2023

Ambiguity Meets Uncertainty: Investigating Uncertainty Estimation for Word Sense Disambiguation

ACL 2023findings

Word sense disambiguation (WSD), which aims to determine an appropriate sense for a target word given its context, is crucial for natural language understanding. Existing supervised methods treat WSD as a classification task and have achieved remarkable performance. However, they ignore uncertainty…

2023

Covariate-informed Representation Learning to Prevent Posterior Collapse of iVAE

AISTATS 2023poster

The recently proposed identifiable variational autoencoder (iVAE) framework provides a promising approach for learning latent independent components (ICs). iVAEs use auxiliary covariates to build an identifiable generation structure from covariates to ICs to observations, and the posterior network a…

2023

Granularity Matters: Pathological Graph-driven Cross-modal Alignment for Brain CT Report Generation

EMNLP 2023long main

The automatic Brain CT reports generation can improve the efficiency and accuracy of diagnosing cranial diseases. However, current methods are limited by 1) coarse-grained supervision: the training data in image-text format lacks detailed supervision for recognizing subtle abnormalities, and 2) coup…

Cited by 0SourceScholar
2022

Cross-modal Contrastive Attention Model for Medical Report Generation

COLING 2022main

Medical report automatic generation has gained increasing interest recently as a way to help radiologists write reports more efficiently. However, this image-to-text task is rather challenging due to the typical data biases: 1) Normal physiological structures dominate the images, with only tiny abno…

2022

Domain Robust Deep Embedding Learning for Speaker Recognition

ICASSP 2022accepted

This paper presents a domain robust deep embedding learning method for speaker verification (SV) tasks. Most recent methods utilize deep neural networks (DNN) to learn compact and discriminative speaker embeddings from large-scale labeled datasets such as VoxCeleb and the NIST SRE corpus. Despite th…

Cited by 0SourceScholar
2021

An Effective Deep Embedding Learning Method Based on Dense-Residual Networks for Speaker Verification

ICASSP 2021accepted

In this paper, we present an effective end-to-end deep embedding learning method based on Dense-Residual networks, which combine the advantages of a densely connected convolutional network (DenseNet) and a residual network (ResNet), for speaker verification (SV). Unlike a model ensemble strategy whi…

Cited by 0SourceScholar
2021

K-PLUG: Knowledge-injected Pre-trained Language Model for Natural Language Understanding and Generation in E-Commerce

EMNLP 2021finding

Existing pre-trained language models (PLMs) have demonstrated the effectiveness of self-supervised learning for a broad range of natural language processing (NLP) tasks. However, most of them are not explicitly aware of domain-specific knowledge, which is essential for downstream tasks in many domai…

2021

Recognition of Dynamic Hand Gesture Based on Mm-Wave Fmcw Radar Micro-Doppler Signatures

ICASSP 2021accepted

Radar-based sensors provide an attractive choice for hand gesture recognition (HGR). The very challenging problems in radar-based HGR are radar echo data preprocessing and recognition accuracy. In this paper, we propose a convolutional neural network (CNN) for dynamic HGR based on a millimeter-wave…

Cited by 0SourceScholar
2019

A Rotation Invariant HOG Descriptor for Tire Pattern Image Classification

ICASSP 2019accepted

Texture feature is important in describing tire pattern image which provides useful clue in solving crime cases and traffic accidents. In this paper, we propose a novel texture feature extraction method based on HOG (Histogram of Oriented Gradient) and dominant gradient (DG) in tire pattern images,…

Cited by 0SourceScholar
2016

Joint-view Kalman-filter recovery of compressed-sensed multiview videos

ICASSP 2016accepted

We develop a novel joint-view Kalman filter for causal reconstruction of compressed-sensed multiview videos. Compressed-sensed multiview video frames are initially reconstructed individually via ℓ <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</inf>…

Cited by 0SourceScholar
2016

Stochastic load scheduling for risk-limiting economic dispatch in smart microgrids

ICASSP 2016accepted

In this work we present a novel scheme for load management in microgrids based on stochastic scheduling of loads under risk-limiting constraints. When trying to enforce adequate power supply in a microgrid, the volatility of renewable resources such as wind energy has to be considered. In the risk o…

Cited by 0SourceScholar
2015

Disparity-compensated total-variation minimization for compressed-sensed multiview image reconstruction

ICASSP 2015accepted

Compressed sensing (CS) is the theory and practice of sub-Nyquist sampling of sparse signals of interest. Perfect reconstruction may then be possible with much fewer than the Nyquist required number of data. In this paper, we consider a distributed multi-view imaging system where each camera at a di…

Cited by 0SourceScholar