← Search

Ying Li

88 accepted papers

2026

4DPC$^2$hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping

ICML 2026poster

Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models (MLLMs). However, existing methods primarily focus on static objects, while understanding dynamic point cloud sequences remains largely unexplored. This…

Cited by 0SourceScholar
2026

Dynamic Momentum Recalibration in Online Gradient Learning

CVPR 2026

Stochastic Gradient Descent (SGD) and its momentum variants form the backbone of deep learning optimization, yet the underlying dynamics of their gradient behavior remain insufficiently understood. In this work, we reinterpret gradient updates through the lens of signal processing and reveal that fi

Cited by 0SourcecodeScholar
2026

EMKG: Embodied Memory Knowledge Graphs for Object-Goal Navigation in Dynamic Open Worlds

RA-L 2026

Object-Goal Navigation (OGN) in complex domestic environments remains challenging due to spatial memory and semantic uncertainties. To address this, we introduce EMKG, an embodied multimodal memory knowledge graph framework that enables open-world navigation. In contrast to conventional vision-langu

Cited by 0SourceScholar
2026

From Manuals to Actions: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation

CVPR 2026

Vision-Language-Action (VLA) models have recently emerged, demonstrating strong generalization in robotic scene understanding and manipulation. However, when confronted with long-horizon tasks that require defined goal states, such as LEGO assembly or object rearrangement, existing VLA models still

Cited by 0SourceScholar
2026

Hierarchical Control for Real-Time 3D Manipulation of Magnetic Bead Using a Single Permanent Magnet

RA-L 2026

Permanent magnet (PM) actuated micro robotics offers significant advantages for minimally invasive medicine, but faces three critical challenges: nonlinear magnetic force relationships, directional control asymmetry between horizontal and vertical motion, and imaging–capturing frequency mismatch. Th

Cited by 0SourceScholar
2026

Learning Whom to Align With: Progressive Anomaly Combination Detection for Partially View-Aligned Clustering

AAAI 2026technical

Partially View-aligned Clustering (PVC) addresses the challenge of partial view alignment in multi-view learning by leveraging complementary and consistent information. While existing PVC methods show promise, most rely on distance-based strategies that are sensitive to view-specific details and noi

Cited by 0SourcePDFScholar
2026

ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory

AAAI 2026technical

Data scarcity continues to be a critical bottleneck in the field of robotic manipulation, limiting the ability to train robust and generalizable models. While diffusion models provide a promising approach to synthesizing realistic robotic manipulation videos, their effectiveness hinges on the availa

Cited by 0SourcePDFScholar
2026

ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance

ICASSP 2026poster

While recent advancements in robotic manipulation video synthesis have shown promise, significant challenges persist in ensuring effective instruction-following and achieving high visual quality. Recent methods, like RoboDreamer, utilize linguistic decomposition to divide instructions into separate…

Cited by 0SourcePDFScholar
2026

MemClaw-RAG: Memory-Driven Navigation and Adaptive Locomotion for Wheeled-Legged Robots in Dynamic Environments

ICRA 2026poster

Object-Goal Navigation in dynamic environments remains challenging because many existing approaches rely primarily on reactive mapping and lack the ability to retain historical experience or establish structured memory associations. To address this limitation, we introduce MemClaw-RAG, an embodied m…

Cited by 0Scholar
2026

Pansharpening for Thin-Cloud Contaminated Remote Sensing Images: A Unified Framework and Benchmark Dataset

AAAI 2026technical

Pansharpening under thin cloudy conditions is a practically significant yet rarely addressed task, challenged by simultaneous spatial resolution degradation and cloud-induced spectral distortions. Existing methods often address cloud removal and pansharpening sequentially, leading to cumulative erro

Cited by 0SourcePDFScholar
2026

ParticleGS: Learning Neural Gaussian Particle Dynamics from Videos for Prior-free Physical Motion Extrapolation

CVPR 2026

The ability to extrapolate dynamic 3D scenes beyond the observed timeframe is fundamental to advancing physical world understanding and predictive modeling. Existing dynamic 3D reconstruction methods have achieved high-fidelity rendering of temporal interpolation, but typically lack physical consist

Cited by 0SourceScholar
2026

Prism-MoE: Efficient Dense-to-MoE Conversion for Visual Autoregressive Generation

ICML 2026poster

Scaling up visual autoregressive models improves generation quality but incurs substantial inference costs. Mixture-of-Experts (MoE) architectures mitigate this issue through sparse activation and have proven effective in large language models. However, training MoE models from scratch remains prohi…

Cited by 0SourceScholar
2026

ReFocusEraser: Refocusing for Small Object Removal with Robust Context-Shadow Repair

ICLR 2026poster

Existing diffusion-based object removal and inpainting methods often fail to recover the fine structural and textural details of small objects. This is primarily due to the VAE encoder’s downsampling, which inevitably compresses small masked regions and causes significant detail loss, while the deco…

Cited by 0SourcecodeScholar
2026

Rethinking Low-Confidence Pseudo Labels: Influence-Aware Semi-Supervised Fine-Tuning for Hyperspectral Change Detection

ICML 2026poster

Hyperspectral image change detection (HSI-CD) suffers from severe annotation scarcity and complex change patterns, which fundamentally limit the effectiveness of directly fine-tuning pre-trained foundation models. Although semi-supervised learning provides a promising direction, existing approaches …

Cited by 0SourceScholar
2026

Revisiting Nonstationary Kernel Design for Multi-Output Gaussian Processes

ICLR 2026poster

Multi-output Gaussian processes (MOGPs) provide a Bayesian framework for modeling non-linear functions with multiple outputs, in which nonstationary kernels are essential for capturing input-dependent variations in observations. However, from a spectral (dual) perspective, existing nonstationary ker…

Cited by 0SourceScholar
2026

UVLM: Benchmarking Video Language Model for Underwater World Understanding

AAAI 2026technical

Recently, video-language models (VidLMs) have gained widespread attention and adoption. However, existing works primarily focus on terrestrial scenarios, overlooking the highly demanding application needs of underwater observation. To overcome this gap, we introduce UVLM, an under water observation

Cited by 0SourcePDFScholar
2026

VeriRole: Verifiable Role-Awareness through Hint-Guided Reinforcement Learning

ICLR 2026poster

Maintaining role-awareness in Role-Playing Conversational Agents (RPCAs) is a significant challenging, largely because the creative nature of role-playing makes it difficult to design verifiable reward signals for reinforcement learning (RL). To address this, we propose VeriRole, a new framework des…

Cited by 0SourcecodeScholar
2025

Anchoring-Guidance Fine-Tuning (AnGFT): Elevating Professional Response Quality in Role-Playing Conversational Agents

EMNLP 2025

Large Language Models (LLMs) have demonstrated significant advancements in various fields, notably in Role-Playing Conversational Agents (RPCAs). However, when confronted with role-specific professional inquiries, LLMs-based RPCAs tend to underperform due to their excessive emphasis on the conversat

2025

Appearance- and Orientation-aware Fine-grained Rotated Ship Detection in High-Resolution Satellite Imagery

ICASSP 2025accepted

Ship detection using remote sensing imagery is a crucial research area with both military and civilian applications. However, it remains challenging due to limitations in current ship datasets, such as insufficient volume, incomplete annotations, and inaccuracies. Additionally, ships often exhibit a…

Cited by 0SourceScholar
2025

Contrastive Auxiliary Learning with Structure Transformation for Heterogeneous Graphs

AAAI 2025technical

In recent years, methods based on heterogeneous graph neural networks (HGNNs) have been widely used for embedding heterogeneous graphs (HGs) due to their ability to effectively encode the rich information from HGs into low-dimensional node embeddings. Existing HGNNs focus on neighbor aggregation and…

2025

Device-aware Optical Adversarial Attack for a Portable Projector-camera System

ICASSP 2025accepted

Deep-learning-based face recognition (FR) systems are susceptible to adversarial examples in both digital and physical domains. Physical attacks present a greater threat to deployed systems as adversaries can easily access the input channel, allowing them to provide malicious inputs to impersonate a…

Cited by 0SourceScholar
2025

Dynamic Syntactic Feature Filtering and Injecting Networks for Cross-lingual Dependency Parsing

AAAI 2025technical

Pre-trained language models enhanced parsers have achieved outstanding performance in rich-resource languages. Cross-lingual dependency parsing aims to learn useful knowledge from high-resource languages to alleviate data scarcity in low-resource languages. However, effectively reducing the syntacti…

2025

EagerLog: Active Learning Enhanced Retrieval Augmented Generation for Log-based Anomaly Detection

ICASSP 2025accepted

Logs record essential information about system operations and serve as a critical source for anomaly detection, which has generated growing research interest. Utilizing large language models (LLMs) within a retrieval-augmented generation (RAG) framework for log-based anomaly detection is an effectiv…

Cited by 0SourceScholar
2025

FinGEAR: Financial Mapping-Guided Enhanced Answer Retrieval

EMNLP 2025

Financial disclosures such as 10-K filings pose challenging retrieval problems because of their length, regulatory section hierarchy, and domain-specific language, which standard retrieval-augmented generation (RAG) models underuse. We present Financial Mapping-Guided Enhanced Answer Retrieval, a re

2025

FreqExit: Enabling Early-Exit Inference for Visual Autoregressive Models via Frequency-Aware Guidance

NeurIPS 2025poster

Visual AutoRegressive (VAR) modeling employs a next-scale decoding paradigm that progresses from coarse structures to fine details. While enhancing fidelity and scalability, this approach challenges two fundamental assumptions of conventional dynamic inference: semantic stability (intermediate outpu…

Cited by 0SourceScholar
2025

FreqMoE: Dynamic Frequency Enhancement for Neural PDE Solvers

IJCAI 2025

Fourier Neural Operators (FNO) have emerged as promising solutions for efficiently solving partial differential equations (PDEs) by learning infinite-dimensional function mappings through frequency domain transformations. However, the sparsity of high-frequency signals limits computational efficienc

Cited by 0SourcePDFScholar
2025

Harnessing the Power of Vibration Motors to Develop Miniature Untethered Robotic Fishes

RA-L 2025

Miniature underwater robots play a crucial role in the exploration and development of marine resources, particularly in confined spaces and high-pressure deep-sea environments. This study presents the design, optimization, and performance of a miniature robotic fish, powered by the oscillation of bi

Cited by 2SourceScholar
2025

HyperDiff: Masked Diffusion Model with High-efficient Transformer for Hyperspectral Image Cross-Scene Classification

ICASSP 2025accepted

Hyperspectral Image (HSI) cross-scene classification is a challenging task in remote sensing, particularly when real-time processing of Target Domain (TD) HSI is required, and data cannot be reused for training. While deep learning methods have shown promising results, the generalization ability of…

Cited by 0SourceScholar
2025

LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation

ICLR 2025poster

Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing visual prompting techniques often pad the prompt parameters around the image, limiting the interaction between the visual p…

2025

MS-UFAD: A Large-Scale Dataset for Real-world Unified Face Attack Detection with Text Descriptions

ICASSP 2025accepted

As deepfake and adversarial attacks evolve, facial recognition systems are encountering increasingly diverse threats. Most existing face liveness detection algorithms focus on single tasks, like spoofing or deepfake attack detection. The corresponding datasets have limited coverage of attack methods…

Cited by 0SourceScholar
2025

Memory-enhanced Large Language Model for Cross-lingual Dependency Parsing via Deep Hierarchical Syntax Understanding

EMNLP 2025

Large language models (LLMs) demonstrate remarkable text generation and syntax parsing capabilities in high-resource languages. However, their performance notably declines in low-resource languages due to memory forgetting stemming from semantic interference across languages. To address this issue,

2025

Multi-View Oriented GPLVM: Expressiveness and Efficiency

NeurIPS 2025poster

The multi-view Gaussian process latent variable model (MV-GPLVM) aims to learn a unified representation from multi-view data but is hindered by challenges such as limited kernel expressiveness and low computational efficiency. To overcome these issues, we first introduce a new duality between the sp…

Cited by 0SourceScholar
2025

OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation

ICCV 2025poster

Retrieval-augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external knowledge to reduce hallucinations and incorporate up-to-date information without retraining. As an essential part of RAG, external knowledge bases are commonly built by extracting structured data from…

2025

OmniArch: Building Foundation Model for Scientific Computing

ICML 2025poster

Foundation models have revolutionized language modeling, while whether this success is replicated in scientific computing remains unexplored. We present OmniArch, the first prototype aiming at solving multi-scale and multi-physics scientific computing problems with physical alignment. We addressed a…

Cited by 0SourcePDFScholar
2025

OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

CVPR 2025poster

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabilities from 2D to full 3D understanding is crucial for real-world applications. To address this challenge, we propose Omn…

2025

Pyramid Attention Enhancement Network for Nighttime UAV Tracking

ICASSP 2025accepted

Whilst Convolutional Neural Network (CNN)-based object tracking methods can achieve promising results on traditional well-lit datasets, it is challenging to accurately locate targets in low-light images taken in nighttime scenes, even for state-of-the-art (SOTA) trackers. Existing solutions often di…

Cited by 0SourceScholar
2025

RAIDEN Benchmark: Evaluating Role-playing Conversational Agents with Measurement-Driven Custom Dialogues

COLING 2025main

As Large-scale Language Models (LLMs) advance, the development of engaging Role-Playing Conversational Agents (RPCAs) has gained prominence. Despite this progress, there is a notable absence of benchmarks designed around dialogues, rather than question-answering formats, to assess the effectiveness…

2025

SLiNT: Structure-aware Language Model with Injection and Contrastive Training for Knowledge Graph Completion

EMNLP 2025

Link prediction in knowledge graphs (KGs) requires integrating structural information and semantic context to infer missing entities. While large language models (LLMs) offer strong generative reasoning capabilities, their limited exploitation of structural signals often results in *structural spars

Cited by 0SourcePDFScholar
2025

Safe Online Convex Optimization with Heavy-Tailed Observation Noises

AAAI 2025technical

We investigate safe online convex optimization (SOCO), where each decision must satisfy a set of unknown linear constraints. Assuming that the unknown constraints can be observed with a sub-Gaussian noise for each chosen decision, previous studies have established a high-probability regret bound of…

Cited by 0SourcePDFScholar
2025

ScalaLog: Scalable Log-Based Failure Diagnosis Using LLM

ICASSP 2025accepted

As Industrial Internet of Things (IIoT) software systems become increasingly complex, precise failure diagnosis has become both essential and challenging. Current log-based failure diagnosis methods lack scalability for different failure types. In IIoT software systems, the number of failure types i…

Cited by 0SourceScholar
2025

Towards Effective and Sparse Adversarial Attack on Spiking Neural Networks via Breaking Invisible Surrogate Gradients

CVPR 2025poster

Spiking neural networks (SNNs) have shown their competence in handling spatial-temporal event-based data with low energy consumption. Similar to conventional artificial neural networks (ANNs), SNNs are also vulnerable to gradient-based adversarial attacks, wherein gradients are calculated by spatial…

2025

URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model

NeurIPS 2025spotlight

Constructing accurate digital twins of articulated objects is essential for robotic simulation training and embodied AI world model building, yet historically requires painstaking manual modeling or multi-stage pipelines. In this work, we propose \textbf{URDF-Anything}, an end-to-end automatic recon…

Cited by 0SourceScholar
2025

Wheel-Legged SLAM: Indoor LiDAR-Inertial SLAM Integrating Kinematic Model of Wheel-Legged Robots

RA-L 2025

SLAM is the key technique for localization and surrounding perception in indoor environments. However, the dynamic posture adjustments of wheel-legged robots cast new challenges that affect the accuracy of localization. Therefore, this letter presents the Wheel-Legged SLAM, a novel indoor SLAM metho

Cited by 3SourceScholar
2024

Any-Size-Diffusion: Toward Efficient Text-Driven Synthesis for Any-Size HD Images

AAAI 2024technical

Stable diffusion, a generative model used in text-to-image synthesis, frequently encounters resolution-induced composition problems when generating images of varying sizes. This issue primarily stems from the model being trained on pairs of single-scale images and their corresponding text descriptio…

2024

BiLLM: Pushing the Limit of Post-Training Quantization for LLMs

ICML 2024poster

Pretrained large language models (LLMs) exhibit exceptional general language processing capabilities but come with significant demands on memory and computational resources. As a powerful compression technology, binarization can extremely reduce model weights to a mere 1 bit, lowering the expensive…

2024

FBLG: A Local Graph Based Approach for Handling Dual Skewed Non-IID Data in Federated Learning

IJCAI 2024poster

In real-world situations, federated learning often needs to process non-IID (non-independent and identically distributed) data with multiple skews, causing inadequate model performance. Existing federated learning methods mainly focus on addressing the problem with a single skew of non-IID, and henc…

2024

HPHS: Hierarchical Planning based on Hybrid Frontier Sampling for Unknown Environments Exploration

IROS 2024poster

Rapid sampling from the environment to acquire available frontier points and timely incorporating them into subsequent planning to reduce fragmented regions are critical to improve the efficiency of autonomous exploration. We propose HPHS, a fast and effective method for the autonomous exploration o…

Cited by 0SourceScholar
2024

Multi-Band Speech Tensor Decomposition for Interactive Feature Extraction in Early Dysphagia Screening

ICASSP 2024accepted

Dysphagia is a prevalent symptom in numerous neurological disorders among older adults. Current dysphagia diagnostic systems either involve invasive procedures or necessitate the ingestion of liquids. Some researchers have devised automatic dysphagia detection methods based on vowels that are easy t…

Cited by 0SourceScholar
2024

Preventing Model Collapse in Gaussian Process Latent Variable Models

ICML 2024poster

Gaussian process latent variable models (GPLVMs) are a versatile family of unsupervised learning models commonly used for dimensionality reduction. However, common challenges in modeling data with GPLVMs include inadequate kernel flexibility and improper selection of the projection noise, leading to…

2024

Representation Alignment and Adversarial Networks for Cross-lingual Dependency Parsing

EMNLP 2024finding

With the strong representational capabilities of pre-trained language models, dependency parsing in resource-rich languages has seen significant advancements. However, the parsing accuracy drops sharply when the model is transferred to low-resource language due to distribution shifts. To alleviate t…

2024

Self-Adaptive Reality-Guided Diffusion for Artifact-Free Super-Resolution

CVPR 2024poster

Artifact-free super-resolution (SR) aims to translate low-resolution images into their high-resolution counterparts with a strict integrity of the original content eliminating any distortions or synthetic details. While traditional diffusion-based SR techniques have demonstrated remarkable abilities…

2024

Surveying the Dead Minds: Historical-Psychological Text Analysis with Contextualized Construct Representation (CCR) for Classical Chinese

EMNLP 2024main

In this work, we develop a pipeline for historical-psychological text analysis in classical Chinese. Humans have produced texts in various languages for thousands of years; however, most of the computational literature is focused on contemporary languages and corpora. The emerging field of historica…

2024

TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling

COLING 2024main

As a cross-modal task, visual storytelling aims to generate a story for an ordered image sequence automatically. Different from the image captioning task, visual storytelling requires not only modeling the relationships between objects in the image but also mining the connections between adjacent im…

Cited by 1SourcePDFScholar
2024

Towards Efficient Modeling and Inference in Multi-Dimensional Gaussian Process State-Space Models

ICASSP 2024accepted

The Gaussian process state-space model (GPSSM) has attracted extensive attention for modeling complex nonlinear dynamical systems. However, the existing GPSSM employs separate Gaussian processes (GPs) for each latent state dimension, leading to escalating computational complexity and parameter proli…

Cited by 0SourceScholar
2023

Boosting Feedback Efficiency of Interactive Reinforcement Learning by Adaptive Learning from Scores

IROS 2023poster

Interactive reinforcement learning has shown promise in learning complex robotic tasks. However, the process can be human-intensive due to the requirement of a large amount of interactive feedback. This paper presents a new method that uses scores provided by humans instead of pairwise preferences t…

Cited by 0SourcecodeScholar
2023

Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection

ICCV 2023poster

In this paper, we propose a long-sequence modeling framework, named StreamPETR, for multi-view 3D object detection. Built upon the sparse query design in the PETR series, we systematically develop an object-centric temporal mechanism. The model is performed in an online manner and the long-term hist…

Cited by 235PDFcodeScholar
2023

HSR-Diff: Hyperspectral Image Super-Resolution via Conditional Diffusion Models

ICCV 2023poster

Despite the proven significance of hyperspectral images (HSIs) in performing various computer vision tasks, its potential is adversely affected by the low-resolution (LR) property in the spatial domain, resulting from multiple physical factors. Inspired by recent advancements in deep generative mode…

Cited by 47PDFScholar
2023

Overcoming Posterior Collapse in Variational Autoencoders Via EM-Type Training

ICASSP 2023accepted

Variational autoencoders (VAE) are one of the most prominent deep generative models for learning the underlying statistical distribution of high-dimensional data. However, training VAEs suffers from a severe issue called posterior collapse; that is, the learned posterior distribution collapses to th…

Cited by 0SourceScholar
2023

SAR Image Despeckling with Residual-in-Residual Dense Generative Adversarial Network

ICASSP 2023accepted

Deep convolutional neural networks have delivered remarkable aptitude in performing Synthetic Aperture Radar (SAR) image speckle removal tasks. Such approaches are nevertheless constrained in balancing speckle removal and preservation of spatial information, particularly with respect to strong speck…

Cited by 0SourceScholar
2023

Semi-attention Partition for Occluded Person Re-identification

AAAI 2023technical

This paper proposes a Semi-Attention Partition (SAP) method to learn well-aligned part features for occluded person re-identification (re-ID). Currently, the mainstream methods employ either external semantic partition or attention-based partition, and the latter manner is usually better than the fo…

Cited by 34SourcePDFScholar
2022

Coarse-To-Fine Unsupervised Change Detection for Remote Sensing Images Via Object-Based MRF and Inception UNET

ICASSP 2022accepted

With the rapid development of various satellite sensor techniques, remote sensing imagery has been an important source of data in change detection applications. This paper aims to propose an unsupervised change detection method based on Object-based Markov Random Filed (OMRF) and Inception UNet (IUN…

Cited by 0SourceScholar
2022

Decoupled Multi-Task Learning With Cyclical Self-Regulation for Face Parsing

CVPR 2022poster

This paper probes intrinsic factors behind typical failure cases (e.g spatial inconsistency and boundary confusion) produced by the existing state-of-the-art method in face parsing. To tackle these problems, we propose a novel Decoupled Multi-task Learning with Cyclical Self-Regulation (DML-CSR) for…

Cited by 43PDFcodeScholar
2022

KSAM: Infusing Multi-Source Knowledge into Dialogue Generation via Knowledge Source Aware Multi-Head Decoding

ACL 2022findings

Knowledge-enhanced methods have bridged the gap between human beings and machines in generating dialogue responses. However, most previous works solely seek knowledge from a single source, and thus they often fail to obtain available knowledge because of the insufficient coverage of a single knowled…

Cited by 6SourcePDFScholar
2022

Section-Aware Commonsense Knowledge-Grounded Dialogue Generation with Pre-trained Language Model

COLING 2022main

In knowledge-grounded dialogue generation, pre-trained language models (PLMs) can be expected to deepen the fusing of dialogue context and knowledge because of their superior ability of semantic understanding. Unlike adopting the plain text knowledge, it is thorny to leverage the structural commonse…

2022

Semi-supervised Domain Adaptation for Dependency Parsing with Dynamic Matching Network

ACL 2022long

Supervised parsing models have achieved impressive results on in-domain texts. However, their performances drop drastically on out-of-domain texts due to the data distribution shift. The shared-private model has shown its promising advantages for alleviating this problem via feature separation, wher…

Cited by 5SourcePDFScholar
2021

A Meta-Learning Framework for Few-Shot Classification of Remote Sensing Scene

ICASSP 2021accepted

While achieving remarkable success in remote sensing (RS) scene classification for the past few years, convolutional neural network (CNN) based methods suffer from the demand for large amounts of training data. The bottleneck in prediction accuracy has shifted from data processing limits toward a la…

Cited by 0SourceScholar
2021

APGN: Adversarial and Parameter Generation Networks for Multi-Source Cross-Domain Dependency Parsing

EMNLP 2021finding

Thanks to the strong representation learning capability of deep learning, especially pre-training techniques with language model loss, dependency parsing has achieved great performance boost in the in-domain scenario with abundant labeled training data for target domains. However, the parsing commun…

Cited by 5SourcePDFScholar
2021

Cirrus: A Long-range Bi-pattern LiDAR Dataset

ICRA 2021poster

In this paper, we introduce Cirrus, a new long-range bi-pattern LiDAR public dataset for autonomous driving tasks such as 3D object detection, critical to highway driving and timely decision making. Our platform is equipped with a high-resolution video camera and a pair of LiDAR sensors with a 250-m…

Cited by 41SourceScholar
2021

Heterogeneous two-Stream Network with Hierarchical Feature Prefusion for Multispectral Pan-Sharpening

ICASSP 2021accepted

Multispectral (MS) pan-sharpening aims at producing a high spatial resolution (HR) MS image by fusing a single-band HR panchromatic (PAN) image and a corresponding MS image with low spatial resolution. In this paper, we propose a heterogeneous two-stream network (HTSNet) with hierarchical feature pr…

Cited by 0SourceScholar
2021

Knowledge-Aware Dialogue Generation via Hierarchical Infobox Accessing and Infobox-Dialogue Interaction Graph Network

IJCAI 2021poster

Due to limited knowledge carried by queries, traditional dialogue systems often face the dilemma of generating boring responses, leading to poor user experience. To alleviate this issue, this paper proposes a novel infobox knowledge-aware dialogue generation approach, HITA-Graph, with three unique f…

2021

More is Better: Enhancing Open-Domain Dialogue Generation via Multi-Source Heterogeneous Knowledge

EMNLP 2021main

Despite achieving remarkable performance, previous knowledge-enhanced works usually only use a single-source homogeneous knowledge base of limited knowledge coverage. Thus, they often degenerate into traditional methods because not all dialogues can be linked with knowledge entries. This paper propo…

2021

Multi Path Training Framework for Data-Driven Open-Domain Conversation System

ICASSP 2021accepted

Nowadays, web data is often used to train a dialogue system. However, noises in web data can disturb the training process, as well as can impact the performance. Consequently, dialogue models tend to be brittle when receiving noisy inputs during the inference. This paper proposes a novel framework,…

Cited by 0SourceScholar
2021

Pushing The Limit of Type I Codebook For Fdd Massive Mimo Beamforming: A Channel Covariance Reconstruction Approach

ICASSP 2021accepted

There is a fundamental trade-off between the channel representation resolution of codebooks and the overheads of feedback communications in the fifth generation new radio (5G NR) frequency division duplex (FDD) massive multiple-input and multiple-output (MIMO) systems. In particular, two types of co…

Cited by 0SourceScholar
2020

Memory-Efficient Hierarchical Neural Architecture Search for Image Denoising

CVPR 2020poster

Recently, neural architecture search (NAS) methods have attracted much attention and outperformed manually designed architectures on a few high-level vision tasks. In this paper, we propose HiNAS (Hierarchical NAS), an effort towards employing NAS to automatically design effective neural network arc…

Cited by 89PDFScholar
2020

Semi-supervised Domain Adaptation for Dependency Parsing via Improved Contextualized Word Representations

COLING 2020main

In recent years, parsing performance is dramatically improved on in-domain texts thanks to the rapid progress of deep neural network models. The major challenge for current parsing research is to improve parsing performance on out-of-domain texts that are very different from the in-domain training d…

2020

TopicKA: Generating Commonsense Knowledge-Aware Dialogue Responses Towards the Recommended Topic Fact

IJCAI 2020poster

Insufficient semantic understanding of dialogue always leads to the appearance of generic responses, in generative dialogue systems. Recently, high-quality knowledge bases have been introduced to enhance dialogue understanding, as well as to reduce the prevalence of boring responses. Although such k…

2019

Asymmetric Local Metric Learning with PSD Constraint for Person Re-identification

ICRA 2019poster

Person re-identification is one of the key issues in both machine learning and video monitor application. In particular, defining an appropriate distance metric between the person images is very important. Existing metric learning approaches used in person re-identification either learn a single mea…

Cited by 1SourceScholar
2019

Exploiting Temporal Consistency for Real-Time Video Depth Estimation

ICCV 2019poster

Accuracy of depth estimation from static images has been significantly improved recently, by exploiting hierarchical features from deep convolutional neural networks (CNNs). Compared with static images, vast information exists among video frames and can be exploited to improve the depth estimation p…

Cited by 142PDFScholar
2016

Classification of voices that elicit soothing effect by applying a voiced vs. unvoiced feature engineering strategy

ICASSP 2016accepted

This paper introduces a novel approach of classifying voices that elicit a soothing effect on listeners from a domain knowledge inspired application of feature engineering. In particular, we utilize the characteristics of voiced vs, unvoiced speech in order to build a more accurate feature set. Larg…

Cited by 0SourceScholar