← Search

Jing Yu

29 accepted papers

2026

Manipulation Intention Understanding for Zero-Shot Composed Image Retrieval

AAAI 2026technical

Zero-shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with varied visual manipulation intents across domains, scenes, objects, and attributes. A key challenge is that existing datasets contain limited intent-relevant annotations, making it hard for models to infer human intent from text

Cited by 0SourcePDFScholar
2025

ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps

ICASSP 2025accepted

Singing voice synthesis (SVS) system is expected to generate high-fidelity singing voice from given music scores (lyrics, duration and pitch). Recently, diffusion models have performed well in this field. However, sacrificing inference speed to exchange with high-quality sample generation limits its…

Cited by 0SourceScholar
2025

MFL-Owner: Ownership Protection for Multi-modal Federated Learning via Orthogonal Transform Watermark

AAAI 2025technical

Multi-modal Federated Learning (MFL) is a distributed machine learning paradigm that enables multiple participants with multi-modal data to collaboratively train a global model for multi-modal tasks without sharing their local data. MFL typically deploys the trained global model as an Embedding-as-a…

2025

Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval

CVPR 2025poster

Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manipulation intent across domain, scene, object, and attribute. The key challenge for ZS-CIR tasks is to modify a reference image according to manipulation text to accurately retrieve a target im…

2025

ProAPO: Progressively Automatic Prompt Optimization for Visual Classification

CVPR 2025poster

Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their performances largely depend on the prompt quality. While recent methods show that visual descriptions generated by large language models (LLMs) enhance the…

2025

Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval

CVPR 2025highlight

Composed Image Retrieval (CIR) aims to retrieve target images that closely resemble a reference image while integrating user-specified textual modifications, thereby capturing user intent more accurately. Existing training-free zero-shot CIR (ZS-CIR) methods often employ a two-stage process: they fi…

2025

SaMer: A Scenario-aware Multi-dimensional Evaluator for Large Language Models

ICLR 2025poster

Evaluating the response quality of large language models (LLMs) for open-ended questions poses a significant challenge, especially given the subjectivity and multi-dimensionality of "quality" in natural language generation. Existing LLM evaluators often neglect that different scenarios require disti…

Cited by 0SourcePDFScholar
2025

Secure and Efficient Watermarking for Latent Diffusion Models in Model Distribution Scenarios

IJCAI 2025

Latent diffusion models have exhibited considerable potential in generative tasks. Watermarking is considered to be an alternative to safeguard the copyright of generative models and prevent their misuse. However, in the context of model distribution scenarios, the accessibility of models to large s

2025

Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization

EMNLP 2025

The widespread dissemination of toxic content on social media poses a serious threat to both online environments and public discourse, highlighting the urgent need for detoxification methods that effectively remove toxicity while preserving the original semantics.However, existing approaches often s

2024

Context-I2W: Mapping Images to Context-Dependent Words for Accurate Zero-Shot Composed Image Retrieval

AAAI 2024technical

Different from the Composed Image Retrieval task that requires expensive labels for training task-specific models, Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manipulation intent that could be related to domain, scene, object, and attribute…

2024

Learning the Uncertainty Sets of Linear Control Systems via Set Membership: A Non-asymptotic Analysis

ICML 2024poster

This paper studies uncertainty set estimation for unknown linear systems. Uncertainty sets are crucial for the quality of robust control since they directly influence the conservativeness of the control design. Departing from the confidence region analysis of least squares estimation, this paper foc…

Cited by 3SourcePDFScholar
2024

PGN: The RNN's New Successor is Effective for Long-Range Time Series Forecasting

NeurIPS 2024poster

Due to the recurrent structure of RNN, the long information propagation path poses limitations in capturing long-term dependencies, gradient explosion/vanishing issues, and inefficient sequential execution. Based on this, we propose a novel paradigm called Parallel Gated Network (PGN) as the new suc…

2024

Risk-Inspired Aerial Active Exploration for Enhancing Autonomous Driving of UGV in Unknown Off-Road Environments

ICRA 2024poster

Unknown area exploration is a crucial but challenging task for autonomous driving of unmanned ground vehicles (UGV) in unknown off-road environments. However, the exploration efficiency of a single UGV is low due to its limited sensing range. To solve this problem, this paper proposes a risk-inspire…

Cited by 1SourceScholar
2023

Achieving Hierarchy-Free Approximation for Bilevel Programs with Equilibrium Constraints

ICML 2023poster

In this paper, we develop an approximation scheme for solving bilevel programs with equilibrium constraints, which are generally difficult to solve. Among other things, calculating the first-order derivative in such a problem requires differentiation across the hierarchy, which is computationally in…

Cited by 8SourcePDFScholar
2022

MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-Based Visual Question Answering

CVPR 2022poster

Knowledge-based visual question answering requires the ability of associating external knowledge for open-ended cross-modal scene understanding. One limitation of existing solutions is that they capture relevant knowledge from text-only knowledge bases, which merely contain facts expressed by first-…

Cited by 136PDFcodeScholar
2022

Wlinker: Modeling Relational Triplet Extraction As Word Linking

ICASSP 2022accepted

Relational triplet extraction (RTE) is a fundamental task for automatically extracting information from unstructured text, which has attracted growing interest in recent years. However, it remains challenging due to the difficulty in extracting the overlapping relational triplets. Existing approache…

Cited by 0SourceScholar
2021

Coarse-To-Careful: Seeking Semantic-Related Knowledge for Open-Domain Commonsense Question Answering

ICASSP 2021accepted

It is prevalent to utilize external knowledge to help machine answer questions that need background commonsense, which faces a problem that unlimited knowledge will transmit noisy and misleading information. Towards the issue of introducing related knowledge, we propose a semantic-driven knowledge-a…

Cited by 0SourceScholar
2021

Evolving Attention with Residual Convolutions

ICML 2021spotlight

Transformer is a ubiquitous model for natural language processing and has attracted wide attentions in computer vision. The attention maps are indispensable for a transformer model to encode the dependencies among input tokens. However, they are learned independently in each layer and sometimes fail…

2021

MCR-NET: A Multi-Step Co-Interactive Relation Network for Unanswerable Questions on Machine Reading Comprehension

ICASSP 2021accepted

Question answering systems usually use keyword searches to retrieve potential passages related to a question, and then extract the answer from passages with the machine reading comprehension methods. However, many questions tend to be unanswerable in the real world. In this case, it is significant a…

Cited by 0SourceScholar
2021

Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-Encoder

AAAI 2021technical

It is important for task-oriented dialogue systems to discover the dialogue structure (i.e. the general dialogue flow) from dialogue corpora automatically. Previous work models dialogue structure by extracting latent states for each utterance first and then calculating the transition probabilities a…

2020

Bi-directional CognitiveThinking Network for Machine Reading Comprehension

COLING 2020main

We propose a novel Bi-directional Cognitive Knowledge Framework (BCKF) for reading comprehension from the perspective of complementary learning systems theory. It aims to simulate two ways of thinking in the brain to answer questions, including reverse thinking and inertial thinking. To validate the…

Cited by 12SourcePDFScholar
2020

DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue

IJCAI 2020poster

Visual Dialogue task requires an agent to be engaged in a conversation with human about an image. The ability of generating detailed and non-repetitive responses is crucial for the agent to achieve human-like conversation. In this paper, we propose a novel generative decoding architecture to generat…

2020

Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering

IJCAI 2020poster

Fact-based Visual Question Answering (FVQA) requires external knowledge beyond the visible content to answer questions about an image. This ability is challenging but indispensable to achieve general VQA. One limitation of existing FVQA solutions is that they jointly embed all kinds of information w…

2017

Blind image deblurring based on sparse representation and structural self-similarity

ICASSP 2017accepted

In this paper, we propose a blind motion deblurring method based on sparse representation and structural self-similarity from a single image. The priors for sparse representation and structural self-similarity are explicitly added into the recovery of the latent image by means of sparse and multi-sc…

Cited by 0SourceScholar
2017

Classification of thyroid nodules in ultrasound images using deep model based transfer learning and hybrid features

ICASSP 2017accepted

Ultrasonography is a valuable diagnosis method for thyroid nodules. Automatically discriminating benign and malignant nodules in the ultrasound images can provide aided diagnosis suggestions, or increase the diagnosis accuracy when lack of experts. The core problem in this issue is how to capture ap…

Cited by 0SourceScholar
2017

Unsupervised feature extraction for hyperspectral images using combined low rank representation and locally linear embedding

ICASSP 2017accepted

Hyperspectral images(HSIs) provide hundreds of narrow spectral bands for the land-covers, thus can provide more powerful discriminative information for the land-cover classification. However, HSIs suffer from the curse of high dimensionality, therefore dimension reduction and feature extraction are…

Cited by 0SourceScholar
2016

Ship wake detection for SAR images with complex backgrounds based on morphological dictionary learning

ICASSP 2016accepted

The ship wake detection of SAR images is useful not only in estimating the speed and the direction of moving ships, but also in finding small ships which are hard to be detected. The traditional ship wake detection methods of SAR images can achieve satisfactory results in simple backgrounds, but har…

Cited by 0SourceScholar