← Search

Jing Jiang

83 accepted papers

2026

A Survey of Personalized Federated Foundation Models for Privacy-Preserving Recommendation

IJCAI 2026

Integrating Foundation Models (FMs) into recommendation systems is an emerging and promising research direction. However, centralized paradigms face growing pressure from privacy concerns and strict regulatory requirements. Federated learning offers a viable solution that enables collaborative model

Cited by 0Scholar
2026

Bi-level Personalization for Federated Foundation Models: A Task-vector Aggregation Approach

AAAI 2026technical

Federated foundation models represent a new paradigm to jointly fine-tune pre-trained foundation models across clients. It is still a challenge to fine-tune foundation models for a small group of new users or specialized scenarios, which typically involve limited data compared to the large-scale dat

Cited by 0SourcePDFScholar
2026

FeDaL: Federated Dataset Learning for General Time Series Foundation Models

ICLR 2026poster

Dataset-level heterogeneity introduces significant domain biases that fundamentally degrade generalization on general Time Series Foundation Models (TSFMs), yet this challenge remains underexplored. This paper rethinks the from-scratch training of TSFMs using the paradigm of federated learning. We p…

Cited by 0SourcecodeScholar
2026

FedMerge: Federated Model Merging for Personalization

AAAI 2026technical

One global model in federated learning (FL) might not be sufficient to serve many clients with non-IID tasks and distributions. Despite recent advances in FL to train multiple global models for better personalization, they only provide limited model choices to clients, so local finetuning of multipl

Cited by 0SourcePDFScholar
2026

Federated Vision-Language-Recommendation with Personalized Fusion

AAAI 2026technical

Applying large pre-trained Vision-Language Models to recommendation is a burgeoning field, a direction we term Vision-Language-Recommendation (VLR). Bringing VLR to user-oriented on-device intelligence within a federated learning framework is a crucial step for enhancing user privacy and delivering

Cited by 0SourcePDFScholar
2026

Learning Disentangled Multi-Agent World Model for Decentralized Control

ICML 2026poster

World models enable learning policies via latent imagination, offering benefits such as history compression and sample efficiency. The primary challenge in applying world models to multi-agent tasks is that modeling multi-agent dynamics in latent space requires integrating information from different…

Cited by 0SourceScholar
2026

Personalized Additive Modeling for Multi-level Federated Learning

ICML 2026poster

Contemporary AI faces the challenge of balancing generality with user-specific personalization. In federated learning (FL), this challenge is amplified by highly heterogeneous client data with complex non-IID patterns beyond standard modeling assumptions. Many existing FL methods are designed for re…

Cited by 0SourceScholar
2025

Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion Adaptation

AAAI 2025technical

Visual emotion recognition (VER), which aims at understanding humans' emotional reactions toward different visual stimuli, has attracted increasing attention. Given the subjective and ambiguous characteristics of emotion, annotating a reliable large-scale dataset is hard. For reducing reliance on da…

2025

CAMI: A Counselor Agent Supporting Motivational Interviewing through State Inference and Topic Exploration

ACL 2025long

Conversational counselor agents have become essential tools for addressing the rising demand for scalable and accessible mental health support. This paper introduces CAMI, a novel automated counselor agent grounded in Motivational Interviewing (MI) – a client-centered counseling approach designed to…

Cited by 0SourcePDFScholar
2025

Cheating Automatic LLM Benchmarks: Null Models Achieve High Win Rates

ICLR 2025oral

Automatic LLM benchmarks, such as AlpacaEval 2.0, Arena-Hard-Auto, and MT-Bench, have become popular for evaluating language models due to their cost-effectiveness and scalability compared to human evaluation. Achieving high win rates on these benchmarks can significantly boost the promotional impac…

2025

Colloquial Singaporean English Style Transfer with Fine-Grained Explainable Control

ACL 2025long

Colloquial Singaporean English (Singlish) is an informal English marked by a unique blend of languages reflecting Singapore’s multicultural identity. Style transfer between Singlish and Standard (formal) English is vital for various applications, yet existing methods often lack explainability and fi…

Cited by 0SourcePDFScholar
2025

Consistent Client Simulation for Motivational Interviewing-based Counseling

ACL 2025long

Simulating human clients in mental health counseling is crucial for training and evaluating counselors (both human or simulated) in a scalable manner. Nevertheless, past research on client simulation did not focus on complex conversation tasks such as mental health counseling. In these tasks, the ch…

Cited by 0SourcePDFScholar
2025

DPC: Dual-Prompt Collaboration for Tuning Vision-Language Models

CVPR 2025poster

The Base-New Trade-off (BNT) problem universally exists during the optimization of CLIP-based prompt tuning, where continuous fine-tuning on base (target) classes leads to a simultaneous decrease of generalization ability on new (unseen) classes. Existing approaches attempt to regulate the prompt tu…

2025

FOCUS: Evaluating Pre-trained Vision-Language Models on Underspecification Reasoning

ACL 2025long

Humans possess a remarkable ability to interpret underspecified ambiguous statements by inferring their meanings from contexts such as visual inputs. This ability, however, may not be as developed in recent pre-trained vision-language models (VLMs). In this paper, we introduce a novel probing datase…

Cited by 0SourcePDFScholar
2025

Federated Foundation Models on Heterogeneous Time Series

AAAI 2025technical

Training a general-purpose time series foundation models with robust generalization capabilities across diverse applications from scratch is still an open challenge. Efforts are primarily focused on fusing cross-domain time series datasets to extract shared subsequences as tokens for training models…

2025

Federated Low-Rank Adaptation for Foundation Models: A Survey

IJCAI 2025

Effectively leveraging private datasets remains a significant challenge in developing foundation models. Federated Learning (FL) has recently emerged as a collaborative framework that enables multiple users to fine-tune these models while mitigating data privacy risks. Meanwhile, Low-Rank Adaptation

2025

Gaussian Constrained Diffeomorphic Deformation Network for Panoramic Semantic Segmentation

ICASSP 2025accepted

Panoramic semantic segmentation has garnered increasing attention due to its ability to provide comprehensive environmental perception. However, it requires a large number of annotated panoramic images to achieve satisfactory performance, which is costly. Recently, Domain Adaptation for Panoramic Se…

Cited by 0SourceScholar
2025

Learning Class Prototypes for Visual Emotion Recognition

ICASSP 2025accepted

Visual emotion recognition (VER), which aims at understanding humans’ emotional reactions toward different visual stimuli, has attracted increasing attention. However, because of the subjectivity and complex nature of emotion, existing VER methods suffer from one or more of the following problems: 1…

Cited by 0SourceScholar
2025

Parameter Selections and Applications for Soft Bellows Actuators (SBAs) with Various Performance Metrics

IROS 2025

Soft bellows actuators (SBAs), a particular type of soft pneumatic actuators (SPAs), are widely used in various applications, such as climbing robots, industrial grippers, and wearable devices. Despite their advantages of uniform motion and high efficiency, the design of SBAs often relies on experie

Cited by 0SourceScholar
2025

Personalized Federated Collaborative Filtering: A Variational AutoEncoder Approach

AAAI 2025technical

Federated Collaborative Filtering (FedCF) is an emerging field focused on developing a new recommendation framework with preserving privacy in a federated setting. Existing FedCF methods typically combine distributed Collaborative Filtering (CF) algorithms with privacy-preserving mechanisms, and the…

2025

RegMix: Data Mixture as Regression for Language Model Pre-training

ICLR 2025spotlight

The data mixture for large language model pre-training significantly impacts performance, yet how to determine an effective mixture remains unclear. We propose RegMix to automatically identify a high-performing data mixture by formulating it as a regression task. RegMix trains many small models on d…

2025

SPASCA: Social Presence and Support with Conversational Agent for Persons Living with Dementia

AAAI 2025technical

We present SPASCA - a conversational AI system that promotes psychological and cognitive well-being of persons living with dementia (PLWD). This system features an AI agent that provides social presence and support to PLWD through verbal communications, without physical presence of human caregivers.…

Cited by 0SourcePDFScholar
2025

Seeing Culture: A Benchmark for Visual Reasoning and Grounding

EMNLP 2025

Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergence of new cultural datasets. However, these datasets frequently fall short of pr

2025

WALL-E: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents

NeurIPS 2025poster

Can we build accurate world models out of large language models (LLMs)? How can world models benefit LLM agents? The gap between the prior knowledge of LLMs and the specified environment's dynamics usually bottlenecks LLMs' performance as world models. To bridge the gap, we propose a training-free "…

Cited by 0SourceScholar
2024

Actively Learn from LLMs with Uncertainty Propagation for Generalized Category Discovery

NAACL 2024long

Generalized category discovery faces a key issue: the lack of supervision for new and unseen data categories. Traditional methods typically combine supervised pretraining with self-supervised learning to create models, and then employ clustering for category identification. However, these approaches…

2024

Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

ICML 2024poster

A multimodal large language model (MLLM) agent can receive instructions, capture images, retrieve histories from memory, and decide which tools to use. Nonetheless, red-teaming efforts have revealed that adversarial images/prompts can jailbreak an MLLM and cause unaligned behaviors. In this work, we…

2024

Dual-Personalizing Adapter for Federated Foundation Models

NeurIPS 2024poster

Recently, foundation models, particularly large language models (LLMs), have demonstrated an impressive ability to adapt to various tasks by fine-tuning diverse instruction data. Notably, federated foundation models (FedFM) emerge as a privacy preservation method to fine-tune models collaboratively…

2024

Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorld

CVPR 2024poster

While large language models (LLMs) excel in a simulated world of texts they struggle to interact with the more realistic world without perceptions of other modalities such as visual or audio signals. Although vision-language models (VLMs) integrate LLM modules (1) aligned with static image features…

2024

Federated Prompt Learning for Weather Foundation Models on Devices

IJCAI 2024poster

On-device intelligence for weather forecasting uses local deep learning models to analyze weather patterns without centralized cloud computing, holds significance for supporting human activates. Federated Learning is a promising solution for such forecasting by enabling collaborative model training…

2024

GliDe with a CaPE: A Low-Hassle Method to Accelerate Speculative Decoding

ICML 2024poster

Speculative decoding is a relatively new decoding framework that leverages small and efficient draft models to reduce the latency of LLMs. In this study, we introduce GliDe and CaPE, two low-hassle modifications to vanilla speculative decoding to further improve the decoding speed of a frozen LLM. S…

Cited by 19SourcePDFScholar
2024

Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses

NeurIPS 2024poster

Recently, Anil et al. (2024) show that many-shot (up to hundreds of) demonstrations can jailbreak state-of-the-art LLMs by exploiting their long-context capability. Nevertheless, is it possible to use few-shot demonstrations to efficiently jailbreak LLMs within limited context sizes? While the vanil…

2024

Intriguing Properties of Data Attribution on Diffusion Models

ICLR 2024poster

Data attribution seeks to trace model outputs back to training data. With the recent development of diffusion models, data attribution has become a desired module to properly assign valuations for high-quality or copyrighted training samples, ensuring that data contributors are fairly compensated or…

2024

Personalized Adapter for Large Meteorology Model on Devices: Towards Weather Foundation Models

NeurIPS 2024poster

This paper demonstrates that pre-trained language models (PLMs) are strong foundation models for on-device meteorological variable modeling. We present LM-Weather, a generic approach to taming PLMs, that have learned massive sequential knowledge from the universe of natural language databases, to ac…

Cited by 7SourcePDFScholar
2024

Speaker Verification in Agent-generated Conversations

ACL 2024long

The recent success of large language models (LLMs) has attracted widespread interest to develop role-playing conversational agents personalized to the characteristics and styles of different speakers to enhance their abilities to perform both general and special purpose dialogue tasks. However, the…

Cited by 2SourcePDFScholar
2024

Synergizing Large Language Models and Pre-Trained Smaller Models for Conversational Intent Discovery

ACL 2024findings

In Conversational Intent Discovery (CID), Small Language Models (SLMs) struggle with overfitting to familiar intents and fail to label newly discovered ones. This issue stems from their limited grasp of semantic nuances and their intrinsically discriminative framework. Therefore, we propose Synergiz…

2024

What Hides behind Unfairness? Exploring Dynamics Fairness in Reinforcement Learning

IJCAI 2024poster

In sequential decision-making problems involving sensitive attributes like race and gender, reinforcement learning (RL) agents must carefully consider long-term fairness while maximizing returns. Recent works have proposed many different types of fairness notions, but how unfairness arises in RL pro…

2023

Adaptive Policy Learning for Offline-to-Online Reinforcement Learning

AAAI 2023technical

Conventional reinforcement learning (RL) needs an environment to collect fresh data, which is impractical when online interactions are costly. Offline RL provides an alternative solution by directly learning from the previously collected dataset. However, it will yield unsatisfactory performance if…

Cited by 27SourcePDFScholar
2023

Continual Task Allocation in Meta-Policy Network via Sparse Prompting

ICML 2023poster

How to train a generalizable meta-policy by continually learning a sequence of tasks? It is a natural human skill yet challenging to achieve by current reinforcement learning: the agent is expected to quickly adapt to new tasks (plasticity) meanwhile retaining the common knowledge from previous task…

2023

Does Continual Learning Equally Forget All Parameters?

ICML 2023poster

Distribution shift (e.g., task or domain shift) in continual learning (CL) usually results in catastrophic forgetting of previously learned knowledge. Although it can be alleviated by repeatedly replaying buffered data, the every-step replay is time-consuming. In this paper, we study which modules i…

Cited by 19SourcePDFScholar
2023

Federated Learning on Non-IID Graphs via Structural Knowledge Sharing

AAAI 2023technical

Graph neural networks (GNNs) have shown their superiority in modeling graph data. Owing to the advantages of federated learning, federated graph learning (FGL) enables clients to train strong GNN models in a distributed manner without sharing their private data. A core challenge in federated systems…

2023

Prompt Federated Learning for Weather Forecasting: Toward Foundation Models on Meteorological Data

IJCAI 2023poster

To tackle the global climate challenge, it urgently needs to develop a collaborative platform for comprehensive weather forecasting on large-scale meteorological data. Despite urgency, heterogeneous meteorological sensors across countries and regions, inevitably causing multivariate heterogeneity an…

2023

ROME: Evaluating Pre-trained Vision-Language Models on Reasoning beyond Visual Common Sense

EMNLP 2023long findings

Humans possess a strong capability for reasoning beyond common sense. For example, given an unconventional image of a goldfish laying on the table next to an empty fishbowl, a human would effortlessly determine that the fish is not inside the fishbowl. The case, however, may be different for a visio…

Cited by 0SourcecodeScholar
2023

Structured Federated Learning through Clustered Additive Modeling

NeurIPS 2023poster

Heterogeneous federated learning without assuming any structure is challenging due to the conflicts among non-identical data distributions of clients. In practice, clients often comprise near-homogeneous clusters so training a server-side model per cluster mitigates the conflicts. However, FL with c…

Cited by 18SourcePDFScholar
2023

Unsupervised Video Domain Adaptation for Action Recognition: A Disentanglement Perspective

NeurIPS 2023poster

Unsupervised video domain adaptation is a practical yet challenging task. In this work, for the first time, we tackle it from a disentanglement view. Our key idea is to handle the spatial and temporal domain divergence separately through disentanglement. Specifically, we consider the generation of c…

2022

Attentional Gated Res2net for Multivariate Time Series Classification

ICASSP 2022accepted

Multivariate time series classification is a critical problem in data mining with broad applications. We design a novel convolutional neural network architecture, Attentional Gated Res2Net, for robust multivariate time series classification. AGRes2Net uses hierarchical residual-like connections to a…

Cited by 0SourceScholar
2022

Context Modeling with Evidence Filter for Multiple Choice Question Answering

ICASSP 2022accepted

Multiple-Choice Question Answering (MCQA) is one of the challenging tasks in machine reading comprehension. The main challenge in MCQA is to extract "evidence" from the given context that supports the correct answer. In OpenbookQA dataset [1], the requirement of extracting "evidence" is particularly…

Cited by 0SourceScholar
2022

EAT-C: Environment-Adversarial sub-Task Curriculum for Efficient Reinforcement Learning

ICML 2022spotlight

Reinforcement learning (RL) is inefficient on long-horizon tasks due to sparse rewards and its policy can be fragile to slightly perturbed environments. We address these challenges via a curriculum of tasks with coupled environments, generated by two policies trained jointly with RL: (1) a co-operat…

2022

Exploring and Adapting Chinese GPT to Pinyin Input Method

ACL 2022long

While GPT has become the de-facto method for text generation tasks, its application to pinyin input method remains unexplored. In this work, we make the first exploration to leverage Chinese GPT for pinyin input method. We find that a frozen GPT achieves state-of-the-art performance on perfect pinyi…

2022

FedProto: Federated Prototype Learning across Heterogeneous Clients

AAAI 2022technical

Heterogeneity across clients in federated learning (FL) usually hinders the optimization convergence and generalization performance when the aggregation of clients' knowledge occurs in the gradient space. For example, clients may differ in terms of data distribution, network latency, input/output sp…

2022

Federated Learning from Pre-Trained Models: A Contrastive Learning Approach

NeurIPS 2022accept

Federated Learning (FL) is a machine learning paradigm that allows decentralized clients to learn collaboratively without sharing their private data. However, excessive computation and communication demands pose challenges to current FL frameworks, especially when training large-scale models. To pre…

Cited by 213SourcePDFScholar
2022

Hierarchical Relation-Guided Type-Sentence Alignment for Long-Tail Relation Extraction with Distant Supervision

NAACL 2022findings

Distant supervision uses triple facts in knowledge graphs to label a corpus for relation extraction, leading to wrong labeling and long-tail problems. Some works use the hierarchy of relations for knowledge transfer to long-tail relations. However, a coarse-grained relation often implies only an att…

Cited by 3SourcePDFScholar
2022

Interventional Training for Out-Of-Distribution Natural Language Understanding

EMNLP 2022main

Out-of-distribution (OOD) settings are used to measure a model’s performance when the distribution of the test data is different from that of the training data. NLU models are known to suffer in OOD. We study this issue from the perspective of causality, which sees confounding bias as the reason for…

2022

Omni-Scale CNNs: a simple and effective kernel size configuration for time series classification

ICLR 2022poster

The size of the receptive field has been one of the most important factors for One Dimensional Convolutional Neural Networks (1D-CNNs) on time series classification tasks. Large efforts have been taken to choose the appropriate receptive field size, for it has a huge influence on the performance and…

2022

Pareto Policy Pool for Model-based Offline Reinforcement Learning

ICLR 2022poster

Online reinforcement learning (RL) can suffer from poor exploration, sparse reward, insufficient data, and overhead caused by inefficient interactions between an immature policy and a complicated environment. Model-based offline RL instead trains an environment model using a dataset of pre-collected…

Cited by 23SourcePDFScholar
2022

Personalized Federated Learning With a Graph

IJCAI 2022poster

Knowledge sharing and model personalization are two key components in the conceptual framework of personalized federated learning (PFL). Existing PFL methods focus on proposing new model personalization mechanisms while simply implementing knowledge sharing by aggregating models from all clients, re…

2022

ngram-OAXE: Phrase-Based Order-Agnostic Cross Entropy for Non-Autoregressive Machine Translation

COLING 2022main

Recently, a new training oaxe loss has proven effective to ameliorate the effect of multimodality for non-autoregressive translation (NAT), which removes the penalty of word order errors in the standard cross-entropy loss. Starting from the intuition that reordering generally occurs between phrases,…

2021

A Low-Complexity Admm-Based Massive Mimo Detectors Via Deep Neural Networks

ICASSP 2021accepted

An alternate direction method of multipliers (ADMM)-based detectors can achieve good performance in both small and large-scale multiple-input multiple-output (MIMO) systems. However, due to the difficulty of choosing the optimal penalty parameters, their performance is limited. This paper presents a…

Cited by 0SourceScholar
2021

A Survey on Complex Knowledge Base Question Answering: Methods, Challenges and Solutions

IJCAI 2021poster

Knowledge base question answering (KBQA) aims to answer a question over a knowledge base (KB). Recently, a large number of studies focus on semantically or syntactically complicated questions. In this paper, we elaborately summarize the typical challenges and solutions for complex KBQA. We begin wi…

Cited by 228SourcePDFScholar
2021

A Universal Representation Transformer Layer for Few-Shot Image Classification

ICLR 2021poster

Few-shot classification aims to recognize unseen classes when presented with only a small number of samples. We consider the problem of multi-domain few-shot image classification, where unseen classes and examples come from diverse data sources. This problem has seen growing interest and has inspire…

2021

CO-PILOT: COllaborative Planning and reInforcement Learning On sub-Task curriculum

NeurIPS 2021poster

Goal-conditioned reinforcement learning (RL) usually suffers from sparse reward and inefficient exploration in long-horizon tasks. Planning can find the shortest path to a distant goal that provides dense reward/guidance but is inaccurate without a precise environment model. We show that RL and plan…

2021

COSY: COunterfactual SYntax for Cross-Lingual Understanding

ACL 2021long

Pre-trained multilingual language models, e.g., multilingual-BERT, are widely used in cross-lingual tasks, yielding the state-of-the-art performance. However, such models suffer from a large performance gap between source and target languages, especially in the zero-shot setting, where the models ar…

2021

Isometric Propagation Network for Generalized Zero-shot Learning

ICLR 2021poster

Zero-shot learning (ZSL) aims to classify images of an unseen class only based on a few attributes describing that class but no access to any training sample. A popular strategy is to learn a mapping between the semantic space of class attributes and the visual space of images based on the seen clas…

Cited by 49SourcePDFScholar
2021

Modeling Transitions of Focal Entities for Conversational Knowledge Base Question Answering

ACL 2021long

Conversational KBQA is about answering a sequence of questions related to a KB. Follow-up questions in conversational KBQA often have missing information referring to entities from the conversation history. In this paper, we propose to model these implied entities, which we refer to as the focal ent…

2021

NOAHQA: Numerical Reasoning with Interpretable Graph Question Answering Dataset

EMNLP 2021finding

While diverse question answering (QA) datasets have been proposed and contributed significantly to the development of deep learning models for QA tasks, the existing datasets fall short in two aspects. First, we lack QA datasets covering complex questions that involve answers as well as the reasonin…

2021

Order-Agnostic Cross Entropy for Non-Autoregressive Machine Translation

ICML 2021oral

We propose a new training objective named order-agnostic cross entropy (OaXE) for fully non-autoregressive translation (NAT) models. OaXE improves the standard cross-entropy loss to ameliorate the effect of word reordering, which is a common source of the critical multimodality problem in NAT. Concr…

2020

Cooperative Heterogeneous Deep Reinforcement Learning

NeurIPS 2020poster

Numerous deep reinforcement learning agents have been proposed, and each of them has its strengths and flaws. In this work, we present a Cooperative Heterogeneous Deep Reinforcement Learning (CHDRL) framework that can learn a policy by integrating the advantages of heterogeneous agents. Specifically…

2020

Effective Search of Logical Forms for Weakly Supervised Knowledge-Based Question Answering

IJCAI 2020poster

Many algorithms for Knowledge-Based Question Answering (KBQA) depend on semantic parsing, which translates a question to its logical form. When only weak supervision is provided, it is usually necessary to search valid logical forms for model training. However, a complex question typically involves…

Cited by 0SourcePDFScholar
2020

Improving Long-Tail Relation Extraction with Collaborating Relation-Augmented Attention

COLING 2020main

Wrong labeling problem and long-tail relations are two main challenges caused by distant supervision in relation extraction. Recent works alleviate the wrong labeling by selective attention via multi-instance learning, but cannot well handle long-tail relations even if hierarchies of the relations a…

2020

MESA: Boost Ensemble Imbalanced Learning with MEta-SAmpler

NeurIPS 2020poster

Imbalanced learning (IL), i.e., learning unbiased models from class-imbalanced data, is a challenging problem. Typical IL methods including resampling and reweighting were designed based on some heuristic assumptions. They often suffer from unstable performance, poor applicability, and high computat…

2020

RatE: Relation-Adaptive Translating Embedding for Knowledge Graph Completion

COLING 2020main

Many graph embedding approaches have been proposed for knowledge graph completion via link prediction. Among those, translating embedding approaches enjoy the advantages of light-weight structure, high efficiency and great interpretability. Especially when extended to complex vector space, they show…

2019

Learning to Propagate for Graph Meta-Learning

NeurIPS 2019poster

Meta-learning extracts the common knowledge from learning different tasks and uses it for unseen tasks. It can significantly improve tasks that suffer from insufficient training data, e.g., few-shot learning. In most meta-learning methods, tasks are implicitly related by sharing parameters or optimize…

2019

Spatio-spectral Modulation Using a Binary Photomask for Compressive Chromotomography

ICASSP 2019accepted

Recent advances in compressive spectral imagers have demonstrated the potential of spatio-spectral modulation (SSM) for improved reconstruction performance. Existing SSM techniques, however, use either a color filter array or a complex optical arrangement, both of which can only provide limited modu…

Cited by 0SourceScholar
2018

Bi-Directional Block Self-Attention for Fast and Memory-Efficient Sequence Modeling

ICLR 2018poster

Recurrent neural networks (RNN), convolutional neural networks (CNN) and self-attention networks (SAN) are commonly used to produce context-aware representations. RNN can capture long-range dependency but is hard to parallelize and not time-efficient. CNN focuses on local dependency but does not per…

2018

Evidence Aggregation for Answer Re-Ranking in Open-Domain Question Answering

ICLR 2018poster

Very recently, it comes to be a popular approach for answering open-domain questions by first searching question-related passages, then applying reading comprehension models to extract answers. Existing works usually extract answers from single passages independently, thus not fully make use of the…