← Search

Xian Wu

85 accepted papers

2026

CMID: Towards Medical Visual Question Answering via Contrastive Mutual Information Decoding

AAAI 2026technical

Medical Visual Question Answering (Med-VQA) aims to generate accurate answers for clinical questions grounded in medical images, which has attracted increasing research attention due to its potential to streamline diagnostics and reduce clinical burden. Recent advances in Large Vision-Language Model

Cited by 0SourcePDFScholar
2026

Copy-Paste to Mitigate Large Language Model Hallucinations

ICLR 2026poster

While Retrieval-Augmented Generation (RAG) enables large language models (LLMs) to generate contextually grounded responses, contextual faithfulness remains challenging as LLMs may not consistently trust provided context, leading to hallucinations that undermine reliability. We observe an inverse co…

Cited by 0SourcecodeScholar
2026

ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation

ICML 2026poster

Electrocardiography (ECG) serves as an indispensable diagnostic tool in clinical practice, yet existing multimodal large language models (MLLMs) remain unreliable for ECG interpretation, often producing plausible but clinically incorrect analyses. To address this, we propose ECG-R1, the first reason…

Cited by 0SourceScholar
2026

From Guessing to Placeholding: A Cost-Theoretic Framework for Uncertainty-Aware Code Completion

ICML 2026poster

While Large Language Models (LLMs) have demonstrated exceptional proficiency in code completion, they typically adhere to a **Hard Completion (HC)** paradigm, compelling the generation of fully concrete code even amidst insufficient context. Our analysis of 3 million real-world interactions exposes …

Cited by 0SourceScholar
2026

HO-SFL: Hybrid-Order Split Federated Learning with Backprop-Free Clients and Dimension-Free Aggregation

ICML 2026poster

Fine-tuning large models on edge devices is severely hindered by the memory-intensive backpropagation (BP) in standard frameworks like federated learning and split learning. While substituting BP with zeroth-order optimization can significantly reduce memory footprints, it typically suffers from pro…

Cited by 0SourceScholar
2026

InfoScan: Information-Efficient Visual Scanning via Resource-Adaptive Walks

ICLR 2026poster

High-resolution visual representation learning remains challenging due to the quadratic complexity of Vision Transformers and the limitations of existing efficient approaches, where fixed scanning patterns in recent Mamba-based models hinder content-adaptive perception. To address these limitations,…

Cited by 0SourceScholar
2026

MGD:Mesh-guided Gaussians with Diffusion Priors for Dynamic Objects Reconstruction from Monocular RGB-D Video

AAAI 2026technical

Reconstructing dynamic objects from monocular RGB-D video is critical for advancing 3D vision applications and enhancing user experience. However, monocular RGB-D video provides limited 3D observations, making the reconstruction of unobserved regions highly under-constrained. Despite recent advanc

Cited by 0SourcePDFScholar
2026

MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current emotional benchmarks remain limited, as it is still unknown: (a)…

Cited by 0SourcecodeScholar
2026

MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning

ICLR 2026poster

Medical Large Vision-Language Models (Med-LVLMs) have shown strong potential in multimodal diagnostic tasks. However, existing single-agent models struggle to generalize across diverse medical specialties, limiting their performance. Recent efforts introduce multi-agent collaboration frameworks insp…

Cited by 0SourceScholar
2026

NeuralOM: Neural Ocean Model for Subseasonal-to-Seasonal Simulation

AAAI 2026technical

Long-term, high-fidelity simulation of slow-changing physical systems, such as the ocean and climate, presents a fundamental challenge in scientific computing. Traditional autoregressive machine learning models often fail in these tasks as minor errors accumulate and lead to rapid forecast degradati

Cited by 0SourcePDFScholar
2026

PnP-Corrector: A Universal Correction Framework for Coupled Spatiotemporal Forecasting

ICML 2026poster

Coupled spatiotemporal forecasting is important for predicting the future evolution of multiple interacting dynamical systems, such as in climate models. However, existing methods are severely constrained by the persistent bottleneck of compounding errors. In coupled systems, errors from each subsys…

Cited by 0SourceScholar
2026

Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception

AAAI 2026technical

Spatio-temporal alignment is crucial for temporal modeling of end-to-end (E2E) perception in autonomous driving (AD), providing valuable structural and textural prior information. Existing methods typically rely on the attention mechanism to align objects across frames, simplifying the motion model

Cited by 0SourcePDFScholar
2026

S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm Detection

AAAI 2026technical

Multimodal sarcasm detection (MSD) aims to identify sarcasm polarity from diverse modalities (i.e., image–text pairs), a task that has received increasing attention. While significant progress has been made, existing approaches still face two major issues: lack of explainability and weak generalizab

Cited by 0SourcePDFScholar
2025

A Layered Debating Multi-Agent System for Similar Disease Diagnosis

NAACL 2025short

Distinguishing between extremely similar diseases is a critical and challenging aspect of clinical decision-making. Traditional classification, contrastive learning, and Large Language Models (LLMs) based methods fail to detect the subtle clues necessary for differentiation. This task demands comple…

Cited by 0SourcePDFScholar
2025

A Multi-Expert Structural-Semantic Hybrid Framework for Unveiling Historical Patterns in Temporal Knowledge Graphs

ACL 2025finding

Temporal knowledge graph reasoning aims to predict future events with knowledge of existing facts and plays a key role in various downstream tasks. Previous methods focused on either graph structure learning or semantic reasoning, failing to integrate dual reasoning perspectives to handle different…

2025

A Survey on Foundation Language Models for Single-cell Biology

ACL 2025long

The recent advancements in language models have significantly catalyzed progress in computational biology. A growing body of research strives to construct unified foundation models for single-cell biology, with language models serving as the cornerstone. In this paper, we systematically review the d…

Cited by 0SourcePDFScholar
2025

A Survey on Multi-modal Intent Recognition: Recent Advances and New Frontiers

EMNLP 2025

Multi-modal intent recognition (MIR) requires integrating non-verbal cues from real-world contexts to enhance human intention understanding, which has attracted substantial research attention in recent years. Despite promising advancements, a comprehensive survey summarizing recent advances and new

2025

AnchorCoT: Anchors Pave the Way for Multi-hop Reasoning

ACL 2025finding

Large Language Models (LLMs) have made substantial strides in a broad array of natural language tasks. Recently, LLMs have demonstrated potential reasoning capabilities through prompt design, such as the Chain of Thought (CoT). Despite their superiority in question answering, LLMs still face challen…

Cited by 0SourcePDFScholar
2025

Can We Trust AI Doctors? A Survey of Medical Hallucination in Large Language and Large Vision-Language Models

ACL 2025finding

Hallucination has emerged as a critical challenge for large language models (LLMs) and large vision-language models (LVLMs), particularly in high-stakes medical applications. Despite its significance, dedicated research on medical hallucination remains unexplored. In this survey, we first provide a…

Cited by 0SourcePDFScholar
2025

CellVerse: Do Large Language Models Really Understand Cell Biology?

NeurIPS 2025poster

Recent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis ta…

Cited by 0SourcecodeScholar
2025

D-VST: Diffusion Transformer for Pathology-Correct Tone-Controllable Cross-Dye Virtual Staining of Whole Slide Images

NeurIPS 2025poster

Diffusion-based virtual staining methods of histopathology images have demonstrated outstanding potential for stain normalization and cross-dye staining (e.g., hematoxylin-eosin to immunohistochemistry). However, achieving pathology-correct cross-dye virtual staining with versatile tone controls pos…

Cited by 0SourceScholar
2025

Degradation-Aware Dynamic Schrödinger Bridge for Unpaired Image Restoration

NeurIPS 2025poster

Image restoration is a fundamental task in computer vision and machine learning, which learns a mapping between the clear images and the degraded images under various conditions (e.g., blur, low-light, haze). Yet, most existing image restoration methods are highly restricted by the requirement of de…

Cited by 0SourceScholar
2025

Guiding Large Language Models for Biomedical Entity Linking via Restrictive and Contrastive Decoding

EMNLP 2025

Biomedical entity linking (BioEL) aims at mapping biomedical mentions to pre-defined entities. While extensive research efforts have been devoted to BioEL, applying large language models (LLMs) for BioEL has not been fully explored. Previous attempts have revealed difficulties when directly applying

Cited by 0SourcePDFScholar
2025

Harnessing Large Language Models for Knowledge Graph Question Answering via Adaptive Multi-Aspect Retrieval-Augmentation

AAAI 2025technical

Large Language Models (LLMs) demonstrate remarkable capabilities, yet struggle with hallucination and outdated knowledge when tasked with complex knowledge reasoning, resulting in factually incorrect outputs. Previous studies have attempted to mitigate it by retrieving factual knowledge from large-s…

2025

LLMEmb: Large Language Model Can Be a Good Embedding Generator for Sequential Recommendation

AAAI 2025technical

Sequential Recommender Systems (SRS), which model a user's interaction history to predict the next item of interest, are widely used in various applications. However, existing SRS often struggle with low-popularity items, a challenge known as the long-tail problem. This issue leads to reduced serend…

2025

Learning a Cross-Modal Schrödinger Bridge for Visual Domain Generalization

NeurIPS 2025poster

Domain generalization aims to train models that perform robustly on unseen target domains without access to target data. The realm of vision-language foundation model has opened a new venue owing to its inherent out-of-distribution generalization capability. However, the static alignment to class-l…

Cited by 0SourceScholar
2025

MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models

EMNLP 2025

Memes have emerged as a popular form of multimodal online communication, where their interpretation heavily depends on the specific context in which they appear. Current approaches predominantly focus on isolated meme analysis, either for harmful content detection or standalone interpretation, overl

Cited by 0SourcePDFScholar
2025

Monte Carlo Tree Search Based Prompt Autogeneration for Jailbreak Attacks against LLMs

COLING 2025main

Jailbreak attacks craft specific prompts or append adversarial suffixes to prompts, thereby inducing language models to generate harmful or unethical content and bypassing the model’s safety guardrails. With the recent blossom of large language models (LLMs), there’s a growing focus on jailbreak att…

2025

NightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene Segmentation

CVPR 2025poster

Night-time scene segmentation is a critical yet challenging task in the real-world applications, primarily due to the complicated lighting conditions. However, existing methods lack sufficient generalization ability to unseen nigh-time scenes with varying illumination.In light of this issue, we focu…

2025

T2: An Adaptive Test-Time Scaling Strategy for Contextual Question Answering

EMNLP 2025

Recent advances in large language models have demonstrated remarkable performance on Contextual Question Answering (CQA). However, prior approaches typically employ elaborate reasoning strategies regardless of question complexity, leading to low adaptability. Recent efficient test-time scaling metho

2025

Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens

NeurIPS 2025spotlight

The recent rise of Large Reasoning Models (LRMs) has significantly improved multi-step reasoning performance, but often at the cost of generating excessively long reasoning chains. This paper revisits the efficiency of such reasoning processes through an information-theoretic lens, revealing a funda…

Cited by 0SourcecodeScholar
2025

Training-free LLM Merging for Multi-task Learning

ACL 2025long

Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse natural language processing (NLP) tasks. The release of open-source LLMs like LLaMA and Qwen has triggered the development of numerous fine-tuned models tailored for various tasks and languages. In this paper, we…

2025

TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice

NAACL 2025industry

Jailbreaking large-language models (LLMs) involves testing their robustness against adversarial prompts and evaluating their ability to withstand prompt attacks that could elicit unauthorized or malicious responses. In this paper, we present TurboFuzzLLM, a mutation-based fuzzing technique for effic…

2024

Alignment before Awareness: Towards Visual Question Localized-Answering in Robotic Surgery via Optimal Transport and Answer Semantics

COLING 2024main

The visual question localized-answering (VQLA) system has garnered increasing attention due to its potential as a knowledgeable assistant in surgical education. Apart from providing text-based answers, VQLA can also pinpoint the specific region of interest for better surgical scene understanding. Al…

2024

Biomedical Entity Linking as Multiple Choice Question Answering

COLING 2024main

Although biomedical entity linking (BioEL) has made significant progress with pre-trained language models, challenges still exist for fine-grained and long-tailed entities. To address these challenges, we present BioELQA, a novel model that treats Biomedical Entity Linking as Multiple Choice Questio…

2024

Can LLMs Replace Clinical Doctors? Exploring Bias in Disease Diagnosis by Large Language Models

EMNLP 2024finding

The bias of disease prediction in Large Language Models (LLMs) is a critical yet underexplored issue, with potential implications for healthcare outcomes and equity. As LLMs increasingly find applications in healthcare, understanding and addressing their biases becomes paramount. This study focuses…

Cited by 0SourcePDFScholar
2024

Cooper: Coordinating Specialized Agents towards a Complex Dialogue Goal

AAAI 2024technical

In recent years, there has been a growing interest in exploring dialogues with more complex goals, such as negotiation, persuasion, and emotional support, which go beyond traditional service-focused dialogue systems. Apart from the requirement for much more sophisticated strategic reasoning and comm…

2024

Fast-Poly: A Fast Polyhedral Algorithm for 3D Multi-Object Tracking

RA-L 2024

3D Multi-Object Tracking (MOT) captures stable and comprehensive motion states of surrounding obstacles, essential for robotic perception. However, current 3D trackers face issues with accuracy and latency consistency. In this letter, we propose Fast-Poly, a fast and effective filter-based method fo

Cited by 16SourceScholar
2024

Guiding Clinical Reasoning with Large Language Models via Knowledge Seeds

IJCAI 2024poster

Clinical reasoning refers to the cognitive process that physicians employ in evaluating and managing patients. This process typically involves suggesting necessary examinations, diagnosing patients’ diseases, and selecting appropriate therapies, etc. Accurate clinical reasoning requires extensive me…

2024

Improving Biomedical Entity Linking with Retrieval-Enhanced Learning

ICASSP 2024accepted

Biomedical entity linking (BioEL) has achieved remarkable progress with the help of pre-trained language models. However, existing BioEL methods usually struggle to handle rare and difficult entities due to long-tailed distribution. To address this limitation, we introduce a new scheme kNN-BioEL, wh…

Cited by 0SourceScholar
2024

JoTR: A Joint Transformer and Reinforcement Learning Framework for Dialogue Policy Learning

COLING 2024main

Dialogue policy learning (DPL) aims to determine an abstract representation (also known as action) to guide what the response should be. Typically, DPL is cast as a sequential decision problem across a series of predefined action candidates. However, such static and narrow actions can limit response…

2024

Knowledge-aware Attention Network for Medication Effectiveness Prediction

COLING 2024main

The first 24 hours’ medication plan is critical to patients with serious or life-threatening illnesses and injuries. An appropriate medication can result in a lower mortality, a shorter length stay and a higher APACHE score. However, in clinical practice, the medication plan is often error-prone, es…

Cited by 0SourcePDFScholar
2024

LLM-ESR: Large Language Models Enhancement for Long-tailed Sequential Recommendation

NeurIPS 2024spotlight

Sequential recommender systems (SRS) aim to predict users' subsequent choices based on their historical interactions and have found applications in diverse fields such as e-commerce and social media. However, in real-world systems, most users interact with only a handful of items, while the majority…

Cited by 10SourcePDFScholar
2024

LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning

EMNLP 2024finding

Chain-of-thought (CoT) prompting is a popular in-context learning (ICL) approach for large language models (LLMs), especially when tackling complex reasoning tasks. Traditional ICL approaches construct prompts using examples that contain questions similar to the input question. However, CoT promptin…

2024

MKeCL: Medical Knowledge-Enhanced Contrastive Learning for Few-shot Disease Diagnosis

COLING 2024main

Artificial intelligence (AI)-aided disease prediction has gained extensive research interest due to its capability to support clinical decision-making. Existing works mainly formulate disease prediction as a multi-label classification problem and use historical Electronic Medical Records (EMR) to tr…

Cited by 2SourcePDFScholar
2024

MedJourney: Benchmark and Evaluation of Large Language Models over Patient Clinical Journey

NeurIPS 2024poster

Large language models (LLMs) have demonstrated remarkable capabilities in language understanding and generation, leading to their widespread adoption across various fields. Among these, the medical field is particularly well-suited for LLM applications, as many medical tasks can be enhanced by LLMs.…

Cited by 1SourcePDFScholar
2024

Mitigating Hallucinations of Large Language Models in Medical Information Extraction via Contrastive Decoding

EMNLP 2024finding

The impressive capabilities of large language models (LLMs) have attracted extensive interests of applying LLMs to medical field. However, the complex nature of clinical environments presents significant hallucination challenges for LLMs, hindering their widespread adoption. In this paper, we addres…

2024

Multi-perspective Improvement of Knowledge Graph Completion with Large Language Models

COLING 2024main

Knowledge graph completion (KGC) is a widely used method to tackle incompleteness in knowledge graphs (KGs) by making predictions for missing links. Description-based KGC leverages pre-trained language models to learn entity and relation representations with their names or descriptions, which shows…

2024

Pre-trained Online Contrastive Learning for Insurance Fraud Detection

AAAI 2024technical

Medical insurance fraud has always been a crucial challenge in the field of healthcare industry. Existing fraud detection models mostly focus on offline learning scenes. However, fraud patterns are constantly evolving, making it difficult for models trained on past data to detect newly emerging frau…

2024

RICE: Breaking Through the Training Bottlenecks of Reinforcement Learning with Explanation

ICML 2024spotlight

Deep reinforcement learning (DRL) is playing an increasingly important role in real-world applications. However, obtaining an optimally performing DRL agent for complex tasks, especially with sparse rewards, remains a significant challenge. The training of a DRL agent can be often trapped in a bottl…

2024

Safeguarding Fraud Detection from Attacks: A Robust Graph Learning Approach

IJCAI 2024poster

Financial fraud is one of the most significant social issues and has caused tremendous property losses. Graph neural networks (GNNs) have been applied to anti-fraud practices and achieved decent results. However, recent researches have discovered flaws in the robustness of fraud-detection models bas…

Cited by 7SourcePDFScholar
2024

Soft-Label Integration for Robust Toxicity Classification

NeurIPS 2024poster

Toxicity classification in textual content remains a significant problem. Data with labels from a single annotator fall short of capturing the diversity of human perspectives. Therefore, there is a growing need to incorporate crowdsourced annotations for training an effective toxicity classifier. Ad…

2024

TFCD: Towards Multi-modal Sarcasm Detection via Training-Free Counterfactual Debiasing

IJCAI 2024poster

Multi-modal sarcasm detection (MSD), which aims to identify whether a given sample with multi-modal information (i.e., text and image) is sarcastic, has garnered widespread attention. Recent approaches focus on designing sophisticated architectures or mechanisms to extract sarcastic cues from entire…

Cited by 10SourcePDFScholar
2023

Dialogue Medical Information Extraction with Medical-Item Graph and Dialogue-Status Enriched Representation

EMNLP 2023long findings

The multi-turn doctor-patient dialogue includes rich medical knowledge, like the symptoms of the patient, the diagnosis and medication suggested by the doctor. If mined and represented properly, such medical knowledge can benefit a large range of clinical applications, including diagnosis assistance…

Cited by 0SourceScholar
2023

MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning

ACL 2023long

Supervised visual captioning models typically require a large scale of images or videos paired with descriptions in a specific language (i.e., the vision-caption pairs) for training. However, collecting and labeling large-scale datasets is time-consuming and expensive for many scenarios and language…

2023

StateMask: Explaining Deep Reinforcement Learning through State Mask

NeurIPS 2023poster

Despite the promising performance of deep reinforcement learning (DRL) agents in many challenging scenarios, the black-box nature of these agents greatly limits their applications in critical domains. Prior research has proposed several explanation techniques to understand the deep learning-based po…

Cited by 11SourcePDFScholar
2022

DeltaNet: Conditional Medical Report Generation for COVID-19 Diagnosis

COLING 2022main

Fast screening and diagnosis are critical in COVID-19 patient treatment. In addition to the gold standard RT-PCR, radiological imaging like X-ray and CT also works as an important means in patient screening and follow-up. However, due to the excessive number of patients, writing reports becomes a he…

2022

Denoising Neural Network for News Recommendation with Positive and Negative Implicit Feedback

NAACL 2022findings

News recommendation is different from movie or e-commercial recommendation as people usually do not grade the news. Therefore, user feedback for news is always implicit (click behavior, reading time, etc). Inevitably, there are noises in implicit feedback. On one hand, the user may exit immediately…

2022

End-to-end Spoken Conversational Question Answering: Task, Dataset and Model

NAACL 2022findings

In spoken question answering, the systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that human seek or test their knowledge is via human conversations. Therefore, we propose a new Spoken Conversational Question An…

Cited by 37SourcePDFScholar
2022

Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations

NeurIPS 2022accept

Most video-and-language representation learning approaches employ contrastive learning, e.g., CLIP, to project the video and text features into a common latent space according to the semantic similarities of text-video pairs. However, such learned shared latent spaces are not often optimal, and the…

2022

Multi-modal Contrastive Representation Learning for Entity Alignment

COLING 2022main

Multi-modal entity alignment aims to identify equivalent entities between two different multi-modal knowledge graphs, which consist of structural triples and images associated with entities. Most previous works focus on how to utilize and encode information from different modalities, while it is not…

2022

Retrieve, Reason, and Refine: Generating Accurate and Faithful Patient Instructions

NeurIPS 2022accept

The "Patient Instruction" (PI), which contains critical instructional information provided both to carers and to the patient at the time of discharge, is essential for the patient to manage their condition outside hospital. An accurate and easy-to-follow PI can improve the self-management of patient…

2022

Towards Joint Intent Detection and Slot Filling via Higher-order Attention

IJCAI 2022poster

Recently, attention-based models for joint intent detection and slot filling have achieved state-of-the-art performance. However, we think the conventional attention can only capture the first-order feature interaction between two tasks and is insufficient. To address this issue, we propose a unifie…

2021

Audio-Oriented Multimodal Machine Comprehension via Dynamic Inter- and Intra-modality Attention

AAAI 2021technical

While Machine Comprehension (MC) has attracted extensive research interests in recent years, existing approaches mainly belong to the category of Machine Reading Comprehension task which mines textual inputs (paragraphs and questions) to predict the answers (choices or text spans). However, there ar…

Cited by 29SourcePDFScholar
2021

Auto-Encoding Knowledge Graph for Unsupervised Medical Report Generation

NeurIPS 2021poster

Medical report generation, which aims to automatically generate a long and coherent report of a given medical image, has been receiving growing research interests. Existing approaches mainly adopt a supervised manner and heavily rely on coupled image-report pairs. However, in the medical domain, bui…

Cited by 135SourcePDFScholar
2021

BACKDOORL: Backdoor Attack against Competitive Reinforcement Learning

IJCAI 2021poster

Recent research has confirmed the feasibility of backdoor attacks in deep reinforcement learning (RL) systems. However, the existing attacks require the ability to arbitrarily modify an agent's observation, constraining the application scope to simple RL systems such as Atari games. In this paper, w…

2021

Competence-based Multimodal Curriculum Learning for Medical Report Generation

ACL 2021long

Medical report generation task, which targets to produce long and coherent descriptions of medical images, has attracted growing research interests recently. Different from the general image captioning tasks, medical report generation is more challenging for data-driven neural models. This is mainly…

2021

Exploring and Distilling Posterior and Prior Knowledge for Radiology Report Generation

CVPR 2021poster

Automatically generating radiology reports can improve current clinical practice in diagnostic radiology. On one hand, it can relieve radiologists from the heavy burden of report writing; On the other hand, it can remind radiologists of abnormalities and avoid the misdiagnosis and missed diagnosis.…

Cited by 393PDFScholar
2021

U-BERT: Pre-training User Representations for Improved Recommendation

AAAI 2021technical

Learning user representation is a critical task for recommendation systems as it can encode user preference for personalized services. User representation is generally learned from behavior data, such as clicking interactions and review comments. However, for less popular domains, the behavior data…

Cited by 152SourcePDFScholar
2020

Automatic Distractor Generation for Multiple Choice Questions in Standard Tests

COLING 2020main

To assess knowledge proficiency of a learner, multiple choice question is an efficient and widespread form in standard tests. However, the composition of the multiple choice question, especially the construction of distractors is quite challenging. The distractors are required to both incorrect and…

2020

Bridging Mixture Density Networks with Meta-Learning for Automatic Speaker Identification

ICASSP 2020accepted

Speaker identification answers the fundamental question "Who is speaking" The identification technology enables various downstream applications to provide a personalized experience. Both the prevalent i-vector based solutions and the state-of-the-art deep learning solutions usually treat all users e…

Cited by 0SourceScholar
2020

Least Squares Regression with Markovian Data: Fundamental Limits and Algorithms

NeurIPS 2020spotlight

We study the problem of least squares linear regression where the datapoints are dependent and are sampled from a Markov chain. We establish sharp information theoretic minimax lower bounds for this problem in terms of $\tmix$, the mixing time of the underlying Markov chain, under different noise se…

Cited by 86SourcePDFScholar
2020

Prophet Attention: Predicting Attention with Future Attention

NeurIPS 2020poster

Recently, attention based models have been used extensively in many sequence-to-sequence learning systems. Especially for image captioning, the attention based models are expected to ground correct image regions with proper generated words. However, for each time step in the decoding process, the at…

Cited by 77SourcePDFScholar
2018

Interpretable Video Captioning via Trajectory Structured Localization

CVPR 2018poster

Automatically describing open-domain videos with natural language are attracting increasing interest in the field of artificial intelligence. Most existing methods simply borrow ideas from image captioning and obtain a compact video representation from an ensemble of global image feature before feed…

Cited by 70SourcePDFScholar
2018

Near-Optimal Time and Sample Complexities for Solving Markov Decision Processes with a Generative Model

NeurIPS 2018poster

In this paper we consider the problem of computing an $\epsilon$-optimal policy of a discounted Markov Decision Process (DMDP) provided we can only access its transition function through a generative sampling model that given any state-action pair samples from the transition function in $O(1)$ time.…

Cited by 269SourcePDFScholar