← Search

Fan Zhou

103 accepted papers

2026

AOEPT: Breaking the Implicit Modality-Reduction Bottleneck in Modality Missing Prompt Tuning

ICML 2026poster

Deploying multimodal systems in real-world environments often entails handling modality-missing scenarios, where one or more modalities are unavailable. While recent studies address this challenge for the general Multimodal Transformer (MT) architecture via prompt tuning, we identify a fundamental l…

Cited by 0SourceScholar
2026

Beyond Graph Priors: A Co-Evolving Framework Under Uncertainty for Enterprise Resilience Assessment

AAAI 2026technical

Assessing enterprise resilience under uncertainty necessitates capturing both intrinsic attributes and evolving inter-enterprise dependencies. However, real-world enterprise systems pose substantial structural challenges: redundant or loosely correlated links can trigger spurious relational inferenc

Cited by 0SourcePDFScholar
2026

Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs

AAAI 2026technical

Object hallucination remains a critical challenge in Large Vision-Language Models (LVLMs), where models generate content inconsistent with visual inputs. Existing language-decoder based mitigation approaches often regulate visual or textual attention independently, overlooking their interaction as t

Cited by 0SourcePDFScholar
2026

Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding

ICML 2026poster

Large Vision-Language Models (LVLMs) exhibit sophisticated reasoning but remain susceptible to object hallucination. Deviating from the prevailing attention intensity assumption, we reveal a deeper dynamic structural misalignment: hallucination is triggered at decision-critical steps where specific …

Cited by 0SourceScholar
2026

JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

ICLR 2026poster

This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Trans- former designed for synchronized audio-video generation (JAVG). Based on the powerful Diffusion Transformer (DiT) architecture, JavisDiT simultaneously generates high-quality audio and video content from open-ended user promp…

Cited by 0SourcecodeScholar
2026

Learning to Curate Context: Jointly Optimizing Retrieval and Prediction for Multimodal Social Media Popularity

AAAI 2026technical

Predicting the popularity of user-generated content (UGC) is a crucial but challenging task in social media analysis. While existing retrieval-augmented models enhance predictions by supplying rich contextual information, they remain limited by a fundamental precision-recall dilemma: enlarging the r

Cited by 0SourcePDFScholar
2026

MS-CRL: Multi-Scale Global Path Planning with Progressive Curriculum Reinforcement Learning

ICRA 2026poster

Global path planning provides high-level guidance for autonomous navigation, supplying reference paths for downstream navigation and control modules. Deep Reinforcement Learning (DRL) has shown strong potential in this domain, but existing methods struggle with multi-scale map inputs. This limitatio…

Cited by 0Scholar
2026

Modality-Balanced Collaborative Distillation for Multi-Modal Domain Generalization

AAAI 2026technical

Weight Averaging (WA) has emerged as a powerful technique for enhancing generalization by promoting convergence to a flat loss landscape, which correlates with stronger out-of-distribution performance. However, applying WA directly to multi-modal domain generalization (MMDG) is challenging: differen

Cited by 0SourcePDFScholar
2026

No More Shortcuts: Network Traffic Anomaly Detection via Bidirectional Prediction

IJCAI 2026

Network Traffic Anomaly Detection (NTAD), particularly under zero-positive settings, is a critical task in cybersecurity. Existing zero-positive NTAD approaches primarily rely on reconstruction-based pipelines. Nevertheless, these methods are susceptible to an identical shortcut issue, where models

Cited by 0Scholar
2026

Self-Consistency Improves the Trustworthiness of Self-Interpretable GNNs

ICLR 2026poster

Graph Neural Networks (GNNs) achieve strong predictive performance but offer limited transparency in their decision-making. Self-Interpretable GNNs (SI-GNNs) address this by generating built-in explanations, yet their training objectives are misaligned with evaluation criteria such as faithfulness.…

Cited by 0SourceScholar
2026

Shedding the Facades, Connecting the Domains: Detecting Shifting Multimodal Hate Video with Test-Time Adaptation

AAAI 2026technical

Hate Video Detection (HVD) is crucial for online ecosystems. Existing methods assume identical distributions between training (source) and inference (target) data. However, hateful content often evolves into irregular and ambiguous forms to evade censorship, resulting in substantial semantic drift a

Cited by 0SourcePDFScholar
2026

The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution

ICLR 2026poster

Real-world language agents must handle complex, multi-step workflows across diverse applications. For instance, an agent may manage emails by coordinating with calendars and file systems, or monitor a production database like BigQuery to detect anomalies and generate reports following a standard ope…

Cited by 0SourcecodeScholar
2025

Borrowing Eyes for the Blind Spot: Overcoming Data Scarcity in Malicious Video Detection via Cross-Domain Retrieval Augmentation

ICCV 2025poster

The rapid proliferation of online video-sharing platforms has accelerated the spread of malicious videos, creating an urgent need for robust detection methods. However, the performance and generalizability of existing detection approaches are severely limited by the scarcity of annotated video data,…

2025

Bridging the Fairness Gap: Enhancing Pre-trained Models with LLM-Generated Sentences

ICASSP 2025accepted

Pre-trained language models (PLMs) are trained on data that inherently contains gender biases, leading to undesirable impacts. Traditional debiasing methods often rely on external corpora, which may lack quality, diversity, or demographic balance, affecting the effectiveness of debiasing. With the r…

Cited by 0SourceScholar
2025

Commonality Augmented Disentanglement for Multimodal Crowdfunding Success Prediction

ICASSP 2025accepted

Online crowdfunding platforms have been gaining increasing popularity due to their convenience in soliciting social capital from the public. These platforms offer valuable opportunities for fundraisers to bring their creative products to life and support pro-social projects. However, the relatively…

Cited by 0SourceScholar
2025

Diving into Self-Evolving Training for Multimodal Reasoning

ICML 2025poster

Self-evolving training—where models iteratively learn from their own outputs—has emerged as a key approach for complex reasoning tasks, addressing the scarcity of high-quality chain-of-thought data. However, its effectiveness in multimodal reasoning, a domain more intricate than text-only reasoning,…

Cited by 0SourcePDFScholar
2025

Enhancing Prediction Performance through Influence Measure

ICLR 2025poster

In the field of machine learning, the pursuit of accurate models is ongoing. A key aspect of improving prediction performance lies in identifying which data points in the training set should be excluded and which high-quality, potentially unlabeled data points outside the training set should be inco…

Cited by 0SourcePDFScholar
2025

Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference

ICCV 2025poster

Video Multimodal Large Language Models (Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the underlying LLMs, existing Video-MLLMs typically exhibit suboptimal performance on long video scenarios. To understand e…

2025

Improving Multimodal Social Media Popularity Prediction via Selective Retrieval Knowledge Augmentation

AAAI 2025technical

Understanding and predicting the popularity of online User-Generated Content (UGC) is critical for various social and recommendation systems. Existing efforts have focused on extracting predictive features and using pre-trained deep models to learn and fuse multimodal UGC representations. However, t…

2025

In-context Prompt-augmented Micro-video Popularity Prediction

AAAI 2025technical

Micro-video popularity prediction (MVPP) plays a crucial role in various downstream applications. Recently, multimodal methods that integrate multiple modalities to predict the popularity have exhibited impressive performance. However, these methods face several unresolved issues: (1) limited contex…

2025

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation

NeurIPS 2025spotlight

This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for Joint Audio-Video (JAV) comprehension and generation. JavisGPT adopts a concise encoder–LLM–decoder architecture, featuring a SyncFusion module for spatio-temporal audio- video fusion and synchrony-aware learn…

Cited by 0SourceScholar
2025

Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale

ICML 2025poster

Large language model pre-training has traditionally relied on human experts to craft heuristics for improving the corpora quality, resulting in numerous rules developed to date. However, these fixed rules lack the flexibility to address the unique characteristics of individual examples, yet crafting…

2025

Redundancy Undermines the Trustworthiness of Self-Interpretable GNNs

ICML 2025poster

This work presents a systematic investigation into the trustworthiness of explanations generated by self-interpretable graph neural networks (GNNs), revealing why models trained with different random seeds yield inconsistent explanations. We identify redundancy—resulting from weak conciseness constr…

2025

Retrieval-Augmented Dynamic Prompt Tuning for Incomplete Multimodal Learning

AAAI 2025technical

Multimodal learning with incomplete modality is practical and challenging. Recently, researchers have focused on enhancing the robustness of pre-trained MultiModal Transformers (MMTs) under missing modality conditions by applying learnable prompts. However, these prompt-based methods face several li…

2025

Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

NeurIPS 2025poster

Reinforcement learning (RL) has shown promise in enhancing large language model (LLM) reasoning, yet progress towards broader capabilities is limited by the availability of high-quality, multi-domain datasets. This work introduces \ours, a 92K RL-for-reasoning dataset designed to address this gap, c…

Cited by 0SourceScholar
2025

Structure-aware Domain Knowledge Injection for Large Language Models

ACL 2025long

This paper introduces a pioneering methodology, termed StructTuning, to efficiently transform foundation Large Language Models (LLMs) into domain specialists. It significantly reduces the training corpus needs to a mere 5% while achieving an impressive 100% of traditional knowledge injection perform…

2025

Versatile Transferable Unlearnable Example Generator

NeurIPS 2025poster

The rapid growth of publicly available data has fueled deep learning advancements but also raises concerns about unauthorized data usage. Unlearnable Examples (UEs) have emerged as a data protection strategy that introduces imperceptible perturbations to prevent unauthorized learning. However, most…

Cited by 0SourcecodeScholar
2025

You Only Query Twice: Multimodal Rumor Detection via Evidential Evaluation from Dual Perspectives

COLING 2025main

Current rumor detectors exhibit limitations in fully exploiting responses to the source tweet as essential public opinions, and in explaining and indicating the reliability of the results obtained. Additionally, the joint utilization of both responses and the multimodal source content for detection…

Cited by 0SourcePDFScholar
2024

Amplifying Diversity and Quality in Commonsense Knowledge Graph Completion (Student Abstract)

AAAI 2024technical

Conventional commonsense knowledge graph completion (CKGC) methods provide inadequate sequence when fine-tuning or generating stages and incorporate full fine-tuning, which fail to align with the autoregressive model's pre-training patterns and have insufficient parameter efficiency. Moreover, decod…

Cited by 2SourcePDFScholar
2024

Biases Mitigation and Expressiveness Preservation in Language Models: A Comprehensive Pipeline (Student Abstract)

AAAI 2024technical

Pre-trained language models (PLMs) have greatly transformed various downstream tasks, yet frequently display social biases from training data, raising fairness concerns. Recent efforts to debias PLMs come with limitations: they either fine-tune the entire parameters in PLMs, which is time-consuming…

Cited by 3SourcePDFScholar
2024

Decoupling User Relationships Guides Information Diffusion Prediction (Student Abstract)

AAAI 2024technical

Information diffusion prediction is a critical task for many social network applications. However, current methods are mainly limited by the following aspects: user relationships behind resharing behaviors are complex and entangled. To address these issues, we propose MHGFormer, a novel multi-channe…

Cited by 0SourcePDFScholar
2024

Deforming Garment Classification With Shallow Temporal Extraction and Tree-Based Fusion

RA-L 2024

A novel RGB-based continuous perception garment classification approach is proposed in this letter, with the aim of identifying the correct category of the garment from a set of categories. It has been observed that treating a video of the continuous deformation of cloth as a set of disordered stati

Cited by 2SourceScholar
2024

Disentanglement-Guided Spatial-Temporal Graph Neural Network for Metro Flow Forecasting (Student Abstract)

AAAI 2024technical

In recent intelligent transportation applications, metro flow forecasting has received much attention from researchers. Most prior arts endeavor to explore spatial or temporal dependencies while ignoring the key characteristic patterns underlying historical flows, e.g., trend and periodicity. Althou…

Cited by 0SourcePDFScholar
2024

EasyTPP: Towards Open Benchmarking Temporal Point Processes

ICLR 2024poster

Continuous-time event sequences play a vital role in real-world domains such as healthcare, finance, online shopping, social networks, and so on. To model such data, temporal point processes (TPPs) have emerged as the most natural and competitive models, making a significant impact in both academic…

2024

Enhancing Fine-Grained Urban Flow Inference via Incremental Neural Operator

IJCAI 2024poster

Fine-grained urban flow inference (FUFI), which involves inferring fine-grained flow maps from their coarse-grained counterparts, is of tremendous interest in the realm of sustainable urban traffic services. To address the FUFI, existing solutions mainly concentrate on investigating spatial dependen…

2024

Enhancing LLM’s Cognition via Structurization

NeurIPS 2024poster

When reading long-form text, human cognition is complex and structurized. While large language models (LLMs) process input contexts through a causal and sequential perspective, this approach can potentially limit their ability to handle intricate and complex inputs effectively. To enhance LLM’s cogn…

2024

Exploring Self-Explainable Street-Level IP Geolocation with Graph Information Bottleneck

ICASSP 2024accepted

Accurate IP geolocation is crucial for location-aware applications. While recent advances in router-centric IP graph methods have garnered attention, they face two persistent challenges: (1) the sparsity problem of IP graphs in rural areas and (2) the limited explainability of current IP geolocation…

Cited by 0SourceScholar
2024

GMP-AR: Granularity Message Passing and Adaptive Reconciliation for Temporal Hierarchy Forecasting

AAAI 2024technical

Time series forecasts of different temporal granularity are widely used in real-world applications, e.g., sales prediction in days and weeks for making different inventory plans. However, these tasks are usually solved separately without ensuring coherence, which is crucial for aligning downstream d…

Cited by 0SourcePDFScholar
2024

Generalizing across Temporal Domains with Koopman Operators

AAAI 2024technical

In the field of domain generalization, the task of constructing a predictive model capable of generalizing to a target domain without access to target data remains challenging. This problem becomes further complicated when considering evolving dynamics between domains. While various approaches have…

Cited by 6SourcePDFScholar
2024

Graph Anomaly Detection with Diffusion Model-Based Graph Enhancement (Student Abstract)

AAAI 2024technical

Graph anomaly detection has gained significant research interest across various domains. Due to the lack of labeled data, contrastive learning has been applied in detecting anomalies and various scales of contrastive strategies have been initiated. However, these methods might force two instances (e…

Cited by 2SourcePDFScholar
2024

Interpreting Temporal Knowledge Graph Reasoning (Student Abstract)

AAAI 2024technical

Temporal knowledge graph reasoning is an essential task that holds immense value in diverse real-world applications. Existing studies mainly focus on leveraging structural and sequential dependencies, excelling in tasks like entity and link prediction. However, they confront a notable interpretabili…

Cited by 2SourcePDFScholar
2024

Lemur: Harmonizing Natural Language and Code for Language Agents

ICLR 2024spotlight

We introduce Lemur and Lemur-Chat, openly accessible language models optimized for both natural language and coding capabilities to serve as the backbone of versatile language agents. The evolution from language chat models to functional language agents demands that models not only master human inte…

2024

MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection

ECCV 2024poster

"Learning from pseudo-labels that generated with VLMs (Vision Language Models) has been shown as a promising solution to assist open vocabulary detection (OVD) in recent studies. However, due to the domain gap between VLM and vision-detection tasks, pseudo-labels produced by the VLMs are prone to be…

2024

Natural Evolution-based Dual-Level Aggregation for Temporal Knowledge Graph Reasoning

EMNLP 2024finding

Temporal knowledge graph (TKG) reasoning aims to predict missing facts based on a given history. Most of the existing methods unifiedly model the evolution process of different events and ignore their inherent asynchronous characteristics, resulting in suboptimal performance. To tackle this challeng…

Cited by 0SourcePDFScholar
2024

OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

NeurIPS 2024poster

The evolution of Artificial Intelligence (AI) has been significantly accelerated by advancements in Large Language Models (LLMs) and Large Multimodal Models (LMMs), gradually showcasing potential cognitive reasoning abilities in problem-solving and scientific discovery (i.e., AI4Science) once exclus…

2024

Rethinking Out-of-Distribution Detection on Imbalanced Data Distribution

NeurIPS 2024poster

Detecting and rejecting unknown out-of-distribution (OOD) samples is critical for deployed neural networks to void unreliable predictions. In real-world scenarios, however, the efficacy of existing OOD detection methods is often impeded by the inherent imbalance of in-distribution (ID) data, which c…

2024

Revitalizing Real Image Deraining via a Generic Paradigm towards Multiple Rainy Patterns

IJCAI 2024poster

Synthetic data-driven methods perform well on image rain removal task, but they still face many challenges in real rainfall scenarios due to the complexity and diversity of rainy patterns. In this paper, we propose a new generic paradigm for real image deraining from the perspective of synthesizing…

Cited by 2SourcePDFScholar
2024

Shallow Diffusion for Fast Speech Enhancement (Student Abstract)

AAAI 2024technical

Recently, the field of Speech Enhancement has witnessed the success of diffusion-based generative models. However, these diffusion-based methods used to take multiple iterations to generate high-quality samples, leading to high computational costs and inefficiency. In this paper, we propose SDFEN (S…

Cited by 0SourcePDFScholar
2024

Spatial-Temporal Augmentation for Crime Prediction (Student Abstract)

AAAI 2024technical

Crime prediction stands as a pivotal concern within the realm of urban management due to its potential threats to public safety. While prior research has predominantly focused on unraveling the intricate dependencies among urban regions and temporal dynamics, the challenges posed by the scarcity and…

Cited by 1SourcePDFScholar
2024

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

NeurIPS 2024poster

This paper studies off-policy evaluation (OPE) in the presence of unmeasured confounders. Inspired by the two-way fixed effects regression model widely used in the panel data literature, we propose a two-way unmeasured confounding assumption to model the system dynamics in causal reinforcement learn…

2023

A Probabilistic Graph Diffusion Model for Source Localization (Student Abstract)

AAAI 2023technical

Source localization, as a reverse problem of graph diffusion, is important for many applications such as rumor tracking, detecting computer viruses, and finding epidemic spreaders. However, it is still under-explored due to the inherent uncertainty of the diffusion process: after a long period of pr…

Cited by 1SourcePDFScholar
2023

CasODE: Modeling Irregular Information Cascade via Neural Ordinary Differential Equations (Student Abstract)

AAAI 2023technical

Predicting information cascade popularity is a fundamental problem for understanding the nature of information propagation on social media. However, existing works fail to capture an essential aspect of information propagation: the temporal irregularity of cascade event -- i.e., users' re-tweetings…

Cited by 1SourcePDFScholar
2023

Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant Learning

ACL 2023long

Demographic biases and social stereotypes are common in pretrained language models (PLMs), and a burgeoning body of literature focuses on removing the unwanted stereotypical associations from PLMs. However, when fine-tuning these bias-mitigated PLMs in downstream natural language processing (NLP) ap…

Cited by 40SourcePDFScholar
2023

DOSE: Diffusion Dropout with Adaptive Prior for Speech Enhancement

NeurIPS 2023poster

Speech enhancement (SE) aims to improve the intelligibility and quality of speech in the presence of non-stationary additive noise. Deterministic deep learning models have traditionally been used for SE, but recent studies have shown that generative approaches, such as denoising diffusion probabilis…

2023

De-biased Teacher: Rethinking IoU Matching for Semi-supervised Object Detection

AAAI 2023technical

Most of the recent research in semi-supervised object detection follows the pseudo-labeling paradigm evolved from the semi-supervised image classification task. However, the training paradigm of the two-stage object detector inevitably makes the pseudo-label learning process for unlabeled images ful…

2023

Debiasing Intrinsic Bias and Application Bias Jointly via Invariant Risk Minimization (Student Abstract)

AAAI 2023technical

Demographic biases and social stereotypes are common in pretrained language models (PLMs), while the fine-tuning in downstream applications can also produce new biases or amplify the impact of the original biases. Existing works separate the debiasing from the fine-tuning procedure, which results in…

Cited by 3SourcePDFScholar
2023

Diffusion Probabilistic Modeling for Fine-Grained Urban Traffic Flow Inference with Relaxed Structural Constraint

ICASSP 2023accepted

Inferring the citywide urban traffic flows is critical for numerous smart city applications such as urban planning, traffic control, and transportation management. Urban traffic flow inference problem aims to generate fine-grained flow maps from the coarse-grained ones. It is still challenging due t…

Cited by 0SourceScholar
2023

Directional diffusion models for graph representation learning

NeurIPS 2023poster

Diffusion models have achieved remarkable success in diverse domains such as image synthesis, super-resolution, and 3D molecule generation. Surprisingly, the application of diffusion models in graph learning has garnered little attention. In this paper, we aim to bridge this gap by exploring the use…

2023

DyCVAE: Learning Dynamic Causal Factors for Non-stationary Series Domain Generalization (Student Abstract)

AAAI 2023technical

Learning domain-invariant representations is a major task of out-of-distribution generalization. To address this issue, recent efforts have taken into accounting causality, aiming at learning the causal factors with regard to tasks. However, extending existing generalization methods for adapting non…

Cited by 0SourcePDFScholar
2023

Enhancing Knowledge Transfer for Task Incremental Learning with Data-free Subnetwork

NeurIPS 2023poster

As there exist competitive subnetworks within a dense network in concert with Lottery Ticket Hypothesis, we introduce a novel neuron-wise task incremental learning method, namely Data-free Subnetworks (DSN), which attempts to enhance the elastic knowledge transfer across the tasks that sequentially…

2023

Exploring Hypergraph of Earnings Call for Risk Prediction (Student Abstract)

AAAI 2023technical

In financial economics, studies have shown that the textual content in the earnings conference call transcript has predictive power for a firm's future risk. However, the conference call transcript is very long and contains diverse non-relevant content, which poses challenges for the text-based risk…

Cited by 2SourcePDFScholar
2023

Foresee What You Will Learn: Data Augmentation for Domain Generalization in Non-stationary Environment

AAAI 2023technical

Existing domain generalization aims to learn a generalizable model to perform well even on unseen domains. For many real-world machine learning applications, the data distribution often shifts gradually along domain indices. For example, a self-driving car with a vision system drives from dawn to du…

2023

Language Models Can Improve Event Prediction by Few-Shot Abductive Reasoning

NeurIPS 2023poster

Large language models have shown astonishing performance on a wide range of reasoning tasks. In this paper, we investigate whether they could reason about real-world events and help improve the prediction performance of event sequence models. We design LAMP, a framework that integrates a large langu…

Cited by 52SourcePDFScholar
2023

Learning Dynamic Temporal Relations with Continuous Graph for Multivariate Time Series Forecasting (Student Abstract)

AAAI 2023technical

The recent advance in graph neural networks (GNNs) has inspired a few studies to leverage the dependencies of variables for time series prediction. Despite the promising results, existing GNN-based models cannot capture the global dynamic relations between variables owing to the inherent limitation…

Cited by 4SourcePDFScholar
2023

Mobility Prediction via Sequential Trajectory Disentanglement (Student Abstract)

AAAI 2023technical

Accurately predicting human mobility is a critical task in location-based recommendation. Most prior approaches focus on fusing multiple semantics trajectories to forecast the future movement of people, and fail to consider the distinct relations in underlying context of human mobility, resulting in…

Cited by 1SourcePDFScholar
2023

Open Anomalous Trajectory Recognition via Probabilistic Metric Learning

IJCAI 2023poster

Typically, trajectories considered anomalous are the ones deviating from usual (e.g., traffic-dictated) driving patterns. However, this closed-set context fails to recognize the unknown anomalous trajectories, resulting in an insufficient self-motivated learning paradigm. In this study, we investiga…

2023

Optimal Treatment Allocation for Efficient Policy Evaluation in Sequential Decision Making

NeurIPS 2023poster

A/B testing is critical for modern technological companies to evaluate the effectiveness of newly developed products against standard baselines. This paper studies optimal designs that aim to maximize the amount of information obtained from online experiments to estimate treatment effects accurately…

Cited by 7SourcePDFScholar
2023

Overcoming Forgetting in Fine-Grained Urban Flow Inference via Adaptive Knowledge Replay

AAAI 2023technical

Fine-grained urban flow inference (FUFI) problem aims at inferring the high-resolution flow maps from the coarse-grained ones, which plays an important role in sustainable and economic urban computing and traffic management. Previous models addressed the FUFI problem from spatial constraint, externa…

2023

Revisiting Denoising Diffusion Probabilistic Models for Speech Enhancement: Condition Collapse, Efficiency and Refinement

AAAI 2023technical

Recent literature has shown that denoising diffusion probabilistic models (DDPMs) can be used to synthesize high-fidelity samples with a competitive (or sometimes better) quality than previous state-of-the-art approaches. However, few attempts have been made to apply DDPM for the speech enhancement…

2023

SLOTH: Structured Learning and Task-Based Optimization for Time Series Forecasting on Hierarchies

AAAI 2023technical

Multivariate time series forecasting with hierarchical structure is widely used in real-world applications, e.g., sales predictions for the geographical hierarchy formed by cities, states, and countries. The hierarchical time series (HTS) forecasting includes two sub-tasks, i.e., forecasting and rec…

Cited by 4SourcePDFScholar
2023

Somali Information Retrieval Corpus: Bridging the Gap between Query Translation and Dedicated Language Resources

EMNLP 2023short main

Despite the growing use of the Somali language in various online domains, research on Somali language information retrieval remains limited and primarily relies on query translation due to the lack of a dedicated corpus. To address this problem, we collaborated with language experts and natural lang…

Cited by 0SourceScholar
2022

Double-Check Soft Teacher for Semi-Supervised Object Detection

IJCAI 2022poster

In the semi-supervised object detection task, due to the scarcity of labeled data and the diversity and complexity of objects to be detected, the quality of pseudo-labels generated by existing methods for unlabeled data is relatively low, which severely restricts the performance of semi-supervised o…

2022

Dynamic Manifold Learning for Land Deformation Forecasting

AAAI 2022technical

Landslides refer to occurrences of massive ground movements due to geological (and meteorological) factors, and can have disastrous impact on property, economy, and even lead to loss of life. The advances of remote sensing provide accurate and continuous terrain monitoring, enabling the study and an…

Cited by 3SourcePDFScholar
2022

Learning Latent Seasonal-Trend Representations for Time Series Forecasting

NeurIPS 2022accept

Forecasting complex time series is ubiquitous and vital in a range of applications but challenging. Recent advances endeavor to achieve progress by incorporating various deep learning techniques (e.g., RNN and Transformer) into sequential models. However, clear patterns are still hard to extract sin…

Cited by 83SourcePDFScholar
2022

Probabilistic Fine-Grained Urban Flow Inference with Normalizing Flows

ICASSP 2022accepted

Fine-grained urban flow inference (FUFI) aims at enhancing the resolution of traffic flow, which plays an important role in intelligent traffic management. Existing FUFI methods are mainly based on techniques from image super-resolution (SR) models, which cannot fully capture the influence of extern…

Cited by 0SourceScholar
2022

Quantification and Analysis of Layer-wise and Pixel-wise Information Discarding

ICML 2022spotlight

This paper presents a method to explain how the information of each input variable is gradually discarded during the forward propagation in a deep neural network (DNN), which provides new perspectives to explain DNNs. We define two types of entropy-based metrics, i.e. (1) the discarding of pixel-wis…

2022

TaCube: Pre-computing Data Cubes for Answering Numerical-Reasoning Questions over Tabular Data

EMNLP 2022main

Existing auto-regressive pre-trained language models (PLMs) like T5 and BART, have been well applied to table question answering by UNIFIEDSKG and TAPEX, respectively, and demonstrated state-of-the-art results on multiple benchmarks. However, auto-regressive PLMs are challenged by recent emerging nu…

2022

Table Pre-training: A Survey on Model Architectures, Pre-training Objectives, and Downstream Tasks

IJCAI 2022poster

Following the success of pre-training techniques in the natural language domain, a flurry of table pre-training frameworks have been proposed and have achieved new state-of-the-arts on various downstream tasks such as table question answering, table type recognition, column relation classification,…

Cited by 71SourcePDFScholar
2021

Non-decreasing Quantile Function Network with Efficient Exploration for Distributional Reinforcement Learning

IJCAI 2021poster

Although distributional reinforcement learning (DRL) has been widely examined in the past few years, there are two open questions people are still trying to address. One is how to ensure the validity of the learned quantile function, the other is how to efficiently utilize the distribution informati…

Cited by 22SourcePDFScholar
2020

Deep Active Learning: Unified and Principled Method for Query and Training

AISTATS 2020poster

In this paper, we are proposing a unified and principled method for both the querying and training processes in deep batch active learning. We are providing theoretical insights from the intuition of modeling the interactive procedure in active learning as distribution matching, by adopting the Wass…

2020

Enhancing Urban Flow Maps via Neural ODEs

IJCAI 2020poster

Flow super-resolution (FSR) enables inferring fine-grained urban flows with coarse-grained observations and plays an important role in traffic monitoring and prediction. The existing FSR solutions rely on deep CNN models (e.g., ResNet) for learning spatial correlation, incurring excessive memory cos…

2020

Non-Crossing Quantile Regression for Distributional Reinforcement Learning

NeurIPS 2020poster

Distributional reinforcement learning (DRL) estimates the distribution over future returns instead of the mean to more efficiently capture the intrinsic uncertainty of MDPs. However, batch-based DRL algorithms cannot guarantee the non-decreasing property of learned quantile curves especially at the…

Cited by 55SourcePDFScholar
2019

Graph-Based Semi-Supervised Learning with Non-ignorable Non-response

NeurIPS 2019poster

Graph-based semi-supervised learning is a very powerful tool in classification tasks, while in most existing literature the labelled nodes are assumed to be randomly sampled. When the labelling status depends on the unobserved node response, ignoring the missingness can lead to significant estimatio…

2017

Illumination insensitive efficient second-order minimization for planar object tracking

ICRA 2017poster

Tracking for planar objects is an important issue to vision-based robotic applications. In direct visual tracking (DVT) methods, the similarity between two images is often measured through the sum of squared differences (SSD) especially with the efficient second-order minimization (ESM) due to its s…

Cited by 30SourceScholar