← Search

Jun Zhou

132 accepted papers

2026

Adaptive Mixture of Disentangled Experts for Dynamic Graphs under Distribution Shifts

ICLR 2026poster

Dynamic graph representation learning under distribution shifts has drawn an increasing amount of attention in the research community, given its wide applicability in real-world scenarios. Existing methods typically employ a fixed-architecture design to extract invariant patterns. However, there may…

Cited by 0SourceScholar
2026

Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation Evaluation

AAAI 2026technical

The rapid development of large language models (LLMs) has highlighted the need for efficient and reliable methods to evaluate their performance. Traditional evaluation methods often face challenges like high costs, limited task formats, dependence on human references, and systematic biases. To addre

Cited by 0SourcePDFScholar
2026

CodeChemist: Test-Time Scaling for Low-Resource Code Generation via Functional Knowledge Transfer

ICML 2026poster

Code Large Language Models (CodeLLMs) have been widely adopted for Natural Language to Programming Language code generation, powering applications with large user bases. Their performance, however, varies sharply across programming languages (PLs) and is particularly suboptimal for low-resource PLs …

Cited by 0SourceScholar
2026

Deterministic Differentiable Structured Pruning for Large Language Models

ICML 2026poster

Structured pruning reduces LLM inference cost by removing low-importance architectural components. This can be viewed as learning a multiplicative gate for each component under an $\ell_0$ sparsity constraint. Due to the discreteness of the $\ell_0$ norm, prior work typically adopts stochastic hard-…

Cited by 0SourceScholar
2026

InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning

ICML 2026poster

Large reasoning models achieve strong performance by scaling inference-time chain-of-thought, but this paradigm suffers from quadratic cost, context length limits, and degraded reasoning due to lost-in-the-middle effects. Iterative reasoning mitigates these issues by periodically summarizing interme…

Cited by 0SourceScholar
2026

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

CVPR 2026

In this work, we introduce LLaDA-V, a purely diffusion-based Multimodal Large Language Model (MLLM) that integrates visual instruction tuning with masked diffusion models, representing a departure from the autoregressive paradigms dominant in current multimodal approaches. Built upon LLaDA, a repres

Cited by 0SourcecodeScholar
2026

Listen and Count: Expanding the Frontier of Zero-Shot Object Counting to Sound-centric Counting

IJCAI 2026

While class-agnostic object counting has recently evolved from image-exemplar to language-guided paradigms, existing methods are limited by text polysemy and the lack of prompts in audio-sensing scenarios. To overcome these challenges, we introduce a sound-centric counting paradigm, enabling models

Cited by 0Scholar
2026

MergeMix: Optimizing Mid-Training Data Mixtures via Learnable Model Merging

ICML 2026poster

Optimizing data mixtures is is essential for unlocking the full potential of of large language models (LLMs), yet identifying the optimal composition remains computationally prohibitive due to reliance on heuristic trials or expensive proxy training. To address this, we introduce MergeMix, a novel a…

Cited by 0SourceScholar
2026

Note2Chat: Improving LLMs for Multi-Turn Clinical History Taking Using Medical Notes

AAAI 2026technical

Effective clinical history taking is a foundational yet underexplored component of clinical reasoning. While large language models (LLMs) have shown promise on static benchmarks, they often fall short in dynamic, multi-turn diagnostic settings that require iterative questioning and hypothesis refine

Cited by 0SourcePDFScholar
2026

Privacy-Aware Video Anomaly Detection: Guided Orthogonal Projection and a Comprehensive Evaluation Framework

ICML 2026oral

Video anomaly detection (VAD) is critical for surveillance systems, but current methods prioritize accuracy while ignoring the ethical risks of encoding sensitive biometric information. This neglect poses significant privacy concerns for real-world deployment. To bridge this gap, we introduce the Gu…

Cited by 0SourceScholar
2026

SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation

ICLR 2026poster

In this paper, we introduce SemHiTok, a unified image Tokenizer via Semantic-Guided Hierarchical codebook (SGHC) that provides consistent discrete representations for multimodal understanding and generation. Recently, unified image tokenizers have sparked exploration within the research community, w…

Cited by 0SourceScholar
2026

Thinker: Training LLMs in Hierarchical Thinking for Deep Search via Multi-Turn Interaction

AAAI 2026technical

Efficient retrieval of external knowledge bases and web pages is crucial for enhancing the reasoning abilities of LLMs. Previous works on training LLMs to leverage external retrievers for solving complex problems have predominantly employed end-to-end reinforcement learning. However, these approache

Cited by 0SourcePDFScholar
2026

Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models

ICLR 2026poster

Mixture-of-Experts (MoE) has become a dominant architecture for scaling Large Language Models (LLMs) efficiently by decoupling total parameters from computational cost. However, this decoupling creates a critical challenge: predicting the model capacity of a given MoE configurations (e.g., expert ac…

Cited by 0SourceScholar
2026

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

ICLR 2026poster

Recent advances in large language models (LLMs) have utilized reinforcement learning with verifiable rewards (RLVR) to improve reasoning capabilities. However, scaling these methods typically requires massive data and extensive rollout computations, leading to high training costs and low data effici…

Cited by 0SourceScholar
2026

Unsupervised Camouflaged Object Detection with Dual-Eigenvector Spectral Pseudo-Labeling and Contrastive Refinement

ICML 2026poster

Unsupervised Camouflaged Object Detection (UCOD) aims to identify objects concealed in their surroundings without relying on pixel-level labels. Existing methods rely solely on simple post-processing of DINO high-dimensional features to generate pseudo labels for training. However, these methods suf…

Cited by 0SourceScholar
2026

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation

CVPR 2026

The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditional generation quality metrics to encompass aesthetic appeal. However, existing benchmarks remain largely focused on technical fidelity, leaving a sig

Cited by 0SourceScholar
2026

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

IJCAI 2026

The rapid advancement of AIGC video generation calls for evaluation frameworks that move beyond technical fidelity and incorporate human-centered aesthetic assessment. Existing benchmarks often overlook fine-grained perceptual qualities such as visual aesthetics, artistic style, and human preference

Cited by 0Scholar
2026

VPHO: Joint Visual-Physical Cue Learning and Aggregation for Hand-Object Pose Estimation

AAAI 2026technical

Estimating the 3D poses of hands and objects from a single RGB image is a fundamental yet challenging problem, with broad applications in augmented reality and human-computer interaction. Existing methods largely rely on visual cues alone, often producing results that violate physical constraints su

Cited by 0SourcePDFScholar
2026

WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training

ICLR 2026oral

Recent advances in learning rate~(LR) scheduling have demonstrated the effectiveness of decay-free approaches that eliminate the traditional decay phase while maintaining competitive performance. Model merging techniques have emerged as particularly promising solutions in this domain. We present War…

Cited by 0SourceScholar
2026

What Really Improves Mathematical Reasoning: Structured Reasoning Signals Beyond Pure Code

ICML 2026poster

Incorporating code into training corpora has become a widely acknowledged practice in the development of modern foundation language models (LMs). Compared with a general Internet corpus, code offers high-quality, well-structured signals that substantially augment the coding proficiency of models. Be…

Cited by 0SourceScholar
2026

Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training

AAAI 2026technical

Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated datasets. Although some zero-shot methods attempt to trajectory control in the l

Cited by 0SourcePDFScholar
2025

ARGenSeg: Image Segmentation with Autoregressive Image Generation Model

NeurIPS 2025poster

We propose a novel AutoRegressive Generation-based paradigm for image Segmentation (ARGenSeg), achieving multimodal understanding and pixel-level perception within a unified framework. Prior works integrating image segmentation into multimodal large language models (MLLMs) typically employ either b…

Cited by 0SourceScholar
2025

Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation

ICCV 2025poster

Talking head synthesis is vital for virtual avatars and human-computer interaction. However, most existing methods are typically limited to accepting control from a single primary modality, restricting their practical utility. To this end, we introduce ACTalker, an end-to-end video diffusion framewo…

2025

BOSE: A Systematic Evaluation Method Optimized for Base Models

ACL 2025finding

This paper poses two critical issues in evaluating base models (without post-training): (1) Unstable evaluation during training: in the early stages of pre-training, the models lack the capability to answer questions as required, leading to unstable evaluation results. This instability makes it diff…

2025

CREST: An Efficient Conjointly-trained Spike-driven Framework for Event-based Object Detection Exploiting Spatiotemporal Dynamics

AAAI 2025technical

Event-based cameras feature high temporal resolution, wide dynamic range, and low power consumption, which are ideal for high-speed and low-light object detection. Spiking neural networks (SNNs) are promising for event-based object recognition and detection due to their spiking nature but lack effic…

2025

Controllable Unlearning for Image-to-Image Generative Models via $\epsilon$-Constrained Optimization

ICLR 2025poster

While generative models have made significant advancements in recent years, they also raise concerns such as privacy breaches and biases. Machine unlearning has emerged as a viable solution, aiming to remove specific training data, e.g., containing private information and bias, from models. In this…

Cited by 1SourcePDFScholar
2025

Dual-Modality Guided Artistic Style Transfer with Pre-trained Diffusion Models

ICASSP 2025accepted

Artistic style transfer aims to replicate an artist’s painting style in a different image. While existing pre-trained model-based methods can generate high-quality stylized images, they often lack precise control over stylistic elements. Recent approaches incorporating textual inversion offer more a…

Cited by 0SourceScholar
2025

Effective and Efficient Masked Image Generation Models

ICML 2025poster

Although masked image generation models and masked diffusion models are designed with different motivations and objectives, we observe that they can be unified within a single framework. Building upon this insight, we carefully explore the design space of training and sampling, identifying key facto…

2025

Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation

EMNLP 2025

Large language models (LLMs) often exhibit limited performance on domain-specific tasks due to the natural disproportionate representation of specialized information in their training data and the static nature of these datasets. Knowledge scarcity and temporal lag create knowledge gaps for domain a

2025

Explore What LLM Does Not Know in Complex Question Answering

AAAI 2025technical

Complex question answering (QA) is a challenging task in artificial intelligence research which requires reasoning based on related knowledge. The retrieval-augmented generation (RAG) based on large language models (LLMs) have become one promising solution in QA. To facilitate RAG more effectively,…

2025

FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model

CVPR 2025poster

Currently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of visual language models (VLMs). However, they still face challenges in three key areas: 1) complex scenarios; 2) semantic consistency; and 3) fine-gra…

Cited by 2SourcePDFScholar
2025

HeMeNet: Heterogeneous Multichannel Equivariant Network for Protein Multi-task Learning

AAAI 2025technical

Understanding and leveraging the 3D structures of proteins is central to a variety of biological and drug discovery tasks. While deep learning has been applied successfully for structure-based protein function prediction tasks, current methods usually employ distinct training for each task. However,…

2025

HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation

CVPR 2025poster

We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the chara…

2025

Improving Natural Language Understanding for LLMs via Large-Scale Instruction Synthesis

AAAI 2025technical

High-quality, large-scale instructions are crucial for aligning large language models (LLMs), however, there is a severe shortage of instruction in the field of natural language understanding (NLU). Previous works on constructing NLU instructions mainly focus on information extraction (IE), neglect…

Cited by 0SourcePDFScholar
2025

InsTaG: Learning Personalized 3D Talking Head from Few-Second Video

CVPR 2025poster

Despite exhibiting impressive performance in synthesizing lifelike personalized 3D talking heads, prevailing methods based on radiance fields suffer from high demands for training data and time for each new identity. This paper introduces InsTaG, a 3D talking head synthesis framework that allows a f…

2025

LLMOPT: Learning to Define and Solve General Optimization Problems from Scratch

ICLR 2025poster

Optimization problems are prevalent across various scenarios. Formulating and then solving optimization problems described by natural language often requires highly specialized human expertise, which could block the widespread application of optimization-based decision making. To automate problem fo…

2025

LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion

ACL 2025long

The safety mechanisms of large language models (LLMs) exhibit notable fragility, as even fine-tuning on datasets without harmful content may still undermine their safety capabilities. Meanwhile, existing safety alignment methods predominantly rely on the fine-tuning process, which inadvertently lead…

2025

LaMP-Val: Large Language Models Empower Personalized Valuation in Auction

EMNLP 2025

Auctions are a vital economic mechanism used to determine the market value of goods or services through competitive bidding within a specific framework. However, much of the current research primarily focuses on the bidding algorithms used within auction mechanisms. This often neglects the potential

2025

Learning Causal Transition Matrix for Instance-dependent Label Noise

AAAI 2025technical

Noisy labels are both inevitable and problematic in machine learning methods, as they negatively impact models' generalization ability by causing overfitting. In the context of learning with noise, the transition matrix plays a crucial role in the design of statistically consistent algorithms. Howev…

Cited by 0SourcePDFScholar
2025

MASS: Mathematical Data Selection via Skill Graphs for Pretraining Large Language Models

ICML 2025poster

High-quality data plays a critical role in the pretraining and fine-tuning of large language models (LLMs), even determining their performance ceiling to some degree. Consequently, numerous data selection methods have been proposed to identify subsets of data that can effectively and efficiently enh…

Cited by 0SourcePDFScholar
2025

Making Large Vision Language Models to Be Good Few-Shot Learners

AAAI 2025technical

Few-shot classification (FSC) is a fundamental yet challenging task in computer vision that involves recognizing novel classes from limited data. While previous methods have focused on enhancing visual features or incorporating additional modalities, Large Vision Language Models (LVLMs) offer a prom…

2025

Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging

NeurIPS 2025poster

Achieving balanced alignment of large language models (LLMs) in terms of Helpfulness, Honesty, and Harmlessness (3H optimization) constitutes a cornerstone of responsible AI. Existing methods like data mixture strategies face limitations, including heavy reliance on expert knowledge and conflicting…

Cited by 0SourceScholar
2025

NATRA: Noise-Agnostic Framework for Trajectory Prediction with Noisy Observations

ICCV 2025poster

Trajectory prediction aims to forecast an agent's future trajectories based on its historical observed trajectories, which is a critical task for various applications such as autonomous driving, robotics, and surveillance systems. Most existing trajectory prediction methods assume that the observed…

Cited by 0SourcePDFScholar
2025

Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing Images

ICASSP 2025accepted

The Direct Segment Anything Model (DirectSAM) excels in class-agnostic contour extraction. In this paper, we explore its use by applying it to optical remote sensing imagery, where semantic contour extraction—such as identifying buildings, road networks, and coastlines-holds significant practical va…

Cited by 0SourceScholar
2025

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution

EMNLP 2025

Reinforcement learning from human feedback (RLHF) offers a promising approach to aligning large language models (LLMs) with human preferences. Typically, a reward model is trained or supplied to act as a proxy for humans in evaluating generated responses during the reinforcement training phase. Howe

2025

Refining Sentence Embedding Model through Ranking Sentences Generation with Large Language Models

ACL 2025finding

Sentence embedding is essential for many NLP tasks, with contrastive learning methods achieving strong performance using annotated datasets like NLI. Yet, the reliance on manual labels limits scalability. Recent studies leverage large language models (LLMs) to generate sentence pairs, reducing annot…

2025

RemoteTrimmer: Adaptive Structural Pruning for Remote Sensing Image Classification

ICASSP 2025accepted

Since high resolution remote sensing image classifi-cation often requires a relatively high computation complexity, lightweight models tend to be practical and efficient. Model pruning is an effective method for model compression. However, existing methods rarely take into account the specificity of…

Cited by 0SourceScholar
2025

Rethinking Causal Ranking: A Balanced Perspective on Uplift Model Evaluation

ICML 2025poster

Uplift modeling is crucial for identifying individuals likely to respond to a treatment in applications like marketing and customer retention, but evaluating these models is challenging due to the inaccessibility of counterfactual outcomes in real-world settings. In this paper, we identify a fundame…

2025

Robust Preference Optimization via Dynamic Target Margins

ACL 2025finding

The alignment of Large Language Models (LLMs) is crucial for ensuring their safety and reliability in practical applications. Direct Preference Optimization (DPO) has emerged as an efficient method that directly optimizes models using preference pairs, significantly reducing resource demands. Howeve…

2025

SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning

NeurIPS 2025poster

Leveraging multimodal large models for image segmentation has become a prominent research direction. However, existing approaches typically rely heavily on manually annotated datasets that include explicit reasoning processes, which are costly and time-consuming to produce. Recent advances suggest t…

Cited by 0SourceScholar
2025

SATA: Spatial Autocorrelation Token Analysis for Enhancing the Robustness of Vision Transformers

CVPR 2025poster

Over the past few years, vision transformers (ViTs) have consistently demonstrated remarkable performance across various visual recognition tasks. However, attempts to enhance their robustness have yielded limited success, mainly focusing on different training strategies, input patch augmentation, o…

2025

STAR: Spatial-Temporal Augmentation with Text-to-Video Models for Real-World Video Super-Resolution

ICCV 2025poster

Image diffusion models have been adapted for real-world video super-resolution to tackle over-smoothing issues in GAN-based methods. However, these models struggle to maintain temporal consistency, as they are trained on static images, limiting their ability to capture temporal dynamics effectively.…

Cited by 0SourcePDFScholar
2025

ShieldHead: Decoding-time Safeguard for Large Language Models

ACL 2025finding

In light of the widespread deployment of Large Language Models (LLMs), the responsibility for safeguarding and regulating LLM-generated content has taken on heightened significance. Recent advancements in LLM-based moderation methods, e.g., LlamaGuard, have demonstrated remarkable promise in identif…

2025

Spatially-variant Blur Degradation Model Based on Depth Estimation

ICASSP 2025accepted

It is well known that the number of aligned images in the single image super-resolution (SISR) models training is limited. Synthesizing data is an effective way to address this issue. However, many degradation models only consider using spatially-invariant blur kernels to blur high-resolution (HR) i…

Cited by 0SourceScholar
2025

TerraFusion: Semi-Supervised Vision-Proprioception Fusion for Robust Terrain Classification

RA-L 2025

Terrain classification is essential for traversability estimation and planning of unmanned ground vehicles (UGVs) in complex environments. Most existing approaches utilize fully supervised learning to classify terrains based on either exteroceptive or proprioceptive sensor modalities. However, visio

Cited by 0SourceScholar
2025

TerraX: Visual Terrain Classification Enhanced by Vision-Language Models

IROS 2025

Visual Terrain Classification (VTC) plays a vital role in enabling unmanned ground vehicles to understand complex environments. Existing research relies on image-label pairs annotated by static label sets, where semantic ambiguity and high annotation costs constrain fine-grained terrain characteriza

Cited by 0SourceScholar
2025

UTC-RS: An Underwater Tracked Cleaning Robot System for Hydraulic Structures

RA-L 2025

During the inspection and maintenance of the underwater part of hydraulic structures, it is often necessary to clean the surface of a certain area for subsequent operations. At present, there are still few robots capable of underwater fine cleaning. Therefore, this letter introduces the design of a

Cited by 4SourceScholar
2025

Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary

EMNLP 2025

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, yet they often refuse to answer legitimate queries—a phenomenon known as overrefusal. Overrefusal typically stems from over-conservative safety alignment, causing models to treat many reasonable prom

2025

Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction Tuning

ICLR 2025oral

Catastrophic forgetting (CF) poses a significant challenge in machine learning, where a model forgets previously learned information upon learning new tasks. Despite the advanced capabilities of Large Language Models (LLMs), they continue to face challenges with CF during continual learning. The ma…

Cited by 1SourcePDFScholar
2025

VADB: A Large-Scale Video Aesthetic Database with Professional and Multi-Dimensional Annotations

NeurIPS 2025poster

Video aesthetic assessment, a vital area in multimedia computing, integrates computer vision with human cognition. Its progress is limited by the lack of standardized datasets and robust models, as the temporal dynamics of video and multimodal fusion challenges hinder direct application of image-bas…

Cited by 0SourcecodeScholar
2024

Backdoor Adjustment via Group Adaptation for Debiased Coupon Recommendations

AAAI 2024technical

Accurate prediction of coupon usage is crucial for promoting user consumption through targeted coupon recommendations. However, in real-world coupon recommendations, the coupon allocation process is not solely determined by the model trained with the history interaction data but is also interfered w…

Cited by 6SourcePDFScholar
2024

ChatUIE: Exploring Chat-based Unified Information Extraction Using Large Language Models

COLING 2024main

Recent advancements in large language models have shown impressive performance in general chat. However, their domain-specific capabilities, particularly in information extraction, have certain limitations. Extracting structured information from natural language that deviates from known schemas or i…

2024

Collaborative Refining for Learning from Inaccurate Labels

NeurIPS 2024poster

This paper considers the problem of learning from multiple sets of inaccurate labels, which can be easily obtained from low-cost annotators, such as rule-based annotators. Previous works typically concentrate on aggregating information from all the annotators, overlooking the significance of data re…

Cited by 0SourcePDFScholar
2024

DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization

CVPR 2024poster

Radiance fields have demonstrated impressive performance in synthesizing novel views from sparse input views yet prevailing methods suffer from high training costs and slow inference speed. This paper introduces DNGaussian a depth-regularized framework based on 3D Gaussian radiance fields offering r…

2024

EasyTPP: Towards Open Benchmarking Temporal Point Processes

ICLR 2024poster

Continuous-time event sequences play a vital role in real-world domains such as healthcare, finance, online shopping, social networks, and so on. To model such data, temporal point processes (TPPs) have emerged as the most natural and competitive models, making a significant impact in both academic…

2024

Efficient Model Stealing Defense with Noise Transition Matrix

CVPR 2024poster

With the escalating complexity and investment cost of training deep neural networks safeguarding them from unauthorized usage and intellectual property theft has become imperative. Especially the rampant misuse of prediction APIs to replicate models without access to the original data or architectur…

Cited by 0SourcePDFScholar
2024

Enhancing Cross-Document Event Coreference Resolution by Discourse Structure and Semantic Information

COLING 2024main

Existing cross-document event coreference resolution models, which either compute mention similarity directly or enhance mention representation by extracting event arguments (such as location, time, agent, and patient), lackingmthe ability to utilize document-level information. As a result, they str…

2024

Enhancing Event Sequence Modeling with Contrastive Relational Inference

ICASSP 2024accepted

Neural temporal point processes(TPPs) have shown promise for modeling continuous-time event sequences. However, capturing the interactions between events is challenging yet critical for performing inference tasks like forecasting on event sequence data. Existing TPP models have focused on parameteri…

Cited by 0SourceScholar
2024

FC-GNN: Recovering Reliable and Accurate Correspondences from Interferences

CVPR 2024poster

Finding correspondences between images is essential for many computer vision tasks and sparse matching pipelines have been popular for decades. However matching noise within and between images along with inconsistent keypoint detection frequently degrades the matching performance. We review these pr…

2024

GMP-AR: Granularity Message Passing and Adaptive Reconciliation for Temporal Hierarchy Forecasting

AAAI 2024technical

Time series forecasts of different temporal granularity are widely used in real-world applications, e.g., sales prediction in days and weeks for making different inventory plans. However, these tasks are usually solved separately without ensuring coherence, which is crucial for aligning downstream d…

Cited by 0SourcePDFScholar
2024

Harvesting Events from Multiple Sources: Towards a Cross-Document Event Extraction Paradigm

ACL 2024findings

Document-level event extraction aims to extract structured event information from unstructured text. However, a single document often contains limited event information and the roles of different event arguments may be biased due to the influence of the information source.This paper addresses the li…

2024

Improving Equivariant Graph Neural Networks on Large Geometric Graphs via Virtual Nodes Learning

ICML 2024poster

Equivariant Graph Neural Networks (GNNs) have made remarkable success in a variety of scientific applications. However, existing equivariant GNNs encounter the efficiency issue for large geometric graphs and perform poorly if the input is reduced to sparse local graph for speed acceleration. In this…

Cited by 5SourcePDFScholar
2024

Keypoint-based Progressive Chain-of-Thought Distillation for LLMs

ICML 2024poster

Chain-of-thought distillation is a powerful technique for transferring reasoning abilities from large language models (LLMs) to smaller student models. Previous methods typically require the student to mimic the step-by-step rationale produced by LLMs, often facing the following challenges: (i) Toke…

Cited by 2SourcePDFScholar
2024

Knowledge-augmented Financial Market Analysis and Report Generation

EMNLP 2024industry

Crafting a convincing financial market analysis report necessitates a wealth of market information and the expertise of financial analysts, posing a highly challenging task. While large language models (LLMs) have enabled the automated generation of financial market analysis text, they still face is…

Cited by 2SourcePDFScholar
2024

LLMRG: Improving Recommendations through Large Language Model Reasoning Graphs

AAAI 2024technical

Recommendation systems aim to provide users with relevant suggestions, but often lack interpretability and fail to capture higher-level semantic relationships between user behaviors and profiles. In this paper, we propose a novel approach that leverages large language models (LLMs) to construct pers…

Cited by 16SourcePDFScholar
2024

Learning to Plan for Retrieval-Augmented Large Language Models from Knowledge Graphs

EMNLP 2024finding

Improving the performance of large language models (LLMs) in complex question-answering (QA) scenarios has always been a research focal point. Recent studies have attempted to enhance LLMs’ performance by combining step-wise planning with external retrieval. While effective for advanced models like…

2024

Lower Bounds of Uniform Stability in Gradient-Based Bilevel Algorithms for Hyperparameter Optimization

NeurIPS 2024poster

Gradient-based bilevel programming leverages unrolling differentiation (UD) or implicit function theorem (IFT) to solve hyperparameter optimization (HO) problems, and is proven effective and scalable in practice. To understand their generalization behavior, existing works establish upper bounds on…

Cited by 0SourcePDFScholar
2024

MDGNN: Multi-Relational Dynamic Graph Neural Network for Comprehensive and Dynamic Stock Investment Prediction

AAAI 2024technical

The stock market is a crucial component of the financial system, but predicting the movement of stock prices is challenging due to the dynamic and intricate relations arising from various aspects such as economic indicators, financial reports, global news, and investor sentiment. Traditional sequent…

Cited by 23SourcePDFScholar
2024

OneGen: Efficient One-Pass Unified Generation and Retrieval for LLMs

EMNLP 2024finding

Despite the recent advancements in Large Language Models (LLMs), which have significantly enhanced the generative capabilities for various NLP tasks, LLMs still face limitations in directly handling retrieval tasks. However, many practical applications demand the seamless integration of both retriev…

2024

Optimizing Language Models with Fair and Stable Reward Composition in Reinforcement Learning

EMNLP 2024main

Reinforcement learning from human feedback (RLHF) and AI-generated feedback (RLAIF) have become prominent techniques that significantly enhance the functionality of pre-trained language models (LMs). These methods harness feedback, sourced either from humans or AI, as direct rewards or to shape rewa…

2024

Rethinking Memory and Communication Costs for Efficient Data Parallel Training of Large Language Models

NeurIPS 2024poster

Recently, various strategies for distributed training of large language models (LLMs) have been proposed. By categorizing them into basic strategies and composite strategies, we have discovered that existing basic strategies provide limited options in specific scenarios, leaving considerable room fo…

Cited by 0SourcePDFScholar
2024

Self-cognitive Denoising in the Presence of Multiple Noisy Label Sources

ICML 2024poster

The strong performance of neural networks typically hinges on the availability of extensive labeled data, yet acquiring ground-truth labels is often challenging. Instead, noisy supervisions from multiple sources, e.g., by multiple well-designed rules, are more convenient to collect. In this paper, w…

Cited by 2SourcePDFScholar
2024

Structural Information Enhanced Graph Representation for Link Prediction

AAAI 2024technical

Link prediction is a fundamental task of graph machine learning, and Graph Neural Network (GNN) based methods have become the mainstream approach due to their good performance. However, the typical practice learns node representations through neighborhood aggregation, lacking awareness of the struct…

Cited by 5SourcePDFScholar
2024

TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting

ECCV 2024poster

"Radiance fields have demonstrated impressive performance in synthesizing lifelike 3D talking heads. However, due to the difficulty in fitting steep appearance changes, the prevailing paradigm that presents facial motions by directly modifying point appearance may lead to distortions in dynamic regi…

2024

TimeMixer: Decomposable Multiscale Mixing for Time Series Forecasting

ICLR 2024poster

Time series forecasting is widely used in extensive applications, such as traffic planning and weather forecasting. However, real-world time series usually present intricate temporal variations, making forecasting extremely challenging. Going beyond the mainstream paradigms of plain decomposition an…

2024

Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations

ICML 2024poster

Bayesian flow networks (BFNs) iteratively refine the parameters, instead of the samples in diffusion models (DMs), of distributions at various noise levels through Bayesian inference. Owing to its differentiable nature, BFNs are promising in modeling both continuous and discrete data, while simultan…

2024

What Factors Influence LLMs’ Judgments? A Case Study on Question Answering

COLING 2024main

Large Language Models (LLMs) are now being considered as judges of high efficiency to evaluate the quality of answers generated by candidate models. However, their judgments may be influenced by complex scenarios and inherent biases, raising concerns about their reliability. This study aims to bridg…

Cited by 3SourcePDFScholar
2024

π-Light: Programmatic Interpretable Reinforcement Learning for Resource-Limited Traffic Signal Control

AAAI 2024technical

The recent advancements in Deep Reinforcement Learning (DRL) have significantly enhanced the performance of adaptive Traffic Signal Control (TSC). However, DRL policies are typically represented by neural networks, which are over-parameterized black-box models. As a result, the learned policies ofte…

2023

An Application of Quantum Mechanics to Attention Methods in Computer Vision

ICASSP 2023accepted

This work proposes the quantum-state-based mapping (QSM) for machine learning. QSM uses wave functions that describe microscopic particle systems as mappings. By QSM, original inputs or features extracted by neural networks are processed as quantum states to train wave function parameters. QSM has a…

Cited by 0SourceScholar
2023

Deep Fusion Transformer Network with Weighted Vector-Wise Keypoints Voting for Robust 6D Object Pose Estimation

ICCV 2023poster

One critical challenge in 6D object pose estimation from a single RGBD image is efficient integration of two different modalities, i.e., color and depth. In this work, we tackle this problem by a novel Deep Fusion Transformer (DFTr) block that can aggregate cross-modality features for improving pose…

Cited by 42PDFcodeScholar
2023

Difference-in-Differences Meets Tree-based Methods: Heterogeneous Treatment Effects Estimation with Unmeasured Confounding

ICML 2023poster

This study considers the estimation of conditional causal effects in the presence of unmeasured confounding for a balanced panel with treatment imposed at the last time point. To address this, we combine Difference-in-differences (DiD) and tree-based methods and propose a new identification assumpti…

Cited by 2SourcePDFScholar
2023

Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis

ICCV 2023poster

This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art performance with small model size. Our idea is to explicitly exploit the unequal…

Cited by 84PDFcodeScholar
2023

FAST: a Fused and Accurate Shrinkage Tree for Heterogeneous Treatment Effects Estimation

NeurIPS 2023poster

This paper proposes a novel strategy for estimating the heterogeneous treatment effect called the Fused and Accurate Shrinkage Tree ($\mathrm{FAST}$). Our approach utilizes both trial and observational data to improve the accuracy and robustness of the estimator. Inspired by the concept of shrinkag…

Cited by 1SourcePDFScholar
2023

Few-shot Classification via Ensemble Learning with Multi-Order Statistics

IJCAI 2023poster

Transfer learning has been widely adopted for few-shot classification. Recent studies reveal that obtaining good generalization representation of images on novel classes is the key to improving the few-shot classification accuracy. To address this need, we prove theoretically that leveraging ensembl…

Cited by 9SourcePDFScholar
2023

Keep Skills in Mind: Understanding and Implementing Skills in Commonsense Question Answering

IJCAI 2023poster

Commonsense Question Answering (CQA) aims to answer questions that require human commonsense. Closed-book CQA, as one of the subtasks, requires the model to answer questions without retrieving external knowledge, which emphasizes the importance of the model's problem-solving ability. Most previous m…

2023

Language Models Can Improve Event Prediction by Few-Shot Abductive Reasoning

NeurIPS 2023poster

Large language models have shown astonishing performance on a wide range of reasoning tasks. In this paper, we investigate whether they could reason about real-world events and help improve the prediction performance of event sequence models. We design LAMP, a framework that integrates a large langu…

Cited by 52SourcePDFScholar
2023

Prompt-augmented Temporal Point Process for Streaming Event Sequence

NeurIPS 2023poster

Neural Temporal Point Processes (TPPs) are the prevalent paradigm for modeling continuous-time event sequences, such as user activities on the web and financial transactions. In real world applications, the event data typically comes in a streaming manner, where the distribution of the patterns may…

Cited by 26SourcePDFScholar
2023

Towards Anytime Fine-tuning: Continually Pre-trained Language Models with Hypernetwork Prompts

EMNLP 2023long findings

Continual pre-training has been urgent for adapting a pre-trained model to a multitude of domains and tasks in the fast-evolving world. In practice, a continually pre-trained model is expected to demonstrate not only greater capacity when fine-tuned on pre-trained domains but also a non-decreasing p…

Cited by 0SourceScholar
2023

Unleashing the Power of Graph Data Augmentation on Covariate Distribution Shift

NeurIPS 2023poster

The issue of distribution shifts is emerging as a critical concern in graph representation learning. From the perspective of invariant learning and stable learning, a recently well-established paradigm for out-of-distribution generalization, stable features of the graph are assumed to causally deter…

2022

A Generalized Kernel Risk Sensitive Loss for Robust Two-Dimensional Singular Value Decomposition

ICASSP 2022accepted

Two-dimensional singular value decomposition (2DSVD) is an important dimensionality reduction algorithm which has inherent advantage in preserving the structure of 2D images. However, 2DSVD algorithm is based on the squared error loss, which may exaggerate the projection errors with the presence of…

Cited by 0SourceScholar
2022

Ask Question First for Enhancing Lifelong Language Learning

COLING 2022main

Lifelong language learning aims to stream learning NLP tasks while retaining knowledge of previous tasks. Previous works based on the language model and following data-free constraint approaches have explored formatting all data as “begin token (B) + context (C) + question (Q) + answer (A)” for diff…

2022

Attribute-Conditioned Face Swapping Network for Low-Resolution Images

ICASSP 2022accepted

Deep learning based face swapping technologies have opened new frontiers for entertainment industries while pose novel threats to identity security. Applying face swapping to real-world products, as well as defending against its misuse, rely on the capacity to generate high quality face swapped imag…

Cited by 0SourceScholar
2022

Debiased Causal Tree: Heterogeneous Treatment Effects Estimation with Unmeasured Confounding

NeurIPS 2022accept

Unmeasured confounding poses a significant threat to the validity of causal inference. Despite that various ad hoc methods are developed to remove confounding effects, they are subject to certain fairly strong assumptions. In this work, we consider the estimation of conditional causal effects in the…

Cited by 13SourcePDFScholar
2022

Learning Mixture of Neural Temporal Point Processes for Multi-dimensional Event Sequence Clustering

IJCAI 2022poster

Multi-dimensional event sequence clustering applies to many scenarios e.g. e-Commerce and electronic health. Traditional clustering models fail to characterize complex real-world processes due to the strong parametric assumption. While Neural Temporal Point Processes (NTPPs) mainly focus on modeling…

Cited by 15SourcePDFScholar
2022

Material-Guided Siamese Fusion Network for Hyperspectral Object Tracking

ICASSP 2022accepted

Hyperspectral videos (HSVs) have more potential in target tracking than color videos thanks to the material identification capability provided by abundant spectral bands. Due to limited HSVs for training, most current hyperspectral trackers are based on hand-crafted features rather than deeply learn…

Cited by 0SourceScholar
2022

Multitask Sparse Neural Network for Hyperspectral Image Denoising

ICASSP 2022accepted

Data-driven deep learning (DL)-based methods directly learn the nonlinear mapping between noisy hyperspectral images (HSIs) and corresponding clean ones. However, DLbased methods neglect the prior knowledge of HSIs embodied by physical models. Consequently, they require complex network architectures…

Cited by 0SourceScholar
2022

Regularizing Graph Neural Networks via Consistency-Diversity Graph Augmentations

AAAI 2022technical

Despite the remarkable performance of graph neural networks (GNNs) in semi-supervised learning, it is criticized for not making full use of unlabeled data and suffering from over-fitting. Recently, graph data augmentation, used to improve both accuracy and generalization of GNNs, has received consid…

Cited by 30SourcePDFScholar
2022

Revisiting Domain Generalized Stereo Matching Networks From a Feature Consistency Perspective

CVPR 2022poster

Despite recent stereo matching networks achieving impressive performance given sufficient training data, they suffer from domain shifts and generalize poorly to unseen domains. We argue that maintaining feature consistency between matching pixels is a vital factor for promoting the generalization ca…

Cited by 79PDFcodeScholar
2022

Robust Heterogeneous Graph Neural Networks against Adversarial Attacks

AAAI 2022technical

Heterogeneous Graph Neural Networks (HGNNs) have drawn increasing attention in recent years and achieved outstanding performance in many tasks. However, despite their wide use, there is currently no understanding of their robustness to adversarial attacks. In this work, we first systematically study…

Cited by 64SourcePDFScholar
2022

SAIL: Self-Augmented Graph Contrastive Learning

AAAI 2022technical

This paper studies learning node representations with graph neural networks (GNNs) for unsupervised scenario. Specifically, we derive a theoretical analysis and provide an empirical demonstration about the non-steady performance of GNNs over different graph datasets, when the supervision signals are…

Cited by 47SourcePDFScholar
2022

Vertically Federated Graph Neural Network for Privacy-Preserving Node Classification

IJCAI 2022poster

Recently, Graph Neural Network (GNN) has achieved remarkable progresses in various real-world tasks on graph data, consisting of node features and the adjacent information between different nodes. High-performance GNN models always depend on both rich features and complete edge information in graph.…

Cited by 132SourcePDFScholar
2022

Where to Focus: Investigating Hierarchical Attention Relationship for Fine-Grained Visual Classification

ECCV 2022poster

"Object categories are often grouped into a multi-granularity taxonomic hierarchy. Classifying objects at coarser-grained hierarchy requires global and common characteristics, while finer-grained hierarchy classification relies on local and discriminative features. Therefore, humans should also subc…

2021

A Bi-Level Framework for Learning to Solve Combinatorial Optimization on Graphs

NeurIPS 2021poster

Combinatorial Optimization (CO) has been a long-standing challenging research topic featured by its NP-hard nature. Traditionally such problems are approximately solved with heuristic algorithms which are usually fast but may sacrifice the solution quality. Currently, machine learning for combinator…

2021

Attention-based Pyramid Dilated Lattice Network for Blind Image Denoising

IJCAI 2021poster

Though convolutional neural networks (CNNs) with residual and dense aggregations have obtained much attention in image denoising, they are incapable of exploiting different levels of contextual information at every convolutional unit in order to infer different levels of noise components with a sing…

Cited by 6SourcePDFScholar
2021

Cross-Domain Recommendation: Challenges, Progress, and Prospects

IJCAI 2021poster

To address the long-standing data sparsity problem in recommender systems (RSs), cross-domain recommendation (CDR) has been proposed to leverage the relatively richer information from a richer domain to improve the recommendation performance in a sparser domain. Although CDR has been extensively stu…

2021

Decomposing Complex Questions Makes Multi-Hop QA Easier and More Interpretable

EMNLP 2021finding

Multi-hop QA requires the machine to answer complex questions through finding multiple clues and reasoning, and provide explanatory evidence to demonstrate the machine’s reasoning process. We propose Relation Extractor-Reader and Comparator (RERC), a three-stage framework based on complex question d…

2021

Goal-Oriented Gaze Estimation for Zero-Shot Learning

CVPR 2021poster

Zero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen classes. Since semantic knowledge is built on attributes shared between different classes, which are highly local, strong prior for localization of object attribute is beneficial f…

Cited by 172PDFcodeScholar
2021

Improving Transferability of Adversarial Patches on Face Recognition With Generative Models

CVPR 2021poster

Face recognition is greatly improved by deep convolutional neural networks (CNNs). Recently, these face recognition models have been used for identity authentication in security sensitive applications. However, deep CNNs are vulnerable to adversarial patches, which are physically realizable and stea…

Cited by 128PDFScholar
2021

MixSeq: Connecting Macroscopic Time Series Forecasting with Microscopic Time Series Data

NeurIPS 2021poster

Time series forecasting is widely used in business intelligence, e.g., forecast stock market price, sales, and help the analysis of data trend. Most time series of interest are macroscopic time series that are aggregated from microscopic data. However, instead of directly modeling the macroscopic ti…

Cited by 21SourcePDFScholar
2021

NMF-SAE: An Interpretable Sparse Autoencoder for Hyperspectral Unmixing

ICASSP 2021accepted

Hyperspectral unmixing is an important tool to learn the material constitution and distribution of a scene. Model-based unmixing methods depend on well-designed iterative optimization algorithms, which is usually time consuming. Learning-based methods perform unmixing in a data-driven manner but hea…

Cited by 0SourceScholar
2020

Bandit Samplers for Training Graph Neural Networks

NeurIPS 2020poster

Several sampling algorithms with variance reduction have been proposed for accelerating the training of Graph Convolution Networks (GCNs). However, due to the intractable computation of optimal sampling distribution, these sampling algorithms are suboptimal for GCNs and are not applicable to more g…

2020

Financial Risk Analysis for SMEs with Graph-based Supply Chain Mining

IJCAI 2020poster

Small and Medium-sized Enterprises (SMEs) are playing a vital role in the modern economy. Recent years, financial risk analysis for SMEs attracts lots of attentions from financial institutions. However, the financial risk analysis for SMEs usually suffers data deficiency problem, especially for the…

Cited by 0SourcePDFScholar
2019

Generalization in Generative Adversarial Networks: A Novel Perspective from Privacy Protection

NeurIPS 2019poster

In this paper, we aim to understand the generalization properties of generative adversarial networks (GANs) from a new perspective of privacy protection. Theoretically, we prove that a differentially private learning algorithm used for training the GAN does not overfit to a certain degree, i.e., the…

Cited by 57SourcePDFScholar