← Search

Wei Wei

141 accepted papers

2026

An Open-Ended Benchmark and Formal Framework for Adjuvant Research with MLLM

ICLR 2026poster

Adjuvants play a critical role in modulating immune responses and are central to the development of vaccines and immunotherapies. Yet progress in this field is constrained by data scarcity and incomplete understanding of mechanisms of action, which limit the transition from experience-based design t…

Cited by 0SourceScholar
2026

Appearance Discrepancy-guided Sequence Hybrid Masking for Robust Scene Text Recognition

AAAI 2026technical

Masked Image Modeling (MIM) has been widely recognized as a powerful self-supervised paradigm for learning general-purpose visual representations. However, standard MIM based on random masking tends to underperform in domain-specific tasks like Scene Text Recognition (STR), due to challenges such as

Cited by 0SourcePDFScholar
2026

DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning

ICLR 2026poster

Unlearning in Large Language Models (LLMs) is crucial for protecting private data and removing harmful knowledge. Most existing approaches rely on fine-tuning to balance unlearning efficiency with general language capabilities. However, these methods typically require training or access to retain da…

Cited by 0SourcecodeScholar
2026

Efficient Offline Reinforcement Learning via Peer-Influenced Constraint

ICLR 2026poster

Offline reinforcement learning (RL) seeks to learn an optimal policy from a fixed dataset, but distributional shift between the dataset and the learned policy often leads to suboptimal real-world performance. Existing methods typically use behavior policy regularization to constrain the learned poli…

Cited by 0SourceScholar
2026

HumanLM: Simulating Users with State Alignment Beats Response Imitation

ICML 2026poster

Large Language Models (LLMs) are increasingly used to simulate how specific users respond to any context, enabling more user-centric applications that rely on user feedback. However, existing user simulators mostly imitate surface-level patterns and language styles, which fails to reflect the underl…

Cited by 0SourceScholar
2026

Inference Time Optimization with Confidence Dynamics

ICML 2026poster

Inference time optimization techniques, such as repeated sampling, have significantly advanced the reasoning capabilities of Large Language Models (LLMs). However, the critical role of model uncertainty remains largely underexplored in these optimization strategies. In this paper, we investigate the…

Cited by 0SourceScholar
2026

JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation Promotion

AAAI 2026technical

Given the inherently costly and time-intensive nature of pixel-level annotation, the generation of synthetic datasets comprising sufficiently diverse synthetic images paired with ground-truth pixel-level annotations has garnered increasing attention recently for training high-performance semantic se

Cited by 0SourcePDFScholar
2026

Language Does Matter for Cross-Domain Few-Shot Visual Feature Enhancement

CVPR 2026

Cross-domain few-shot image interpretation (CD-FSII) has been significantly advanced by fine-tuning pre-trained visual feature models using limited labeled samples in target domains. However, profound cross-domain distribution discrepancies, along with inherent conflicts between extensive object vis

Cited by 0SourcecodeScholar
2026

ProRefine: Inference-Time Prompt Refinement with Textual Feedback (Student Abstract)

AAAI 2026technical

Agentic workflows, where multiple AI agents collaborate to accomplish complex tasks like reasoning or planning, play a substantial role in many cutting-edge commercial applications. These workflows depend critically on the prompts used to provide the roles models play in such workflows. Poorly desig

Cited by 0SourcePDFScholar
2026

Q-SAM: Unlocking Sharpness-Aware Minimization for Generalization in Offline Reinforcement Learning

ICML 2026poster

Generalization remains a central challenge in offline reinforcement learning (RL), where policies are trained solely from static datasets and must perform reliably under distribution shift. While most existing offline RL methods focus on reducing training loss using standard optimizers such as Adam,…

Cited by 0SourceScholar
2026

RDF-MIG: A Robust Diffusion Framework for Masked Image Generation to Augment Semantic Segmentation and Change Detection

CVPR 2026

Change detection and semantic segmentation are key techniques for satellite image analysis in remote sensing. However, acquiring high-quality labeled data is costly and time-consuming. Although recent studies have explored generative models to ease data scarcity, a unified framework supporting both

Cited by 0SourceScholar
2026

RIVS: Mitigating Hallucination in Large Vision-Language Models via Representation Intervention on Visual Grounding Shift

IJCAI 2026

Large Vision-Language Models (LVLMs) demonstrate powerful generative capabilities yet remain prone to object hallucinations. Most existing methods mitigate this issue through training or decoding strategies, but provide limited exploration of how hallucinations arise from internal representations du

Cited by 0Scholar
2026

Representation Alignment for Diffusion Transformers without External Components

ICLR 2026poster

Recent studies have demonstrated that learning a meaningful internal represen- tation can accelerate generative training. However, existing approaches necessi- tate to either introduce an off-the-shelf external representation task or rely on a large-scale, pre-trained external representation encoder…

Cited by 0SourcecodeScholar
2026

Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding

AAAI 2026technical

Multimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering f

Cited by 0SourcePDFScholar
2026

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet their performance degrades significantly under real-world visual corruptions. While existing robustness enhancement approaches exist, they are limited: black-box feature alignment lacks interpr…

Cited by 0SourceScholar
2026

ScaleADFG: Affordance-Based Dexterous Functional Grasping via Scalable Dataset

RA-L 2026

Dexterous functional tool-use grasping is essential for effective robotic manipulation of tools. However, existing approaches face significant challenges in efficiently constructing large-scale datasets and ensuring generalizability to everyday object scales. These issues primarily arise from size m

Cited by 1SourcecodeScholar
2026

TANGO: Text-Anchored Guided Optimization for Robust Fine-tuning Vision-Language Models under Label Noise

CVPR 2026

Fine-tuning large-scale Vision-Language Models (VLMs) is crucial for specialized tasks, but their performance is often undermined by the label noise prevalent in real-world datasets. Traditional approaches to learning with noisy labels typically rely on a self-referential loop, using a model's own p

Cited by 0SourceScholar
2026

VA-p: Variational Policy Alignment for Pixel-Aware Autoregressive Generation

CVPR 2026

Autoregressive (AR) visual generation relies on tokenizers to map images to and from discrete sequences. However, tokenizers are trained to reconstruct clean images from ground-truth tokens, while AR generators are optimized only for token likelihood. This misalignment leads to generated token seque

Cited by 0SourcecodeScholar
2025

Avoiding Knowledge Edit Skipping in Multi-hop Question Answering with Guided Decomposition

EMNLP 2025

In a rapidly evolving world where information updates swiftly, knowledge in large language models (LLMs) becomes outdated quickly. Retraining LLMs is not a cost-effective option, making knowledge editing (KE) without modifying parameters particularly necessary. We find that although existing retriev

Cited by 0SourcePDFScholar
2025

Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive Injection

ICCV 2025poster

Controllable diffusion models have been widely applied in image stylization. However, existing methods often treat the style in the reference image as a single, indivisible entity, which makes it difficult to transfer specific stylistic attributes. To address this issue, we propose a fine-grained co…

2025

CoMIF: Modeling of Complex Multiple Interaction Factors for Conversation Generation

COLING 2025main

Highly realistic human-machine interaction is challenging for open-domain dialogue systems. Although existing methods have achieved notable progress by leveraging various interaction factors (e.g., emotion, personality, topic) for delivering human-like (e.g., empathetic, personalized and semanticall…

Cited by 2SourcePDFScholar
2025

Cooperative or Competitive? Understanding the Interaction between Attention Heads From A Game Theory Perspective

ACL 2025long

Despite the remarkable success of attention-based large language models (LLMs), the precise interaction mechanisms between attention heads remain poorly understood. In contrast to prevalent methods that focus on individual head contributions, we rigorously analyze the intricate interplay among atten…

2025

Dynamic Uncertainty Estimation for Offline Reinforcement Learning

AAAI 2025technical

Offline reinforcement learning confronts the distributional shift challenge, a consequence of learning policy from static datasets. Current methods primarily handle this issue by aligning the learned policy with the behavior policy or conservatively estimating Q-values for out-of-distribution (OOD)…

Cited by 0SourcePDFScholar
2025

Enhanced Sample Selection with Confidence Tracking: Identifying Correctly Labeled Yet Hard-to-Learn Samples in Noisy Data

AAAI 2025technical

We propose a novel sample selection method for image classification in the presence of noisy labels. Existing methods typically consider small-loss samples as correctly labeled. However, some correctly labeled samples are inherently difficult for the model to learn and can exhibit high loss similar…

2025

EvoPrompt: Evolving Prompts for Enhanced Zero-Shot Named Entity Recognition with Large Language Models

COLING 2025main

Large language models (LLMs) possess extensive prior knowledge and powerful in-context learning (ICL) capabilities, presenting significant opportunities for low-resource tasks. Though effective, several key issues still have not been well-addressed when focusing on zero-shot named entity recognition…

2025

Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think

CVPR 2025highlight

Image-to-Video (I2V) generation aims to synthesize a video clip according to a given image and condition (e.g., text). The key challenge of this task lies in simultaneously generating natural motions while preserving the original appearance of the images. However, current I2V diffusion models (I2V-D…

2025

FANS: A Flatness-Aware Network Structure for Generalization in Offline Reinforcement Learning

NeurIPS 2025poster

Offline reinforcement learning (RL) aims to learn optimal policies from static datasets while enhancing generalization to out-of-distribution (OOD) data. To mitigate overfitting to suboptimal behaviors in offline datasets, existing methods often relax constraints on policy and data or extract inform…

Cited by 0SourceScholar
2025

From Isolated Conversations to Hierarchical Schemas: Dynamic Tree Memory Representation for LLMs

ICLR 2025poster

Recent advancements in large language models have significantly improved their context windows, yet challenges in effective long-term memory management remain. We introduce MemTree, an algorithm that leverages a dynamic, tree-structured memory representation to optimize the organization, retrieval,…

Cited by 3SourcePDFScholar
2025

GraspAgent 1.0: Adversarial Continual Dexterous Grasp Learning

RA-L 2025

Grasp is at the core of robotic manipulation tasks. Nonetheless, most 6-DOF methods resort to a one-time setup via intensive analytics and targeting a predetermined domain. On the other hand, learning and adapting in real environments is of great promise to robotics yet challenging. In this context,

Cited by 0SourceScholar
2025

Improving Data Efficiency via Curating LLM-Driven Rating Systems

ICLR 2025poster

Instruction tuning is critical for adapting large language models (LLMs) to downstream tasks, and recent studies have demonstrated that small amounts of human-curated data can outperform larger datasets, challenging traditional data scaling laws. While LLM-based data quality rating systems offer a c…

Cited by 3SourcePDFScholar
2025

Improving Generalization in Offline Reinforcement Learning via Latent Distribution Representation Learning

AAAI 2025technical

Dealing with the distribution shift is a significant challenge when building offline reinforcement learning (RL) models that can generalize from a static dataset to out-of-distribution (OOD) scenarios. Previous approaches have employed pessimism or conservatism strategies. More recently, data-driven…

Cited by 0SourcePDFScholar
2025

Indirect Alignment and Relationship Preservation for Domain Generalization

IJCAI 2025

Domain generalization (DG) aims to train models on multiple source domains to generalize effectively to unseen target domains, addressing performance degradation caused by domain shifts. Many existing methods rely on direct feature alignment, which disrupts natural sequence relationships, causes mis

Cited by 0SourcePDFScholar
2025

KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse

NeurIPS 2025poster

We describe KVLink, an approach for efficient key-value (KV) cache reuse in large language models (LLMs). In many LLM applications, different inputs can share overlapping context, such as the same retrieved document appearing in multiple queries. However, the LLMs still need to encode the entire con…

Cited by 0SourcecodeScholar
2025

LLM Unlearning via Loss Adjustment with Only Forget Data

ICLR 2025poster

Unlearning in Large Language Models (LLMs) is essential for ensuring ethical and responsible AI use, especially in addressing privacy leak, bias, safety, and evolving regulations. Existing approaches to LLM unlearning often rely on retain data or a reference LLM, yet they struggle to adequately bala…

Cited by 2SourcePDFScholar
2025

Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning

COLING 2025main

Recently, Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multi-modal context comprehension. However, they still suffer from hallucination problems referring to generating inconsistent outputs with the image content. To mitigate hallucinations, previous studies main…

2025

Low-Biased General Annotated Dataset Generation

CVPR 2025poster

Pre-training backbone networks on a general annotated dataset (e.g., ImageNet) that comprises numerous manually collected images with category annotations has proven to be indispensable for enhancing the generalization capacity of downstream visual tasks. However, those manually collected images oft…

2025

MMSciBench: Benchmarking Language Models on Chinese Multimodal Scientific Problems

ACL 2025finding

Recent advances in large language models (LLMs) and vision-language models (LVLMs) have shown promise across many tasks, yet their scientific reasoning capabilities remain untested, particularly in multimodal settings. We present MMSciBench, a benchmark for evaluating mathematical and physical reaso…

2025

Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment

ICML 2025poster

While Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning for Large Language Models (LLMs), its performance often falls short of Full Fine-Tuning (Full FT). Current methods optimize LoRA by initializing with static singular value decomposition (SVD) subsets, leading to suboptimal lev…

2025

MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation

IROS 2025

Human action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action recognition, the mask-based reconstruction paradigm learns the spatial structure and motion patterns of the skeleton by

Cited by 0SourcecodeScholar
2025

Modal Feature Optimization Network with Prompt for Multimodal Sentiment Analysis

COLING 2025main

Multimodal sentiment analysis(MSA) is mostly used to understand human emotional states through multimodal. However, due to the fact that the effective information carried by multimodal is not balanced, the modality containing less effective information cannot fully play the complementary role betwee…

2025

Multi-granularity Knowledge Transfer for Continual Reinforcement Learning

IJCAI 2025

Continual reinforcement learning (CRL) empowers RL agents with the ability to learn a sequence of tasks, accumulating knowledge learned in the past and using the knowledge for problemsolving or future task learning. However, existing methods often focus on transferring fine-grained knowledge across

Cited by 0SourcePDFScholar
2025

Multi-level Association Refinement Network for Dialogue Aspect-based Sentiment Quadruple Analysis

ACL 2025long

Dialogue Aspect-based Sentiment Quadruple (DiaASQ) analysis aims to identify all quadruples (i.e., target, aspect, opinion, sentiment) from the dialogue. This task is challenging as different elements within a quadruple may manifest in different utterances, requiring precise handling of associations…

Cited by 0SourcePDFScholar
2025

Prompt-Free Conditional Diffusion for Multi-object Image Augmentation

IJCAI 2025

Diffusion model has underpinned much recent advances of dataset augmentation in various computer vision tasks. However, when involving generating multi-object images as real scenarios, most existing methods either rely entirely on text condition, resulting in a deviation between the generated object

2025

RhythmGuassian: Repurposing Generalizable Gaussian Model For Remote Physiological Measurement

ICCV 2025poster

Remote Photoplethysmography (rPPG) enables non-contact extraction of physiological signals, providing significant advantages in medical monitoring, emotion recognition, and face anti-spoofing. However, the extraction of reliable rPPG signals is hindered by motion variations in real-world environment…

2025

Risk-aware Direct Preference Optimization under Nested Risk Measure

NeurIPS 2025poster

When fine-tuning pre-trained Large Language Models (LLMs) to align with human values and intentions, maximizing the estimated reward can lead to superior performance, but it also introduces potential risks due to deviations from the reference model's intended behavior. Most existing methods typicall…

Cited by 0SourcecodeScholar
2025

Selecting and Merging: Towards Adaptable and Scalable Named Entity Recognition with Large Language Models

ACL 2025long

Supervised fine-tuning (SFT) is widely used to align large language models (LLMs) with information extraction (IE) tasks, such as named entity recognition (NER). However, annotating such fine-grained labels and training domain-specific models is costly. Existing works typically train a unified model…

2025

Semantic Enhanced Heterogeneous Hypergraph Network for Collaborative Filtering

AAAI 2025technical

Collaborative Filtering (CF) based on graph neural networks (GNNs) has yielded immense success for recommendation systems by capturing high-order dependencies from implicit feedback. Recently, the outstanding text comprehension ability of the Large Language Models (LLMs) has shown promising potentia…

Cited by 0SourcePDFScholar
2025

Towards Effective Foundation Model Adaptation for Extreme Cross-Domain Few-Shot Learning

ICCV 2025poster

Large-scale pre-trained foundation models have demonstrated remarkable generalization capabilities across diverse computer vision tasks through fine-tuning. However, existing fine-tuning approaches often encounter challenges in extreme cross-domain few-shot learning scenarios, primarily due to the s…

2025

Towards Physically Realizable Adversarial Attacks in Embodied Vision Navigation

IROS 2025

The significant advancements in embodied vision navigation have raised concerns about its susceptibility to adversarial attacks exploiting deep neural networks. Investigating the adversarial robustness of embodied vision navigation is crucial, especially given the threat of 3D physical attacks that

Cited by 7SourcecodeScholar
2025

Zero-Shot Cross-Domain Aspect-Based Sentiment Analysis via Domain-Contextualized Chain-of-Thought Reasoning

EMNLP 2025

Cross-domain aspect-based sentiment analysis (ABSA) aims at learning specific knowledge from a source domain to perform various ABSA tasks on a target domain. Recent works mainly focus on how to use domain adaptation techniques to transfer the domain-agnostic features from the labeled source domain

Cited by 0SourcePDFScholar
2024

4D Gaussian Splatting for Real-Time Dynamic Scene Rendering

CVPR 2024poster

Representing and rendering dynamic scenes has been an important but challenging task. Especially to accurately model complex motions high efficiency is usually hard to guarantee. To achieve real-time dynamic scene rendering while also enjoying high training and storage efficiency we propose 4D Gauss…

2024

BBox-Adapter: Lightweight Adapting for Black-Box Large Language Models

ICML 2024spotlight

Adapting state-of-the-art Large Language Models (LLMs) like GPT-4 and Gemini for specific tasks is challenging. Due to the opacity in their parameters, embeddings, and even output probabilities, existing fine-tuning adaptation methods are inapplicable. Consequently, adapting these black-box LLMs is…

2024

Confidence is not Timeless: Modeling Temporal Validity for Rule-based Temporal Knowledge Graph Forecasting

ACL 2024long

Recently, Temporal Knowledge Graph Forecasting (TKGF) has emerged as a pivotal domain for forecasting future events. Unlike black-box neural network methods, rule-based approaches are lauded for their efficiency and interpretability. For this line of work, it is crucial to correctly estimate the pre…

Cited by 8SourcePDFScholar
2024

Detection-Based Intermediate Supervision for Visual Question Answering

AAAI 2024technical

Recently, neural module networks (NMNs) have yielded ongoing success in answering compositional visual questions, especially those involving multi-hop visual and logical reasoning. NMNs decompose the complex question into several sub-tasks using instance-modules from the reasoning paths of that ques…

2024

Enhancing Low-Resource Relation Representations through Multi-View Decoupling

AAAI 2024technical

Recently, prompt-tuning with pre-trained language models (PLMs) has demonstrated the significantly enhancing ability of relation extraction (RE) tasks. However, in low-resource scenarios, where the available training data is scarce, previous prompt-based methods may still perform poorly for prompt-…

2024

FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained Understanding

NeurIPS 2024poster

Contrastive Language-Image Pre-training (CLIP) achieves impressive performance on tasks like image classification and image-text retrieval by learning on large-scale image-text datasets. However, CLIP struggles with dense prediction tasks due to the poor grasp of the fine-grained details. Although e…

Cited by 3SourcePDFScholar
2024

HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models

CVPR 2024poster

Personalization has emerged as a prominent aspect within the field of generative AI enabling the synthesis of individuals in diverse contexts and styles while retaining high-fidelity to their identities. However the process of personalization presents inherent challenges in terms of time and memory…

Cited by 191SourcePDFScholar
2024

Improving Generalization in Offline Reinforcement Learning via Adversarial Data Splitting

ICML 2024poster

Offline Reinforcement Learning (RL) commonly suffers from the out-of-distribution (OOD) overestimation issue due to the distribution shift. Prior work gradually shifts their focus from suppressing OOD overestimation to avoiding overly conservative learning from suboptimal behavior policies to improv…

2024

Improving Pseudo Labels with Global-Local Denoising Framework for Cross-lingual Named Entity Recognition

IJCAI 2024poster

Cross-lingual named entity recognition (NER) aims to train an NER model for the target language leveraging only labeled source language data and unlabeled target language data. Prior approaches either perform label projection on translated source language data or employ a source model to assign pseu…

2024

Joint Multi-Facts Reasoning Network for Complex Temporal Question Answering Over Knowledge Graph

ICASSP 2024accepted

Temporal Knowledge Graph (TKG) is an extension of regular knowledge graph by attaching the time scope. Existing temporal knowledge graph question answering (TKGQA) models solely approach simple questions, owing to the prior assumption that each question only contains a single temporal fact with expl…

Cited by 0SourceScholar
2024

Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph Reasoning

NeurIPS 2024poster

Temporal Knowledge Graph Reasoning (TKGR) is the process of utilizing temporal information to capture complex relations within a Temporal Knowledge Graph (TKG) to infer new knowledge. Conventional methods in TKGR typically depend on deep learning algorithms or temporal logical rules. However, deep l…

2024

Learning Realistic and Reasonable Grasps for Anthropomorphic Hand in Cluttered Scenes

ICRA 2024poster

Grasping is one of the most fundamental skills for humans to interact with objects. However, it remains a challenging problem for anthropomorphic hands, due to the lack of object affordance understanding and high-dimensional grasp planning. In this work, we propose an anthropomorphic hand grasping f…

Cited by 2SourceScholar
2024

Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot Learning

NeurIPS 2024poster

Meta-learning offers a promising avenue for few-shot learning (FSL), enabling models to glean a generalizable feature embedding through episodic training on synthetic FSL tasks in a source domain. Yet, in practical scenarios where the target task diverges from that in the source domain, meta-learnin…

Cited by 1SourcePDFScholar
2024

Mitigating Boundary Ambiguity and Inherent Bias for Text Classification in the Era of Large Language Models

ACL 2024findings

Text classification is a crucial task encountered frequently in practical scenarios, yet it is still under-explored in the era of large language models (LLMs). This study shows that LLMs are vulnerable to changes in the number and arrangement of options in text classification. Our extensive empirica…

2024

On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits Fusion

NeurIPS 2024poster

Efficient fine-tuning of large language models for task-specific applications is imperative, yet the vast number of parameters in these models makes their training increasingly challenging. Despite numerous proposals for effective methods, a substantial memory overhead remains for gradient computati…

Cited by 4SourcePDFScholar
2024

PRDP: Proximal Reward Difference Prediction for Large-Scale Reward Finetuning of Diffusion Models

CVPR 2024poster

Reward finetuning has emerged as a promising approach to aligning foundation models with downstream objectives. Remarkable success has been achieved in the language domain by using reinforcement learning (RL) to maximize rewards that reflect human preference. However in the vision domain existing RL…

Cited by 15SourcePDFScholar
2024

Personalized Topic Selection Model for Topic-Grounded Dialogue

ACL 2024findings

Recently, the topic-grounded dialogue (TGD) system has become increasingly popular as its powerful capability to actively guide users to accomplish specific tasks through topic-guided conversations. Most existing works utilize side information (e.g. topics or personas) in isolation to enhance the to…

2024

Position Debiasing Fine-Tuning for Causal Perception in Long-Term Dialogue

IJCAI 2024poster

The core of the dialogue system is to generate relevant, informative, and human-like responses based on extensive dialogue history. Recently, dialogue generation domain has seen mainstream adoption of large language models (LLMs), due to its powerful capability in generating utterances. However, the…

Cited by 2SourcePDFScholar
2024

Reinforcement Learning with Token-level Feedback for Controllable Text Generation

NAACL 2024findings

To meet the requirements of real-world applications, it is essential to control generations of large language models (LLMs). Prior research has tried to introduce reinforcement learning (RL) into controllable text generation while most existing methods suffer from overfitting issues (finetuning-base…

2024

Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement Learning

NeurIPS 2024poster

A challenging problem in seeking to bring multi-agent reinforcement learning (MARL) techniques into real-world applications, such as autonomous driving and drone swarms, is how to control multiple agents safely and cooperatively to accomplish tasks. Most existing safe MARL methods learn the centrali…

Cited by 1SourcePDFScholar
2024

Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging

NeurIPS 2024poster

In the era of large language models, model merging is a promising way to combine multiple task-specific models into a single multitask model without extra training. However, two challenges remain: (a) interference between different models and (b) heterogeneous data during testing. Traditional model…

2023

An Empirical Study on the Language Modal in Visual Question Answering

IJCAI 2023poster

Generalization beyond in-domain experience to out-of-distribution data is of paramount significance in the AI domain. Of late, state-of-the-art Visual Question Answering (VQA) models have shown impressive performance on in-domain data, partially due to the language prior bias which, however, hinders…

Cited by 7SourcePDFScholar
2023

AttenWalker: Unsupervised Long-Document Question Answering via Attention-based Graph Walking

ACL 2023findings

Annotating long-document question answering (long-document QA) pairs is time-consuming and expensive. To alleviate the problem, it might be possible to generate long-document QA pairs via unsupervised question answering (UQA) methods. However, existing UQA tasks are based on short documents, and can…

2023

Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework

ACL 2023findings

Despite recent success on various tasks, deep learning techniques still perform poorly on adversarial examples with small perturbations. While optimization-based methods for adversarial attacks are well-explored in the field of computer vision, it is impractical to directly apply them in natural lan…

2023

Glocal Energy-Based Learning for Few-Shot Open-Set Recognition

CVPR 2023poster

Few-shot open-set recognition (FSOR) is a challenging task of great practical value. It aims to categorize a sample to one of the pre-defined, closed-set classes illustrated by few examples while being able to reject the sample from unknown classes. In this work, we approach the FSOR task by proposi…

2023

Mind the Gap: Polishing Pseudo Labels for Accurate Semi-supervised Object Detection

AAAI 2023technical

Exploiting pseudo labels (e.g., categories and bounding boxes) of unannotated objects produced by a teacher detector have underpinned much of recent progress in semi-supervised object detection (SSOD). However, due to the limited generalization capacity of the teacher detector caused by the scarce a…

2023

Miracle: Towards Personalized Dialogue Generation with Latent-Space Multiple Personal Attribute Control

EMNLP 2023long findings

Personalized dialogue systems aim to endow the chatbot agent with more anthropomorphic traits for human-like interactions. Previous approaches have explored explicitly user profile modeling using text descriptions, implicit derivation of user embeddings, or utilizing handicraft prompts for ChatGPT-…

Cited by 0SourcecodeScholar
2023

On Task-personalized Multimodal Few-shot Learning for Visually-rich Document Entity Retrieval

EMNLP 2023long findings

Visually-rich document entity retrieval (VDER), which extracts key information (e.g. date, address) from document images like invoices and receipts, has become an important topic in industrial NLP applications. The emergence of new document types at a constant pace, each with its unique entity types…

Cited by 0SourceScholar
2023

Revisiting Prototypical Network for Cross Domain Few-Shot Learning

CVPR 2023poster

Prototypical Network is a popular few-shot solver that aims at establishing a feature metric generalizable to novel few-shot classification (FSC) tasks using deep neural networks. However, its performance drops dramatically when generalizing to the FSC tasks in new domains. In this study, we revisit…

2023

STAGE: Span Tagging and Greedy Inference Scheme for Aspect Sentiment Triplet Extraction

AAAI 2023technical

Aspect Sentiment Triplet Extraction (ASTE) has become an emerging task in sentiment analysis research, aiming to extract triplets of the aspect term, its corresponding opinion term, and its associated sentiment polarity from a given sentence. Recently, many neural networks based models with differen…

2023

Semi-Implicit Denoising Diffusion Models (SIDDMs)

NeurIPS 2023poster

Despite the proliferation of generative models, achieving fast sampling during inference without compromising sample diversity and quality remains challenging. Existing models such as Denoising Diffusion Probabilistic Models (DDPM) deliver high-quality, diverse samples but are slowed by an inherentl…

2023

Set-membership Belief State-based Reinforcement Learning for POMDPs

ICML 2023poster

Reinforcement learning (RL) has made significant progress in areas such as Atari games and robotic control, where the agents have perfect sensing capabilities. However, in many real-world sequential decision-making tasks, the observation data could be noisy or incomplete due to the intrinsic low qua…

Cited by 0SourcePDFScholar
2023

TREA: Tree-Structure Reasoning Schema for Conversational Recommendation

ACL 2023long

Conversational recommender systems (CRS) aim to timely trace the dynamic interests of users through dialogues and generate relevant responses for item recommendations. Recently, various external knowledge bases (especially knowledge graphs) are incorporated into CRS to enhance the understanding of c…

2023

Towards Hierarchical Policy Learning for Conversational Recommendation with Hypergraph-based Reinforcement Learning

IJCAI 2023poster

Conversational recommendation systems (CRS) aim to timely and proactively acquire user dynamic preferred attributes through conversations for item recommendation. In each turn of CRS, there naturally have two decision-making processes with different roles that influence each other: 1) director, whic…

2022

Balancing between Forgetting and Acquisition in Incremental Subpopulation Learning

ECCV 2022poster

"The subpopulation shifting challenge, known as some subpopulations of a category that are not seen during training, severely limits the classification performance of the state-of-the-art convolutional neural networks. Thus, to mitigate this practical issue, we explore incremental subpopulation lear…

2022

BiSyn-GAT+: Bi-Syntax Aware Graph Attention Network for Aspect-based Sentiment Analysis

ACL 2022findings

Aspect-based sentiment analysis (ABSA) is a fine-grained sentiment analysis task that aims to align aspects and corresponding sentiments for aspect-specific sentiment polarity inference. It is challenging because a sentence may contain multiple aspects or complicated (e.g., conditional, coordinating…

2022

Capturing Global Structural Information in Long Document Question Answering with Compressive Graph Selector Network

EMNLP 2022main

Long document question answering is a challenging task due to its demands for complex reasoning over long text. Previous works usually take long documents as non-structured flat texts or only consider the local structure in long documents. However, these methods usually ignore the global structure o…

2022

Controlling Underestimation Bias in Reinforcement Learning via Quasi-median Operation

AAAI 2022technical

How to get a good value estimation is one of the key problems in reinforcement learning (RL). Current off-policy methods, such as Maxmin Q-learning, TD3 and TADD, suffer from the underestimation problem when solving the overestimation problem. In this paper, we propose the Quasi-Median Operation, a…

Cited by 16SourcePDFScholar
2022

DVGG: Deep Variational Grasp Generation for Dextrous Manipulation

RA-L 2022

Grasping with anthropomorphic robotic hands involves much more hand-object interactions compared to parallel-jaw grippers. Modeling hand-object interactions is essential to the study of multi-finger hand dextrous manipulation. This work presents DVGG, an efficient grasp generation network that takes

Cited by 64SourceScholar
2022

Declaration-based Prompt Tuning for Visual Question Answering

IJCAI 2022poster

In recent years, the pre-training-then-fine-tuning paradigm has yielded immense success on a wide spectrum of cross-modal tasks, such as visual question answering (VQA), in which a visual-language (VL) model is first optimized via self-supervised task objectives, e.g., masked language modeling (MLM)…

2022

HCL-TAT: A Hybrid Contrastive Learning Method for Few-shot Event Detection with Task-Adaptive Threshold

EMNLP 2022finding

Event detection has been suffering from constantly emerging event types with lack of sufficient data. Existing works formulate the new problem as few-shot event detection (FSED), and employ two-stage or unified models based on meta-learning to address the problem. However, these methods fall far sho…

2022

HGC-Net: Deep Anthropomorphic Hand Grasping in Clutter

ICRA 2022poster

Grasping in cluttered environments is one of the most fundamental skills in robotic manipulation. Most of the current works focus on estimating grasp poses for parallel-jaw or suction-cup end effectors. However, the study for dexterous anthropomorphic hand grasping in clutter remains a great challen…

Cited by 20SourcecodeScholar
2022

Incorporating Causal Analysis into Diversified and Logical Response Generation

COLING 2022main

Although the Conditional Variational Auto-Encoder (CVAE) model can generate more diversified responses than the traditional Seq2Seq model, the responses often have low relevance with the input words or are illogical with the question. A causal analysis is carried out to study the reasons behind, and…

Cited by 7SourcePDFScholar
2022

Multi-View Intent Disentangle Graph Networks for Bundle Recommendation

AAAI 2022technical

Bundle recommendation aims to recommend the user a bundle of items as a whole. Previous models capture user’s preferences on both items and the association of items. Nevertheless, they usually neglect the diversity of user’s intents on adopting items and fail to disentangle user’s intents in represe…

2022

Sequential Topic Selection Model with Latent Variable for Topic-Grounded Dialogue

EMNLP 2022finding

Recently, topic-grounded dialogue system has attracted significant attention due to its effectiveness in predicting the next topic to yield better responses via the historical context and given topic sequence. However, almost all existing topic prediction solutions focus on only the current conversa…

Cited by 1SourcePDFScholar
2021

A Student-Teacher Architecture for Dialog Domain Adaptation Under the Meta-Learning Setting

AAAI 2021technical

Numerous new dialog domains are being created every day while collecting data for these domains is extremely costly since it involves human interactions. Therefore, it is essential to develop algorithms that can adapt to different domains efficiently when building data-driven dialog models. Most rec…

Cited by 7SourcePDFScholar
2021

Context-Aware Biaffine Localizing Network for Temporal Sentence Grounding

CVPR 2021poster

This paper addresses the problem of temporal sentence grounding (TSG), which aims to identify the temporal boundary of a specific segment from an untrimmed video by a sentence query. Previous works either compare pre-defined candidate segments with the query and select the best one by ranking, or di…

Cited by 176PDFcodeScholar
2021

GPR: Grasp Pose Refinement Network for Cluttered Scenes

ICRA 2021poster

Object grasping in cluttered scenes is a widely investigated field of robot manipulation. Most of the current works focus on estimating grasp pose from point clouds based on an efficient single-shot grasp detection network. However, due to the lack of geometry awareness of the local grasping area, i…

Cited by 39SourceScholar
2021

Overcoming Catastrophic Forgetting by Bayesian Generative Regularization

ICML 2021spotlight

In this paper, we propose a new method to over-come catastrophic forgetting by adding generative regularization to Bayesian inference frame-work. Bayesian method provides a general frame-work for continual learning. We could further construct a generative regularization term for all given classifica…

2021

Reinforced History Backtracking for Conversational Question Answering

AAAI 2021technical

To model the context history in multi-turn conversations has become a critical step towards a better understanding of the user query in question answering systems. To utilize the context history, most existing studies treat the whole context as input, which will inevitably face the following two cha…

2021

Spatiotemporal Graph Neural Network based Mask Reconstruction for Video Object Segmentation

AAAI 2021technical

This paper addresses the task of segmenting class-agnostic objects in semi-supervised setting. Although previous detection based methods achieve relatively good performance, these approaches extract the best proposal by a greedy strategy, which may lose the local patch details outside the chosen can…

Cited by 27SourcePDFScholar
2021

Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation Approach

EMNLP 2021main

Reliable automatic evaluation of dialogue systems under an interactive environment has long been overdue. An ideal environment for evaluating dialog systems, also known as the Turing test, needs to involve human interaction, which is usually not affordable for large-scale experiments. Though researc…

2020

Differentiable Top-k with Optimal Transport

NeurIPS 2020poster

Finding the k largest or smallest elements from a collection of scores, i.e., top-k operation, is an important model component widely used in information retrieval, machine learning, and data mining. However, if the top-k operation is implemented in an algorithmic way, e.g., using bubble algorithm,…

2020

MLS3RDUH: Deep Unsupervised Hashing via Manifold based Local Semantic Similarity Structure Reconstructing

IJCAI 2020poster

Most of the unsupervised hashing methods usually map images into semantic similarity-preserving hash codes by constructing local semantic similarity structure as guiding information, i.e., treating each point similar to its k nearest neighbours. However, for an image, some of its k nearest neighbour…

Cited by 0SourcePDFScholar
2020

Mitigating Forgetting in Online Continual Learning via Instance-Aware Parameterization

NeurIPS 2020poster

Online continual learning is a challenging scenario where a model needs to learn from a continuous stream of data without revisiting any previously encountered data instances. The phenomenon of catastrophic forgetting is worsened since the model should not only address the forgetting at the task-lev…

Cited by 50SourcePDFScholar
2020

Unsupervised Adaptation Learning for Hyperspectral Imagery Super-Resolution

CVPR 2020poster

The key for fusion based hyperspectral image (HSI) super-resolution (SR) is to infer the posteriori of a latent HSI using appropriate image prior and likelihood that depends on degeneration. However, in practice the priors of high-dimensional HSIs can be extremely complicated and the degeneration is…

Cited by 130PDFcodeScholar
2019

COCO-GAN: Generation by Parts via Conditional Coordinating

ICCV 2019oral

Humans can only interact with part of the surrounding environment due to biological restrictions. Therefore, we learn to reason the spatial relationships across a series of observations to piece together the surrounding environment. Inspired by such behavior and the fact that machines also have comp…

Cited by 170PDFcodeScholar
2019

Complement Objective Training

ICLR 2019poster

Learning with a primary objective, such as softmax cross entropy for classification and sequence generation, has been the norm for training deep neural networks for years. Although being a widely-adopted approach, using cross entropy as the primary objective exploits mostly the information from the…

2019

Improving Adversarial Robustness via Guided Complement Entropy

ICCV 2019poster

Adversarial robustness has emerged as an important topic in deep learning as carefully crafted attack samples can significantly disturb the performance of a model. Many recent methods have proposed to improve adversarial robustness by utilizing adversarial training or model distillation, which adds…

Cited by 73PDFScholar
2019

Policy Certificates: Towards Accountable Reinforcement Learning

ICML 2019oral

The performance of a reinforcement learning algorithm can vary drastically during learning because of exploration. Existing algorithms provide little information about the quality of their current policy before executing it, and thus have limited use in high-stakes applications like healthcare. We a…

Cited by 176SourcePDFScholar
2019

Vehicle Re-Identification in Aerial Imagery: Dataset and Approach

ICCV 2019poster

In this work, we construct a large-scale dataset for vehicle re-identification (ReID), which contains 137k images of 13k vehicle instances captured by UAV-mounted cameras. To our knowledge, it is the largest UAV-based vehicle ReID dataset. To increase intra-class variation, each vehicle is captured…

Cited by 78PDFScholar
2018

DPP-Net: Device-aware Progressive Search for Pareto-optimal Neural Architectures

ECCV 2018poster

Recent breakthroughs in Neural Architectural Search (NAS) have achieved state-of-the-art performances in applications such as image classification and language modeling. However, these techniques typically ignore device-related objectives such as inference time, memory usage, and power consumption.…

Cited by 272SourcePDFScholar
2018

Escaping from Collapsing Modes in a Constrained Space

ECCV 2018poster

Generative adversarial networks (GANs) often suffer from unpredictable mode-collapsing during training. We study the issue of mode collapse of Boundary Equilibrium Generative Adversarial Network (BEGAN), which is one of the state-of-the-art generative models. Despite its potential of generating high…

2018

Thoracic Disease Identification and Localization With Limited Supervision

CVPR 2018poster

Accurate identification and localization of abnormalities from radiology images play an integral part in clinical diagnosis and treatment planning. Building a highly accurate prediction model for these tasks usually requires a large number of images manually annotated with labels and finding sites o…

Cited by 455SourcePDFScholar
2018

Video Rain Streak Removal by Multiscale Convolutional Sparse Coding

CVPR 2018poster

Videos captured by outdoor surveillance equipments sometimes contain unexpected rain streaks, which brings difficulty in subsequent video processing tasks. Rain streak removal from a video is thus an important topic in recent computer vision research. In this paper, we raise two intrinsic characte…

Cited by 228SourcePDFScholar
2017

Should We Encode Rain Streaks in Video as Deterministic or Stochastic?

ICCV 2017poster

Videos taken in the wild sometimes contain unexpected rain streaks, which brings difficulty in subsequent video processing tasks. Rain streak removal in a video (RSRV) is thus an important issue and has been attracting much attention in computer vision. Different from previous RSRV methods formulati…

Cited by 162PDFScholar
2017

When Unsupervised Domain Adaptation Meets Tensor Representations

ICCV 2017poster

Domain adaption (DA) allows machine learning methods trained on data sampled from one distribution to be applied to data sampled from another. It is thus of great practical importance to the application of such methods. Despite the fact that tensor representations are widely used in Computer Vision…

Cited by 89PDFcodeScholar
2016

Pairwise Matching Through Max-Weight Bipartite Belief Propagation

CVPR 2016poster

Feature matching is a key problem in computer vision and pattern recognition. One way to encode the essential interdependence between potential feature matches is to cast the problem as inference in a graphical model, though recently alternatives such as spectral methods, or approaches based on the…

Cited by 66PDFScholar
2015

Hyperspectral Compressive Sensing Using Manifold-Structured Sparsity Prior

ICCV 2015poster

To reconstruct hyperspectral image (HSI) accurately from a few noisy compressive measurements, we present a novel manifold-structured sparsity prior based hyperspectral compressive sensing (HCS) method in this study. A matrix based hierarchical prior is first proposed to represent the spectral struc…

Cited by 18PDFScholar
2015

Reweighted Laplace Prior Based Hyperspectral Compressive Sensing for Unknown Sparsity

CVPR 2015poster

Compressive sensing(CS) has been exploited for hypespectral image(HSI) compression in recent years. Though it can greatly reduce the costs of computation and storage, the reconstruction of HSI from a few linear measurements is challenging. The underlying sparsity of HSI is crucial to improve the rec…

Cited by 39SourcePDFScholar