← Search

Weili Guan

29 accepted papers

2026

Amplifying Discrepancies: Exploiting Macro and Micro Inconsistencies for Image Manipulation Localization

AAAI 2026technical

The rapid development of image manipulation technologies poses significant challenges to multimedia forensics, especially in accurate localization of manipulated regions. Existing methods often fail to fully explore the intrinsic discrepancies between manipulated and authentic regions, resulting in

Cited by 0SourcePDFScholar
2026

Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question Answering

AAAI 2026technical

Multi-hop question answering (MHQA) requires integrating knowledge scattered across multiple passages to derive the correct answer. Traditional retrieval-augmented generation (RAG) methods primarily focus on coarse-grained textual semantic similarity and ignore structural associations among disperse

Cited by 5SourcePDFScholar
2026

EnergyAction: Unimanual to Bimanual Composition with Energy-Based Models

CVPR 2026

Recent advances in unimanual manipulation policies have achieved remarkable success across diverse robotic tasks through abundant training data and well-established model architectures. However, extending these capabilities to bimanual manipulation remains challenging due to the lack of bimanual dem

Cited by 0SourcecodeScholar
2026

EnsembleVLA: Ensemble Learning for Vision-Language Action Models

ICML 2026poster

Diverse Vision-language-action (VLA) models have been proposed and demonstrated remarkable capabilities in robotic manipulation. However, how to effectively ensemble VLAs to further enhance performance remains largely unexplored, as conventional ensemble techniques designed for discriminative tasks …

Cited by 0SourceScholar
2026

HATS: Hardness-Aware Trajectory Synthesis for GUI Agents

CVPR 2026

Graphical user interface (GUI) agents powered by large vision-language models (VLMs) have shown remarkable potential in automating digital tasks, highlighting the need for high-quality trajectory data to support effective agent training. Yet existing trajectory synthesis pipelines often yield agents

Cited by 0SourcecodeScholar
2026

SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting

ICLR 2026poster

Continual Learning (CL) requires a model to learn multiple tasks in sequence while maintaining both stability—preserving knowledge from previously learned tasks, and plasticity—effectively learning new tasks. Orthogonal projection has emerged as an effective and popular paradigm in CL, where it part…

Cited by 0SourcecodeScholar
2026

TGV-KV: Text-Grounded KV Eviction for Vision-Language Models

ICML 2026poster

Vision-Language Models (VLMs) inherit the auto-regressive generation paradigm and cache the keys and values (KV) of all previous tokens to accelerate inference, resulting in memory consumption that scales linearly with context length. This issue is particularly pronounced in VLMs due to substantial …

Cited by 0SourceScholar
2025

Breakthrough Sensor-Limited Single View: Towards Implicit Temporal Dynamics for Time Series Domain Adaptation

NeurIPS 2025poster

Unsupervised domain adaptation has emerged as a pivotal paradigm for mitigating distribution shifts in time series analysis. The fundamental challenge in time series domain adaptation arises from the entanglement of domain shifts and intricate temporal patterns. Crucially, the latent continuous-time…

Cited by 0SourcecodeScholar
2025

Content-aware Balanced Spectrum Encoding in Masked Modeling for Time Series Classification

AAAI 2025technical

Due to the superior ability of global dependency, transformer and its variants have become the primary choice in Masked Time-series Modeling (MTM) towards time-series classification task. In this paper, we experimentally analyze that existing transformer-based MTM methods encounter with two under-ex…

Cited by 0SourcePDFScholar
2025

Curriculum Coarse-to-Fine Selection for High-IPC Dataset Distillation

CVPR 2025poster

Dataset distillation (DD) excels in synthesizing a small number of images per class (IPC) but struggles to maintain its effectiveness in high-IPC settings. Recent works on dataset distillation demonstrate that combining distilled and real data can mitigate the effectiveness decay. However, our analy…

2025

Debiased Curriculum Adaptation for Safe Transfer Learning in Chest X-ray Classification

ICCV 2025poster

Chest X-ray classification is extensively utilized within the field of medical image analysis. However, manually labeling chest X-ray images is time-consuming and costly. Domain adaptation, which is designed to transfer knowledge from related domains, could offer a promising solution. Existing metho…

2025

ENCODER: Entity Mining and Modification Relation Binding for Composed Image Retrieval

AAAI 2025technical

The objective of Composed Image Retrieval (CIR) is to identify a target image that meets the requirement based on a multimodal query (including the reference image and the modification text) provided by the user. Despite the notable success of existing approaches, they fail to adequately address the…

Cited by 2SourcePDFScholar
2025

Enhancing GUI Agent with Uncertainty-Aware Self-Trained Evaluator

NeurIPS 2025poster

Benefiting from the availability of extensive navigation trajectories, both manually and automatically annotated, current graphical user interface (GUI) agents have achieved remarkable advancements in performance. However, these annotated datasets often contain substantial noise, which impedes effec…

Cited by 0SourceScholar
2025

FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers

ICCV 2025poster

The incorporation of high-resolution visual input equips multimodal large language models (MLLMs) with enhanced visual perception capabilities for real-world tasks. However, most existing high-resolution MLLMs rely on a cropping-based approach to process images, which leads to fragmented visual enco…

2025

Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin

ICML 2025poster

Adapting vision-language models (VLMs) to downstream tasks with pseudolabels has gained increasing attention. A major obstacle is that the pseudolabels generated by VLMs tend to be imbalanced, leading to inferior performance. While existing methods have explored various strategies to address this,…

2025

Meta Guidance: Incorporating Inductive Biases into Deep Time Series Imputers

NeurIPS 2025poster

Missing values, frequently encountered in time series data, can significantly impair the effectiveness of analytical methods. While deep imputation models have emerged as the predominant approach due to their superior performance, explicitly incorporating inductive biases aligned with time-series ch…

Cited by 0SourceScholar
2025

Object-Shot Enhanced Grounding Network for Egocentric Video

CVPR 2025poster

Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentric and exocentric videos but often neglect key characteristics of egocentric vid…

2025

PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models

ACL 2025long

Large Language Models (LLMs) suffer severe performance degradation when facing extremely low-bit (sub 2-bit) quantization. Several existing sub 2-bit post-training quantization (PTQ) methods utilize a mix-precision scheme by leveraging an unstructured fine-grained mask to explicitly distinguish sali…

2025

Social Debiasing for Fair Multi-modal LLMs

ICCV 2025poster

Multi-modal Large Language Models (MLLMs) have dramatically advanced the research field and delivered powerful vision-language understanding capabilities. However, these models often inherit deep-rooted social biases from their training data, leading to uncomfortable responses with respect to attrib…

Cited by 0SourcePDFScholar
2025

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data

NeurIPS 2025spotlight

Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied interaction. However, existing methods face spatial uncertainty and data scarcity, limiting the 3D spatial reasoning capab…

Cited by 0SourceScholar
2025

Towards Stable and Storage-efficient Dataset Distillation: Matching Convexified Trajectory

CVPR 2025poster

The rapid evolution of deep learning and large language models has led to an exponential growth in the demand for training data, prompting the development of Dataset Distillation methods to address the challenges of managing large datasets. Among these, Matching Training Trajectories (MTT) has been…

Cited by 2SourcePDFScholar
2025

Unified Transferability Metrics for Time Series Foundation Models

NeurIPS 2025poster

With the increasing number of time series pre-trained models, designing transferability evaluation metrics for time series has become an urgent problem to address. While transferability evaluation has been extensively studied in computer vision, we aim to address a critical gap by developing tailor…

Cited by 0SourceScholar
2024

Boosting Transferability and Discriminability for Time Series Domain Adaptation

NeurIPS 2024poster

Unsupervised domain adaptation excels in transferring knowledge from a labeled source domain to an unlabeled target domain, playing a critical role in time series applications. Existing time series domain adaptation methods either ignore frequency features or treat temporal and frequency features eq…

2024

MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models

NeurIPS 2024poster

Multimodal large language models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, a generalist MLLM typically underperforms compared with a specialist MLLM on most VL tasks, which can be attributed to task interference. In this paper, we propose a mixt…

2023

Knowledge-Aware Prompt Tuning for Generalizable Vision-Language Models

ICCV 2023poster

Pre-trained vision-language models, e.g., CLIP, working with manually designed prompts have demonstrated great effectiveness in transfer learning. Recently, learnable prompts achieve state-of-the-art performance, which however are prone to overfit to seen classes while failing to generalize to unsee…

Cited by 37PDFScholar
2023

Semi-Supervised Video Inpainting With Cycle Consistency Constraints

CVPR 2023poster

Deep learning-based video inpainting has yielded promising results and gained increasing attention from researchers. Generally, these methods usually assume that the corrupted region masks of each frame are known and easily obtained. However, the annotation of these masks are labor-intensive and exp…

Cited by 18SourcePDFScholar
2023

Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training Models

ICCV 2023oral

Vision-language pre-training (VLP) models have shown vulnerability to adversarial examples in multimodal tasks. Furthermore, malicious adversaries can be deliberately transferred to attack other black-box models. However, existing work has mainly focused on investigating white-box attacks. In this p…

Cited by 65PDFcodeScholar