← Search

Daizong Liu

52 accepted papers

2026

Attacking Gray-Box Large Vision-Language Models with Adaptive SVD-Structured Adversarial Alignment

ICML 2026poster

Large vision-language models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal reasoning tasks. However, recent research shows that they are susceptible to adversarial examples. Existing LVLM attack methods are generally deployed in the white- or black-box setting, …

Cited by 0SourceScholar
2026

CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image Retrieval

CVPR 2026

Interactive Text-to-Image Retrieval (I-TIR) aims to refine image retrieval results through natural language dialogues, which allows users to progressively supplement or correct their search intention across multiple rounds, enabling a more precise and user-aligned visual search experience.However, e

Cited by 0SourcecodeScholar
2026

FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting

ICLR 2026poster

While Large Vision-Language Models (LVLMs) have achieved substantial progress in video understanding, their application to long video reasoning is hindered by uniform frame sampling and static textual reasoning, which are inefficient and struggle to handle visually intensive video tasks. To overcom…

Cited by 0SourceScholar
2026

Newton-coupled Dual-Teacher Semi-supervised Learning Framework

ICML 2026poster

Most semi-supervised learning frameworks rely on a single teacher that transfers zero-order supervision through pseudo-labels, constraining the student to imitate categorical outputs without perceiving the loss geometry. This design often leads to unstable optimization and limited generalization und…

Cited by 0SourceScholar
2026

Rethinking Video-Language Model from the Language Input Perspective

AAAI 2026technical

Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between videos and texts. Although previous VLM works have made significant progress, almost all of them implicitly assume that all the texts are predefine

Cited by 0SourcePDFScholar
2026

Spatial-Spectral Homogeneous Attacks on Physical-World Large Vision-Language Models

AAAI 2026technical

Although large vision-language models (LVLMs) have demonstrated promising versatile capabilities on various downstream tasks, they are shown to be susceptible to adversarial examples. Existing LVLM attackers simply implement adversarial patterns in an impracticable setting: i) add digital global per

Cited by 0SourcePDFScholar
2026

Spotlight on Token Perception for Multimodal Reinforcement Learning

ICLR 2026poster

While Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Vision-Language Models (LVLMs), most existing methods in multimodal reasoning neglect the critical role of visual perception within the RLVR optimization process. In this paper, we undertake…

Cited by 0SourcecodeScholar
2026

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

AAAI 2026technical

Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these VLMs are task-specific and assume that both video and language inputs are complete. However, real-world VLM applications might face challenges due

Cited by 0SourcePDFScholar
2026

Understanding and Exploiting Phase Sensitivity for Attacking Large Vision–Language Models

IJCAI 2026

Although Large Vision-Language Models (LVLMs) have demonstrated remarkable reasoning capabilities across various downstream multimodal tasks, they are proven to be vulnerable to carefully designed adversarial examples. Existing LVLM attackers show that exploring external components of adversarial gu

Cited by 0Scholar
2026

VideoSSR: Video Self-Supervised Reinforcement Learning

CVPR 2026

Reinforcement Learning with Verifiable Reward (RLVR) has substantially advanced the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, the rapid progress of MLLMs is outpacing the complexity of existing video datasets, while the manual annotation of new, high-qual

Cited by 0SourcecodeScholar
2025

Cooperative or Competitive? Understanding the Interaction between Attention Heads From A Game Theory Perspective

ACL 2025long

Despite the remarkable success of attention-based large language models (LLMs), the precise interaction mechanisms between attention heads remain poorly understood. In contrast to prevalent methods that focus on individual head contributions, we rigorously analyze the intricate interplay among atten…

2025

Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language Models

NeurIPS 2025poster

Although Multimodal Large Language Models (MLLMs) have demonstrated remarkable achievements in recent years, they remain vulnerable to adversarial examples that result in harmful responses. Existing attacks typically focus on optimizing adversarial perturbations for a certain multimodal image-prompt…

Cited by 0SourceScholar
2025

Imperceptible 3D Point Cloud Attacks on Lattice-based Barycentric Coordinates

AAAI 2025technical

Imperceptible adversarial attacks on 3D point clouds rely on effective constraints. While manifold constraints have notable advantages over Euclidean ones, the global parameterization used in current methods often fails to fully preserve manifold properties. In this paper, we propose to constrain la…

Cited by 1SourcePDFScholar
2025

LLM-assisted Entropy-based Adaptive Distillation for Unsupervised Fine-grained Visual Representation Learning

ICCV 2025poster

Unsupervised Fine-grained Visual Represent Learning (FVRL) aims to learn discriminative features to distinguish subtle differences among visually similar categories without using labeled fine-grained data. Existing works, which typically learn representation from target data, often struggle to captu…

2025

Learning from Few Samples: A Novel Approach for High-Quality Malcode Generation

EMNLP 2025

Intrusion Detection Systems (IDS) play a crucial role in network security defense. However, a significant challenge for IDS in training detection models is the shortage of adequately labeled malicious samples. To address these issues, this paper introduces a novel semi-supervised framework GANGRL-LL

Cited by 0SourcePDFScholar
2025

Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning

COLING 2025main

Recently, Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multi-modal context comprehension. However, they still suffer from hallucination problems referring to generating inconsistent outputs with the image content. To mitigate hallucinations, previous studies main…

2025

Misalignment Attack on Text-to-Image Models via Text Embedding Optimization and Inversion

EMNLP 2025

Text embedding serves not only as a core component of modern NLP models but also plays a pivotal role in multimodal systems such as text-to-image (T2I) models, significantly facilitating user-friendly image generation through natural language instructions. However, with the convenience being brought

Cited by 0SourcePDFScholar
2025

Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer Network

AAAI 2025technical

Given some video-query pairs with untrimmed videos and sentence queries, temporal sentence grounding (TSG) aims to locate query-relevant segments in these videos. Although previous respectable TSG methods have achieved remarkable success, they train each video-query pair separately and ignore the re…

Cited by 4SourcePDFScholar
2025

Seeing is Not Believing: Adversarial Natural Object Optimization for Hard-Label 3D Scene Attacks

CVPR 2025poster

Deep learning models for 3D data have shown to be vulnerable to adversarial attacks, which have received increasing attention in various safety-critical applications such as autonomous driving and robotic navigation. Existing 3D attackers mainly put effort into attacking the simple 3D classification…

Cited by 0SourcePDFScholar
2025

Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language Models

NeurIPS 2025spotlight

Although Large Vision-Language Models (LVLMs) exhibit impressive multimodal capabilities, their vulnerability to adversarial examples has raised serious security concerns. Existing LVLM attackers simply optimize adversarial images that easily overfit a certain model/prompt, making them ineffective o…

Cited by 0SourceScholar
2024

Explicitly Perceiving and Preserving the Local Geometric Structures for 3D Point Cloud Attack

AAAI 2024technical

Deep learning models for point clouds have shown to be vulnerable to adversarial attacks, which have received increasing attention in various safety-critical applications such as autonomous driving, robotics, and surveillance. Existing 3D attack methods generally employ global distance losses to imp…

Cited by 10SourcePDFScholar
2024

FLAT: Flux-aware Imperceptible Adversarial Attacks on 3D Point Clouds

ECCV 2024poster

"Adversarial attacks on point clouds play a vital role in assessing and enhancing the adversarial robustness of 3D deep learning models. While employing a variety of geometric constraints, existing adversarial attack solutions often display unsatisfactory imperceptibility due to inadequate considera…

Cited by 5SourcePDFScholar
2024

Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using Language

AAAI 2024technical

Given an untrimmed video and a sentence query, video moment retrieval using language (VMR) aims to locate a target query-relevant moment. Since the untrimmed video is overlong, almost all existing VMR methods first sparsely down-sample each untrimmed video into multiple fixed-length video clips and…

Cited by 19SourcePDFScholar
2024

Hiding Imperceptible Noise in Curvature-Aware Patches for 3D Point Cloud Attack

ECCV 2024poster

"With the maturity of depth sensors, point clouds have received increasing attention in various 3D safety-critical applications, while deep point cloud learning models have been shown to be vulnerable to adversarial attacks. Most existing 3D attackers rely on implicit global distance losses to pertu…

Cited by 6SourcePDFScholar
2024

Manifold Constraints for Imperceptible Adversarial Attacks on Point Clouds

AAAI 2024technical

Adversarial attacks on 3D point clouds often exhibit unsatisfactory imperceptibility, which primarily stems from the disregard for manifold-aware distortion, i.e., distortion of the underlying 2-manifold surfaces. In this paper, we develop novel manifold constraints to reduce such distortion, aiming…

Cited by 11SourcePDFScholar
2024

Pandora's Box: Towards Building Universal Attackers against Real-World Large Vision-Language Models

NeurIPS 2024poster

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities across a wide range of multimodal understanding tasks. Nevertheless, these models are susceptible to adversarial examples. In real-world applications, existing LVLM attackers generally rely on the detailed prior knowledge…

Cited by 7SourcePDFScholar
2024

Temporal Sentence Grounding with Relevance Feedback in Videos

NeurIPS 2024poster

As a widely explored multi-modal task, Temporal Sentence Grounding in videos (TSG) endeavors to retrieve a specific video segment matched with a given query text from a video. The traditional paradigm for TSG generally assumes that relevant segments always exist within a given video. However, this a…

2024

Towards Robust Temporal Activity Localization Learning with Noisy Labels

COLING 2024main

This paper addresses the task of temporal activity localization (TAL). Although recent works have made significant progress in TAL research, almost all of them implicitly assume that the dense frame-level correspondences in each video-query pair are correctly annotated. However, in reality, such an…

Cited by 6SourcePDFScholar
2024

Unsupervised Domain Adaptative Temporal Sentence Localization with Mutual Information Maximization

AAAI 2024technical

Temporal sentence localization (TSL) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on abundant yet expensive manual annotations for training. Moreover, these trained data-depen…

Cited by 7SourcePDFScholar
2023

3DHacker: Spectrum-based Decision Boundary Generation for Hard-label 3D Point Cloud Attack

ICCV 2023poster

With the maturity of depth sensors, the vulnerability of 3D point cloud models has received increasing attention in various applications such as autonomous driving and robot navigation. Previous 3D adversarial attackers either follow the white-box setting to iteratively update the coordinate perturb…

Cited by 21PDFScholar
2023

Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding

EMNLP 2023long findings

This paper addresses the task of temporal sentence grounding (TSG). Although many respectable works have made decent achievements in this important topic, they severely rely on massive expensive video-query paired annotations, which require a tremendous amount of human effort to collect in real-worl…

Cited by 0SourceScholar
2023

Density-Insensitive Unsupervised Domain Adaption on 3D Object Detection

CVPR 2023poster

3D object detection from point clouds is crucial in safety-critical autonomous driving. Although many works have made great efforts and achieved significant progress on this task, most of them suffer from expensive annotation cost and poor transferability to unknown data due to the domain gap. Recen…

2023

Distantly-Supervised Named Entity Recognition with Adaptive Teacher Learning and Fine-Grained Student Ensemble

AAAI 2023technical

Distantly-Supervised Named Entity Recognition (DS-NER) effectively alleviates the data scarcity problem in NER by automatically generating training samples. Unfortunately, the distant supervision may induce noisy labels, thus undermining the robustness of the learned models and restricting the pract…

2023

Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video Retrieval

ICCV 2023poster

Almost all previous text-to-video retrieval works assume that videos are pre-trimmed with short durations. However, in practice, videos are generally untrimmed containing much background content. In this work, we investigate the more practical but challenging Partially Relevant Video Retrieval (PRVR…

Cited by 19PDFcodeScholar
2023

Hypotheses Tree Building for One-Shot Temporal Sentence Localization

AAAI 2023technical

Given an untrimmed video, temporal sentence localization (TSL) aims to localize a specific segment according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on dense video frame annotations, which require a tremendous amount of human…

Cited by 20SourcePDFScholar
2023

Jointly Visual- and Semantic-Aware Graph Memory Networks for Temporal Sentence Localization in Videos

ICASSP 2023accepted

Temporal sentence localization in videos (TSLV) aims to retrieve the most interested segment in an untrimmed video according to a given sentence query. However, almost of existing TSLV approaches suffer from the same limitations: (1) They only focus on either frame-level or object-level visual repre…

Cited by 0SourceScholar
2023

Tracking Objects and Activities with Attention for Temporal Sentence Grounding

ICASSP 2023accepted

Temporal sentence grounding (TSG) aims to localize the temporal segment which is semantically aligned with a natural language query in an untrimmed video. Most existing methods extract frame-grained features or object-grained features by 3D ConvNet or detection network under a conventional TSG frame…

Cited by 0SourceScholar
2023

You Can Ground Earlier Than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed Videos

CVPR 2023poster

Given an untrimmed video, temporal sentence grounding (TSG) aims to locate a target moment semantically according to a sentence query. Although previous respectable works have made decent success, they only focus on high-level visual features extracted from the consecutive decoded frames and fail to…

Cited by 54SourcePDFScholar
2022

Exploring Motion and Appearance Information for Temporal Sentence Grounding

AAAI 2022technical

This paper addresses temporal sentence grounding. Previous works typically solve this task by learning frame-level video features and align them with the textual information. A major limitation of these works is that they fail to distinguish ambiguous video frames with subtle appearance differences…

Cited by 41SourcePDFScholar
2022

Exploring the Devil in Graph Spectral Domain for 3D Point Cloud Attacks

ECCV 2022poster

"With the maturity of depth sensors, point clouds have received increasing attention in various applications such as autonomous driving, robotics, surveillance, \etc., while deep point cloud learning models have shown to be vulnerable to adversarial attacks. Existing attack methods generally add/del…

2022

Memory-Guided Semantic Learning Network for Temporal Sentence Grounding

AAAI 2022technical

Temporal sentence grounding (TSG) is crucial and fundamental for video understanding. Although existing methods train well-designed deep networks with large amount of data, we find that they can easily forget the rarely appeared cases during training due to the off-balance data distribution, which i…

Cited by 66SourcePDFScholar
2022

Rethinking the Video Sampling and Reasoning Strategies for Temporal Sentence Grounding

EMNLP 2022finding

Temporal sentence grounding (TSG) aims to identify the temporal boundary of a specific segment from an untrimmed video by a sentence query. All existing works first utilize a sparse sampling strategy to extract a fixed number of video frames and then interact them with query for reasoning.However, w…

Cited by 22SourcePDFScholar
2022

Unsupervised Temporal Video Grounding with Deep Semantic Clustering

AAAI 2022technical

Temporal video grounding (TVG) aims to localize a target segment in a video according to a given sentence query. Though respectable works have made decent achievements in this task, they severely rely on abundant video-query paired data, which is expensive to collect in real-world scenarios. In this…

Cited by 59SourcePDFScholar
2021

Adaptive Proposal Generation Network for Temporal Sentence Localization in Videos

EMNLP 2021main

We address the problem of temporal sentence localization in videos (TSLV). Traditional methods follow a top-down framework which localizes the target segment with pre-defined segment proposals. Although they have achieved decent performance, the proposals are handcrafted and redundant. Recently, bot…

Cited by 61SourcePDFScholar
2021

Context-Aware Biaffine Localizing Network for Temporal Sentence Grounding

CVPR 2021poster

This paper addresses the problem of temporal sentence grounding (TSG), which aims to identify the temporal boundary of a specific segment from an untrimmed video by a sentence query. Previous works either compare pre-defined candidate segments with the query and select the best one by ranking, or di…

Cited by 176PDFcodeScholar
2021

F2Net: Learning to Focus on the Foreground for Unsupervised Video Object Segmentation

AAAI 2021technical

Although deep learning based methods have achieved great progress in unsupervised video object segmentation, difficult scenarios (e.g., visual similarity, occlusions, and appearance changing) are still no well-handled. To alleviate these issues, we propose a novel Focus on Foreground Network (F2Net…

Cited by 51SourcePDFScholar
2021

Progressively Guide to Attend: An Iterative Alignment Framework for Temporal Sentence Grounding

EMNLP 2021main

A key solution to temporal sentence grounding (TSG) exists in how to learn effective alignment between vision and language features extracted from an untrimmed video and a sentence description. Existing methods mainly leverage vanilla soft attention to perform the alignment in a single-step process.…

Cited by 49SourcePDFScholar
2021

Spatiotemporal Graph Neural Network based Mask Reconstruction for Video Object Segmentation

AAAI 2021technical

This paper addresses the task of segmenting class-agnostic objects in semi-supervised setting. Although previous detection based methods achieve relatively good performance, these approaches extract the best proposal by a greedy strategy, which may lose the local patch details outside the chosen can…

Cited by 27SourcePDFScholar
2020

Reasoning Step-by-Step: Temporal Sentence Localization in Videos via Deep Rectification-Modulation Network

COLING 2020main

Temporal sentence localization in videos aims to ground the best matched segment in an untrimmed video according to a given sentence query. Previous works in this field mainly rely on attentional frameworks to align the temporal boundaries by a soft selection. Although they focus on the visual conte…

Cited by 36SourcePDFScholar
2019

MHP-VOS: Multiple Hypotheses Propagation for Video Object Segmentation

CVPR 2019oral

We address the problem of semi-supervised video object segmentation (VOS), where the masks of objects of interests are given in the first frame of an input video. To deal with challenging cases where objects are occluded or missing, previous work relies on greedy data association strategies that mak…

Cited by 67PDFcodeScholar