← Search

Yunjian Zhang

19 accepted papers

2026

EventFlash: Towards Efficient MLLMs for Event-Based Vision

ICLR 2026poster

Event-based multimodal large language models (MLLMs) enable robust perception in high-speed and low-light scenarios, addressing key limitations of frame-based MLLMs. However, current event-based MLLMs often rely on dense image-like processing paradigms, overlooking the spatiotemporal sparsity of eve…

Cited by 0SourcecodeScholar
2026

Mitigating Error Propagation in Low-Rank Approximation of Large Models via Distribution-Aware Whitening

ICML 2026poster

Low-rank approximation has emerged as a cornerstone technique for model compression and parameter-efficient fine-tuning, enabling substantial reductions in computation and memory without altering model architectures. However, existing approaches often overlook the shifts in feature distributions ind…

Cited by 0SourceScholar
2026

Pruning as a Cooperative Game: Surrogate-Assisted Layer Contribution Estimation for Large Language Models

ICLR 2026poster

While large language models (LLMs) demonstrate impressive performance across various tasks, their deployment in real-world scenarios is still constrained by high computational demands. Layer-wise pruning, a commonly employed strategy to mitigate inference costs, can partially address this challenge.…

Cited by 0SourceScholar
2026

Resource-Efficient Reinforcement for Reasoning Large Language Models via Dynamic One-Shot Policy Refinement

ICML 2026poster

Large language models (LLMs) have exhibited remarkable performance on complex reasoning tasks, with reinforcement learning under verifiable rewards (RLVR) emerging as a principled framework for aligning model behavior with reasoning chains. Despite its promise, RLVR remains prohibitively resource-in…

Cited by 0SourceScholar
2026

TFD-Net: Towards Intelligent Time-Frequency Mode Decomposition with Practical Applications

AAAI 2026technical

Time-frequency analysis (TFA) and mode decomposition for non-stationary signals are research hotspots in the field of signal processing. Current optimization-based decomposition methods require a good initial IF estimate. However, due to the Heisenberg uncertainty, achieving accurate ridge extractio

Cited by 0SourcePDFScholar
2025

Bridging the Gap Between Ideal and Real-world Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios

ICCV 2025poster

With the rapid advancement of generative models, highly realistic image synthesis has posed new challenges to digital security and media credibility. Although AI-generated image detection methods have partially addressed these concerns, a substantial research gap remains in evaluating their performa…

Cited by 0SourcePDFScholar
2025

ELMoD-Net: Exploring Pattern Similarities for Robust Event-Based Lane Detection

RA-L 2025

Lane detection plays a critical role in autonomous driving, requiring low-latency response and robustness to varying lighting conditions to support downstream algorithms effectively. Event cameras, with their low latency and high dynamic range, are well-suited for such tasks. However, current event-

Cited by 0SourceScholar
2025

Enhanced Event-based Dense Stereo via Cross-Sensor Knowledge Distillation

ICCV 2025poster

Accurate stereo matching under fast motion and extreme lighting conditions is a challenge for many vision applications. Event cameras have the advantages of low latency and high dynamic range, thus providing a reliable solution to this challenge. However, since events are sparse, this makes it an il…

Cited by 0SourcePDFScholar
2025

EventGPT: Event Stream Understanding with Multimodal Large Language Models

CVPR 2025poster

Event cameras capture visual information as asynchronous pixel change streams, excelling in challenging lighting and high-dynamic scenarios. Existing multimodal large language models (MLLMs) concentrate on natural RGB images, failing in scenarios where event data fits better. In this paper, we intro…

Cited by 3SourcePDFScholar
2025

SHIFT: Smoothing Hallucinations by Information Flow Tuning for Multimodal Large Language Models

ICCV 2025poster

Large Language Models (LLMs) are prone to hallucinations, which pose significant risks in their applications. Most existing hallucination detection methods rely on internal probabilities or external knowledge, and they are limited to identifying hallucinations at the sentence or passage level. In th…

Cited by 0SourcePDFScholar
2025

Towards Annotation-Free Evaluation: KPAScore for Human Keypoint Detection

ICCV 2025poster

Human keypoint detection is fundamental in computer vision, with applications in pose estimation and action recognition. However, existing evaluation metrics (e.g., OKS, PCP, PDJ) rely on human-annotated ground truth, a labor-intensive process that increases costs, limits scalability. To address thi…

Cited by 0SourcePDFScholar
2025

Towards Understanding How Knowledge Evolves in Large Vision-Language Models

CVPR 2025poster

Large Vision-Language Models (LVLMs) are gradually becoming the foundation for many artificial intelligence applications. However, understanding their internal working mechanisms has continued to puzzle researchers, which in turn limits the further enhancement of their capabilities. In this paper, w…

2024

Unleashing the Potential of Large Language Models through Spectral Modulation

EMNLP 2024finding

Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, garnering significant attention from both academia and industry. However, enhancing the performance of LLMs typically requires scaling up model sizes or fine-tuning with additional datasets, which results…

2022

360-Attack: Distortion-Aware Perturbations From Perspective-Views

CVPR 2022poster

The application of deep neural networks (DNNs) on 360-degree images has achieved remarkable progress in the recent years. However, DNNs have been demonstrated to be vulnerable to well-crafted adversarial examples, which may trigger severe safety problems in the real-world applications based on 360-d…

Cited by 6PDFScholar
2022

Mitigating the Inconsistency Between Word Saliency and Model Confidence with Pathological Contrastive Training

ACL 2022findings

Neural networks are widely used in various NLP tasks for their remarkable performance. However, the complexity makes them difficult to interpret, i.e., they are not guaranteed right for the right reason. Besides the complexity, we reveal that the model pathology - the inconsistency between word sali…

Cited by 6SourcePDFScholar
2022

PARSE: An Efficient Search Method for Black-box Adversarial Text Attacks

COLING 2022main

Neural networks are vulnerable to adversarial examples. The adversary can successfully attack a model even without knowing model architecture and parameters, i.e., under a black-box scenario. Previous works on word-level attacks widely use word importance ranking (WIR) methods and complex search met…

Cited by 9SourcePDFScholar
2022

SP Attack: Single-Perspective Attack for Generating Adversarial Omnidirectional Images

ICASSP 2022accepted

The safety of Deep Neural Networks (DNNs) processing omnidirectional images (ODIs) is an under-researched topic. In this paper, we propose a novel sparse attack, named Single-Perspective (SP) Attack, towards fooling these models by perturbing only one perspective image (PI) rendered from the target…

Cited by 0SourceScholar