← Search

Jingyi Wang

11 accepted papers

2026

ATTENTION TO DETAILS, LOGITS TO TRUTH: VISUAL-AWARE ATTENTION AND LOGITS ENHANCEMENT TO MITIGATE HALLUCINATIONS IN LVLMS

ICASSP 2026oral

Existing Large Vision-Language Models (LVLMs) exhibit insufficient visual attention, leading to hallucinations. To alleviate this problem, some previous studies adjust and amplify visual attention. These methods present a limitation that boosting attention for all visual tokens inevitably increases…

Cited by 0SourcePDFScholar
2026

OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models

CVPR 2026

Multimodal large language models (MLLMs) have achieved remarkable performance across a wide range of vision-language tasks. However, their ability in low-level visual perception, particularly in detecting fine-grained visual discrepancies, remains underexplored and lacks systematic analysis.In this

Cited by 0SourcecodeScholar
2025

An Online Adaptive Sampling Algorithm for Stochastic Difference-of-convex Optimization with Time-varying Distributions

ICML 2025oral

We propose an online adaptive sampling algorithm for solving stochastic nonsmooth difference-of-convex (DC) problems under time-varying distributions. At each iteration, the algorithm relies solely on data generated from the current distribution and employs distinct adaptive sampling rates for the c…

Cited by 0SourcePDFScholar
2025

Convergence Rates of Constrained Expected Improvement

NeurIPS 2025spotlight

Constrained Bayesian optimization (CBO) methods have seen significant success in black-box optimization with constraints. One of the most commonly used CBO methods is the constrained expected improvement (CEI) algorithm. CEI is a natural extension of expected improvement (EI) when constraints are i…

Cited by 0SourceScholar
2025

LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models

ICASSP 2025accepted

Recent advances in large vision-language models (LVLMs) typically employ vision encoders based on the Vision Transformer (ViT) architecture. The division of the images into patches by ViT results in a fragmented perception, thereby hindering the visual understanding capabilities of LVLMs. In this pa…

Cited by 0SourceScholar
2025

TableEval: A Real-World Benchmark for Complex, Multilingual, and Multi-Structured Table Question Answering

EMNLP 2025

LLMs have shown impressive progress in natural language processing. However, they still face significant challenges in TableQA, where real-world complexities such as diverse table structures, multilingual data, and domain-specific reasoning are crucial. Existing TableQA benchmarks are often limited

2025

VisNumBench: Evaluating Number Sense of Multimodal Large Language Models

ICCV 2025poster

Can Multimodal Large Language Models (MLLMs) develop an intuitive number sense similar to humans? Targeting this problem, we introduce Visual Number Benchmark (VisNumBench) to evaluate the number sense abilities of MLLMs across a wide range of visual numerical tasks. VisNumBench consists of about 1,…

2024

Continuous Review and Timely Correction: Enhancing the Resistance to Noisy Labels via Self-Not-True Distillation

ICASSP 2024accepted

Deep neural networks possess substantial learning capacities and robust expressive power, making them prone to overfitting mislabeled data. Fortunately, the memorization effect shows that the networks tend to memorize the clean data first, and then gradually memorize the mislabeled data. Correspondi…

Cited by 0SourceScholar
2023

Boosting Adversarial Training in Safety-Critical Systems Through Boundary Data Selection

RA-L 2023

AI-enabled collaborative robots are designed to be used in close collaboration with humans, thus requiring stringent safety standards and quick response times. Adversarial attacks pose a significant threat to the deep learning models of these systems, making it crucial to develop methods to improve

Cited by 3SourceScholar
2023

Cross-Modality Time-Variant Relation Learning for Generating Dynamic Scene Graphs

ICRA 2023poster

Dynamic scene graphs generated from video clips could help enhance the semantic visual understanding in a wide range of challenging tasks such as environmental perception, autonomous navigation, and task planning of self-driving vehicles and mobile robots. In the process of temporal and spatial mode…

Cited by 9SourcecodeScholar