← Search

Yiwei Shi

7 accepted papers

2026

Belief-Contraction-Driven Active Inverse Source Localization and Characterization

IJCAI 2026

Active inverse source localization and characterization (ISLC) in dynamic fields requires sequential decision making under partial observability, where a mobile sensor must infer latent source parameters from sparse, noisy readings. We introduce a belief-contraction-driven approach that unifies infe

Cited by 0Scholar
2026

Markovian Scale Prediction: A New Era of Visual Autoregressive Generation

CVPR 2026

Visual AutoRegressive modeling (VAR) based on next-scale prediction has revitalized autoregressive visual generation. Although its full-context dependency, i.e., modeling all previous scales for next-scale prediction, facilitates more stable and comprehensive representation learning by leveraging co

Cited by 0SourceScholar
2026

Swordsman: Entropy-Driven Adaptive Block Partition for Efficient Diffusion Language Models

ICML 2026poster

Block-wise decoding effectively improves the inference speed and quality in diffusion language models (DLMs) by combining inter-block sequential denoising and intra-block parallel unmasking. However, existing block-wise decoding methods typically partition blocks in a rigid and fixed manner, which i…

Cited by 0SourceScholar
2025

Autonomous Goal Detection and Cessation in Reinforcement Learning: A Case Study on Source Term Estimation

AAAI 2025technical

Reinforcement Learning has revolutionized decision-making processes in dynamic environments, yet it often struggles with autonomously detecting and achieving goals without clear feedback signals. For example, in a Source Term Estimation problem, the lack of precise environmental information makes it…

Cited by 3SourcePDFScholar
2025

Reidentify: Context-Aware Identity Generation for Contextual Multi-Agent Reinforcement Learning

ICML 2025poster

Generalizing multi-agent reinforcement learning (MARL) to accommodate variations in problem configurations remains a critical challenge in real-world applications, where even subtle differences in task setups can cause pre-trained policies to fail. To address this, we propose Context-Aware Identity…

Cited by 0SourcePDFScholar
2024

MLIP: Efficient Multi-Perspective Language-Image Pretraining with Exhaustive Data Utilization

ICML 2024poster

Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success, leading to rapid advancements in multimodal studies. However, CLIP faces a notable challenge in terms of *inefficient data utilization*. It relies on a single contrastive supervision for each image-text pair during repres…

Cited by 3SourcePDFScholar
2023

MG-ViT: A Multi-Granularity Method for Compact and Efficient Vision Transformers

NeurIPS 2023poster

Vision Transformer (ViT) faces obstacles in wide application due to its huge computational cost. Almost all existing studies on compressing ViT adopt the manner of splitting an image with a single granularity, with very few exploration of splitting an image with multi-granularity. As we know, import…

Cited by 12SourcePDFScholar