← Search

Shida Wang

5 accepted papers

2026

DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Models

CVPR 2026

Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning remain limited. In this work, we introduce DiG (Differential Grounding), a novel proxy task framework where MLLMs learn

Cited by 0SourceScholar
2026

TableMix: Enhancing Multimodal Table Reasoning in MLLMs from a Data-Centric Perspective

CVPR 2026

Recent advances in Multimodal Large Language Models (MLLMs) have enabled promising progress in table reasoning from visual table inputs. Despite their ability to capture rich visual cues such as color and layout, MLLMs still underperform compared to text-only models.We argue that a major limitation

Cited by 0SourceScholar
2024

Inverse Approximation Theory for Nonlinear Recurrent Neural Networks

ICLR 2024spotlight

We prove an inverse approximation theorem for the approximation of nonlinear sequence-to-sequence relationships using recurrent neural networks (RNNs). This is a so-called Bernstein-type result in approximation theory, which deduces properties of a target function under the assumption that it can be…

2024

StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization

ICML 2024poster

In this paper, we investigate the long-term memory learning capabilities of state-space models (SSMs) from the perspective of parameterization. We prove that state-space models without any reparameterization exhibit a memory limitation similar to that of traditional RNNs: the target relationships th…

Cited by 9SourcePDFScholar
2023

State-space models with layer-wise nonlinearity are universal approximators with exponential decaying memory

NeurIPS 2023poster

State-space models have gained popularity in sequence modelling due to their simple and efficient network structures. However, the absence of nonlinear activation along the temporal direction limits the model's capacity. In this paper, we prove that stacking state-space models with layer-wise nonlin…

Cited by 28SourcePDFScholar