← Search

Bingyang Wang

3 accepted papers

2026

Vision Language Models Cannot Reason About Physical Transformation

ICML 2026poster

Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they genuinely understand physical transformations remains unclear. We introduce ***ConservationBench*** evaluating ***conse…

Cited by 0SourceScholar
2025

Core Knowledge Deficits in Multi-Modal Language Models

ICML 2025poster

While Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falling short on tasks that are intuitive and effortless for humans. We examine the hypothesis that these deficiencies stem f…

Cited by 0SourcePDFScholar
2025

TrackFusion: Enhancing Multi-Object Tracking With Temporal Trajectory Modeling and Frame-Integrated Detection

ICASSP 2025accepted

Although MOTIP is the SOTA multi-object tracking method, there are still some issues that limit its performance. First, MOTIP still has defects in temporal information modeling, which leads to the failure to fully utilize the historical information of the tracked target and affects the correlation p…

Cited by 0SourceScholar