← Search

Lixiang Liu

4 accepted papers

2026

Test-Time Perturbation Tuning with Delayed Feedback for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action models (VLAs) achieve remarkable performance in sequential decision-making but remain fragile to subtle environmental shifts, such as small changes in object pose. We attribute this brittleness to trajectory overfitting, where VLAs over-attend to the spurious correlation betwe

Cited by 0SourcecodeScholar
2025

TFS: Revisiting Temporal Language Grounding from Frequency Spiking Perspective

ICASSP 2025accepted

Temporal Language Grounding (TLG) aims to localize moments in untrimmed videos that are most relevant to natural language queries. While existing weakly-supervised methods have achieved significant success in exploring cross-modal relationships, they still face a critical bottleneck: the interferenc…

Cited by 0SourceScholar
2024

Rethinking Misalignment in Vision-Language Model Adaptation from a Causal Perspective

NeurIPS 2024poster

Foundational Vision-Language models such as CLIP have exhibited impressive generalization in downstream tasks. However, CLIP suffers from a two-level misalignment issue, i.e., task misalignment and data misalignment, when adapting to specific tasks. Soft prompt tuning has mitigated the task misalign…

Cited by 5SourcePDFScholar