← Search

Bicheng Xu

3 accepted papers

2025

Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?

NeurIPS 2025poster

Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI assistants. Effective guidance requires not only delivering instructions but also detecting their successful execution, a…

Cited by 0SourcecodeScholar
2025

OCCAM: Towards Cost-Efficient and Accuracy-Aware Classification Inference

ICLR 2025poster

Classification tasks play a fundamental role in various applications, spanning domains such as healthcare, natural language processing and computer vision. With the growing popularity and capacity of machine learning models, people can easily access trained classifiers as a service online or offline…

Cited by 0SourcePDFScholar
2019

Watch, Listen and Tell: Multi-Modal Weakly Supervised Dense Event Captioning

ICCV 2019poster

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of the research has been limited to approaches that either do no…

Cited by 115PDFcodeScholar