DialCLIP: Empowering Clip As Multi-Modal Dialog Retriever
Recently, substantial advancements in pre-trained visionlanguage models have greatly enhanced the capabilities of multi-modal dialog systems. These models have demonstrated significant improvements by fine-tuning on downstream tasks. However, the existing pre-trained models primarily focus on effect…