← Search

Cong Cao

10 accepted papers

2026

OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval

AAAI 2026technical

Recent advances in large language models (LLMs) and dense retrievers have driven significant progress in retrieval-augmented generation (RAG). However, existing approaches face significant challenges in complex reasoning-oriented multi-hop retrieval tasks: 1) Ineffective reasoning-oriented planning:

Cited by 0SourcePDFScholar
2026

SpiralDiff: Spiral Diffusion with LoRA for RGB-to-RAW Conversion Across Cameras

CVPR 2026

RAW images preserve superior fidelity and rich scene information compared to RGB, making them essential for tasks in challenging imaging conditions. To alleviate the high cost of data collection, recent RGB-to-RAW conversion methods aim to synthesize RAW images from RGB. However, they overlook two k

Cited by 0SourcecodeScholar
2025

Emotion Transfer with Enhanced Prototype for Unseen Emotion Recognition in Conversation

EMNLP 2025

Current Emotion Recognition in Conversation (ERC) research follows a closed-domain assumption. However, there is no clear consensus on emotion classification in psychology, which presents a challenge for models when it comes to recognizing previously unseen emotions in real-world applications. To br

2025

Multi-View Incongruity Learning for Multimodal Sarcasm Detection

COLING 2025main

Multimodal sarcasm detection (MSD) is essential for various downstream tasks. Existing MSD methods tend to rely on spurious correlations. These methods often mistakenly prioritize non-essential features yet still make correct predictions, demonstrating poor generalizability beyond training environme…

Cited by 1SourcePDFScholar
2025

ReTD: Reconstruction-Based Traceability Detection for Generated Images

ICASSP 2025accepted

The objective of generated image traceability is to accurately identify and locate the source models. In this paper, we propose ReTD (Reconstruction-Based Traceability Detection), a generalized model for generated image traceability detection. Firstly, we use VAE to reconstruct images which are comp…

Cited by 0SourceScholar
2025

T-T: Table Transformer for Tagging-based Aspect Sentiment Triplet Extraction

IJCAI 2025

Aspect sentiment triplet extraction (ASTE) aims to extract triplets composed of aspect terms, opinion terms, and sentiment polarities from given sentences. The table tagging method is a popular approach to addressing this task, which encodes a sentence into a 2-dimensional table, allowing for the ta

2025

Zero-shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model

AAAI 2025technical

Diffusion-based zero-shot image restoration and enhancement models have achieved great success in various tasks of image restoration and enhancement. However, directly applying them to video restoration and enhancement results in severe temporal flickering artifacts. In this paper, we propose the fi…

2023

Mulan: A Multi-Level Alignment Model for Video Question Answering

EMNLP 2023long findings

Video Question Answering (VideoQA) aims to answer questions about the visual content of a video. Current methods mainly focus on improving joint representations of video and text. However, these methods pay little attention to the fine-grained semantic interaction between video and text. In this pap…

Cited by 0SourceScholar
2021

HSAN: A Hierarchical Self-Attention Network for Multi-Turn Dialogue Generation

ICASSP 2021accepted

In the multi-turn dialogue system, response generation is not only related to the sentences in context but also relies on the words in each utterance. Although there are lots of methods that pay attention to model words and utterances, there still exist problems such as tending to generate common re…

Cited by 0SourceScholar
2020

Supervised Raw Video Denoising With a Benchmark Dataset on Dynamic Scenes

CVPR 2020poster

In recent years, the supervised learning strategy for real noisy image denoising has been emerging and has achieved promising results. In contrast, realistic noise removal for raw noisy videos is rarely studied due to the lack of noisy-clean pairs for dynamic scenes. Clean video frames for dynamic s…

Cited by 142PDFcodeScholar