← Search

Rong Fu

11 accepted papers

2026

CACR: Reinforcing Temporal Answer Grounding in Instructional Video via Candidate-Aware Causal Reasoning

ICML 2026poster

The task of temporal answer grounding in instructional videos (TAGV), which aims to locate precise video segments that respond to natural language queries, is increasingly important for direct video answer retrieval. This task remains challenging due to the need to comprehend semantically complex qu…

Cited by 0SourceScholar
2026

CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning

CVPR 2026

Multimodal learning aims to capture both shared and private information from multiple modalities. However, existing methods that project all modalities into a single latent space for fusion often overlook the asynchronous, multi-level semantic structure of multimodal data. This oversight induces sem

Cited by 0SourceScholar
2026

Escaping the Likelihood Trap: Geometric Diversity Optimization for Long-Form Image Captioning

ICML 2026poster

The utility of Vision-Language Models (VLMs) in reasoning and auditing tasks hinges on their ability to exhaustively describe visual scenes. However, current models exhibit a pathology we term the Likelihood Trap: standard alignment objectives, specifically MLE and KL-regularization, drive generatio…

Cited by 0SourceScholar
2026

FINSENTLLM: MULTI-LLM AND STRUCTURED SEMANTIC SIGNALS FOR ENHANCED FINANCIAL SENTIMENT FORECASTING

ICASSP 2026poster

Financial sentiment analysis (FSA) has attracted significant attention, and recent studies increasingly explore large language models (LLMs) for this field. Yet most work evaluates only classification metrics, leaving unclear whether sentiment signals align with market behavior. We propose FinSentLL…

Cited by 0SourcePDFScholar
2026

Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents Collaboration

ICML 2026poster

Centralized multimodal learning commonly compresses language, acoustic, and visual signals into a single fused representation for prediction. While effective, this paradigm suffers from two limitations: modality dominance, where optimization gravitates towards the path of least resistance, ignoring …

Cited by 0SourceScholar
2026

MULTIMODAL MULTI-AGENT EMPOWERED LEGAL JUDGMENT PREDICTION

ICASSP 2026poster

Legal Judgment Prediction (LJP) aims to predict the outcomes of legal cases based on factual descriptions, serving as a fundamental task to advance the development of legal systems. Traditional methods often rely on statistical analyses or role-based simulations but face challenges with multiple all…

Cited by 0SourcePDFScholar
2026

Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis

CVPR 2026

Multimodal Sentiment Analysis (MSA) integrates language, visual, and acoustic modalities to infer human sentiment. Most existing methods either focus on globally shared representations or modality-specific features, while overlooking signals that are shared only by certain modality pairs. This limit

Cited by 0SourceScholar
2025

FlashGS: Efficient 3D Gaussian Splatting for Large-scale and High-resolution Rendering

CVPR 2025poster

Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated significant potential over traditional rendering techniques, attracting widespread attention from both industry and academia. However, real-time rendering with 3DGS remains a challenging problem, particularly in large-scale, high-reso…

2022

InterFusion: Interaction-based 4D Radar and LiDAR Fusion for 3D Object Detection

IROS 2022poster

Many recent works detect 3D objects by several sensor modalities for autonomous driving, where high-resolution cameras and high-line LiDARs are mostly used but relatively expensive. To achieve a balance between overall cost and detection accuracy, many multi-modal fusion techniques have been suggest…

Cited by 26SourceScholar