← Search

Nong Xiao

8 accepted papers

2026

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning

CVPR 2026

When video reasoning requires external knowledge, many systems with large multimodal models (LMMs) adopt retrieval augmentation to supply the missing context. Appending textual or multi-clip evidence, however, forces heterogeneous signals into a single attention space. We observe diluted attention a

Cited by 0SourceScholar
2026

Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video Reasoning

CVPR 2026

Video reasoning has advanced with large multimodal models (LMMs), yet their inference is often a single pass that returns an answer without verifying whether the reasoning is evidence-aligned. We introduce **Reinforce to Learn, Elect to Reason (RLER)**, a dual paradigm that decouples learning to pro

Cited by 0SourceScholar
2025

FedCSR: A Federated Framework for Multi-Platform Cross-Domain Sequential Recommendation with Dual Contrastive Learning

COLING 2025main

Cross-domain sequential recommendation (CSR) has garnered significant attention. Current federated frameworks for CSR leverage information across multiple domains but often rely on user alignment, which increases communication costs and privacy risks. In this work, we propose FedCSR, a novel federat…

2023

Efficient Personalized Federated Learning on Selective Model Training

ICASSP 2023accepted

Personalized Federated Learning (FL) handles the data heterogeneous problem by tailoring local models for each distributed data owner. Previous studies first train a highly-adaptable global model and then transfer it for personalization. However, the additional training aggravates burden of resource…

Cited by 0SourceScholar
2021

Improving Math Word Problems with Pre-trained Knowledge and Hierarchical Reasoning

EMNLP 2021main

The recent algorithms for math word problems (MWP) neglect to use outside knowledge not present in the problems. Most of them only capture the word-level relationship and ignore to build hierarchical reasoning like the human being for mining the contextual structure between words and sentences. In t…

Cited by 46SourcePDFScholar
2021

Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering

NeurIPS 2021poster

Recent advances in the video question answering (i.e., VideoQA) task have achieved strong success by following the paradigm of fine-tuning each clip-text pair independently on the pretrained transformer-based model via supervised learning. Intuitively, multiple samples (i.e., clips) should be interd…

Cited by 48SourcePDFScholar
2019

Heterogeneous Graph Learning for Visual Commonsense Reasoning

NeurIPS 2019spotlight

Visual commonsense reasoning task aims at leading the research field into solving cognition-level reasoning with the ability to predict correct answers and meanwhile providing convincing reasoning paths, resulting in three sub-tasks i.e., Q->A, QA->R and Q->AR. It poses great challenges over the pro…

2019

Layout-Graph Reasoning for Fashion Landmark Detection

CVPR 2019poster

Detecting dense landmarks for diverse clothes, as a fundamental technique for clothes analysis, has attracted increasing research attention due to its huge application potential. However, due to the lack of modeling underlying semantic layout constraints among landmarks, prior works often detect amb…

Cited by 52PDFScholar