← Search

Jianjun Li

11 accepted papers

2026

OPERA: A Reinforcement Learning--Enhanced Orchestrated Planner-Executor Architecture for Reasoning-Oriented Multi-Hop Retrieval

AAAI 2026technical

Recent advances in large language models (LLMs) and dense retrievers have driven significant progress in retrieval-augmented generation (RAG). However, existing approaches face significant challenges in complex reasoning-oriented multi-hop retrieval tasks: 1) Ineffective reasoning-oriented planning:

Cited by 0SourcePDFScholar
2026

SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema

AAAI 2026technical

Although Large language Model (LLM)-powered information extraction (IE) systems have shown impressive capabilities, current fine-tuning paradigms face two major limitations: high training costs and difficulties in aligning with LLM preferences. To address these issues, we propose a novel universal

Cited by 0SourcePDFScholar
2024

LGMRec: Local and Global Graph Learning for Multimodal Recommendation

AAAI 2024technical

The multimodal recommendation has gradually become the infrastructure of online media platforms, enabling them to provide personalized service to users through a joint modeling of user historical behaviors (e.g., purchases, clicks) and item various modalities (e.g., visual and textual). The majority…

2024

LMD: Faster Image Reconstruction with Latent Masking Diffusion

AAAI 2024technical

As a class of fruitful approaches, diffusion probabilistic models (DPMs) have shown excellent advantages in high-resolution image reconstruction. On the other hand, masked autoencoders (MAEs), as popular self-supervised vision learners, have demonstrated simpler and more effective image reconstructi…

2023

HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question Answering

AAAI 2023technical

Visual Question Answering (VQA) aims to answer the natural language question about a given image by understanding multimodal content. However, the answer quality of most existing visual-language pre-training (VLP) methods is still limited, mainly due to: (1) Incompatibility. Upstream pre-training ta…

2023

Multi-Aspect Interest Neighbor-Augmented Network for Next-Basket Recommendation

ICASSP 2023accepted

Next-basket recommendation (NBR) is a type of recommendation task that focuses on mining user interests based on the sequential basket records in which users purchase multiple items at a time. Limited by the sparsity brought by short-term user interaction behavior, existing NBR methods typically fai…

Cited by 0SourceScholar
2022

GLAF: Global-to-Local Aggregation and Fission Network for Semantic Level Fact Verification

COLING 2022main

Accurate fact verification depends on performing fine-grained reasoning over crucial entities by capturing their latent logical relations hidden in multiple evidence clues, which is generally lacking in existing fact verification models. In this work, we propose a novel Global-to-Local Aggregation a…

2022

UniTranSeR: A Unified Transformer Semantic Representation Framework for Multimodal Task-Oriented Dialog System

ACL 2022long

As a more natural and intelligent interaction manner, multimodal task-oriented dialog system recently has received great attention and many remarkable progresses have been achieved. Nevertheless, almost all existing studies follow the pipeline to first learn intra-modal features separately and then…

2021

Intention Reasoning Network for Multi-Domain End-to-end Task-Oriented Dialogue

EMNLP 2021main

Recent years has witnessed the remarkable success in end-to-end task-oriented dialog system, especially when incorporating external knowledge information. However, the quality of most existing models’ generated response is still limited, mainly due to their lack of fine-grained reasoning on determin…

2020

Weakly Supervised Fine-Grained Image Classification via Guassian Mixture Model Oriented Discriminative Learning

CVPR 2020oral

Existing weakly supervised fine-grained image recognition (WFGIR) methods usually pick out the discriminative regions from the high-level feature maps directly. We discover that due to the operation of stacking local receptive filed, Convolutional Neural Network causes the discriminative region diff…

Cited by 104PDFScholar