← Search

Lihong Wang

11 accepted papers

2026

RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models

AAAI 2026technical

Pre-trained Vision-Language Models (VLMs), e.g. CLIP, have become essential tools in multimodal transfer learning. However, fine-tuning VLMs in few-shot scenarios poses significant challenges in balancing task-specific adaptation and generalization in the obtained model. Meanwhile, current researc

Cited by 0SourcePDFScholar
2026

ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking

CVPR 2026

CoT has significantly enhanced the reasoning ability of LLMs while it faces challenges when extended to multimodal domains, particularly in mathematical tasks.Existing MLLMs typically perform textual reasoning solely from a single static mathematical image, overlooking dynamic visual acquisition dur

Cited by 0SourcecodeScholar
2025

4KAgent: Agentic Any Image to 4K Super-Resolution

NeurIPS 2025poster

We present 4KAgent, a unified agentic super-resolution generalist system designed to universally upscale any image to 4K resolution (and even higher, if applied iteratively). Our system can transform images from extremely low resolutions with severe degradations, for example, highly distorted inputs…

Cited by 0SourcecodeScholar
2025

Diff3DS: Generating View-Consistent 3D Sketch via Differentiable Curve Rendering

ICLR 2025poster

3D sketches are widely used for visually representing the 3D shape and structure of objects or scenes. However, the creation of 3D sketch often requires users to possess professional artistic skills. Existing research efforts primarily focus on enhancing the ability of interactive sketch generation…

2025

Towards S²-Challenges Underlying LLM-Based Augmentation for Personalized News Recommendation

AAAI 2025technical

Personalized news recommendation aims to recommend candidate news to the target user. Since the data and knowledge involved in traditional recommender systems are restricted, recent studies utilize large language models (LLMs) to generate news articles and augment the original dataset. However, desp…

Cited by 0SourcePDFScholar
2023

Dual-Gated Fusion with Prefix-Tuning for Multi-Modal Relation Extraction

ACL 2023findings

Multi-Modal Relation Extraction (MMRE) aims at identifying the relation between two entities in texts that contain visual clues. Rich visual content is valuable for the MMRE task, but existing works cannot well model finer associations among different modalities, failing to capture the truly helpful…

2023

Multi-Modal Knowledge Graph Transformer Framework for Multi-Modal Entity Alignment

EMNLP 2023long findings

Multi-Modal Entity Alignment (MMEA) is a critical task that aims to identify equivalent entity pairs across multi-modal knowledge graphs (MMKGs). However, this task faces challenges due to the presence of different types of information, including neighboring entities, multi-modal attributes, and ent…

Cited by 0SourcecodeScholar
2023

Prototype-Guided Pseudo Labeling for Semi-Supervised Text Classification

ACL 2023long

Semi-supervised text classification (SSTC) aims at text classification with few labeled data and massive unlabeled data. Recent works achieve this task by pseudo-labeling methods, with the belief that the unlabeled and labeled data have identical data distribution, and assign the unlabeled data with…

2022

Explicitly Modeling Importance and Coherence for Timeline Summarization

ICASSP 2022accepted

Timeline summarization (TLS) identifies major events and generates short summaries on how the event evolves in a period of time. Existing timeline summarization methods generate summaries by considering the coverage and diversity of the content and temporized information but ignore the importance an…

Cited by 0SourceScholar
2021

Enhancing Deep Paraphrase Identification via Leveraging Word Alignment Information

ICASSP 2021accepted

Recent deep learning based methods have achieved impressive performance on paraphrase identification (PI), a fundamental NLP task, judging whether two sentences are semantically equivalent or not. However, their success heavily relies on massive labeled samples, which are time-consuming and expensiv…

Cited by 0SourceScholar
2021

TextGTL: Graph-based Transductive Learning for Semi-supervised Text Classification via Structure-Sensitive Interpolation

IJCAI 2021poster

Compared with traditional sequential learning models, graph-based neural networks exhibit excellent properties when encoding text, such as the capacity of capturing global and local information simultaneously. Especially in the semi-supervised scenario, propagating information along the edge can eff…

Cited by 31SourcePDFScholar