← Search

Xingjiao Wu

13 accepted papers

2026

Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models

ICLR 2026poster

Large Vision-Language Models (LVLMs) exhibit outstanding performance on vision-language tasks but struggle with hallucination problems. Through in-depth analysis of LVLM activation patterns, we reveal two key findings: 1) truthfulness and visual perception capabilities predominantly engage different…

Cited by 4SourceScholar
2026

VecDesigner: Exploring Visual Guidance and Structural Consistency for Semantic Typography

ICML 2026poster

Semantic Typography aims to visualize the meaning of an input word through the form of a character, while preserving its legibility. Existing vector-based methods, which primarily rely on text-driven optimization like Score Distillation Sampling (SDS), often produce glyphs that lack rich semantic de…

Cited by 0SourceScholar
2025

An Exemplar-based Framework for Chinese Text Recognition

AAAI 2025technical

This paper introduces a novel exemplar-based framework for reading Chinese texts in natural scene or document images. We present the Deep Exemplar-based Chinese Text Recognizer, which is structured to first identify candidate characters as exemplars from each text-line, and subsequently recognize th…

Cited by 0SourcePDFScholar
2025

CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering

CVPR 2025highlight

Multimodal large language models (MLLMs) have garnered widespread attention from researchers due to their remarkable understanding and generation capabilities in visual language tasks (e.g., visual question answering). However, the rapid pace of knowledge updates in the real world makes offline trai…

Cited by 2SourcePDFScholar
2025

Lark: Low-Rank Updates After Knowledge Localization for Few-shot Class-Incremental Learning

ICCV 2025poster

For Few-Shot Class-Incremental Learning (FSCIL), direct fine-tuning causes significant parameter shifts, resulting in catastrophic forgetting and increased resource consumption. While, freezing the pre-trained backbone exacerbates the inconsistency between the backbone and the evolving classifier. T…

Cited by 0SourcePDFScholar
2025

Multi-Type Preference Learning: Empowering Preference-Based Reinforcement Learning with Equal Preferences

ICRA 2025

Preference-Based reinforcement learning (PBRL) learns directly from the preferences of human teachers regarding agent behaviors without needing meticulously designed reward functions. However, existing PBRL methods often learn primarily from explicit preferences, neglecting the possibility that teac

Cited by 1SourcecodeScholar
2025

RMoA: Optimizing Mixture-of-Agents through Diversity Maximization and Residual Compensation

ACL 2025finding

Although multi-agent systems based on large language models show strong capabilities on multiple tasks, they are still limited by high computational overhead, information loss, and robustness. Inspired by ResNet’s residual learning, we propose Residual Mixture-of-Agents (RMoA), integrating residual…

2025

Unleashing the Semantic Adaptability of Controlled Diffusion Model for Image Colorization

IJCAI 2025

Recent data-driven image colorization methods have leveraged pre-trained Text-to-Image (T2I) diffusion models as generative prior, while still suffering from unsatisfactory and inaccurate semantic-level color control. To address these issues, we propose a Semantic Adaptation method (SeAda) that enha

2023

LoGoNet: Towards Accurate 3D Object Detection With Local-to-Global Cross-Modal Fusion

CVPR 2023poster

LiDAR-camera fusion methods have shown impressive performance in 3D object detection. Recent advanced multi-modal methods mainly perform global fusion, where image features and point cloud features are fused across the whole scene. Such practice lacks fine-grained region-level information, yielding…

2022

Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection

ECCV 2022poster

"Multi-modal 3D object detection has been an active research topic in autonomous driving. Nevertheless, it is non-trivial to explore the cross-modal feature fusion between sparse 3D points and dense 2D pixels. Recent approaches either fuse the image features with the point cloud features that are pr…

2022

Multi-Channel Attentive Graph Convolutional Network with Sentiment Fusion for Multimodal Sentiment Analysis

ICASSP 2022accepted

Nowadays, with the explosive growth of multimodal reviews on social media platforms, multimodal sentiment analysis has recently gained popularity because of its high relevance to these social media posts. Although most previous studies design various fusion frameworks for learning an interactive rep…

Cited by 0SourceScholar
2020

Scene Text Recognition with Temporal Convolutional Encoder

ICASSP 2020accepted

Texts from scene images typically consist of several characters and exhibit a characteristic sequence structure. Existing methods capture the structure with the sequence-to-sequence models by an encoder to have the visual representations and then a decoder to translate the features into the label se…

Cited by 0SourceScholar