← Search

Qing Du

10 accepted papers

2026

FAM: Fine-Grained Alignment Matters in Multimodal Embedding Learning with Large Vision-Language Models

AAAI 2026technical

Learning multimodal representation is a fundamental task that supports a wide range of applications such as visual-text retrieval. While pioneering approaches e.g., CLIP paves the way by learning separated encoders for different modalities, they struggle to model complex interactions between modalit

Cited by 0SourcePDFScholar
2026

Future-Gain Guided Test-Time Learning for Large Language Models

ICML 2026poster

Large language models (LLMs) inevitably encounter distribution shifts during real-world deployment, leading to performance degradation. Although test-time learning (TTL) adapts LLMs from unlabeled test streams, applying entropy minimization to autoregressive generation faces two challenges: (i) earl…

Cited by 0SourceScholar
2026

NaVLA$^2$: A Vision-Language-Audio-Action Model for Multimodal Instruction Navigation

AAAI 2026technical

Embodied navigation is a fundamental capability for intelligent agents, yet remains challenging in partially observable environments where navigation instructions can be difficult to interpret. However, existing tasks only provide unimodal instructions, which are ambiguous in complex multimodal envi

Cited by 0SourcePDFScholar
2026

Tavatar: Topology-Aware Gaussian Attribute Derivation for Animatable Human Avatars

CVPR 2026

Reconstructing high-fidelity, animatable human avatars from monocular videos remains a critical challenge. Existing 3DGS-based human animation methods constrain Gaussian parameters but exclude scale, which we argue is crucial for adapting human poses to challenging out-of-distribution poses. To achi

Cited by 0SourceScholar
2024

AlphaFin: Benchmarking Financial Analysis with Retrieval-Augmented Stock-Chain Framework

COLING 2024main

The task of financial analysis primarily encompasses two key areas: stock trend prediction and the corresponding financial question answering. Currently, machine learning and deep learning algorithms (ML&DL) have been widely applied for stock trend predictions, leading to significant progress. Howev…

2023

Disentangling Writer and Character Styles for Handwriting Generation

CVPR 2023poster

Training machines to synthesize diverse handwritings is an intriguing task. Recently, RNN-based methods have been proposed to generate stylized online Chinese characters. However, these methods mainly focus on capturing a person's overall writing style, neglecting subtle style inconsistencies betwee…

2022

Beyond Fixation: Dynamic Window Visual Transformer

CVPR 2022poster

Recently, a surge of interest in visual transformers is to reduce the computational cost by limiting the calculation of self-attention to a local window. Most current work uses a fixed single-scale window for modeling by default, ignoring the impact of window size on model performance. However, this…

Cited by 41PDFcodeScholar
2021

Source-free Domain Adaptation via Avatar Prototype Generation and Adaptation

IJCAI 2021poster

We study a practical domain adaptation task, called source-free unsupervised domain adaptation (UDA) problem, in which we cannot access source domain data due to data privacy issues but only a pre-trained source model and unlabeled target data are available. This task, however, is very difficult du…

2021

Towards Accurate Text-Based Image Captioning With Content Diversity Exploration

CVPR 2021poster

Text-based image captioning (TextCap) which aims to read and reason images with texts is crucial for a machine to understand a detailed and complex scene environment, considering that texts are omnipresent in daily life. This task, however, is very challenging because an image often contains complex…

Cited by 84PDFcodeScholar