← Search

Anxiang Zeng

7 accepted papers

2026

CCCaption: Dual-Reward Reinforcement Learning for Complete and Correct Image Captioning

CVPR 2026

Image captioning remains a fundamental task for vision-language understanding, yet ground-truth supervision still relies predominantly on human-annotated references.Because human annotations reflect subjective preferences and expertise, ground-truth captions are often incomplete or even incorrect, w

Cited by 0SourcecodeScholar
2026

SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia

CVPR 2026

Multilingual document and scene text understanding plays an important role in applications such as search, finance, and public services. However, most existing benchmarks focus on high-resource languages and fail to evaluate models in realistic multilingual environments. In Southeast Asia, the diver

Cited by 0SourcecodeScholar
2026

Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training

ICML 2026poster

Supervised fine-tuning (SFT) is computationally efficient but often yields inferior generalization compared to reinforcement learning (RL). This gap is primarily driven by RL’s use of on-policy data. We propose a framework to bridge this chasm by enabling On-Policy SFT. We first present ***Distribut…

Cited by 0SourceScholar
2026

TreeBridge: Aligning LLM Embeddings in Industrial Recommender Systems

AAAI 2026technical

Large language models (LLMs) have shown great potential in enhancing search and recommender systems by providing rich semantic representations from unstructured texts. However, directly integrating LLM embeddings into industrial recommendation pipelines often results in subpar performance due to the

Cited by 0SourcePDFScholar
2025

Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization

ACL 2025long

Direct Preference Optimization (DPO) has emerged as a promising framework for aligning Large Language Models (LLMs) with human preferences by directly optimizing the log-likelihood difference between chosen and rejected responses. However, existing methods assign equal importance to all tokens in th…

2024

Transformer-Empowered Multi-Modal Item Embedding for Enhanced Image Search in E-commerce

AAAI 2024technical

Over the past decade, significant advances have been made in the field of image search for e-commerce applications. Traditional image-to-image retrieval models, which focus solely on image details such as texture, tend to overlook useful semantic information contained within the images. As a result,…

Cited by 1SourcePDFScholar
2023

Recurrent Temporal Revision Graph Networks

NeurIPS 2023poster

Temporal graphs offer more accurate modeling of many real-world scenarios than static graphs. However, neighbor aggregation, a critical building block of graph networks, for temporal graphs, is currently straightforwardly extended from that of static graphs. It can be computationally expensive when…

Cited by 2SourcePDFScholar