← Search

Jihai Zhang

11 accepted papers

2026

Less Is More: Vision Representation Compression for Efficient Video Generation with Large Language Models

AAAI 2026technical

Video generation using Large Language Models (LLMs) has shown promising potential, effectively leveraging the extensive LLM infrastructure to provide a unified framework for multimodal understanding and content generation. However, these methods face critical challenges, i.e., token redundancy and i

Cited by 0SourcePDFScholar
2025

CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling

EMNLP 2025

Contrastive Language-Image Pre-training (CLIP) has become a cornerstone in multimodal intelligence. However, recent studies discovered that CLIP can only encode one aspect of the feature space, leading to substantial information loss and indistinctive features. To mitigate this issue, this paper int

2025

Filter-then-Generate: Large Language Models with Structure-Text Adapter for Knowledge Graph Completion

COLING 2025main

Large Language Models (LLMs) present massive inherent knowledge and superior semantic comprehension capability, which have revolutionized various tasks in natural language processing. Despite their success, a critical gap remains in enabling LLMs to perform knowledge graph completion (KGC). Empirica…

2025

Scale Down to Speed Up: Dynamic Data Selection for Reinforcement Learning

EMNLP 2025

Optimizing data utilization remains a central challenge in applying Reinforcement Learning (RL) to Large Language Models (LLMs), directly impacting sample efficiency, training stability, and final model performance.Current approaches often rely on massive static datasets, leading to computational in

Cited by 0SourcePDFScholar
2025

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model

NeurIPS 2025poster

While large language models (LLMs) demonstrate strong reasoning capabilities utilizing reinforcement learning (RL) with verifiable reward, whether large vision-language models (VLMs) can directly inherit such capabilities through similar post-training strategies remains underexplored. In this work,…

Cited by 0SourcecodeScholar
2024

BC-Prover: Backward Chaining Prover for Formal Theorem Proving

EMNLP 2024main

Despite the remarkable progress made by large language models in mathematical reasoning, interactive theorem proving in formal logic still remains a prominent challenge. Previous methods resort to neural models for proofstep generation and search. However, they suffer from exploring possible proofst…

Cited by 0SourcePDFScholar
2024

Learning the Unlearned: Mitigating Feature Suppression in Contrastive Learning

ECCV 2024poster

"Self-Supervised Contrastive Learning has proven effective in deriving high-quality representations from unlabeled data. However, a major challenge that hinders both unimodal and multimodal contrastive learning is feature suppression, a phenomenon where the trained model captures only a limited port…

2024

SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information

EMNLP 2024main

Large Vision-Language Models (LVLMs) have become pivotal at the intersection of computer vision and natural language processing. However, the full potential of LVLMs’ Retrieval-Augmented Generation (RAG) capabilities remains underutilized. Existing works either focus solely on the text modality or a…

2024

Solving General Natural-Language-Description Optimization Problems with Large Language Models

NAACL 2024industry

Optimization problems seek to find the best solution to an objective under a set of constraints, and have been widely investigated in real-world applications. Modeling and solving optimization problems in a specific domain typically require a combination of domain knowledge, mathematical skills, and…

2022

Neighbor-Augmented Transformer-Based Embedding for Retrieval

ICASSP 2022accepted

With rapid evolution of e-commerce, it is essential but challenging to quickly provide a recommending service for users. The recommender system can be divided into two stages: retrieval and ranking. However, most recent academic research has focused on the second stage for datasets with limited size…

Cited by 0SourceScholar
2021

Synergetic Learning of Heterogeneous Temporal Sequences for Multi-Horizon Probabilistic Forecasting

AAAI 2021technical

Time-series is ubiquitous across applications, such as transportation, finance and healthcare. Time-series is often influenced by external factors, especially in the form of asynchronous events, making forecasting difficult. However, existing models are mainly designated for either synchronous time-…

Cited by 17SourcePDFScholar