← Search

Yipeng Zhang

16 accepted papers

2026

Advancing LLM Reasoning with Natural Language and Numerical Feedback

ICML 2026spotlight

Recent advances in reinforcement learning (RL) using numerical rewards have significantly enhanced the complex reasoning capabilities of large language models (LLMs). However, we identify three fundamental limitations of purely numerical feedback: performance plateaus, ineffective spontaneous self-r…

Cited by 0SourceScholar
2026

Cross-Scale Collaboration between LLMs and Lightweight Sequential Recommenders with Domain-Specific Latent Reasoning

AAAI 2026technical

Sequential recommendation aims to predict the next item based on historical interactions. To further enhance the reasoning capability in sequential recommendation, LLMs are employed to predict the next item or generate semantic IDs for item representation, given LLMs

Cited by 0SourcePDFScholar
2026

LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic Pyramid

AAAI 2026technical

Vision transformers (ViTs) are widely employed in multimodal large language models (MLLMs) for visual encoding. However, they exhibit inferior performance on tasks regarding fine-grained visual perception. We attribute this to the inner limitations of ViTs in capturing diverse visual semantic level

Cited by 0SourcePDFScholar
2026

Omni-iEEG: A Large-Scale, Comprehensive iEEG Dataset and Benchmark for Epilepsy Research

ICLR 2026poster

Epilepsy affects over 50 million people worldwide, and one-third of patients suffer drug-resistant seizures where surgery offers the best chance of seizure freedom. Accurate localization of the epileptogenic zone (EZ) relies on intracranial EEG (iEEG). Clinical workflows, however, remain constrained…

Cited by 0SourcecodeScholar
2026

Reasoning Diffusion for Unpaired Test Time Out-of-distribution Text-Image to Video Generation

CVPR 2026

Text-image to video generation aims to synthesize a video conditioned on the given text-image inputs. Nevertheless, existing methods generally assume that the semantic information carried in the input text and image tends to be perfectly paired and temporally aligned, occurring simultaneously in the

Cited by 0SourceScholar
2026

Self-Supervised Learning from Structural Invariance

ICLR 2026poster

Joint-embedding self-supervised learning (SSL), the key paradigm for unsupervised representation learning from visual data, learns from invariances between semantically-related data pairs. We study the one-to-many mapping problem in SSL, where each datum may be mapped to multiple valid targets. Thi…

Cited by 0SourcecodeScholar
2026

Temporal-aware Flow Matching for Video Generation with Temporally Coherent Motion

ICML 2026poster

Despite rapid advances in text-to-video generation, state-of-the-art generative models still suffer from producing temporally incoherent and unrealistic motion for videos. The key weakness of existing works is that they commonly treat videos as frame sequences and directly adopt Flow Matching object…

Cited by 0SourcecodeScholar
2025

Modular-Cam: Modular Dynamic Camera-view Video Generation with LLM

AAAI 2025technical

Text-to-Video generation, which utilizes the provided text prompt to generate high-quality videos, has drawn increasing attention and achieved great success due to the development of diffusion models recently. Existing methods mainly rely on a pre-trained text encoder to capture the semantic informa…

2025

Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching

ACL 2025finding

Large language models (LLMs) often struggle to provide up-to-date information due to their one-time training and the constantly evolving nature of the world. To keep LLMs current, existing approaches typically involve continued pre-training on new documents. However, they frequently face difficultie…

2024

Coarse-to-Fine Detection of Multiple Seams for Robotic Welding

IROS 2024poster

Efficiently detecting target weld seams while ensuring sub-millimeter accuracy has always been an important challenge in autonomous welding, which has significant application in industrial practice. Previous works mostly focused on recognizing and localizing welding seams one by one, leading to infe…

Cited by 0SourceScholar
2024

DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation

ICLR 2024poster

Subject-driven text-to-image generation aims to generate customized images of the given subject based on the text descriptions, which has drawn increasing attention. Existing methods mainly resort to finetuning a pretrained generative model, where the identity-relevant information (e.g., the boy) an…

2020

E3SN: Efficient End-to-End Siamese Network for Video Object Segmentation

IJCAI 2020poster

In the semi-supervised video object segmentation (VOS) field, SiamMask has achieved competitive accuracy and the fastest running speed. However, the two-stage training procedure requires additional manual intervention, and using only single-level features does not maximize the rich hierarchical feat…

Cited by 0SourcePDFScholar
2020

Towards A Friendly Online Community: An Unsupervised Style Transfer Framework for Profanity Redaction

COLING 2020main

Offensive and abusive language is a pressing problem on social media platforms. In this work, we propose a method for transforming offensive comments, statements containing profanity or offensive language, into non-offensive ones. We design a Retrieve, Generate and Edit unsupervised style transfer p…

Cited by 35SourcePDFScholar