← Search

Yuqi Zhang

24 accepted papers

2026

MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action Agent

CVPR 2026

Recent Vision-Language-Action (VLA) models reformulate vision-language models by tuning them with millions of robotic demonstrations. While they perform well when fine-tuned for a single embodiment or task family, extending them to multi-skill settings remains challenging: directly merging VLA exper

Cited by 0SourcecodeScholar
2025

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders

EMNLP 2025

We introduce AMIA, a lightweight, inference-only defense for Large Vision–Language Models (LVLMs) that (1) Automatically Masks a small set of text-irrelevant image patches to disrupt adversarial perturbations, and (2) conducts joint Intention Analysis to uncover and mitigate hidden harmful intents b

2025

Goku: Flow Based Video Generative Foundation Models

CVPR 2025highlight

This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model ar…

Cited by 15SourcePDFScholar
2025

Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

CVPR 2025poster

We present Infinity, a Bitwise Visual AutoRegressive Modeling capable of generating high-resolution, photorealistic images following language instruction. Infinity refactors visual autoregressive model under a bitwise token prediction framework with an infinite-vocabulary classifier and bitwise self…

2025

Intention Analysis Makes LLMs A Good Jailbreak Defender

COLING 2025main

Aligning large language models (LLMs) with human values, particularly when facing complex and stealthy jailbreak attacks, presents a formidable challenge. Unfortunately, existing methods often overlook this intrinsic nature of jailbreaks, which limits their effectiveness in such complex scenarios. I…

2025

RobustSplat: Decoupling Densification and Dynamics for Transient-Free 3DGS

ICCV 2025poster

3D Gaussian Splatting (3DGS) has gained significant attention for its real-time, photo-realistic rendering in novel-view synthesis and 3D modeling. However, existing methods struggle with accurately modeling scenes affected by transient objects, leading to artifacts in the rendered images. We identi…

2025

The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking

ICML 2025poster

This work identifies the *Energy Loss Phenomenon* in Reinforcement Learning from Human Feedback (RLHF) and its connection to reward hacking. Specifically, energy loss in the final layer of a Large Language Model (LLM) gradually increases during the RL process, with an *excessive* increase in energy…

Cited by 0SourcePDFScholar
2025

Tree-of-AdEditor: Heuristic Tree Reasoning for Automated Video Advertisement Editing with Large Language Model

IJCAI 2025

Video advertising has become a popular marketing strategy on e-commerce platforms, requiring high-level semantic reasoning like selling point discovery, narrative organization. Previous rule-based methods struggle with these complex tasks, and learning-based approaches demand large datasets and high

2024

Aerial Lifting: Neural Urban Semantic and Building Instance Lifting from Aerial Imagery

CVPR 2024poster

We present a neural radiance field method for urban-scale semantic and building-level instance segmentation from aerial images by lifting noisy 2D labels to 3D. This is a challenging problem due to two primary reasons. Firstly objects in urban aerial images exhibit substantial variations in size inc…

2024

Towards CLIP-driven Language-free 3D Visual Grounding via 2D-3D Relational Enhancement and Consistency

CVPR 2024poster

3D visual grounding plays a crucial role in scene understanding with extensive applications in AR/VR. Despite the significant progress made in recent methods the requirement of dense textual descriptions for each individual object which is time-consuming and costly hinders their scalability. To miti…

2023

Disambiguated Lexically Constrained Neural Machine Translation

ACL 2023findings

Lexically constrained neural machine translation (LCNMT), which controls the translation generation with pre-specified constraints, is important in many practical applications. Current approaches to LCNMT typically assume that the pre-specified lexicon constraints are contextually appropriate. This…

Cited by 4SourcePDFScholar
2023

Easy Guided Decoding in Providing Suggestions for Interactive Machine Translation

ACL 2023long

Machine translation technology has made great progress in recent years, but it cannot guarantee error-free results. Human translators perform post-editing on machine translations to correct errors in the scene of computer aided translation. In favor of expediting the post-editing process, many works…

2023

Improving Neural Machine Translation by Multi-Knowledge Integration with Prompting

EMNLP 2023long findings

Improving neural machine translation (NMT) systems with prompting has achieved significant progress in recent years. In this work, we focus on how to integrate multi-knowledge, multiple types of knowledge, into NMT models to enhance the performance with prompting. We propose a unified framework, whi…

Cited by 0SourceScholar
2022

Ada-NETS: Face Clustering via Adaptive Neighbour Discovery in the Structure Space

ICLR 2022poster

Face clustering has attracted rising research interest recently to take advantage of massive amounts of face images on the web. State-of-the-art performance has been achieved by Graph Convolutional Networks (GCN) due to their powerful representation capacity. However, existing GCN-based methods buil…

2022

Adaptive Matching Strategy for Multi-Target Multi-Camera Tracking

ICASSP 2022accepted

Multi-Target Multi-Camera Tracking has a wide range of applications and is the basis for many high-level inference and prediction tasks. How to make the system perform efficiently on a large number of cameras is a crucial research issue. Previous works have proposed many matching strategies to reduc…

Cited by 0SourceScholar
2022

Domain Adaptation via Mutual Information Maximization for Handwriting Recognition

ICASSP 2022accepted

Deep learning models for handwriting recognition have been developed in recent years. To improve the model’s generalization ability for sequence modeling task, this paper proposes to use domain adaptation with statistical distribution alignment and entropy regularization. For statistical distributio…

Cited by 0SourceScholar
2022

Graph Convolution for Re-Ranking in Person Re-Identification

ICASSP 2022accepted

Nowadays, deep learning is widely applied to extract features for similarity computation in person re-identification (re-ID). However, the difference between the training data and testing data makes the performance of learned feature degraded during testing. Hence, re-ranking is proposed to mitigate…

Cited by 0SourceScholar
2022

Learning to Incorporate Texture Saliency Adaptive Attention to Image Cartoonization

ICML 2022spotlight

Image cartoonization is recently dominated by generative adversarial networks (GANs) from the perspective of unsupervised image-to-image translation, in which an inherent challenge is to precisely capture and sufficiently transfer characteristic cartoon styles (e.g., clear edges, smooth color shadin…

2022

Third-Party Aligner for Neural Word Alignments

EMNLP 2022finding

Word alignment is to find translationally equivalent words between source and target sentences. Previous work has demonstrated that self-training can achieve competitive word alignment results. In this paper, we propose to use word alignments generated by a third-party word aligner to supervise the…

2021

Beyond Glass-Box Features: Uncertainty Quantification Enhanced Quality Estimation for Neural Machine Translation

EMNLP 2021finding

Quality Estimation (QE) plays an essential role in applications of Machine Translation (MT). Traditionally, a QE system accepts the original source text and translation from a black-box MT system as input. Recently, a few studies indicate that as a by-product of translation, QE benefits from the mod…

Cited by 5SourcePDFScholar
2020

Placepedia: Comprehensive Place Understanding with Multi-Faceted Annotations

ECCV 2020poster

Place is an important element in visual understanding. Given a photo of a building, people can often tell its functionality, e.g. a restaurant or a shop, its cultural style, e.g. Asian or European, as well as its economic type, e.g. industry oriented or tourism oriented. While place recognition has…

Cited by 8SourcePDFScholar