← Search

Linlin Zhang

8 accepted papers

2025

OmniTry: Virtual Try-On Anything without Masks

NeurIPS 2025poster

Virtual Try-ON (VTON) is a practical and widely-applied task, for which most of existing works focus on clothes. This paper presents OmniTry, a unified framework that extends VTON beyond garment to encompass any wearable objects, e.g., jewelries and accessories, with mask-free setting for more pract…

Cited by 0SourcecodeScholar
2025

ThinkAnswer Loss: Balancing Semantic Similarity and Exact Matching for LLM Reasoning Enhancement

EMNLP 2025

Knowledge distillation for large language models often uses Chain-of-Thought (CoT) and answer pairs, but existing methods struggle with appropriate supervision signals. Uniform constraints (e.g., cross-entropy) on CoT can enforce literal, verbose reasoning and suppress expressive diversity, while so

Cited by 0SourcePDFScholar
2024

Divergence-Guided Simultaneous Speech Translation

AAAI 2024technical

To achieve high-quality translation with low latency, a Simultaneous Speech Translation (SimulST) system relies on a policy module to decide whether to translate immediately or wait for additional streaming input, along with a translation model capable of effectively handling partial speech input. P…

2024

Simple but Effective Compound Geometric Operations for Temporal Knowledge Graph Completion

ACL 2024long

Temporal knowledge graph completion aims to infer the missing facts in temporal knowledge graphs. Current approaches usually embed factual knowledge into continuous vector space and apply geometric operations to learn potential patterns in temporal knowledge graphs. However, these methods only adopt…

2023

A Simple Concatenation can Effectively Improve Speech Translation

ACL 2023short

A triple speech translation data comprises speech, transcription, and translation. In the end-to-end paradigm, text machine translation (MT) usually plays the role of a teacher model for the speech translation (ST) via knowledge distillation. Parameter sharing with the teacher is often adopted to co…

2023

Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens Strategy

EMNLP 2023long main

Simultaneous Speech Translation (SimulST) is a task focused on ensuring high-quality translation of speech in low-latency situations. Despite this, the modality gap (\emph{e.g.}, unknown word boundaries) between audio and text presents a challenge. This gap hinders the effective application of pol…

Cited by 0SourceScholar
2022

Context-Adaptive Document-Level Neural Machine Translation

ICASSP 2022accepted

Document-level translation models are still far from perfect. Most existing document-level neural machine translation (NMT) models leverage a fixed number of the previous or all global sentences to handle the context-independent problem in standard NMT. However, the translating of each source senten…

Cited by 0SourceScholar
2020

GFNet: A Lightweight Group Frame Network for Efficient Human Action Recognition

ICASSP 2020accepted

Human action recognition aims at assigning an action label to a well-segmented video. Recent work using two-stream or 3D convolutional neural networks achieves high recognition rates at the cost of huge computation complexity, memory footprint, and parameters. In this paper, we propose a lightweight…

Cited by 0SourceScholar