← Search

Yaping Zhang

16 accepted papers

2026

MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation

CVPR 2026

End-to-end text-image machine translation (TIMT), which directly translates textual content in images across languages, is crucial for real-world multilingual scene understanding. Despite advances in vision-language large models (VLLMs), robustness across diverse visual scenes and low-resource langu

Cited by 0SourceScholar
2025

A Query-Response Framework for Whole-Page Complex-Layout Document Image Translation with Relevant Regional Concentration

ACL 2025finding

Document Image Translation (DIT), which aims at translating documents in images from source language to the target, plays an important role in Document Intelligence. It requires a comprehensive understanding of document multi-modalities and a focused concentration on relevant textual regions during…

Cited by 0SourcePDFScholar
2025

From Chaotic OCR Words to Coherent Document: A Fine-to-Coarse Zoom-Out Network for Complex-Layout Document Image Translation

COLING 2025main

Document Image Translation (DIT) aims to translate documents in images from one language to another. It requires visual layouts and textual contents understanding, as well as document coherence capturing. However, current methods often rely on the quality of OCR output, which, particularly in comple…

2025

From Generic Empathy to Personalized Emotional Support: A Self-Evolution Framework for User Preference Alignment

EMNLP 2025

Effective emotional support hinges on understanding users’ emotions and needs to provide meaningful comfort during multi-turn interactions. Large Language Models (LLMs) show great potential for expressing empathy; however, they often deliver generic responses that fail to address users’ specific nee

2025

Improving MLLM’s Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency

ACL 2025finding

Multimodal Large Language Models (MLLMs) have shown strong performance in document image tasks, especially Optical Character Recognition (OCR). However, they struggle with Document Image Machine Translation (DIMT), which requires handling both cross-modal and cross-lingual challenges. Previous effor…

Cited by 0SourcePDFScholar
2025

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation

ACL 2025long

Document Image Machine Translation (DIMT) aims to translate text within document images, facing generalization challenges due to limited training data and the complex interplay between visual and textual information. To address these challenges, we introduce M4Doc, a novel single-to-mix Modality ali…

Cited by 0SourcePDFScholar
2025

SweetieChat: A Strategy-Enhanced Role-playing Framework for Diverse Scenarios Handling Emotional Support Agent

COLING 2025main

Large Language Models (LLMs) have demonstrated promising potential in providing empathetic support during interactions. However, their responses often become verbose or overly formulaic, failing to adequately address the diverse emotional support needs of real-world scenarios. To tackle this challen…

Cited by 4SourcePDFScholar
2024

Born a BabyNet with Hierarchical Parental Supervision for End-to-End Text Image Machine Translation

COLING 2024main

Text image machine translation (TIMT) aims at translating source language texts in images into another target language, which has been proven successful by bridging text image recognition encoder and text translation decoder. However, it is still an open question of how to incorporate fine-grained k…

2024

Document Image Machine Translation with Dynamic Multi-pre-trained Models Assembling

NAACL 2024long

Text image machine translation (TIMT) is a task that translates source texts embedded in the image to target translations. The existing TIMT task mainly focuses on text-line-level images. In this paper, we extend the current TIMT task and propose a novel task, **D**ocument **I**mage **M**achine **T*…

2024

Vector Quantization Knowledge Transfer for End-to-End Text Image Machine Translation

ICASSP 2024accepted

End-to-end text image machine translation (TIMT) aims at translating source language embedded in images into target language without recognizing intermediate texts in images. However, the data scarcity of end-to-end TIMT task limits the translation performance. Existing research explores aligning co…

Cited by 0SourceScholar
2023

CCIM: Cross-modal Cross-lingual Interactive Image Translation

EMNLP 2023short findings

Text image machine translation (TIMT) which translates source language text images into target language texts has attracted intensive attention in recent years. Although the end-to-end TIMT model directly generates target translation from encoded text image features with an efficient architecture, i…

Cited by 0SourceScholar
2023

LayoutDIT: Layout-Aware End-to-End Document Image Translation with Multi-Step Conductive Decoder

EMNLP 2023long findings

Document image translation (DIT) aims to translate text embedded in images from one language to another. It is a challenging task that needs to understand visual layout with text semantics simultaneously. However, existing methods struggle to capture the crucial visual layout in real-world complex d…

Cited by 0SourceScholar
2019

Efficient Belief Propagation Detection Based on Channel Hardening for Massive MIMO

ICASSP 2019accepted

For massive multiple-input multiple-output (MIMO) detection, belief propagation (BP) based on graphical models has become a popular detection algorithm since it provides a good tradeoff between performance and complexity. To further lower the complexity of BP detection, an efficient BP detection bas…

Cited by 0SourceScholar
2019

Loss and Double-edge-triggered Detector for Robust Small-footprint Keyword Spotting

ICASSP 2019accepted

Keyword spotting (KWS) system constitutes a critical component of human-computer interfaces, which detects the specific keyword from a continuous stream of audio. The goal of KWS is providing a high detection accuracy at a low false alarm rate while having small memory and computation requirements.…

Cited by 17SourceScholar
2019

Sequence-To-Sequence Domain Adaptation Network for Robust Text Image Recognition

CVPR 2019poster

Domain adaptation has shown promising advances for alleviating domain shift problem. However, recent visual domain adaptation works usually focus on non-sequential object recognition with a global coarse alignment, which is inadequate to transfer effective knowledge for sequence-like text images wit…

Cited by 163PDFScholar
2018

Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training

ICASSP 2018accepted

In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a well-designed speech enhancement approach as the front-end of…

Cited by 0SourceScholar