← Search

Jun-Yan He

15 accepted papers

2026

ViType: High-Fidelity Visual Text Rendering via Glyph-Aware Multimodal Diffusion

AAAI 2026technical

Current text-to-image models face challenges in visual text rendering: text encoders like CLIP and T5 lack glyph-level understanding and often struggle to distinguish between the specific words to be rendered and their intended semantic meaning within prompts. In addition, inconsistencies between th

Cited by 0SourcePDFScholar
2025

MetaDesigner: Advancing Artistic Typography through AI-Driven, User-Centric, and Multilingual WordArt Synthesis

ICLR 2025poster

MetaDesigner introduces a transformative framework for artistic typography synthesis, powered by Large Language Models (LLMs) and grounded in a user-centric design paradigm. Its foundation is a multi-agent system comprising the Pipeline, Glyph, and Texture agents, which collectively orchestrate the…

Cited by 2SourcePDFScholar
2025

POPoS: Improving Efficient and Robust Facial Landmark Detection with Parallel Optimal Position Search

AAAI 2025technical

Achieving a balance between accuracy and efficiency is a critical challenge in facial landmark detection (FLD). This paper introduces Parallel Optimal Position Search (POPoS), a high-precision encoding-decoding framework designed to address the limitations of traditional FLD methods. POPoS employs t…

2025

UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts

ICASSP 2025accepted

Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of human emotions, primarily relying on oversimplified emotional…

Cited by 0SourceScholar
2024

AnyText: Multilingual Visual Text Generation and Editing

ICLR 2024spotlight

Diffusion model based Text-to-Image has achieved impressive achievements recently. Although current technology for synthesizing images is highly advanced and capable of generating images with high fidelity, it is still possible to give the show away when focusing on the text area in the generated im…

2024

DCPT: Darkness Clue-Prompted Tracking in Nighttime UAVs

ICRA 2024poster

Existing nighttime unmanned aerial vehicle (UAV) trackers follow an "Enhance-then-Track" architecture - first using a light enhancer to brighten the nighttime video, then employing a daytime tracker to locate the object. This separate enhancement and tracking fails to build an end-to-end trainable v…

Cited by 17SourcecodeScholar
2024

Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning

NeurIPS 2024poster

Accurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling. However, traditional single-modality approaches often fail to capture the complexity of real-world emotional expressions, which are inherently multimodal. Moreover, exist…

2024

Human-Aware Vision-and-Language Navigation: Bridging Simulation to Reality with Dynamic Human Interactions

NeurIPS 2024spotlight

Vision-and-Language Navigation (VLN) aims to develop embodied agents that navigate based on human instructions. However, current VLN frameworks often rely on static environments and optimal expert supervision, limiting their real-world applicability. To address this, we introduce Human-Aware Vision-…

2024

Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception

CVPR 2024highlight

Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabilities. However there still remains a gap in providing fine-grained pixel-level…

2023

DAMO-StreamNet: Optimizing Streaming Perception in Autonomous Driving

IJCAI 2023poster

In the realm of autonomous driving, real-time perception or streaming perception remains under-explored. This research introduces DAMO-StreamNet, a novel framework that merges the cutting-edge elements of the YOLO series with a detailed examination of spatial and temporal perception techniques. DAMO…

2023

HDFormer: High-order Directed Transformer for 3D Human Pose Estimation

IJCAI 2023poster

Human pose estimation is a challenging task due to its structured data sequence nature. Existing methods primarily focus on pair-wise interaction of body joints, which is insufficient for scenarios involving overlapping joints and rapidly changing poses. To overcome these issues, we introduce a nove…

2023

Longshortnet: Exploring Temporal and Semantic Features Fusion In Streaming Perception

ICASSP 2023accepted

Streaming perception is a fundamental task in autonomous driving that requires a careful balance between the latency and accuracy of the autopilot system. However, current methods for streaming perception are limited as they rely only on the current and adjacent two frames to learn movement patterns…

Cited by 0SourceScholar
2023

Optimal Proposal Learning for Deployable End-to-End Pedestrian Detection

CVPR 2023poster

End-to-end pedestrian detection focuses on training a pedestrian detection model via discarding the Non-Maximum Suppression (NMS) post-processing. Though a few methods have been explored, most of them still suffer from longer training time and more complex deployment, which cannot be deployed in the…

Cited by 19SourcePDFScholar
2023

Procontext: Exploring Progressive Context Transformer for Tracking

ICASSP 2023accepted

Existing Visual Object Tracking (VOT) only takes the target area in the first frame as a template. This causes tracking to inevitably fail in fast-changing and crowded scenes, as it cannot account for changes in object appearance between frames. To this end, we revamped the tracking framework with P…

Cited by 0SourceScholar
2023

Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning

ICCV 2023oral

Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pursue unified frameworks to tackle this challenge but mostly still treat it as two individual learning tasks, which limits…

Cited by 12PDFcodeScholar