← Search

Junhao Chen

24 accepted papers

2026

$\textit{MADFormer}$: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation

ICLR 2026poster

Recent progress in multimodal generation has increasingly combined autoregressive (AR) and diffusion-based approaches, leveraging their complementary strengths: AR models capture long-range dependencies and produce fluent, context-aware outputs, while diffusion models operate in continuous latent sp…

Cited by 0SourceScholar
2026

Animator-Centric Skeleton Generation on Objects with Fine-Grained Details

CVPR 2026

Skeleton generation is essential for animating 3D assets, but current deep learning methods remain limited: they cannot handle the growing structural complexity of modern models and offer minimal controllability, creating a major bottleneck for real-world animation workflows. To address this, we pro

Cited by 0SourceScholar
2026

DanceTogether: Generating Interactive Multi-Person Video without Identity Drifting

ICLR 2026poster

Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first end-to-end diffusion framework that turns a single reference image…

Cited by 0SourceScholar
2026

GarmentGPT: Compositional Garment Pattern Generation via Discrete Latent Tokenization

ICLR 2026poster

Apparel is a fundamental component of human appearance, making garment digitalization critical for digital human creation. However, sewing pattern creation traditionally relies on the intuition and extensive experience of skilled artisans. This manual bottleneck significantly hinders the scalability…

Cited by 0SourcecodeScholar
2026

HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

CVPR 2026

Recent methods have made notable progress in the visual quality of hand-object interaction video synthesis. However, most approaches rely on 2D control signals that lack spatial expressiveness and limit the utilization of synthetic 3D conditional data. To address these limitations, we propose HVG-3D

Cited by 0SourceScholar
2026

Interpreting Fedspeak with Confidence: A LLM-Based Uncertainty-Aware Framework Guided by Monetary Policy Transmission Paths

AAAI 2026technical

"Fedspeak", the stylized and often nuanced language used by the U.S. Federal Reserve, encodes implicit policy signals and strategic stances. The Federal Open Market Committee strategically employs Fedspeak as a communication tool to shape market expectations and influence both domestic and global e

Cited by 0SourcePDFScholar
2026

LottieGPT: Tokenizing Vector Animation for Autoregressive Generation

CVPR 2026

Despite rapid progress in video generation, existing models are incapable of producing vector animation, a dominant and highly expressive form of multimedia on the Internet. Vector animations offer resolution-independence, compactness, semantic structure, and editable parametric motion representatio

Cited by 0SourceScholar
2025

DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters

CVPR 2025poster

Recent advances in generative models have enabled high-quality 3D character reconstruction from multi-modal. However, animating these generated characters remains a challenging task, especially for complex elements like garments and hair, due to the lack of large-scale datasets and effective rigging…

2025

FLIQA-AD: a Fusion Model with Large Language Model for Better Diagnose and MMSE Prediction of Alzheimer’s Disease

NAACL 2025short

Tracking a patient’s cognitive status early in the onset of the disease provides an opportunity to diagnose and intervene in Alzheimer’s disease (AD). However, relying solely on magnetic resonance imaging (MRI) images with traditional classification and regression models may not fully extract finer-…

Cited by 0SourcePDFScholar
2025

IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web

ACL 2025finding

Recently, advancements in large multimodal models have led to significant strides in image comprehension capabilities. Despite these advancements, there is a lack of a robust benchmark specifically for assessing the image‐to‐web conversion proficiency of these large models. It is essential to ensure…

2025

Idea23D: Collaborative LMM Agents Enable 3D Model Generation from Interleaved Multimodal Inputs

COLING 2025main

With the success of 2D diffusion models, 2D AIGC content has already transformed our lives. Recently, this success has been extended to 3D AIGC, with state-of-the-art methods generating textured 3D models from single images or text. However, we argue that current 3D AIGC methods still don’t fully un…

2025

LITE: A Learning-Integrated Topological Explorer for Multi-Floor Indoor Environments

IROS 2025

This work focuses on multi-floor indoor exploration, which remains an open area of research. Compared to traditional methods, recent learning-based explorers have demonstrated significant potential due to their robust environmental learning and modeling capabilities, but most are restricted to 2D en

Cited by 0SourceScholar
2025

LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts

EMNLP 2025

As large language models (LLMs) advance across diverse tasks, the need for comprehensive evaluation beyond single metrics becomes increasingly important.To fully assess LLM intelligence, it is crucial to examine their interactive dynamics and strategic behaviors.We present LLMsPark, a game theory–ba

2025

Putting People in LLMs’ Shoes: Generating Better Answers via Question Rewriter

AAAI 2025technical

Large Language Models (LLMs) have demonstrated significant capabilities, particularly in the domain of question answering (QA). However, their effectiveness in QA is often undermined by the vagueness of user questions. To address this issue, we introduce single-round instance-level prompt optimizat…

2025

ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training

COLING 2025main

Recent approaches for visually-rich document understanding (VrDU) uses manually annotated semantic groups, where a semantic group encompasses all semantically relevant but not obviously grouped words. As OCR tools are unable to automatically identify such grouping, we argue that current VrDU approac…

2024

DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models

NeurIPS 2024poster

Large language models (LLMs) have recently showcased remarkable capabilities, spanning a wide range of tasks and applications, including those in the medical domain. Models like GPT-4 excel in medical question answering but may face challenges in the lack of interpretability when handling complex ta…

2024

MMAD:Multi-modal Movie Audio Description

COLING 2024main

Audio Description (AD) aims to generate narrations of information that is not accessible through unimodal hearing in movies to aid the visually impaired in following film narratives. Current solutions rely heavily on manual work, resulting in high costs and limited scalability. While automatic metho…

2024

MaxQ: Multi-Axis Query for N:M Sparsity Network

CVPR 2024poster

N:M sparsity has received increasing attention due to its remarkable performance and latency trade-off compared with structured and unstructured sparsity. However existing N:M sparsity methods do not differentiate the relative importance of weights among blocks and leave important weights underappre…

2024

SDSTrack: Self-Distillation Symmetric Adapter Learning for Multi-Modal Visual Object Tracking

CVPR 2024poster

Multimodal Visual Object Tracking (VOT) has recently gained significant attention due to its robustness. Early research focused on fully fine-tuning RGB-based trackers which was inefficient and lacked generalized representation due to the scarcity of multimodal data. Therefore recent studies have ut…

2024

∞Bench: Extending Long Context Evaluation Beyond 100K Tokens

ACL 2024long

Processing and reasoning over long contexts is crucial for many practical applications of Large Language Models (LLMs), such as document comprehension and agent construction. Despite recent strides in making LLMs process contexts with more than 100K tokens, there is currently a lack of a standardize…

2023

IAST: Instance Association Relying on Spatio-Temporal Features for Video Instance Segmentation

ICASSP 2023accepted

Most offline video instance segmentation (VIS) methods lack consideration for multi-scale spatio-temporal features, which leads to unstable instance association across frames. To address this problem, we propose IAST that builds Instance Association relying on Spatio-Temporal features for video inst…

Cited by 0SourceScholar
2023

SQA: Strong Guidance Query with Self-Selected Attention for Human-Object Interaction Detection

ICASSP 2023accepted

The attention mechanism in Transformer-based HOI models plays important role in the comprehension of human and object interaction. However, most previous Transformer-based models ignore the guidance on the query and attention, which leads to a poor understanding of interaction behaviour. In this pap…

Cited by 0SourceScholar
2023

VLKP:Video Instance Segmentation with Visual-Linguistic Knowledge Prompts

ICASSP 2023accepted

Most video instance segmentation(VIS) models only focused on visual knowledge and ignored intrinsic linguistic knowledge. Based on the observation that incorporating linguistic knowledge can significantly improve the model’s contextual understanding of the video, in this paper, we present a Video In…

Cited by 0SourceScholar