← Search

Mingyang Zhang

24 accepted papers

2026

DcSplat: Dual-Constraint Human Gaussian Splatting with Latent Multi-View Consistency

AAAI 2026technical

Human Novel View Synthesis (HNVS) aims to synthesize photorealistic human images from novel viewpoints given observations from known views. Despite significant advances achieved by existing methods such as NeRF, diffusion models, and 3DGS, they still face substantial challenges in achieving stable m

Cited by 0SourcePDFScholar
2026

Distributed Image Compression with Multimodal Side Information at Extremely Low Bitrates

CVPR 2026

Distributed Image Compression (DIC) is crucial for multi-view transmission, especially when operating at extremely low bitrates (< 0.1 bpp). Its core challenge is effectively utilizing side information to achieve high-quality reconstruction under strict bitrate budgets. However, existing DIC approac

Cited by 0SourceScholar
2026

TG-RAG: A Retrieval-Augmented Framework for Reasoning Guidance in Specialized Domains

ICML 2026oral

Enhancing Large Reasoning Models (LRMs) for specialized domains remains a critical challenge. While recent industrial frameworks attempt to encapsulate Standard Operating Procedures into modular "skills" for dynamic retrieval, utilizing them via context engineering often proves insufficient for comp…

Cited by 0SourceScholar
2026

Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training

AAAI 2026technical

Trajectory-Guided image-to-video (I2V) generation aims to synthesize videos that adhere to user-specified motion instructions. Existing methods typically rely on computationally expensive fine-tuning on scarce annotated datasets. Although some zero-shot methods attempt to trajectory control in the l

Cited by 0SourcePDFScholar
2025

AdvDisplay: Adversarial Display Assembled by Thermoelectric Cooler for Fooling Thermal Infrared Detectors

AAAI 2025technical

When the current physical adversarial patches cannot deceive thermal infrared detectors, the existing techniques implement adversarial attacks from scratch, such as digital patch generation, material production, and physical deployment. Besides, it is difficult to finely regulate infrared radiation.…

2025

Channel Merging: Preserving Specialization for Merged Experts

AAAI 2025technical

Lately, the practice of utilizing task-specific fine-tuning has been implemented to improve the performance of large language models (LLM) in subsequent tasks. Through the integration of diverse LLMs, the overall competency of LLMs is significantly boosted. Nevertheless, traditional ensemble methods…

2025

Consistency Trajectory Matching for One-Step Generative Super-Resolution

ICCV 2025poster

Current diffusion-based super-resolution (SR) approaches achieve commendable performance at the cost of high inference overhead. Therefore, distillation techniques are utilized to accelerate the multi-step teacher model into one-step student model. Nevertheless, these methods significantly raise tra…

2025

DRANet: Dual-threshold Guided Reliability Aware Network for Semi-Supervised Image Semantic Segmentation

ICASSP 2025accepted

Different from supervised semantic segmentation task, semi-supervised semantic segmentation (SSSS) aims to alleviate the burden of time-consuming pixel-wise manual labeling. Although existing methods have achieved the promising performance with a small amount of labeled images, they still suffer fro…

Cited by 0SourceScholar
2025

Decoupling Scattering: Pseudo-Label Guided NeRF for Scenes with Scattering Media

AAAI 2025technical

Neural Radiance Fields (NeRF) has been widely used in computer vision and graphics, achieving impressive results in novel view synthesis and multi-view 3D reconstruction. However, despite its excellent performance under ideal conditions, NeRF struggles in challenging environments such as hazy, foggy…

2025

MUCD: Unsupervised Point Cloud Change Detection via Masked Consistency

AAAI 2025technical

3D Change Detection (3DCD) has gradually become another research hotspot after image change detection. Recent works focus on using artificial labels for supervised or weakly-supervised training of siamese networks to segment changed points. However, labeling every points of multi-temporal point clou…

Cited by 0SourcePDFScholar
2025

PointTruss: K-Truss for Point Cloud Registration

NeurIPS 2025poster

Point cloud registration is a fundamental task in 3D computer vision. Recent advances have shown that graph-based methods are effective for outlier rejection in this context. However, existing clique-based methods impose overly strict constraints and are NP-hard, making it difficult to achieve both…

Cited by 0SourceScholar
2024

Bridging the Preference Gap between Retrievers and LLMs

ACL 2024long

Large Language Models (LLMs) have demonstrated superior results across a wide range of tasks, and Retrieval-augmented Generation (RAG) is an effective way to enhance the performance by locating relevant information and placing it into the context window of the LLM. However, the relationship between…

Cited by 30SourcePDFScholar
2024

Enhancing Hyperspectral Images via Diffusion Model and Group-Autoencoder Super-resolution Network

AAAI 2024technical

Existing hyperspectral image (HSI) super-resolution (SR) methods struggle to effectively capture the complex spectral-spatial relationships and low-level details, while diffusion models represent a promising generative model known for their exceptional performance in modeling complex relations and l…

2024

LoRAPrune: Structured Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning

ACL 2024findings

Large Language Models (LLMs), such as LLaMA and T5, have shown exceptional performance across various tasks through fine-tuning. Although low-rank adaption (LoRA) has emerged to cheaply fine-tune these LLMs on downstream tasks, their deployment is still hindered by the vast model scale and computati…

2024

On provable privacy vulnerabilities of graph representations

NeurIPS 2024poster

Graph representation learning (GRL) is critical for extracting insights from complex network structures, but it also raises security concerns due to potential privacy vulnerabilities in these representations. This paper investigates the structural vulnerabilities in graph neural models where sensiti…

Cited by 2SourcePDFScholar
2024

PRewrite: Prompt Rewriting with Reinforcement Learning

ACL 2024short

Prompt engineering is critical for the development of LLM-based applications. However, it is usually done manually in a “trial and error” fashion that can be time consuming, ineffective, and sub-optimal. Even for the prompts which seemingly work well, there is always a lingering question: can the pr…

Cited by 10SourcePDFScholar
2024

PointMC: Multi-instance Point Cloud Registration based on Maximal Cliques

ICML 2024poster

Multi-instance point cloud registration is the problem of estimating multiple rigid transformations between two point clouds. Existing solutions rely on global spatial consistency of ambiguity and the time-consuming clustering of highdimensional correspondence features, making it difficult to handle…

Cited by 1SourcePDFScholar
2024

Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach

EMNLP 2024industry

Retrieval Augmented Generation (RAG) has been a powerful tool for Large Language Models (LLMs) to efficiently process overly lengthy contexts. However, recent LLMs like Gemini-1.5 and GPT-4 show exceptional capabilities to understand long contexts directly. We conduct a comprehensive comparison betw…

Cited by 33SourcePDFScholar
2024

Transfer the Linguistic Representations from TTS to Accent Conversion with Non-Parallel Data

ICASSP 2024accepted

Accent conversion aims to convert the accent of a source speech to a target accent, meanwhile preserving the speaker’s identity. This paper introduces a novel non-autoregressive framework for accent conversion that learns accent-agnostic linguistic representations and employs them to convert the acc…

Cited by 0SourceScholar
2023

A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation From a Single RGB Image

CVPR 2023poster

3D interacting hand pose estimation from a single RGB image is a challenging task, due to serious self-occlusion and inter-occlusion towards hands, confusing similar appearance patterns between 2 hands, ill-posed joint position mapping from 2D to 3D, etc.. To address these, we propose to extend A2J-…

2022

C3P: Cross-Domain Pose Prior Propagation for Weakly Supervised 3D Human Pose Estimation

ECCV 2022poster

"This paper first proposes and solves weakly supervised 3D human pose estimation (HPE) problem in point cloud, via propagating the pose prior within unlabelled RGB-point cloud sequence to 3D domain. Our approach termed C3P does not require any labor-consuming 3D keypoint annotation for training. To…

2022

Visualtts: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over

ICASSP 2022accepted

In this paper, we formulate a novel task to synthesize speech in sync with a silent pre-recorded video, denoted as automatic voice over (AVO). Unlike traditional speech synthesis, AVO seeks to generate not only human-sounding speech, but also perfect lip-speech synchronization. A natural solution to…

Cited by 0SourceScholar
2020

Multi-View Joint Graph Representation Learning for Urban Region Embedding

IJCAI 2020poster

The increasing amount of urban data enable us to investigate urban dynamics, assist urban planning, and eventually, make our cities more livable and sustainable. In this paper, we focus on learning an embedding space from urban data for urban regions. For the first time, we propose a multi-view join…