← Search

HE Zhang

72 accepted papers

2026

Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models

ICLR 2026poster

In this work, we propose aligning pretrained visual encoders to serve as tokenizers for latent diffusion models in image generation. Unlike training a variational autoencoder (VAE) from scratch, which primarily emphasizes low-level details, our approach leverages the rich semantic structure of found…

Cited by 0SourceScholar
2026

Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

ICML 2026poster

Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation enc…

Cited by 0SourceScholar
2026

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

ICLR 2026oral

Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified frameworks, video generation and editing remain fragmented due…

Cited by 0SourcecodeScholar
2026

From Subtle to Significant: Prompt-Driven Self-Improving Optimization in Test-Time Graph OOD Detection

AAAI 2026technical

Graph Out-of-Distribution (OOD) detection aims to identify whether a test graph deviates from the distribution of graphs observed during training, which is critical for ensuring the reliability of Graph Neural Networks (GNNs) when deployed in open-world scenarios. Recent advances in graph OOD detect

Cited by 0SourcePDFScholar
2026

HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation

CVPR 2026

Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanced methods such as BAGEL and LMFusion follow the Mixture-of-Transformers (MoT) paradigm, adopting a symmetric design tha

Cited by 0SourceScholar
2026

Imprint of the Forgotten: Stealthy Membership Inference in Unlearned Graph Neural Networks

AAAI 2026technical

Graphs effectively model interactions in real-world applications such as social and trade networks, where Graph Neural Networks (GNNs) excel at tasks such as link prediction to enhance user experiences. Despite these benefits, users raise privacy concerns as user data can be exploited to improve GNN

Cited by 0SourcePDFScholar
2026

MetricHMSR: Metric Human Mesh and Scene Recovery from Monocular Images

CVPR 2026

We introduce MetricHMSR (Metric Human Mesh and Scene Recovery), a novel approach for metric human mesh and scene recovery from monocular images. Due to unrealistic assumptions in the camera model and inherent challenges in metric perception, existing approaches struggle to achieve human pose and met

Cited by 0SourcecodeScholar
2026

Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action

ICML 2026poster

We introduce SOMA, the Spatial Memory framework for Out-of-Vision Manipulation in Vision-Language-Action (VLA) models. Most existing VLAs implicitly assume that task-relevant objects are always visible, leading to brittle and reactive behaviors when targets fall outside the camera’s field of view. S…

Cited by 0SourceScholar
2026

TokenLight: Precise Lighting Control in Images using Attribute Tokens

CVPR 2026

This paper presents a method for image relighting that enables precise and continuous control over multiple illumination attributes in a photograph. We formulate relighting as a conditional image generation task and introduce attribute tokens to encode distinct lighting factors such as intensity, co

Cited by 0SourceScholar
2026

Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

RSS 2026poster

Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely on VLM backbones optimized for Visual Question Answering (VQA…

Cited by 0SourceScholar
2025

Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction

ICCV 2025poster

Existing feedforward image-to-3D methods mainly rely on 2D multi-view diffusion models that cannot guarantee 3D consistency. These methods easily collapse when changing the prompt view direction and mainly handle object-centric cases. In this paper, we propose a novel single-stage 3D diffusion model…

2025

BiMark: Unbiased Multilayer Watermarking for Large Language Models

ICML 2025poster

Recent advances in Large Language Models (LLMs) have raised urgent concerns about LLM-generated text authenticity, prompting regulatory demands for reliable identification mechanisms. Although watermarking offers a promising solution, existing approaches struggle to simultaneously achieve three cri…

Cited by 0SourcePDFScholar
2025

Comprehensive Relighting: Generalizable and Consistent Monocular Human Relighting and Harmonization

CVPR 2025poster

This paper introduces Comprehensive Relighting, the first all-in-one approach that can both control and harmonize the lighting from an image or video of humans with arbitrary body parts from any scene. Building such a generalizable model is extremely challenging due to the lack of dataset, restricti…

Cited by 0SourcePDFScholar
2025

DIVE: Taming DINO for Subject-Driven Video Editing

ICCV 2025poster

Building on the success of diffusion models in image generation and editing, video editing has recently gained substantial attention. However, maintaining temporal consistency and motion alignment still remains challenging. To address these issues, this paper proposes DINO-guided Video Editing (DIVE…

2025

Hybrid Contrastive Learning Decoupling Speech Emotion Recognition

ICASSP 2025accepted

Speech signals contain rich information, such as textual content, emotion, and speaker identity. To extract these features more efficiently, researchers are investigating joint training across multiple tasks, like Speech Emotion Recognition (SER) and Speaker Verification (SV), aiming to improve perf…

Cited by 0SourceScholar
2025

Learning to Hang Crumpled Garments with Confidence-Guided Grasping and Active Perception

IROS 2025

Accurately recognizing the structural regions of targeted objects is crucial for successful manipulation. In this study, we concentrate on the task of hanging crumpled garments on a rack, a common scenario in household environments. This context presents two primary challenges: (1) perceiving and gr

Cited by 0SourceScholar
2025

MotionPRO: Exploring the Role of Pressure in Human MoCap and Beyond

CVPR 2025highlight

Existing human Motion Capture (MoCap) methods mostly focus on the visual similarity while neglecting the physical plausibility. As a result, downstream tasks such as driving virtual human in 3D scene or humanoid robots in real world suffer from issues such as timing drift and jitter, spatial problem…

2025

Multitwine: Multi-Object Compositing with Text and Layout Control

CVPR 2025highlight

We introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects within a scene, capturing a range of interactions from simple positional relations (e.g., next to, in front of) to complex…

Cited by 1SourcePDFScholar
2025

OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions

NeurIPS 2025poster

Existing feedforward subject-driven video customization methods mainly study single-subject scenarios due to the difficulty of constructing multi-subject training data pairs. Another challenging problem that how to use the signals such as depth, mask, camera, and text prompts to control and edit the…

Cited by 0SourcecodeScholar
2025

Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment

ICLR 2025poster

Personalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is…

Cited by 1SourcePDFScholar
2025

RelitLRM: Generative Relightable Radiance for Large Reconstruction Models

ICLR 2025spotlight

We propose RelitLRM, a Large Reconstruction Model (LRM) for generating high-quality Gaussian splatting representations of 3D objects under novel illuminations from sparse (4-8) posed images captured under unknown static lighting. Unlike prior inverse rendering methods requiring dense captures and sl…

2025

Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion

ACL 2025long

This paper addresses the critical need for democratizing large language models (LLM) in the Arab world, a region that has seen slower progress in developing models comparable to state-of-the-art offerings like GPT-4 or GPT-3.5, due to a predominant focus on mainstream languages (e.g., English and Ch…

2025

Self-Prompting Driven SAM2 for 3D Medical Image Segmentation

ICASSP 2025accepted

The latest advancement in large foundational model, SAM2, has demonstrated significant potential in 3D medical image segmentation due to their capability to effectively segment video streams. However, its application in medical image segmentation presents challenges, requiring extensive training on…

Cited by 0SourceScholar
2025

TDMER: A Task-Driven Method for Multimodal Emotion Recognition

ICASSP 2025accepted

In multimodal emotion recognition, disentangled representation learning method effectively address the inherent heterogeneity among modalities. To facilitate the flexible integration of enhanced disentangled features into multimodal emotional features, we propose a task-driven multimodal emotion rec…

Cited by 0SourceScholar
2025

Text2Relight: Creative Portrait Relighting with Text Guidance

AAAI 2025technical

We present a lighting-aware image editing pipeline that, given a portrait image and a text prompt, performs single image relighting. Our model modifies the lighting and color of both the foreground and background to align with the provided text description. The unbounded nature in creativeness of a…

Cited by 1SourcePDFScholar
2025

TransPixeler: Advancing Text-to-Video Generation with Transparency

CVPR 2025poster

Text-to-video generative models have made significant strides, enabling diverse applications in entertainment, advertising, and education. However, generating RGBA video, which includes alpha channels for transparency, remains a challenge due to limited datasets and the difficulty of adapting existi…

2025

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

CVPR 2025highlight

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation…

2024

COMPOSE: Comprehensive Portrait Shadow Editing

ECCV 2024poster

"Existing portrait relighting methods struggle with precise control over facial shadows, particularly when faced with challenges such as handling hard shadows from directional light sources or adjusting shadows while remaining in harmony with existing lighting conditions. In many situations, complet…

Cited by 3SourcePDFScholar
2024

GOODAT: Towards Test-Time Graph Out-of-Distribution Detection

AAAI 2024technical

Graph neural networks (GNNs) have found widespread application in modeling graph data across diverse domains. While GNNs excel in scenarios where the testing data shares the distribution of their training counterparts (in distribution, ID), they often exhibit incorrect predictions when confronted wi…

2024

Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single Image

CVPR 2024poster

At the core of portrait photography is the search for ideal lighting and viewpoint. The process often requires advanced knowledge in photography and an elaborate studio setup. In this work we propose Holo-Relighting a volumetric relighting method that is capable of synthesizing novel viewpoints and…

Cited by 12SourcePDFScholar
2024

IMPRINT: Generative Object Compositing by Learning Identity-Preserving Representation

CVPR 2024poster

Generative object compositing emerges as a promising new avenue for compositional image editing. However the requirement of object identity preservation poses a significant challenge limiting practical usage of most existing methods. In response this paper introduces IMPRINT a novel diffusion-based…

Cited by 29SourcePDFScholar
2024

Infusing Self-Consistency into Density Functional Theory Hamiltonian Prediction via Deep Equilibrium Models

NeurIPS 2024poster

In this study, we introduce a unified neural network architecture, the Deep Equilibrium Density Functional Theory Hamiltonian (DEQH) model, which incorporates Deep Equilibrium Models (DEQs) for predicting Density Functional Theory (DFT) Hamiltonians. The DEQH model inherently captures the self-consi…

2024

Learning Highly Dynamic Behaviors for Quadrupedal Robots

ICRA 2024poster

Learning highly dynamic behaviors for robots has been a longstanding challenge. Traditional approaches have demonstrated robust locomotion, but the exhibited behaviors lack diversity and agility. They employ approximate models, which lead to compromises in performance. Data-driven approaches have be…

Cited by 5SourceScholar
2024

MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors

CVPR 2024poster

Foot contact is an important cue for human motion capture understanding and generation. Existing datasets tend to annotate dense foot contact using visual matching with thresholding or incorporating pressure signals. However these approaches either suffer from low accuracy or are only designed for s…

2024

Relightful Harmonization: Lighting-aware Portrait Background Replacement

CVPR 2024poster

Portrait harmonization aims to composite a subject into a new background adjusting its lighting and color to ensure harmony with the background scene. Existing harmonization techniques often only focus on adjusting the global color and brightness of the foreground and ignore crucial illumination cue…

Cited by 15SourcePDFScholar
2024

Self-Consistency Training for Density-Functional-Theory Hamiltonian Prediction

ICML 2024poster

Predicting the mean-field Hamiltonian matrix in density functional theory is a fundamental formulation to leverage machine learning for solving molecular science problems. Yet, its applicability is limited by insufficient labeled data for training. In this work, we highlight that Hamiltonian predict…

Cited by 5SourcePDFScholar
2024

SumSurvey: An Abstractive Dataset of Scientific Survey Papers for Long Document Summarization

ACL 2024findings

With the popularity of large language models (LLMs) and their ability to handle longer input documents, there is a growing need for high-quality long document summarization datasets. Although many models already support 16k input, current lengths of summarization datasets are inadequate, and salient…

2024

SwapAnything: Enabling Arbitrary Object Swapping in Personalized Image Editing

ECCV 2024poster

"Effective editing of personal content holds a pivotal role in enabling individuals to express their creativity, weaving captivating narratives within their visual stories, and elevate the overall quality and impact of their visual content. Therefore, in this work, we introduce , a novel framework t…

Cited by 16SourcePDFScholar
2023

Demystifying Uneven Vulnerability of Link Stealing Attacks against Graph Neural Networks

ICML 2023poster

While graph neural networks (GNNs) dominate the state-of-the-art for exploring graphs in real-world applications, they have been shown to be vulnerable to a growing number of privacy attacks. For instance, link stealing is a well-known membership inference attack (MIA) on edges that infers the prese…

Cited by 26SourcePDFScholar
2023

Finding the Missing-half: Graph Complementary Learning for Homophily-prone and Heterophily-prone Graphs

ICML 2023poster

Real-world graphs generally have only one kind of tendency in their connections. These connections are either homophilic-prone or heterophily-prone. While graphs with homophily-prone edges tend to connect nodes with the same class (i.e., intra-class nodes), heterophily-prone edges tend to build rela…

2023

Interactive Portrait Harmonization

ICLR 2023poster

Current image harmonization methods consider the entire background as the guidance for harmonization. However, this may limit the capability for user to choose any specific object/person in the background to guide the harmonization. To enable flexible interaction between user and harmonization, we i…

Cited by 17SourcePDFScholar
2023

LightPainter: Interactive Portrait Relighting With Freehand Scribble

CVPR 2023poster

Recent portrait relighting methods have achieved realistic results of portrait lighting effects given a desired lighting representation such as an environment map. However, these methods are not intuitive for user interaction and lack precise lighting control. We introduce LightPainter, a scribble-b…

Cited by 14SourcePDFScholar
2023

PHOTOSWAP: Personalized Subject Swapping in Images

NeurIPS 2023poster

In an era where images and visual content dominate our digital landscape, the ability to manipulate and personalize these images has become a necessity. Envision seamlessly substituting a tabby cat lounging on a sunlit window sill in a photograph with your own playful puppy, all while preserving the…

Cited by 36SourcePDFScholar
2023

Perceptual Artifacts Localization for Image Synthesis Tasks

ICCV 2023poster

Recent advancements in deep generative models have facilitated the creation of photo-realistic images across various tasks. However, these generated images often exhibit perceptual artifacts in specific regions, necessitating manual correction. In this study, we present a comprehensive empirical exa…

Cited by 23PDFcodeScholar
2023

PixHt-Lab: Pixel Height Based Light Effect Generation for Image Compositing

CVPR 2023highlight

Lighting effects such as shadows or reflections are key in making synthetic images realistic and visually appealing. To generate such effects, traditional computer graphics uses a physically-based renderer along with 3D geometry. To compensate for the lack of geometry in 2D Image compositing, recent…

Cited by 22SourcePDFScholar
2023

RD-Suite: A Benchmark for Ranking Distillation

NeurIPS 2023poster

The distillation of ranking models has become an important topic in both academia and industry. In recent years, several advanced methods have been proposed to tackle this problem, often leveraging ranking information from teacher rankers that is absent in traditional classification settings. To dat…

Cited by 7SourcePDFScholar
2023

Semi-Supervised Parametric Real-World Image Harmonization

CVPR 2023poster

Learning-based image harmonization techniques are usually trained to undo synthetic global transformations, applied to a masked foreground in a single ground truth photo. This simulated data does not model many important appearance mismatches (illumination, object boundaries, etc.) between foregroun…

2023

Temporal Knowledge Graph Completion: A Survey

IJCAI 2023poster

Knowledge graph completion (KGC) predicts missing links and is crucial for real-life knowledge graphs, which widely suffer from incompleteness. KGC methods assume a knowledge graph is static, but that may lead to inaccurate prediction results because many facts in the knowledge graphs change over…

Cited by 134SourcePDFScholar
2022

Boosting Robustness of Image Matting With Context Assembling and Strong Data Augmentation

CVPR 2022poster

Deep image matting methods have achieved increasingly better results on benchmarks (e.g., Composition-1k/alphamatting.com). However, the robustness, including robustness to trimaps and generalization to images from different domains, is still under-explored. Although some works propose to either ref…

Cited by 38PDFScholar
2022

Controllable Shadow Generation Using Pixel Height Maps

ECCV 2022poster

"Shadows are essential for realistic image compositing. Physics based shadow rendering methods require 3D geometries, which are not always available. Deep learning-based shadow synthesis methods learn a mapping from the light information to an object’s shadow without explicitly modeling the shadow g…

Cited by 30SourcePDFScholar
2022

DoubleField: Bridging the Neural Surface and Radiance Fields for High-Fidelity Human Reconstruction and Rendering

CVPR 2022poster

We introduce DoubleField, a novel framework combining the merits of both surface field and radiance field for high-fidelity human reconstruction and rendering. Within DoubleField, the surface field and radiance field are associated together by a shared feature embedding and a surface-guided sampling…

Cited by 186PDFScholar
2022

How Far are We from Robust Long Abstractive Summarization?

EMNLP 2022main

Abstractive summarization has made tremendous progress in recent years. In this work, we perform fine-grained human annotations to evaluate long document abstractive summarization systems (i.e., models and metrics) with the aim of implementing them to generate reliable summaries. For long document a…

2022

Lite Vision Transformer With Enhanced Self-Attention

CVPR 2022poster

Despite the impressive representation capacity of vision transformer models, current light-weight vision transformer models still suffer from inconsistent and incorrect dense predictions at local regions. We suspect that the power of their self-attention mechanism is limited in shallower and thinner…

Cited by 151PDFcodeScholar
2022

Needle Tip Pose Estimation for Ultrasound-Guided Steerable Flexible Needle With a Complicated Trajectory in Soft Tissue

RA-L 2022

Visualization of the surgical needle is critical for an image guided insertion. However, it is difficult to obtain a 3D tip pose of a steerable flexible needle under 2D US images. In this letter, an image processing method is used to extract the needle shaft radial cross-sectional centroid coordinat

Cited by 8SourceScholar
2022

SE(3) Equivariant Graph Neural Networks with Complete Local Frames

ICML 2022spotlight

Group equivariance (e.g. SE(3) equivariance) is a critical physical symmetry in science, from classical and quantum physics to computational biology. It enables robust and accurate prediction under arbitrary reference transformations. In light of this, great efforts have been put on encoding this sy…

2022

Vision Assisted Control of Lower Extremity Exoskeleton for Obstacle Avoidance With Dynamic Constraint Based Piecewise Nonlinear MPC

RA-L 2022

This article proposed a humanoid obstacle passability strategy (OPS) to enhance human-exoskeleton integrated system to cross over multi-obstacle in sagittal plane. A hybrid bounding box integrated with closeness regression in L-shape section and synchronous convergence in convex hull search (HBB-LC)

Cited by 7SourceScholar
2021

Co-evolution Transformer for Protein Contact Prediction

NeurIPS 2021poster

Proteins are the main machinery of life and protein functions are largely determined by their 3D structures. The measurement of the pairwise proximity between amino acids of a protein, known as inter-residue contact map, well characterizes the structural information of a protein. Protein contact pre…

2021

Mask Guided Matting via Progressive Refinement Network

CVPR 2021poster

We propose Mask Guided (MG) Matting, a robust matting framework that takes a general coarse mask as guidance. MG Matting leverages a network (PRN) design which encourages the matting model to provide self-guidance to progressively refine the uncertain regions through the decoding process. A series o…

Cited by 153PDFcodeScholar
2021

Progressive Co-Teaching for Ambiguous Speech Emotion Recognition

ICASSP 2021accepted

Speech emotion recognition is a challenging task due to the ambiguity of emotion, which makes it difficult to learn the features of emotion data using machine learning algorithms. However, previous studies conventionally ignore the ambiguity of emotion and treat the emotion data as the same difficul…

Cited by 0SourceScholar
2021

SSH: A Self-Supervised Framework for Image Harmonization

ICCV 2021poster

Image harmonization aims to improve the quality of image compositing by matching the "appearance"" (e.g., color tone, brightness and contrast) between foreground and background images. However, collecting large-scale annotated datasets for this task requires complex professional retouching. Instead,…

Cited by 96PDFcodeScholar
2020

The VCU-RVI Benchmark: Evaluating Visual Inertial Odometry for Indoor Navigation Applications with an RGB-D Camera

IROS 2020poster

This paper presents VCU-RVI, a new visual inertial odometry (VIO) benchmark with a set of diverse data sequences in different indoor scenarios. The benchmark was captured using an Structure Core (SC) sensor, consisting of an RGB-D camera and an IMU. It provides aligned color and depth images with 64…

Cited by 16SourceScholar