← Search

Ke ZHANG

47 accepted papers

2026

DialogueVPR: Towards Conversational Visual Place Recognition

CVPR 2026

Inspired by how humans communicate spatial information, language-guided geo-localization has gained significant traction for its intuitive and practical value. Despite this progress, most methods still rely on a static, one-shot retrieval paradigm, which fails to handle the ambiguity and incompleten

Cited by 0SourcecodeScholar
2026

FreeViS: Training-free Video Stylization with Inconsistent References

ICLR 2026poster

Video stylization plays a key role in content creation, but it remains a challenging problem. Naïvely applying image stylization frame-by-frame hurts temporal consistency and reduces style richness. Alternatively, training a dedicated video stylization model typically requires paired video data and…

Cited by 0SourcecodeScholar
2026

GRASP: Awakening Latent Spatial Reasoning in LVLMs via Training-free Geometric Rectification

ICML 2026poster

Large Vision-Language Models (LVLMs) exhibit remarkable general capabilities but struggle significantly with spatial reasoning tasks. In this paper, we uncover a critical representation-output misalignment via linear probing: LVLMs correctly encode spatial features internally, but generate incorrect…

Cited by 0SourceScholar
2026

ProCrop: Learning Aesthetic Image Cropping from Professional Compositions

AAAI 2026technical

Image cropping is crucial for enhancing the visual appeal and narrative impact of photographs, yet existing rule-based and data-driven approaches often lack diversity or require annotated training data. We introduce ProCrop, a retrieval-based method that leverages professional photography to guide c

Cited by 0SourcePDFScholar
2026

Revisiting Model Stitching In the Foundation Model Era

CVPR 2026

Model stitching, connecting early layers of one model (source) to later layers of another (target) via a light stitch layer, has served as a probe of representational compatibility. Prior work finds that models trained on the same dataset remain stitchable (negligible accuracy drop) despite differen

Cited by 0SourceScholar
2026

Towards Understanding Generalization of Federated Adversarial Learning: Perspective of Algorithmic Stability

ICML 2026poster

Federated Adversarial Learning (FAL) enhances model robustness by integrating adversarial training into the federated learning framework. Despite recent advances proposing efficient FAL algorithms, existing work has mainly focused on convergence properties, with limited understanding of their genera…

Cited by 0SourceScholar
2025

An Abnormal Audio Generation Method for Fault Diagnosis of Power Transformers

ICASSP 2025accepted

Existing deep learning-based models can achieve a prompt diagnosis of operational anomalies by analyzing the audios emitted from power transformers. However, the practical abnormal data are insufficient for model training, resulting in limited diagnostic performance. To address this problem, we prop…

Cited by 0SourceScholar
2025

Beyond Logits: Aligning Feature Dynamics for Effective Knowledge Distillation

ACL 2025long

Knowledge distillation (KD) compresses large language models (LLMs), known as teacher models, into lightweight versions called student models, enabling efficient inference and downstream applications. However, prevailing approaches accomplish this by predominantly focusing on matching the final outp…

2025

Details Matter for Indoor Open-vocabulary 3D Instance Segmentation

ICCV 2025poster

Unlike closed-vocabulary 3D instance segmentation that is often trained end-to-end, open-vocabulary 3D instance segmentation (OV-3DIS) often leverages vision-language models (VLMs) to generate 3D instance proposals and classify them. While various concepts have been proposed from existing research,…

Cited by 0SourcePDFScholar
2025

Drama: Mamba-Enabled Model-Based Reinforcement Learning Is Sample and Parameter Efficient

ICLR 2025poster

Model-based reinforcement learning (RL) offers a solution to the data inefficiency that plagues most model-free RL algorithms. However, learning a robust world model often requires complex and deep architectures, which are computationally expensive and challenging to train. Within the world model, s…

2025

Fine-grained Adaptive Visual Prompt for Generative Medical Visual Question Answering

AAAI 2025technical

Medical Visual Question Answering (MedVQA) serves as an automated medical assistant, capable of answering patient queries and aiding physician diagnoses based on medical images and questions. Recent advancements have shown that incorporating Large Language Models (LLMs) into MedVQA tasks significant…

2025

HarmonySeg: Tubular Structure Segmentation with Deep-Shallow Feature Fusion and Growth-Suppression Balanced Loss

ICCV 2025poster

Accurate segmentation of tubular structures in medical images, such as vessels and airway trees, is crucial for computer-aided diagnosis, radiotherapy, and surgical planning. However, significant challenges exist in algorithm design when faced with diverse sizes, complex topologies, and (often) inco…

Cited by 0SourcePDFScholar
2025

IMDPrompter: Adapting SAM to Image Manipulation Detection by Cross-View Automated Prompt Learning

ICLR 2025poster

Using extensive training data from SA-1B, the Segment Anything Model (SAM) has demonstrated exceptional generalization and zero-shot capabilities, attracting widespread attention in areas such as medical image segmentation and remote sensing image segmentation. However, its performance in the field…

Cited by 0SourcePDFScholar
2025

Less is More: Efficient Image Vectorization with Adaptive Parameterization

CVPR 2025poster

Image vectorization aims to convert raster images to vector ones, allowing for easy scaling and editing.Existing works mainly rely on preset parameters (i.e., a fixed number of paths and control points), ignoring the complexity of the image and posing significant challenges to practical applications…

Cited by 0SourcePDFScholar
2025

Multi-Level Speaker Representation for Target Speaker Extraction

ICASSP 2025accepted

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of speakers may suffer from confusion of speaker identity. In th…

Cited by 0SourceScholar
2025

OpenM3D: Open Vocabulary Multi-view Indoor 3D Object Detection without Human Annotations

ICCV 2025poster

Open-vocabulary (OV) 3D object detection is an emerging field, yet its exploration through image-based methods remains limited compared to 3D point cloud-based methods. We introduce OpenM3D, a novel open-vocabulary multi-view indoor 3D object detector trained without human annotations. In particular…

Cited by 0SourcePDFScholar
2025

Rethinking Pseudo-Label Guided Learning for Weakly Supervised Temporal Action Localization from the Perspective of Noise Correction

AAAI 2025technical

Pseudo-label learning methods have been widely applied in weakly-supervised temporal action localization. Existing works directly utilize weakly-supervised base model to generate instance-level pseudo-labels for training the fully-supervised detection head. We argue that the noise in pseudo-labels w…

Cited by 1SourcePDFScholar
2025

Super Deep Contrastive Information Bottleneck for Multi-modal Clustering

ICML 2025poster

In an era of increasingly diverse information sources, multi-modal clustering (MMC) has become a key technology for processing multi-modal data. It can apply and integrate the feature information and potential relationships of different modalities. Although there is a wealth of research on MMC, due…

Cited by 0SourcePDFScholar
2025

TANDEM: Bi-Level Data Mixture Optimization with Twin Networks

NeurIPS 2025poster

The capabilities of large language models (LLMs) significantly depend on training data drawn from various domains. Optimizing domain-specific mixture ratios can be modeled as a bi-level optimization problem, which we simplify into a single-level penalized form and solve with twin networks: a proxy m…

Cited by 0SourceScholar
2025

Ultrasound-Guided Robotic Blood Drawing and In Vivo Studies on Submillimetre Vessels of Rats

ICRA 2025

Billions of vascular access procedures are performed annually worldwide, serving as a crucial first step in various clinical diagnostic and therapeutic procedures. For pediatric or elderly individuals, whose vessels are small in size (typically 2 to 3 mm in diameter for adults and <1 mm in children)

Cited by 2SourceScholar
2025

Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models

CVPR 2025poster

Recent breakthroughs in Multimodal Large Language Models (MLLMs) have gained significant recognition within the deep learning community, where the fusion of the Video Foundation Models (VFMs) and Large Language Models(LLMs) has proven instrumental in constructing robust video understanding systems,…

Cited by 0SourcePDFScholar
2024

HumVI: A Multilingual Dataset for Detecting Violent Incidents Impacting Humanitarian Aid

EMNLP 2024finding

Humanitarian organizations can enhance their effectiveness by analyzing data to discover trends, gather aggregated insights, manage their security risks, support decision-making, and inform advocacy and funding proposals. However, data about violent incidents with direct impact and relevance for hum…

2024

Nearest Neighbour Score Estimators for Diffusion Generative Models

ICML 2024poster

Score function estimation is the cornerstone of both training and sampling from diffusion generative models. Despite this fact, the most commonly used estimators are either biased neural network approximations or high variance Monte Carlo estimators based on the conditional score. We introduce a nov…

2024

Prompt-based Generation of Natural Language Explanations of Synthetic Lethality for Cancer Drug Discovery

COLING 2024main

Synthetic lethality (SL) offers a promising approach for targeted anti-cancer therapy. Deeply understanding SL gene pair mechanisms is vital for anti-cancer drug discovery. However, current wet-lab and machine learning-based SL prediction methods lack user-friendly and quantitatively evaluable expla…

2024

Typos Correction Training against Misspellings from Text-to-Text Transformers

COLING 2024main

Dense retrieval (DR) has become a mainstream approach to information seeking, where a system is required to return relevant information to a user query. In real-life applications, typoed queries resulting from the users’ mistyping words or phonetic typing errors exist widely in search behaviors. Cur…

2023

AGD: an Auto-switchable Optimizer using Stepwise Gradient Difference for Preconditioning Matrix

NeurIPS 2023poster

Adaptive optimizers, such as Adam, have achieved remarkable success in deep learning. A key component of these optimizers is the so-called preconditioning matrix, providing enhanced gradient information and regulating the step size of each gradient direction. In this paper, we propose a novel approa…

2023

BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics

ACL 2023long

The proliferation of automatic faithfulness metrics for summarization has produced a need for benchmarks to evaluate them. While existing benchmarks measure the correlation with human judgements of faithfulness on model-generated summaries, they are insufficient for diagnosing whether metrics are: 1…

2023

Graph Contrastive Learning with Learnable Graph Augmentation

ICASSP 2023accepted

Graph contrastive learning has gained popularity due to its success in self-supervised graph representation learning. Augmented views in contrastive learning greatly determine the quality of the learned representations. Handcrafted data augmentations in previous work require tedious trial-and- error…

Cited by 0SourceScholar
2023

ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object Detection

ICCV 2023poster

We propose ImGeoNet, a multi-view image-based 3D object detection framework that models a 3D space by an image-induced geometry-aware voxel representation. Unlike previous methods which aggregate 2D features into 3D voxels without considering geometry, ImGeoNet learns to induce geometry from multi-…

Cited by 11PDFcodeScholar
2023

Multi-Head Uncertainty Inference for Adversarial Attack Detection

ICASSP 2023accepted

Deep neural networks (DNNs) are sensitive and susceptible to tiny perturbations by adversarial attacks which cause erroneous predictions. Various methods, including adversarial defense and uncertainty inference (UI), have been developed to overcome adversarial attacks in recent years. In this paper,…

Cited by 0SourceScholar
2023

RH-BrainFS: Regional Heterogeneous Multimodal Brain Networks Fusion Strategy

NeurIPS 2023poster

Multimodal fusion has become an important research technique in neuroscience that completes downstream tasks by extracting complementary information from multiple modalities. Existing multimodal research on brain networks mainly focuses on two modalities, structural connectivity (SC) and functional…

2023

Topgformer: Topological-Based Graph Transformer for Mapping Brain Structural Connectivity to Functional Connectivity

ICASSP 2023accepted

Exploring the mapping between structural connectivity (SC) and functional connectivity (FC) is of essential importance to understanding the working mechanism of the human brain. Traditional methods are difficult to represent the complex relationship of high-order interaction between SC and FC. Recen…

Cited by 0SourceScholar
2023

Topology Uncertainty Modeling For Imbalanced Node Classification on Graphs

ICASSP 2023accepted

Most existing graph neural networks work under a class-balanced assumption, while ignoring class-imbalanced scenarios that widely exist in real-world graphs. Although there are many methods in other fields that can alleviate this issue, they do not consider the special topology of the non-Euclidean…

Cited by 0SourceScholar
2022

An Exploration of Post-Editing Effectiveness in Text Summarization

NAACL 2022long

Automatic summarization methods are efficient but can suffer from low quality. In comparison, manual summarization is expensive but produces higher quality. Can humans and AI collaborate to improve summarization performance? In similar text generation tasks (e.g., machine translation), human-AI coll…

2022

CrisisLTLSum: A Benchmark for Local Crisis Event Timeline Extraction and Summarization

EMNLP 2022finding

Social media has increasingly played a key role in emergency response: first responders can use public posts to better react to ongoing crisis events and deploy the necessary resources where they are most needed. Timeline extraction and abstractive summarization are critical technical tasks to lever…

2022

Feature Space Message Passing Network for Medical Image Semantic Segmentation

ICASSP 2022accepted

Accurate semantic segmentation of medical images is of significant importance for subsequent processing and analysis. The encoder-decoder deep learning framework has been widely applied for numerous medical image segmentation tasks. However, most existing approaches are restricted by the limited rec…

Cited by 0SourceScholar
2022

Hierarchical Diffusion Scattering Graph Neural Network

IJCAI 2022poster

Graph neural network (GNN) is popular now to solve the tasks in non-Euclidean space and most of them learn deep embeddings by aggregating the neighboring nodes. However, these methods are prone to some problems such as over-smoothing because of the single-scale perspective field and the nature of lo…

2022

Mapping the Design Space of Human-AI Interaction in Text Summarization

NAACL 2022long

Automatic text summarization systems commonly involve humans for preparing data or evaluating model performance, yet, there lacks a systematic understanding of humans’ roles, experience, and needs when interacting with or being assisted by AI. From a human-centered perspective, we map the design opp…

Cited by 34SourcePDFScholar
2021

Olá, Bonjour, Salve! XFORMAL: A Benchmark for Multilingual Formality Style Transfer

NAACL 2021long

We take the first step towards multilingual style transfer by creating and releasing XFORMAL, a benchmark of multiple formal reformulations of informal text in Brazilian Portuguese, French, and Italian. Results on XFORMAL suggest that state-of-the-art style transfer approaches perform close to simpl…

2021

Secure Deep Graph Generation with Link Differential Privacy

IJCAI 2021poster

Many data mining and analytical tasks rely on the abstraction of networks (graphs) to summarize relational structures among individuals (nodes). Since relational data are often sensitive, we aim to seek effective approaches to generate utility-preserved yet privacy-protected structured data. In thi…

Cited by 47SourcePDFScholar
2021

Subgraph Federated Learning with Missing Neighbor Generation

NeurIPS 2021spotlight

Graphs have been widely used in data mining and machine learning due to their unique representation of real-world objects and their interactions. As graphs are getting bigger and bigger nowadays, it is common to see their subgraphs separately collected and stored in multiple local systems. Therefore…

2016

Summary Transfer: Exemplar-Based Subset Selection for Video Summarization

CVPR 2016poster

Video summarization has unprecedented importance to help us digest, browse, and search today's ever-growing video collections. We propose a novel subset selection technique that leverages supervision in the form of human-created summaries to perform automatic keyframe-based video summarization. The…

Cited by 271PDFScholar