← Search

Qin Zhang

45 accepted papers

2026

ChartR: Evaluating Reasoning Accuracy and Robustness in Chart Question Answering

CVPR 2026

Chart Question Answering (CQA) benchmarks are critical for evaluating Multimodal Large Language Models (MLLMs) on visual data reasoning. Existing benchmarks focus mainly on final-answer correctness, ignoring intermediate reasoning steps and the propagation of errors in multi-step processes. To addre

Cited by 0SourceScholar
2026

G-Merging: Graph Models Merging for Parameter-Efficient Multi-Task Knowledge Consolidation

ICLR 2026poster

The pretrain-finetuning paradigm has achieved notable success in graph learning. Moreover, merging models fine-tuned on different tasks to enable a parameter-efficient model with multi-task capabilities is gaining increasing attention for its practicality. However, existing model merging methods, su…

Cited by 0SourcecodeScholar
2026

LR-AdaInSeg:Adaptive Instance Segmentation of Incomplete 3D Scenes Driven by Low-Rank Networks

AAAI 2026technical

3D full-scene segmentation technology has demonstrated great potential driven by large models, but it often faces challenges of incomplete scenes and identification of invisible classes in practical applications. To address this, we propose the LR-AdaInSeg method, which significantly enhances the mo

Cited by 0SourcePDFScholar
2026

LangRef3DGS: Natural Language-Guided 3D Referential Segmentation from Partial Observations via 3D Gaussian Splatting

CVPR 2026

Language-guided 3D segmentation is crucial for linking 3D perception with semantic understanding, yet it remains vulnerable to the sparse and occluded views common in real-world RGB-D data. To overcome this, we present a real-time framework that leverages 3D Gaussian Splatting (3DGS) to build a sema

Cited by 0SourcecodeScholar
2026

Learn to Merge: Meta-Learning for Adaptive Multi-Task Model Merging

ICML 2026poster

Model merging in the pretrain-finetune paradigm has proven effective by combining multiple finetuned models into one with multi-task capabilities. However, existing methods rely on fix or manually tuned merging coefficients, making the unified model sensitive to the initial merging strategy and subo…

Cited by 0SourceScholar
2026

VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models

CVPR 2026

Advances in large reasoning models have shown strong performance on complex reasoning tasks by scaling test-time compute through extended inference-time thinking. However, recent studies observe that in vision-dependent tasks, extended textual reasoning at inference time can often degrade performanc

Cited by 0SourceScholar
2025

Breaking the $\log(1/\Delta_2)$ Barrier: Better Batched Best Arm Identification with Adaptive Grids

ICLR 2025poster

We investigate the problem of batched best arm identification in multi-armed bandits, where we want to find the best arm from a set of $n$ arms while minimizing both the number of samples and batches. We introduce an algorithm that achieves near-optimal sample complexity and features an instance-sen…

Cited by 0SourcePDFScholar
2025

CSC-PA: Cross-image Semantic Correlation via Prototype Attentions for Single-network Semi-supervised Breast Tumor Segmentation

CVPR 2025poster

Accurate automatic breast ultrasound (BUS) image segmentation is essential for early breast cancer screening and diagnosis. However, it remains challenging owing to (1) breast lesions of various scale and shape, (2) ambiguous boundaries caused by speckle noise and artifacts, and (3) the scarcity of…

2025

Faster Approximation Algorithms for k-Center via Data Reduction

ICML 2025poster

We study efficient algorithms for the Euclidean $k$-Center problem, focusing on the regime of large $k$. We take the approach of data reduction by considering $\alpha$-coreset, which is a small subset $S$ of the dataset $P$ such that any $\beta$-approximation on $S$ is an $(\alpha + \beta)$-approxim…

Cited by 0SourcePDFScholar
2025

Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels

CVPR 2025poster

This work presents a simple yet effective workflow for automatically scaling instruction-following data to elicit pixel-level grounding capabilities of VLMs under complex instructions. In particular, we address five critical real-world challenges in text-instruction-based grounding: hallucinated ref…

Cited by 0SourcePDFScholar
2025

Model Diagnosis and Correction via Linguistic and Implicit Attribute Editing

CVPR 2025poster

How can we troubleshoot a deep visual model, i.e., understand why it makes certain mistakes and further take action to correct its behavior? We design a Model Diagnosis and Correction system (MDC), an automated framework that analyzes the pattern of errors, proposes candidate causes of attributes, c…

Cited by 0SourcePDFScholar
2025

Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing

CVPR 2025poster

Developing a face anti-spoofing model that meets the security requirements of clients worldwide is challenging due to the domain gap between training datasets and the diverse end-user test data. Moreover, for security and privacy reasons, it is undesirable for clients to share a large amount of thei…

Cited by 0SourcePDFScholar
2025

TGLsta: Low-resource Textual Graph Learning with Semantic and Topological Awareness via LLMs

AAAI 2025technical

Textual Graphs (TGs) present a graph-based representation of textual data and find wide applications in real-world scenarios, such as citation networks, knowledge graphs, and social networks. While the traditional "pre-train, fine-tune" framework effectively addresses tasks requiring abundant labele…

Cited by 0SourcePDFScholar
2025

Uncertainty-Aware Contrastive Learning with Hard Negative Sampling for Code Search Tasks

AAAI 2025technical

Code search is a highly required technique for software development. In recent years, the rapid development of transformer-based language models has made it increasingly more popular to adapt a pre-trained language model to a code search task, where contrastive learning is typically adopted to seman…

Cited by 0SourcePDFScholar
2025

Understanding Large Language Model Vulnerabilities to Social Bias Attacks

ACL 2025long

Large Language Models (LLMs) have become foundational in human-computer interaction, demonstrating remarkable linguistic capabilities across various tasks. However, there is a growing concern about their potential to perpetuate social biases present in their training data. In this paper, we comprehe…

Cited by 0SourcePDFScholar
2024

CHAmbi: A New Benchmark on Chinese Ambiguity Challenges for Large Language Models

EMNLP 2024finding

Ambiguity is an inherent feature of language, whose management is crucial for effective communication and collaboration. This is particularly true for Chinese, a language with extensive lexical-morphemic ambiguity. Despite the wide use of large language models (LLMs) in numerous domains and their gr…

2024

CONC: Complex-noise-resistant Open-set Node Classification with Adaptive Noise Detection

IJCAI 2024poster

As a popular task in graph learning, node classification seeks to assign labels to nodes, taking into account both their features and connections. However, an important challenge for its application in real-world scenarios is the presence of newly-emerged out-of-distribution samples and noisy sample…

Cited by 1SourcePDFScholar
2024

EGonc : Energy-based Open-Set Node Classification with substitute Unknowns

NeurIPS 2024poster

Open-set Classification (OSC) is a critical requirement for safely deploying machine learning models in the open world, which aims to classify samples from known classes and reject samples from out-of-distribution (OOD). Existing methods exploit the feature space of trained network and attempt at e…

Cited by 0SourcePDFScholar
2024

Learning for Transductive Threshold Calibration in Open-World Recognition

CVPR 2024poster

In deep metric learning for visual recognition the calibration of distance thresholds is crucial for achieving desired model performance in the true positive rates (TPR) or true negative rates (TNR). However calibrating this thresh- old presents challenges in open-world scenarios where the test clas…

Cited by 0SourcePDFScholar
2024

Open-World Dynamic Prompt and Continual Visual Representation Learning

ECCV 2024poster

"The open world is inherently dynamic, characterized by ever-evolving concepts and distributions. Continual learning (CL) in this dynamic open-world environment presents a significant challenge in effectively generalizing to unseen test-time classes. To address this challenge, we introduce a new pra…

Cited by 2SourcePDFScholar
2024

PH-Net: Semi-Supervised Breast Lesion Segmentation via Patch-wise Hardness

CVPR 2024poster

We present a novel semi-supervised framework for breast ultrasound (BUS) image segmentation which is a very challenging task owing to (1) large scale and shape variations of breast lesions and (2) extremely ambiguous boundaries caused by massive speckle noise and artifacts in BUS images. While exist…

2024

ROG_PL: Robust Open-Set Graph Learning via Region-Based Prototype Learning

AAAI 2024technical

Open-set graph learning is a practical task that aims to classify the known class nodes and to identify unknown class samples as unknowns. Conventional node classification methods usually perform unsatisfactorily in open-set scenarios due to the complex data they encounter, such as out-of-distributi…

Cited by 2SourcePDFScholar
2024

Threshold-Consistent Margin Loss for Open-World Deep Metric Learning

ICLR 2024poster

Existing losses used in deep metric learning (DML) for image retrieval often lead to highly non-uniform intra-class and inter-class representation structures across test classes and data distributions. When combined with the common practice of using a fixed threshold to declare a match, this gives r…

Cited by 4SourcePDFScholar
2023

A Survey for Efficient Open Domain Question Answering

ACL 2023long

Open domain question answering (ODQA) is a longstanding task aimed at answering factual questions from a large knowledge corpus without any explicit evidence in natural language processing (NLP). Recent works have predominantly focused on improving the answering accuracy and have achieved promising…

2023

Design, Modeling, and Control of a Low-Cost and Rapid Response Soft-Growing Manipulator for Orchard Operations

IROS 2023poster

Tree fruit growers around the world are facing labor shortages for critical operations, including harvest and pruning. There is a great interest in developing robotic solutions for these labor-intensive tasks, but current efforts have been prohibitively costly, slow, or require a reconfiguration of…

Cited by 3SourceScholar
2023

Evaluating Model-Free Reinforcement Learning toward Safety-Critical Tasks

AAAI 2023technical

Safety comes first in many real-world applications involving autonomous agents. Despite a large number of reinforcement learning (RL) methods focusing on safety-critical tasks, there is still a lack of high-quality evaluation of those algorithms that adheres to safety constraints at each decision st…

Cited by 31SourcePDFScholar
2023

G2Pxy: Generative Open-Set Node Classification on Graphs with Proxy Unknowns

IJCAI 2023poster

Node classification is the task of predicting the labels of unlabeled nodes in a graph. State-of-the-art methods based on graph neural networks achieve excellent performance when all labels are available during training. But in real-life, models are of ten applied on data with new classes, which…

2022

Deep Unsupervised Hashing with Latent Semantic Components

AAAI 2022technical

Deep unsupervised hashing has been appreciated in the regime of image retrieval. However, most prior arts failed to detect the semantic components and their relationships behind the images, which makes them lack discriminative power. To make up the defect, we propose a novel Deep Semantic Component…

Cited by 26SourcePDFScholar
2022

Improved Representation Learning For Acoustic Event Classification Using Tree-Structured Ontology

ICASSP 2022accepted

Acoustic events have a hierarchical structure analogous to a tree (or a directed acyclic graph). In this work, we propose a structure-aware semi-supervised learning framework for acoustic event classification (AEC). Our hypothesis is that the audio label structure contains useful information that is…

Cited by 0SourceScholar
2022

Wikitag: Wikipedia-Based Knowledge Embeddings Towards Improved Acoustic Event Classification

ICASSP 2022accepted

Acoustic event classification (AEC) is the task of determining whether certain events occur in an audio clip. Inspired by previous research [1], [2], [3] that embeddings from event labels can be leveraged to facilitate the learning of new detectors with no or limited audio samples, we introduce Wiki…

Cited by 0SourceScholar
2018

A Practical Algorithm for Distributed Clustering and Outlier Detection

NeurIPS 2018poster

We study the classic k-means/median clustering, which are fundamental problems in unsupervised learning, in the setting where data are partitioned across multiple sites, and where we are allowed to discard a small portion of the data by labeling them as outliers. We propose a simple approach based…

Cited by 31SourcePDFScholar
2016

Proof-of-concept of a robotic apple harvester

IROS 2016poster

There are no mechanical harvesters for the fresh market apple industry commercially available. The absence of automated harvesting technology is a critical problem because of rising production costs and increasing uncertainty about future labor availability. This paper presents the preliminary desig…

Cited by 81SourceScholar