← Search

Qi Jia

31 accepted papers

2026

ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art

ICLR 2026poster

Perceiving visual semantics embedded within consecutive characters is a crucial yet under-explored capability for both Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs). In this work, we select ASCII art as a representative artifact. It depicts concepts through careful arran…

Cited by 0SourcecodeScholar
2026

Can VLMs Diagnose and Recover from VLA Manipulation Faults?

ICML 2026poster

Existing VLA models frequently fail in robotic manipulation tasks, with poorly structured fault types that often require expert diagnosis.While VLMs offer strong explanatory capabilities, their effectiveness in assisting VLAs is limited by their unclear role in diagnostics and inadequate collaborati…

Cited by 0SourceScholar
2026

Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm Detection

AAAI 2026technical

Multimodal sarcasm detection is a complex task that requires distinguishing subtle complementary signals across modalities while filtering out irrelevant information. Many advanced methods rely on learning shortcuts from datasets rather than extracting intended sarcasm-related features. However, our

Cited by 0SourcePDFScholar
2026

Dual-Level Hypergraph Generation for Addressing Feature Scarcity in Whole-Slide Image Classification

CVPR 2026

Lymph node metastasis diagnosis in pathological images is a highly challenging four-class classification task, comprising macrometastasis, micrometastasis, isolated tumor cells (ITC), and negative lesions.Unlike conventional classification settings, this four-class scenario simultaneously suffers fr

Cited by 0SourcecodeScholar
2026

Improve MLLM Benchmark Efficiency through Interview

ICASSP 2026poster

The rapid development of Multimodal Large Language Models (MLLM) has led to a wide range of MLLM applications, and a number of benchmark datasets have sprung up in order to assess MLLM abilities. However, full-coverage Q&A testing on large-scale data is resource-intensive and time-consuming. To addr…

Cited by 0SourcePDFScholar
2026

RSOD: Reliability-Guided Sonar Image Object Detection with Extremely Limited Labels

AAAI 2026technical

Object detection in sonar images is a key technology in underwater detection systems. Compared to natural images, sonar images contain fewer texture details and are more susceptible to noise, making it difficult for non-experts to distinguish subtle differences between classes. This leads to their i

Cited by 0SourcePDFScholar
2026

SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond

ICML 2026poster

The success of large language models (LLMs) in scientific domains has heightened safety concerns, prompting numerous benchmarks to evaluate their scientific safety. Existing benchmarks often suffer from limited risk coverage and a reliance on subjective evaluation. To address thess problems, we intr…

Cited by 0SourceScholar
2026

Time Series Class-Incremental Learning via Confidence-guided Mask Distillation and Prototype-guided Contrastive Learning

AAAI 2026technical

Class-incremental learning (CIL) has recently gained great attention in the field of time series classification. Existing CIL methods based on knowledge distillation exhibit impressive ability to retain prior knowledge and overcome catastrophic forgetting, however, their effectiveness faces major c

Cited by 0SourcePDFScholar
2025

As Pseudo-Label Free as Possible: Leveraging Adaptive Feature Generation for Sparsely Annotated Object Detection

AAAI 2025technical

Compared to fully supervised object detection, training with sparse annotations typically leads to a decline in performance due to insufficient feature diversity. Existing sparsely annotated object detection (SAOD) methods often rely on pseudo-labeling strategies, but these pseudo-labels tend to int…

2025

DeMAC: Enhancing Multi-Agent Coordination with Dynamic DAG and Manager-Player Feedback

EMNLP 2025

Multi-agent systems (MAS) powered by large language models (LLMs) have shown potential in tackling multifaceted problems through advanced understanding and reasoning. However, they struggle to adapt to evolving task dependencies and to handle uncertainties, such as shifting priorities or unpredictab

Cited by 0SourcePDFScholar
2025

DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation

ICCV 2025poster

Spatio-temporal consistency is a critical topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying viewpoints. Prior research, especially in open-source projects, primarily foc…

2025

Info-Coevolution: An Efficient Framework for Data Model Coevolution

ICML 2025poster

Machine learning relies heavily on data, yet the continuous growth of real-world data poses challenges for efficient dataset construction and training. A fundamental yet unsolved question is: given our current model and data, does a new data (sample/batch) need annotation/learning? Conventional appr…

2025

SimulBench: Evaluating Language Models with Creative Simulation Tasks

NAACL 2025findings

We introduce SimulBench, a benchmark designed to evaluate large language models (LLMs) across a diverse collection of creative simulation tasks, such as acting as a Linux terminal or playing text games with users. While these simulation tasks serve as effective measures of an LLM’s general intellige…

2024

CSCNet: Class-Specified Cascaded Network for Compositional Zero-Shot Learning

ICASSP 2024accepted

Attribute and object (A-O) disentanglement is a fundamental and critical problem for Compositional Zero-shot Learning (CZSL), whose aim is to recognize novel A-O compositions based on foregone knowledge. Existing methods based on disentangled representation learning lose sight of the contextual depe…

Cited by 0SourceScholar
2024

Depth-Guided Dominant Plane Perception for Unsupervised Homography Estimation

ICASSP 2024accepted

Homography describes the mapping relations of the same plane across views. In scenarios with multiple planes, single homography estimation aims to obtain the optimal solution generated by the largest consistent plane to obey the coplanar constraints. However, existing methods typically consider all…

Cited by 0SourceScholar
2024

Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline

NeurIPS 2024poster

Existing video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watching the videos. Induced sentiment of viewers is essential for inferring the public response to videos and has broad appl…

2024

Not Just Object, But State: Compositional Incremental Learning without Forgetting

NeurIPS 2024poster

Most incremental learners excessively prioritize object classes while neglecting various kinds of states (e.g. color and material) attached to the objects. As a result, they are limited in the ability to model state-object compositionality accurately. To remedy this limitation, we propose a novel ta…

2024

Novel Class Discovery for Ultra-Fine-Grained Visual Categorization

CVPR 2024highlight

Ultra-fine-grained visual categorization (Ultra-FGVC) aims at distinguishing highly similar sub-categories within fine-grained objects such as different soybean cultivars. Compared to traditional fine-grained visual categorization Ultra-FGVC encounters more hurdles due to the small inter-class and l…

2024

Sketch-Based 3D Shape Retrieval With Multi-View Fusion Transformer

ICASSP 2024accepted

Sketch-based 3D shape retrieval aims to retrieve similar 3D shapes given a 2D sketch query. Although this task has been studied for years, the inherent cross-modal gap and data imbalance between 2D sketches and 3D shapes remain challenging. To address the problems, we propose a simple and effective…

Cited by 0SourceScholar
2023

Global and Local Mixture Consistency Cumulative Learning for Long-Tailed Visual Recognitions

CVPR 2023poster

In this paper, our goal is to design a simple learning paradigm for long-tail visual recognition, which not only improves the robustness of the feature extractor but also alleviates the bias of the classifier towards head classes while reducing the training skills and overhead. We propose an efficie…

2023

In-sample Curriculum Learning by Sequence Completion for Natural Language Generation

ACL 2023long

Curriculum learning has shown promising improvements in multiple domains by training machine learning models from easy samples to hard ones. Previous works which either design rules or train models for scoring the difficulty highly rely on task-specific expertise, and cannot generalize. Inspired by…

2023

Reducing Sensitivity on Speaker Names for Text Generation from Dialogues

ACL 2023findings

Changing speaker names consistently throughout a dialogue should not affect its meaning and corresponding outputs for text generation from dialogues. However, pre-trained language models, serving as the backbone for dialogue-processing tasks, have shown to be sensitive to nuances. This may result in…

2023

Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language Model

EMNLP 2023long main

Despite tremendous improvements in natural language generation, summarization models still suffer from the unfaithfulness issue. Previous work evaluates faithfulness either using models trained on the other tasks or in-domain synthetic data, or prompting a large model such as ChatGPT. This paper pro…

Cited by 0SourcecodeScholar
2022

Length Control in Abstractive Summarization by Pretraining Information Selection

ACL 2022long

Previous length-controllable summarization models mostly control lengths at the decoding stage, whereas the encoding or the selection of information from the source document is not sensitive to the designed length. They also tend to generate summaries as long as those in the training data. In this p…

2022

Post-Training Dialogue Summarization using Pseudo-Paraphrasing

NAACL 2022findings

Previous dialogue summarization techniques adapt large language models pretrained on the narrative text by injecting dialogue-specific features into the models. These features either require additional knowledge to recognize or make the resulting models harder to tune. To bridge the format gap betwe…

2022

Reference-free Summarization Evaluation via Semantic Correlation and Compression Ratio

NAACL 2022long

A document can be summarized in a number of ways. Reference-based evaluation of summarization has been criticized for its inflexibility. The more sufficient the number of abstracts, the more accurate the evaluation results. However, it is difficult to collect sufficient reference summaries. In this…

2022

Segment, Magnify and Reiterate: Detecting Camouflaged Objects the Hard Way

CVPR 2022poster

It is challenging to accurately detect camouflaged objects from their highly similar surroundings. Existing methods mainly leverage a single-stage detection fashion, while neglecting small objects with low-resolution fine edges requires more operations than the larger ones. To tackle camouflaged obj…

Cited by 207PDFcodeScholar
2021

DDRel: A New Dataset for Interpersonal Relation Classification in Dyadic Dialogues

AAAI 2021technical

Interpersonal language style shifting in dialogues is an interesting and almost instinctive ability of human. Understanding interpersonal relationship from language content is also a crucial step toward further understanding dialogues. Previous work mainly focuses on relation extraction between name…

2021

Leveraging Line-Point Consistence To Preserve Structures for Wide Parallax Image Stitching

CVPR 2021poster

Generating high-quality stitched images with natural structures is a challenging task in computer vision. In this paper, we succeed in preserving both local and global geometric structures for wide parallax images, while reducing artifacts and distortions. A projective invariant, Characteristic Numb…

Cited by 129PDFcodeScholar