← Search

Hanjiang Lai

21 accepted papers

2025

Answering Complex Geographic Questions by Adaptive Reasoning with Visual Context and External Commonsense Knowledge

ACL 2025long

This paper focuses on a new task of answering geographic reasoning questions based on the given image (called GeoVQA). Unlike traditional VQA tasks, GeoVQA asks for details about the image-related culture, landscape, etc. This requires not only the identification of the objects in the image, their p…

Cited by 0SourcePDFScholar
2025

Detecting Emotional Incongruity of Sarcasm by Commonsense Reasoning

COLING 2025main

This paper focuses on sarcasm detection, which aims to identify whether given statements convey criticism, mockery, or other negative sentiment opposite to the literal meaning. To detect sarcasm, humans often require a comprehensive understanding of the semantics in the statement and even resort to…

Cited by 1SourcePDFScholar
2025

Enhancing Few-Shot Out-of-Distribution Detection with Gradient Aligned Context Optimization

ICASSP 2025accepted

Few-shot out-of-distribution (OOD) detection aims to detect OOD images from unseen classes with only a few labeled in-distribution (ID) images. To detect OOD images and classify ID samples, prior methods have been proposed by regarding the background regions of ID samples as the OOD knowledge and pe…

Cited by 0SourceScholar
2025

Generating Commonsense Reasoning Questions with Controllable Complexity through Multi-step Structural Composition

COLING 2025main

This paper studies the task of generating commonsense reasoning questions (QG) with desired difficulty levels. Compared to traditional shallow questions that can be solved by simple term matching, ours are more challenging. Our answering process requires reasoning over multiple contextual and common…

Cited by 1SourcePDFScholar
2025

On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free Approach

CVPR 2025poster

Pre-trained Vision-Language Models (VLMs) like CLIP, have demonstrated strong zero-shot generalization capabilities. Despite their effectiveness on various downstream tasks, they remain vulnerable to adversarial samples. Existing methods fine-tune VLMs to improve their performance via performing adv…

Cited by 0SourcePDFScholar
2024

Hierarchical Topic Modeling via Contrastive Learning and Hyperbolic Embedding

COLING 2024main

Hierarchical topic modeling, which can mine implicit semantics in the corpus and automatically construct topic hierarchical relationships, has received considerable attention recently. However, the current hierarchical topic models are mainly based on Euclidean space, which cannot well retain the im…

2024

MimicDiffusion: Purifying Adversarial Perturbation via Mimicking Clean Diffusion Model

CVPR 2024poster

Deep neural networks (DNNs) are vulnerable to adversarial perturbation where an imperceptible perturbation is added to the image that can fool the DNNs. Diffusion-based adversarial purification uses the diffusion model to generate a clean image against such adversarial attacks. Unfortunately the gen…

2023

Counterfactual Multihop QA: A Cause-Effect Approach for Reducing Disconnected Reasoning

ACL 2023long

Multi-hop QA requires reasoning over multiple supporting facts to answer the question. However, the existing QA models always rely on shortcuts, e.g., providing the true answer by only one fact, rather than multi-hop reasoning, which is referred as disconnected reasoning problem. To alleviate this i…

2023

Deep Hashing With Minimal-Distance-Separated Hash Centers

CVPR 2023poster

Deep hashing is an appealing approach for large-scale image retrieval. Most existing supervised deep hashing methods learn hash functions using pairwise or triple image similarities in randomly sampled mini-batches. They suffer from low training efficiency, insufficient coverage of data distribution…

Cited by 47SourcePDFScholar
2023

From Parse-Execute to Parse-Execute-Refine: Improving Semantic Parser for Complex Question Answering over Knowledge Base

EMNLP 2023long main

Parsing questions into executable logical forms has showed impressive results for knowledge-base question answering (KBQA). However, complex KBQA is a more challenging task that requires to perform complex multi-step reasoning. Recently, a new semantic parser called KoPL has been proposed to explici…

Cited by 0SourceScholar
2022

Towards Better Plasticity-Stability Trade-Off in Incremental Learning: A Simple Linear Connector

CVPR 2022poster

Plasticity-stability dilemma is a main problem for incremental learning, where plasticity is referring to the ability to learn new knowledge, and stability retains the knowledge of previous tasks. Many methods tackle this problem by storing previous samples, while in some applications, training data…

Cited by 71PDFcodeScholar
2022

You Never Stop Dancing: Non-freezing Dance Generation via Bank-constrained Manifold Projection

NeurIPS 2022accept

One of the most overlooked challenges in dance generation is that the auto-regressive frameworks are prone to freezing motions due to noise accumulation. In this paper, we present two modules that can be plugged into the existing models to enable them to generate non-freezing and high fidelity dance…

Cited by 27SourcePDFScholar
2020

Learning Expensive Coordination: An Event-Based Deep RL Approach

ICLR 2020poster

Existing works in deep Multi-Agent Reinforcement Learning (MARL) mainly focus on coordinating cooperative agents to complete certain tasks jointly. However, in many cases of the real world, agents are self-interested such as employees in a company and clubs in a league. Therefore, the leader, i.e.,…

Cited by 11SourceScholar
2019

Towards Multi-Pose Guided Virtual Try-On Network

ICCV 2019poster

Virtual try-on systems under arbitrary human poses have significant application potential, yet also raise extensive challenges, such as self-occlusions, heavy misalignment among different poses, and complex clothes textures. Existing virtual try-on methods can only transfer clothes given a fixed hum…

Cited by 252PDFScholar
2018

Soft-Gated Warping-GAN for Pose-Guided Person Image Synthesis

NeurIPS 2018poster

Despite remarkable advances in image synthesis research, existing works often fail in manipulating images under the context of large geometric transformations. Synthesizing person images conditioned on arbitrary poses is one of the most representative examples where the generation quality largely re…

Cited by 205SourcePDFScholar
2015

Simultaneous Feature Learning and Hash Coding With Deep Neural Networks

CVPR 2015poster

Similarity-preserving hashing is a widely-used method for nearest neighbour search in large-scale image retrieval tasks. For most existing hashing methods, an image is first encoded as a vector of hand-engineering visual features, followed by another separate projection or quantization step that gen…

Cited by 1028SourcePDFScholar