← Search

Haotian Liu

33 accepted papers

2025

AnomalyNCD: Towards Novel Anomaly Class Discovery in Industrial Scenarios

CVPR 2025poster

Recently, multi-class anomaly classification has garnered increasing attention. Previous methods directly cluster anomalies but often struggle due to the lack of anomaly-prior knowledge. Acquiring this knowledge faces two issues: the non-prominent and weak-semantics anomalies. In this paper, we prop…

2025

DTRT: Enhancing Human Intent Estimation and Role Allocation for Physical Human-Robot Collaboration

ICRA 2025

In physical Human-Robot Collaboration (pHRC), accurate human intent estimation and rational human-robot role allocation are crucial for safe and efficient assistance. Existing methods that rely on short-term motion data for intention estimation lack multi-step prediction capabilities, hindering thei

Cited by 5SourceScholar
2025

Fantastic Copyrighted Beasts and How (Not) to Generate Them

ICLR 2025poster

Recent studies show that image and video generation models can be prompted to reproduce copyrighted content from their training data, raising serious legal con- cerns about copyright infringement. Copyrighted characters (e.g., Mario, Batman) present a significant challenge: at least one lawsuit has…

Cited by 12SourcePDFScholar
2025

GPS: A Probabilistic Distributional Similarity with Gumbel Priors for Set-to-Set Matching

ICLR 2025poster

Set-to-set matching aims to identify correspondences between two sets of unordered items by minimizing a distance metric or maximizing a similarity measure. Traditional metrics, such as Chamfer Distance (CD) and Earth Mover’s Distance (EMD), are widely used for this purpose but often suffer from lim…

2025

IDAGC: Adaptive Generalized Human-Robot Collaboration via Human Intent Estimation and Multimodal Policy Learning

IROS 2025

In Human-Robot Collaboration (HRC), which encompasses physical interaction and remote cooperation, accurate estimation of human intentions and seamless switching of collaboration modes to adjust robot behavior remain paramount challenges. To address these issues, we propose an Intent-Driven Adaptive

Cited by 6SourceScholar
2025

Match Policy: A Simple Pipeline from Point Cloud Registration to Manipulation Policies

ICRA 2025

Many manipulation tasks require the robot to rearrange objects relative to one another. Such tasks can be described as a sequence of relative poses between parts of a set of rigid bodies. In this work, we propose Match Policy, a simple but novel pipeline for solving high-precision pick and place tas

Cited by 5SourcecodeScholar
2025

SeaS: Few-shot Industrial Anomaly Image Generation with Separation and Sharing Fine-tuning

ICCV 2025poster

We introduce SeaS, a unified industrial generative model for automatically creating diverse anomalies, authentic normal products, and precise anomaly masks. While extensive research exists, most efforts either focus on specific tasks, i.e., anomalies or normal products only, or require separate mode…

2024

Aligning Large Multimodal Models with Factually Augmented RLHF

ACL 2024findings

Large Multimodal Models (LMM) are built across modalities and the misalignment between two modalities can result in “hallucination”, generating textual outputs that are not grounded by the multimodal information in context. To address the multimodal misalignment issue, we adapt the Reinforcement Lea…

2024

Bridging The Gap between Low-rank and Orthogonal Adaptation via Householder Reflection Adaptation

NeurIPS 2024spotlight

While following different technical routes, both low-rank and orthogonal adaptation techniques can efficiently adapt large-scale pre-training models in specific tasks or domains based on a small piece of trainable parameters. In this study, we bridge the gap between these two techniques, proposing a…

2024

CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

NeurIPS 2024poster

Chart understanding plays a pivotal role when applying Multimodal Large Language Models (MLLMs) to real-world tasks such as analyzing scientific papers or financial reports. However, existing datasets often focus on oversimplified and homogeneous charts with template-based questions, leading to an o…

2024

Edit One for All: Interactive Batch Image Editing

CVPR 2024poster

In recent years image editing has advanced remarkably. With increased human control it is now possible to edit an image in a plethora of ways; from specifying in text what we want to change to straight up dragging the contents of the image in an interactive point-based manner. However most of the fo…

Cited by 4SourcePDFScholar
2024

Generalizable Face Landmarking Guided by Conditional Face Warping

CVPR 2024poster

As a significant step for human face modeling editing and generation face landmarking aims at extracting facial keypoints from images. A generalizable face landmarker is required in practice because real-world facial images e.g. the avatars in animations and games are often stylized in various ways.…

2024

IMAGINATION POLICY: Using Generative Point Cloud Models for Learning Manipulation Policies

CoRL 2024poster

Humans can imagine goal states during planning and perform actions to match those goals. In this work, we propose IMAGINATION POLICY, a novel multi-task key-frame policy network for solving high-precision pick and place tasks. Instead of learning actions directly, IMAGINATION POLICY generates point…

Cited by 7SourceScholar
2024

Inferring Iterated Function Systems Approximately from Fractal Images

IJCAI 2024poster

As an important mathematical concept, fractals commonly appear in nature and inspire the design of many artistic works. Although we can generate various fractal images easily based on different iterated function systems (IFSs), inferring an IFS from a given fractal image is still a challenging inve…

2024

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

ECCV 2024poster

"This paper presents (), a general-purpose multimodal assistant trained using an end-to-end approach that systematically expands the capabilities of large multimodal models (LMMs). maintains a skill repository that contains a wide range of vision and vision-language pre-trained models (tools), and i…

2024

Loss Distillation via Gradient Matching for Point Cloud Completion with Weighted Chamfer Distance

IROS 2024poster

3D point clouds enhanced the robot’s ability to perceive the geometrical information of the environments, making it possible for many downstream tasks such as grasp pose detection and scene understanding. The performance of these tasks, though, heavily relies on the quality of data input, as incompl…

Cited by 1SourcecodeScholar
2024

PCDepth: Pattern-based Complementary Learning for Monocular Depth Estimation by Best of Both Worlds

IROS 2024poster

Event cameras can record scene dynamics with high temporal resolution, providing rich scene details for monocular depth estimation (MDE) even at low-level illumination. Therefore, existing complementary learning approaches for MDE fuse intensity information from images and scene details from event d…

Cited by 3SourceScholar
2024

Removing Distributional Discrepancies in Captions Improves Image-Text Alignment

ECCV 2024poster

"In this paper, we introduce a model designed to improve the prediction of image-text alignment, targeting the challenge of compositional understanding in current visual-language models. Our approach focuses on generating high-quality training datasets for the alignment task by producing mixed-type…

2024

UNICORN: A Unified Causal Video-Oriented Language-Modeling Framework for Temporal Video-Language Tasks

EMNLP 2024main

The great success of large language models has encouraged the development of large multimodal models, with a focus on image-language interaction. Despite promising results in various image-language downstream tasks, it is still challenging and unclear how to extend the capabilities of these models t…

2024

ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts

CVPR 2024poster

While existing large vision-language multimodal models focus on whole image understanding there is a prominent gap in achieving region-specific comprehension. Current approaches that use textual coordinates or spatial encodings often fail to provide a user-friendly interface for visual prompting. To…

Cited by 106SourcePDFScholar
2024

Yo'LLaVA: Your Personalized Language and Vision Assistant

NeurIPS 2024poster

Large Multimodal Models (LMMs) have shown remarkable capabilities across a variety of tasks (e.g., image captioning, visual question answering). While broad, their knowledge remains generic (e.g., recognizing a dog), and they are unable to handle personalized subjects (e.g., recognizing a user's pet…

2023

A Novel Transformer-Based Pipeline for Lung Cytopathological Whole Slide Image Classification

ICASSP 2023accepted

We propose a novel three-stage Transformer-based methodology for entire cytopathological whole slide image (WSI) classification. The key idea is to leverage Transformer to extract the fine-grained lesion-level features and then progressively aggregate them into intermediate-grained patch-level featu…

Cited by 0SourceScholar
2023

Data-Efficient Image Quality Assessment with Attention-Panel Decoder

AAAI 2023technical

Blind Image Quality Assessment (BIQA) is a fundamental task in computer vision, which however remains unresolved due to the complex distortion conditions and diversified image contents. To confront this challenge, we in this paper propose a novel BIQA pipeline based on the Transformer architecture,…

2023

GLIGEN: Open-Set Grounded Text-to-Image Generation

CVPR 2023poster

Large-scale text-to-image diffusion models have made amazing advances. However, the status quo is to use text input alone, which can impede controllability. In this work, we propose GLIGEN: Open-Set Grounded Text-to-Image Generation, a novel approach that builds upon and extends the functionality of…

2023

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

NeurIPS 2023spotlight

Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversational AI has seen rapid progress by leveraging billions of image-text pairs from the public web, but such general-domain vi…

Cited by 828SourcePDFScholar
2023

Learning Customized Visual Models With Retrieval-Augmented Knowledge

CVPR 2023highlight

Image-text contrastive learning models such as CLIP have demonstrated strong task transfer ability. The high generality and usability of these visual models is achieved via a web-scale data collection process to ensure broad concept coverage, followed by expensive pre-training to feed all the knowle…

2023

TMA: Temporal Motion Aggregation for Event-based Optical Flow

ICCV 2023poster

Event cameras have the ability to record continuous and detailed trajectories of objects with high temporal resolution, thereby providing intuitive motion cues for optical flow estimation. Nevertheless, most existing learning-based approaches for event optical flow estimation directly remould the pa…

Cited by 31PDFcodeScholar
2022

ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models

NeurIPS 2022accept

Learning visual representations from natural language supervision has recently shown great promise in a number of pioneering works. In general, these language-augmented visual models demonstrate strong transferability to a variety of datasets/tasks. However, it remains challenging to evaluate the tr…

Cited by 159SourcePDFScholar
2021

YolactEdge: Real-time Instance Segmentation on the Edge

ICRA 2021poster

We propose YolactEdge, the first competitive instance segmentation approach that runs on small edge devices at real-time speeds. Specifically, YolactEdge runs at up to 30.8 FPS on a Jetson AGX Xavier (and 172.7 FPS on an RTX 2080 Ti) with a ResNet-101 backbone on 550x550 resolution images. To achiev…

Cited by 104SourcecodeScholar
2019

Identity From Here, Pose From There: Self-Supervised Disentanglement and Generation of Objects Using Unlabeled Videos

ICCV 2019poster

We propose a novel approach that disentangles the identity and pose of objects for image generation. Our model takes as input an ID image and a pose image, and generates an output image with the identity of the ID image and the pose of the pose image. Unlike most previous unsupervised work which rel…

Cited by 21PDFScholar