← Search

Chaoyi Zhang

14 accepted papers

2026

INT vs. FP: A Comprehensive Study of Fine-Grained Low-bit Quantization Formats

ICML 2026poster

Modern AI hardware, such as Nvidia's Blackwell architecture, is increasingly embracing low-precision floating-point (FP) formats to handle the pervasive activation outliers in Large Language Models (LLMs). Despite this industry trend, a unified comparison of FP and integer (INT) quantization across …

Cited by 0SourceScholar
2025

Model Merging in Pre-training of Large Language Models

NeurIPS 2025poster

Model merging has emerged as a promising technique for enhancing large language models, though its application in large-scale pre-training remains relatively unexplored. In this paper, we present a comprehensive investigation of model merging techniques during the pre-training process. Through exten…

Cited by 0SourceScholar
2025

Multimodal Causal Reasoning Benchmark: Challenging Multimodal Large Language Models to Discern Causal Links Across Modalities

ACL 2025finding

Multimodal Large Language Models (MLLMs) have showcased exceptional Chain-of-Thought (CoT) reasoning ability in complex textual inference tasks including causal reasoning. However, will these causalities remain straightforward when crucial hints hide in visual details? If not, what factors might inf…

Cited by 0SourcePDFScholar
2024

A Robot Humanoid Control Framework Through Human Arm Active Endpoint Stiffness and Direction Adaptive Compensation

RA-L 2024

In the process of human-robot interaction (HRI), the controleffect cannot meet the needs of HRI tasks if a control strategy is developed solely from the robot's point of view. It's necessary to take into account the characteristics of the human operator. In this letter, a novel HRI framework is deve

Cited by 8SourceScholar
2024

Controllable Contextualized Image Captioning: Directing the Visual Narrative through User-Defined Highlights

ECCV 2024poster

"(CIC) evolves traditional image captioning into a more complex domain, necessitating the ability for multimodal reasoning. It aims to generate image captions given specific contextual information. This paper further introduces a novel domain of (). Unlike CIC, which solely relies on broad context,…

2024

Enhancing Advanced Visual Reasoning Ability of Large Language Models

EMNLP 2024main

Recent advancements in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models’ advanced reasoning ability. Traditional Vision-Language models (VLMs) perform well in visual perception tasks while struggling with complex reasoning scenarios. Converse…

Cited by 7SourcePDFScholar
2024

MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning

CVPR 2024highlight

We present MM-Narrator a novel system leveraging GPT-4 with multimodal in-context learning for the generation of audio descriptions (AD). Unlike previous methods that primarily focused on downstream fine-tuning with short video clips MM-Narrator excels in generating precise audio descriptions for vi…

Cited by 27SourcePDFScholar
2023

PaRot: Patch-Wise Rotation-Invariant Network via Feature Disentanglement and Pose Restoration

AAAI 2023technical

Recent interest in point cloud analysis has led rapid progress in designing deep learning methods for 3D models. However, state-of-the-art models are not robust to rotations, which remains an unknown prior to real applications and harms the model performance. In this work, we introduce a novel Patch…

2023

SQUID: Deep Feature In-Painting for Unsupervised Anomaly Detection

CVPR 2023poster

Radiography imaging protocols focus on particular body regions, therefore producing images of great similarity and yielding recurrent anatomical structures across patients. To exploit this structured information, we propose the use of Space-aware Memory Queues for In-painting and Detecting anomalies…

2022

Spatiality-guided Transformer for 3D Dense Captioning on Point Clouds

IJCAI 2022poster

Dense captioning in 3D point clouds is an emerging vision-and-language task involving object-level 3D scene understanding. Apart from coarse semantic class prediction and bounding box regression as in traditional 3D object detection, 3D dense captioning aims at producing a further and finer instance…

2021

Exploiting Edge-Oriented Reasoning for 3D Point-Based Scene Graph Analysis

CVPR 2021poster

Scene understanding is a critical problem in computer vision. In this paper, we propose a 3D point-based scene graph generation (SGGpoint) framework to effectively bridge perception and reasoning to achieve scene understanding via three sequential stages, namely scene graph construction, reasoning,…

Cited by 65PDFScholar
2021

Walk in the Cloud: Learning Curves for Point Clouds Shape Analysis

ICCV 2021poster

Discrete point cloud objects lack sufficient shape descriptors of 3D geometries. In this paper, we present a novel method for aggregating hypothetical curves in point clouds. Sequences of connected points (curves) are initially grouped by taking guided walks in the point clouds, and then subsequentl…

Cited by 367PDFcodeScholar