← Search

Hongyu Sun

9 accepted papers

2025

Beyond Film Subtitles: Is YouTube the Best Approximation of Spoken Vocabulary?

COLING 2025main

Word frequency is a key variable in psycholinguistics, useful for modeling human familiarity with words even in the era of large language models (LLMs). Frequency in film subtitles has proved to be a particularly good approximation of everyday language exposure. For many languages, however, film sub…

2025

LoCt-Instruct: An Automatic Pipeline for Constructing Datasets of Logical Continuous Instructions

EMNLP 2025

Continuous instruction following closely mirrors real-world tasks by requiring models to solve sequences of interdependent steps, yet existing multi-step instruction datasets suffer from three key limitations: (1) lack of logical coherence across turns, (2) narrow topical breadth and depth, and (3)

2025

Point-Cache: Test-time Dynamic and Hierarchical Cache for Robust and Generalizable Point Cloud Analysis

CVPR 2025poster

This paper proposes a general solution to enable point cloud recognition models to handle distribution shifts at test time. Unlike prior methods, which rely heavily on training data (often inaccessible during online inference) and are limited to recognizing a fixed set of point cloud classes predefi…

2024

Applying Contrastive Learning to Code Vulnerability Type Classification

EMNLP 2024main

Vulnerability classification is a crucial task in software security analysis, essential for identifying and mitigating potential security risks. Learning-based methods often perform poorly due to the long-tail distribution of vulnerability classification datasets. Recent approaches try to address th…

Cited by 0SourcePDFScholar
2024

Parameter-efficient Prompt Learning for 3D Point Cloud Understanding

ICRA 2024poster

This paper presents a parameter-efficient prompt tuning method, named PPT, to adapt a large multi-modal model for 3D point cloud understanding. Existing strategies are quite expensive in computation and storage, and depend on timeconsuming prompt engineering. We address the problems from three aspec…

Cited by 7SourcecodeScholar
2024

Point-PRC: A Prompt Learning Based Regulation Framework for Generalizable Point Cloud Analysis

NeurIPS 2024poster

This paper investigates the 3D domain generalization (3DDG) ability of large 3D models based on prevalent prompt learning. Recent works demonstrate the performances of 3D point cloud recognition can be boosted remarkably by parameter-efficient prompt tuning. However, we observe that the improvement…

2024

Practical Measurements of Translucent Materials with Inter-Pixel Translucency Prior

CVPR 2024poster

Material appearance is a key component of photorealism with a pronounced impact on human perception. Although there are many prior works targeting at measuring opaque materials using light-weight setups (e.g. consumer-level cameras) little attention is paid on acquiring the optical properties of tra…

Cited by 1SourcePDFScholar
2023

Lightweight Fisher Vector Transfer Learning for Video Deduplication

ICASSP 2023accepted

Video deduplication in cloud and on devices is a key challenge for storage and communication efficiency. The lifetime of video content creation, communication/sharing, and consumption can generate multiple versions of the same content with variations in coding and editing effects. In this work, we d…

Cited by 0SourceScholar
2023

ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud Understanding

ICRA 2023poster

Recently, a growing number of work design unsupervised paradigms for point cloud processing to alleviate the limitation of expensive manual annotation and poor transferability of supervised methods. Among them, CrossPoint follows the contrastive learning framework and exploits image and point cloud…

Cited by 11SourcecodeScholar