← Search

Bohan Yu

15 accepted papers

2026

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

CVPR 2026

Omnimodal large language models (OmniLLMs) have attracted increasing research attention of late towards unified audio-video understanding. However, the high computational cost of processing longer joint audio-video token sequences has become a key bottleneck. Existing token compression methods have

Cited by 0SourcecodeScholar
2026

TextShield-R1: Reinforced Reasoning for Tampered Text Detection

AAAI 2026technical

The growing prevalence of tampered images poses serious security threats, highlighting the urgent need for reliable detection methods. Multimodal large language models (MLLMs) demonstrate strong potential in analyzing tampered images and generating interpretations. However, they still struggle with

Cited by 0SourcePDFScholar
2025

Active Hyperspectral Imaging Using an Event Camera

CVPR 2025highlight

Hyperspectral imaging plays a critical role in numerous scientific and industrial fields. Conventional hyperspectral imaging systems often struggle with the trade-off between capture speed, spectral resolution, and bandwidth, particularly in dynamic environments. In this work, we present a novel eve…

Cited by 0SourcePDFScholar
2025

CaDRL: Document-level Relation Extraction via Context-aware Differentiable Rule Learning

COLING 2025main

Document-level Relation Extraction (DocRE) aims to extract relations from documents. Compared with sentence-level relation extraction, it is necessary to extract long-distance dependencies. Existing methods enhance the output of trained DocRE models either by learning logical rules or by extracting…

2025

Development of an Efficient Stiffness Modulation Mechanism in Fish-like Robots for Enhanced Swimming Performance

IROS 2025

Drawing inspiration from the ability of fish to maintain efficient swimming over a wide range of speeds by tuning the stiffness of their tails, researchers have explored stiffness adjustment mechanisms in fish-like robots. Typically, existing mechanisms require extra actuators or power sources only

Cited by 0SourceScholar
2025

EventPSR: Surface Normal and Reflectance Estimation from Photometric Stereo Using an Event Camera

CVPR 2025highlight

Simultaneously acquisition of the surface normal and reflectance parameters is a crucial but challenging technique in the field of computer vision and graphics. It requires capturing multiple high dynamic range (HDR) images in existing methods using frame-based cameras. In this paper, we propose Eve…

Cited by 0SourcePDFScholar
2025

EventUPS: Uncalibrated Photometric Stereo Using an Event Camera

ICCV 2025poster

We present EventUPS, the first uncalibrated photometric stereo (UPS) method using an event camera--a neuromorphic sensor that asynchronously detects brightness changes with microsecond resolution. Traditional frame-based UPS methods are hindered by high bandwidth demands and limited use in dynamic s…

Cited by 0SourcePDFScholar
2025

TableEval: A Real-World Benchmark for Complex, Multilingual, and Multi-Structured Table Question Answering

EMNLP 2025

LLMs have shown impressive progress in natural language processing. However, they still face significant challenges in TableQA, where real-world complexities such as diverse table structures, multilingual data, and domain-specific reasoning are crucial. Existing TableQA benchmarks are often limited

2024

EventPS: Real-Time Photometric Stereo Using an Event Camera

CVPR 2024poster

Photometric stereo is a well-established technique to estimate the surface normal of an object. However the requirement of capturing multiple high dynamic range images under different illumination conditions limits the speed and real-time applications. This paper introduces EventPS a novel approach…

Cited by 12SourcePDFScholar
2024

Latency Correction for Event-guided Deblurring and Frame Interpolation

CVPR 2024poster

Event cameras with their high temporal resolution dynamic range and low power consumption are particularly good at time-sensitive applications like deblurring and frame interpolation. However their performance is hindered by latency variability especially under low-light conditions and with fast-mov…

Cited by 9SourcePDFScholar
2024

Towards an Interpretable Representation of Speaker Identity via Perceptual Voice Qualities

ICASSP 2024accepted

Unlike other data modalities such as text and vision, speech does not lend itself to easy interpretation. While lay people can understand how to describe an image or sentence via perception, non-expert descriptions of speech often end at high-level demographic information, such as gender or age. In…

Cited by 0SourceScholar
2023

ReLeaPS : Reinforcement Learning-based Illumination Planning for Generalized Photometric Stereo

ICCV 2023poster

Illumination planning in photometric stereo aims to find a balance between tween surface normal estimation accuracy and image capturing efficiency by selecting optimal light configurations. It depends on factors such as the unknown shape and general reflectance of the target object, global illuminat…

Cited by 2PDFScholar