← Search

Bowen Chen

22 accepted papers

2026

Falsdo: Benchmarking Artifact-Controlled Multimodal Fake News Verification via Failure-Aligned Auditing

IJCAI 2026

Recent generative AI renders multimodal misinformation structurally harder to detect, making reliable detection dependent on semantic verification grounded in verifiable evidence. However, current benchmarks often fail to isolate true semantic checking from superficial shortcut exploitation. We intr

Cited by 0Scholar
2026

ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution

IJCAI 2026

Autonomous agents are increasingly expected to support scientific research, and recent benchmarks report progress in code repair and autonomous experimentation. However, these evaluations typically assume a pre-configured execution environment, which requires resolving complex software dependencies,

Cited by 0Scholar
2026

SPICE: Submodular Penalized Information–Conflict Selection for Efficient Large Language Model Training

ICLR 2026poster

Information-based data selection for instruction tuning is compelling: maximizing the log-determinant of the Fisher information yields a monotone submodular objective, enabling greedy algorithms to achieve a $(1-1/e)$ approximation under a cardinality budget. In practice, however, we identify allevi…

Cited by 0SourceScholar
2026

Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos

CVPR 2026

High Dynamic Range (HDR) user-generated (UGC) videos are rapidly proliferating across social platforms, yet most perceptual video quality assessment (VQA) systems remain tailored to Standard Dynamic Range (SDR). HDR's higher bit depth, wide color gamut, and elevated luminance range expose distortion

Cited by 0SourcecodeScholar
2026

Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action Segmentation

CVPR 2026

Skeleton-based Temporal Action Segmentation (STAS) seeks to densely segment and classify diverse actions within long, untrimmed skeletal motion sequences. However, existing STAS methodologies face challenges of limited inter-class discriminability and blurred segmentation boundaries, primarily due t

Cited by 0SourcecodeScholar
2025

A Statistical and Multi-Perspective Revisiting of the Membership Inference Attack in Large Language Models

ACL 2025long

The lack of data transparency in Large Language Models (LLMs) has highlighted the importance of Membership Inference Attack (MIA), which differentiates trained (member) and untrained (non-member) data. Though it shows success in previous studies, recent research reported a near-random performance in…

2025

Connectome-Based Modelling Reveals Orientation Maps in the Drosophila Optic Lobe

NeurIPS 2025poster

The ability to extract oriented edges from visual input is a core computation across animal vision systems. Orientation maps, long associated with the layered architecture of the mammalian visual cortex, systematically organise neurons by their preferred edge orientation. Despite lacking cortical st…

Cited by 0SourceScholar
2025

GloPER: Unsupervised Animal Pattern Extraction from Local Reconstruction

ICCV 2025poster

Traditional image segmentation methods struggle with fine-grained pattern extraction, especially in an unsupervised setting without labeled data. Shallow and deep learning approaches either lack structural coherence or focus on object-level segmentation rather than internal textures. Additionally, e…

2025

OMS: On-the-fly, Multi-Objective, Self-Reflective Ad Keyword Generation via LLM Agent

EMNLP 2025

Keyword decision in Sponsored Search Advertising is critical to the success of ad campaigns. While LLM-based methods offer automated keyword generation, they face three major limitations: reliance on large-scale query–keyword pair data, lack of online multi-objective performance monitoring and optim

Cited by 0SourcePDFScholar
2025

XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation

NeurIPS 2025poster

Achieving fine-grained control over subject identity and semantic attributes (pose, style, lighting) in text-to-image generation, particularly for multiple subjects, often undermines the editability and coherence of Diffusion Transformers (DiTs). Many approaches introduce artifacts or suffer from at…

Cited by 0SourcecodeScholar
2024

AvatarVerse: High-Quality & Stable 3D Avatar Creation from Text and Pose

AAAI 2024technical

Creating expressive, diverse and high-quality 3D avatars from highly customized text descriptions and pose guidance is a challenging task, due to the intricacy of modeling and texturing in 3D that ensure details and various styles (realistic, fictional, etc). We present AvatarVerse, a stable pipelin…

2024

Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action Segmentation

ECCV 2024poster

"Skeleton-based Temporal Action Segmentation (STAS) aims to densely segment and classify human actions in long, untrimmed skeletal motion sequences. Existing STAS methods primarily model spatial dependencies among joints and temporal relationships among frames to generate frame-level one-hot classif…

2023

Visual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology Images

CVPR 2023poster

Contrastive visual language pretraining has emerged as a powerful method for either training new language-aware image encoders or augmenting existing pretrained models with zero-shot visual recognition capabilities. However, existing works typically train on large datasets of image-text pairs and ha…

2023

WSCFER: Improving Facial Expression Representations by Weak Supervised Contrastive Learning

IROS 2023poster

The major challenge of Facial Expression Recog-nition (FER) is to learn class discriminative representations, and the existing works mainly address it by designing various classification networks from class level. However, learning representations at class level is limited due to the inconspicuous c…

Cited by 3SourceScholar
2022

Mitigating Reporting Bias in Semi-supervised Temporal Commonsense Inference with Probabilistic Soft Logic

AAAI 2022technical

Acquiring high-quality temporal common sense (TCS) knowledge from free-form text is a crucial but challenging problem for event-centric natural language understanding, due to the language reporting bias problem: people rarely report the commonly observed events but highlight the special cases. For e…

2022

Syntactic and Semantic Uniformity for Semantic Parsing and Task-Oriented Dialogue Systems

EMNLP 2022finding

This paper proposes a data representation framework for semantic parsing and task-oriented dialogue systems, aiming to achieve a uniform representation for syntactically and semantically diverse machine-readable formats.Current NLP systems heavily rely on adapting pre-trained language models to spec…

Cited by 4SourcePDFScholar
2021

Adaptive RF Fingerprint Decomposition in Micro UAV Detection based on Machine Learning

ICASSP 2021accepted

Radio frequency (RF) signal classification has significantly been used for detecting and identifying the features of unknown unmanned aerial vehicles (UAVs). This paper proposes a method using empirical mode decomposition (EMD) and ensemble empirical mode decomposition (EEMD) on extracting the commu…

Cited by 0SourceScholar
2021

Fusion-DHL: WiFi, IMU, and Floorplan Fusion for Dense History of Locations in Indoor Environments

ICRA 2021poster

The paper proposes a multi-modal sensor fusion algorithm that fuses WiFi, IMU, and floorplan information to infer an accurate and dense location history in indoor environments. The algorithm uses 1) an inertial navigation algorithm to estimate a relative motion trajectory from IMU sensor data; 2) a…

Cited by 29SourcecodeScholar
2020

ScribbleBox: Interactive Annotation Framework for Video Object Segmentation

ECCV 2020poster

Manually labeling video datasets for segmentation tasks is extremely time consuming. We introduce ScribbleBox, an interactive framework for annotating object instances with masks in videos with a significant boost in efficiency. In particular, we split annotation into two steps: annotating objects w…

Cited by 23SourcePDFScholar