← Search

Wynne Hsu

27 accepted papers

2026

LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical Supervision

ICLR 2026poster

Although LLMs exhibit strong reasoning capabilities, existing training methods largely depend on outcome-based feedback, which can produce correct answers with flawed reasoning. Prior work introduces supervision on intermediate steps but still lacks guarantees of logical soundness, which is crucial…

Cited by 0SourcecodeScholar
2026

Orthogonal Spatial-temporal Distributional Transfer for 4D Generation

AAAI 2026technical

In the AIGC era, generating high-quality 4D content has garnered increasing research attention. Unfortunately, current 4D synthesis research is severely constrained by the lack of large-scale 4D datasets, preventing models from adequately learning the critical spatial-temporal features necessary for

Cited by 0SourcePDFScholar
2026

UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark

CVPR 2026

In real-world multimodal applications, systems usually need to comprehend arbitrarily combined and interleaved multimodal inputs from users, while also generating outputs in any interleaved multimedia form. This capability defines the goal of any-to-any interleaved multimodal learning under a unifie

Cited by 0SourceScholar
2026

Unveiling the Cognitive Compass: Theory-of-Mind–Guided Multimodal Emotion Reasoning

ICLR 2026poster

Despite rapid progress in multimodal large language models (MLLMs), their capability for deep emotional understanding remains limited. We argue that genuine affective intelligence requires explicit modeling of Theory of Mind (ToM), the cognitive substrate from which emotions arise. To this end, we i…

Cited by 0SourceScholar
2025

Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework

ACL 2025long

In the context of large language models (LLMs), current advanced reasoning methods have made impressive strides in various reasoning tasks. However, when it comes to logical reasoning tasks, significant challenges remain in both efficacy and efficiency. This is rooted in the fact that these systems…

2025

From Personas to Talks: Revisiting the Impact of Personas on LLM-Synthesized Emotional Support Conversations

EMNLP 2025

The rapid advancement of Large Language Models (LLMs) has revolutionized the generation of emotional support conversations (ESC), offering scalable solutions with reduced costs and enhanced data privacy. This paper explores the role of personas in the creation of ESC by LLMs. Our research utilizes e

Cited by 0SourcePDFScholar
2025

MuSLR: Multimodal Symbolic Logical Reasoning

NeurIPS 2025poster

Multimodal symbolic logical reasoning, which aims to deduce new facts from multimodal input via formal logic, is critical in high-stakes applications such as autonomous driving and medical diagnosis, as its rigorous, deterministic reasoning helps prevent serious consequences. To evaluate such capabi…

Cited by 0SourceScholar
2025

TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection

EMNLP 2025

Multimodal misinformation, encompassing textual, visual, and cross-modal distortions, poses an increasing societal threat that is amplified by generative AI. Existing methods typically focus on a single type of distortion and struggle to generalize to unseen scenarios. In this work, we observe that

Cited by 0SourcePDFScholar
2025

Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models

ICML 2025poster

Multi-Modal Large Language Models (MLLMs) have exhibited remarkable performance on various vision-language tasks such as Visual Question Answering (VQA). Despite accumulating evidence of privacy concerns associated with task-relevant content, it remains unclear whether MLLMs inadvertently memorize p…

2024

Cross-Domain Feature Augmentation for Domain Generalization

IJCAI 2024poster

Domain generalization aims to develop models that are robust to distribution shifts. Existing methods focus on learning invariance across domains to enhance model robustness, and data augmentation has been widely used to learn invariant predictors, with most methods performing augmentation in the in…

2024

Faithful Logical Reasoning via Symbolic Chain-of-Thought

ACL 2024long

While the recent Chain-of-Thought (CoT) technique enhances the reasoning ability of large language models (LLMs) with the theory of mind, it might still struggle in handling logical reasoning that relies much on symbolic expressions and rigid deducing rules. To strengthen the logical reasoning capab…

2024

SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection

CVPR 2024poster

Misinformation is a prevalent societal issue due to its potential high risks. Out-Of-Context (OOC) misinformation where authentic images are repurposed with false text is one of the easiest and most effective ways to mislead audiences. Current methods focus on assessing image-text consistency but la…

Cited by 50SourcePDFScholar
2024

Time Matters: An End-to-End Solution for Temporal Claim Verification

EMNLP 2024industry

Automated claim verification plays an essential role in fostering trust in the digital space. Despite the growing interest, the verification of temporal claims has not received much attention in the community. Temporal claim verification brings new challenges where cues of the temporal information n…

Cited by 0SourcePDFScholar
2024

Towards Robust Out-of-Distribution Generalization Bounds via Sharpness

ICLR 2024spotlight

Generalizing to out-of-distribution (OOD) data or unseen domain, termed OOD generalization, still lacks appropriate theoretical guarantees. Canonical OOD bounds focus on different distance measurements between source and target domains but fail to consider the optimization property of the learned mo…

Cited by 7SourcePDFScholar
2024

Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition

ICML 2024oral

Existing research of video understanding still struggles to achieve in-depth comprehension and reasoning in complex videos, primarily due to the under-exploration of two key bottlenecks: fine-grained spatial-temporal perceptive understanding and cognitive-level video scene comprehension. This paper…

Cited by 99SourcePDFScholar
2023

Label-Efficient Online Continual Object Detection in Streaming Video

ICCV 2023poster

Humans can watch a continuous video stream and effortlessly perform continual acquisition and transfer of new knowledge with minimal supervision yet retaining previously learnt experiences. In contrast, existing continual learning (CL) methods require fully annotated labels to effectively learn from…

Cited by 18PDFcodeScholar
2023

Leveraging Old Knowledge to Continually Learn New Classes in Medical Images

AAAI 2023technical

Class-incremental continual learning is a core step towards developing artificial intelligence systems that can continuously adapt to changes in the environment by learning new concepts without forgetting those previously learned. This is especially needed in the medical domain where continually lea…

2023

Multi-Object Representation Learning via Feature Connectivity and Object-Centric Regularization

NeurIPS 2023spotlight

Discovering object-centric representations from images has the potential to greatly improve the robustness, sample efficiency and interpretability of machine learning algorithms. Current works on multi-object images typically follow a generative approach that optimizes for input reconstruction and f…

Cited by 2SourcePDFScholar
2023

REFINE: A Fine-Grained Medication Recommendation System Using Deep Learning and Personalized Drug Interaction Modeling

NeurIPS 2023poster

Patients with co-morbidities often require multiple medications to manage their conditions. However, existing medication recommendation systems only offer class-level medications and regard all interactions among drugs to have the same level of severity. This limits their ability to provide personal…

Cited by 14SourcePDFScholar
2023

Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation

ICCV 2023poster

To replicate the success of text-to-image (T2I) generation, recent works employ large-scale video datasets to train a text-to-video (T2V) generator. Despite their promising results, such paradigm is computationally expensive. In this work, we propose a new T2V generation setting--One-Shot Video Tuni…

Cited by 853PDFcodeScholar
2022

Chronic Disease Management with Personalized Lab Test Response Prediction

IJCAI 2022poster

Chronic disease management involves frequent administration of invasive lab procedures in order for clinicians to determine the best course of treatment regimes for these patients. However, patients are often put off by these invasive lab procedures and do not follow the appointment schedules. T…

Cited by 4SourcePDFScholar
2021

Improving Evidence Retrieval for Automated Explainable Fact-Checking

NAACL 2021system demonstrations

Automated fact-checking on a large-scale is a challenging task that has not been studied systematically until recently. Large noisy document collections like the web or news articles make the task more difficult. We describe a three-stage automated fact-checking system, named Quin+, using evidence r…

2020

Towards Maximizing the Representation Gap between In-Domain & Out-of-Distribution Examples

NeurIPS 2020poster

Among existing uncertainty estimation approaches, Dirichlet Prior Network (DPN) distinctly models different predictive uncertainty types. However, for in-domain examples with high data uncertainties among multiple classes, even a DPN model often produces indistinguishable representations from the o…