← Search

Zhiyu Wu

8 accepted papers

2026

Not Just What’s There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-Tuning

AAAI 2026technical

Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing methods refine negation understanding via fine-tuning CLIP’s text encoder, risking overfitting. In this work, we propose C

Cited by 0SourcePDFScholar
2025

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

CVPR 2025poster

We introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing levels of information granularity required by multimodal understanding and gen…

2025

JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

CVPR 2025poster

We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model.JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling.Our key finding demonstrate…

2025

RecNet: Optimization for Dense Object Detection in Retail Scenarios Based on View Rectification

ICASSP 2025accepted

High-precision dense object detection in retail is crucial for automation, inventory management, and sales optimization. Our experiments revealed that detection models perform significantly better with frontal views than with oblique views, motivating the development of RecNet. RecNet utilizes a Rec…

Cited by 0SourceScholar
2025

The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization

NeurIPS 2025spotlight

As the adoption of Generative AI in real-world services grow explosively, energy has emerged as a critical bottleneck resource. However, energy remains a metric that is often overlooked, under-explored, or poorly understood in the context of building ML systems. We present the ML.ENERGY Benchmark, a…

Cited by 0SourcecodeScholar
2023

BERT-ERC: Fine-Tuning BERT Is Enough for Emotion Recognition in Conversation

AAAI 2023technical

Previous works on emotion recognition in conversation (ERC) follow a two-step paradigm, which can be summarized as first producing context-independent features via fine-tuning pretrained language models (PLMs) and then analyzing contextual information and dialogue structure information among the ext…

Cited by 38SourcePDFScholar