← Search

Muli Yang

19 accepted papers

2026

Editing Is a Bargaining Game: Balanced Knowledge Editing in Large Language Models

AAAI 2026technical

Large Language Models (LLMs) are prone to generating incorrect or outdated information, thereby necessitating efficient and precise mechanisms for knowledge updates. Existing knowledge editing approaches, however, often encounter conflicts between two competing objectives: maintaining existing knowl

Cited by 0SourcePDFScholar
2026

Multimodal Semantic Bias Mitigation for Diverse Text-To-3D Generation

CVPR 2026

The latest progress in text-to-3D generative models makes it possible to generate high-quality 3D content. Recent text-to-3D large model have achieved remarkable breakthroughs in multi-view consistency. However, their effectiveness is often affected by inherent biases, resulting in sensitivity to de

Cited by 0SourceScholar
2026

Next-Generation Metalens Vision System: Powered by AI and Applied to AI

AAAI 2026technical

Metalenses have been widely recognized as a key building block of next-generation optical systems, offering unprecedented advantages in compactness, lightweight design, and scalable manufacturing compared to traditional refractive optics. Despite this promise, practical use is limited by optical abe

Cited by 0SourcePDFScholar
2026

Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles

CVPR 2026

Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an intuitive interface, yet translating passenger open-ended instructions into control signals--without sacrificing interpretability and traceability--remai

Cited by 0SourceScholar
2026

Towards Illumination-Aware Restoration of Metalens-Captured Images: A New Dataset and a Strong Baseline

AAAI 2026technical

Metalenses offer compelling advantages such as lightweight and ultra-thin design, making them promising alternatives to conventional lenses. However, their widespread adoption is hindered by image quality degradation caused by chromatic and angular aberrations. To mitigate this, restoration processe

Cited by 0SourcePDFScholar
2026

Your AI-Generated Image Detector Can Secretly Achieve SOTA Accuracy, If Calibrated

AAAI 2026technical

Despite being trained on balanced datasets, existing AI-generated image detectors often exhibit systematic bias at test time, frequently misclassifying fake images as real. We hypothesize that this behavior stems from distributional shift in fake samples and implicit priors learned during training.

Cited by 0SourcePDFScholar
2025

Detecting Open World Objects via Partial Attribute Assignment

CVPR 2025poster

Despite being trained on massive data, today's vision foundation models still fall short in detecting open world objects. Apart from recognizing known objects from training, a successful Open World Object Detection (OWOD) system must also be able to detect unknown objects never seen before, without…

2025

Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance Grounding

CVPR 2025poster

This paper pays attention to Weakly Supervised Affordance Grounding (WSAG) task that aims to train model to identify affordance regions using human-object interaction images and egocentric images without the need for costly pixel-level annotations. Most existing methods usually consider the affordan…

Cited by 0SourcePDFScholar
2025

Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative Modeling

NeurIPS 2025poster

In dance performances, choreographers define the visual expression of movement, while cinematographers shape its final presentation through camera work. Consequently, the synthesis of camera movements informed by both music and dance has garnered increasing research interest. While recent advancemen…

Cited by 0SourceScholar
2025

Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization

ICLR 2025poster

Recently, the comprehensive understanding of human motion has been a prominent area of research due to its critical importance in many fields. However, existing methods often prioritize specific downstream tasks and roughly align text and motion features within a CLIP-like framework. This results in…

Cited by 1SourcePDFScholar
2025

Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph Generation

ICCV 2025poster

To promote the deployment of scenario understanding in the real world, Open-Vocabulary Scene Graph Generation (OV-SGG) has attracted much attention recently, aiming to generalize beyond the limited number of relation categories labeled during training and detect those unseen relations during inferen…

2024

LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures

ACL 2024long

In response to the escalating demand for digital human representations, progress has been made in the generation of realistic human gestures from given speeches. Despite the remarkable achievements of recent research, the generation process frequently includes unintended, meaningless, or non-realist…

Cited by 3SourcePDFScholar
2023

Bootstrap Your Own Prior: Towards Distribution-Agnostic Novel Class Discovery

CVPR 2023poster

Novel Class Discovery (NCD) aims to discover unknown classes without any annotation, by exploiting the transferable knowledge already learned from a base set of known classes. Existing works hold an impractical assumption that the novel class distribution prior is uniform, yet neglect the imbalanced…

2023

Hierarchical Prompt Learning for Compositional Zero-Shot Recognition

IJCAI 2023poster

Compositional Zero-Shot Learning (CZSL) aims to imitate the powerful generalization ability of human beings to recognize novel compositions of known primitive concepts that correspond to a state and an object, e.g., purple apple. To fully capture the intra- and inter-class correlations between compo…

Cited by 23SourcePDFScholar
2022

Divide and Conquer: Compositional Experts for Generalized Novel Class Discovery

CVPR 2022poster

In response to the explosively-increasing requirement of annotated data, Novel Class Discovery (NCD) has emerged as a promising alternative to automatically recognize unknown classes without any annotation. To this end, a model makes use of a base set to learn basic semantic discriminability that ca…

Cited by 51PDFcodeScholar
2020

Fewer is More: A Deep Graph Metric Learning Perspective Using Fewer Proxies

NeurIPS 2020spotlight

Deep metric learning plays a key role in various machine learning tasks. Most of the previous works have been confined to sampling from a mini-batch, which cannot precisely characterize the global geometry of the embedding space. Although researchers have developed proxy- and classification-based me…

2020

Learning Unseen Concepts via Hierarchical Decomposition and Composition

CVPR 2020poster

Composing and recognizing new concepts from known sub-concepts has been a fundamental and challenging vision task, mainly due to 1) the diversity of sub-concepts and 2) the intricate contextuality between sub-concepts and their corresponding visual features. However, most of the current methods simp…

Cited by 66PDFScholar
2020

Progressive Domain-Independent Feature Decomposition Network for Zero-Shot Sketch-Based Image Retrieval

IJCAI 2020poster

Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is a specific cross-modal retrieval task for searching natural images given free-hand sketches under the zero-shot scenario. Most existing methods solve this problem by simultaneously projecting visual features and semantic supervision into a low-dime…

Cited by 0SourcePDFScholar
2019

Adversarial Fine-Grained Composition Learning for Unseen Attribute-Object Recognition

ICCV 2019poster

Recognizing unseen attribute-object pairs never appearing in the training data is a challenging task, since an object often refers to a specific entity while an attribute is an abstract semantic description. Besides, attributes are highly correlated to objects, i.e., an attribute tends to describe d…

Cited by 116PDFScholar