← Search

Tianyu Yang

33 accepted papers

2026

Active Intelligence in Video Avatars via Closed-loop World Modeling

CVPR 2026

Current video avatar generation methods excel at identity preservation and motion alignment but lack genuine agency--they cannot autonomously pursue long-term goals through adaptive environmental interaction. We address this by introducing L-IVA (Long-horizon Interactive Visual Avatar), a task and b

Cited by 0SourceScholar
2026

Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory

ICML 2026poster

We propose **Infinite-World**, a robust interactive world model capable of maintaining coherent visual memory over **1000+ frames** in complex real-world environments. While existing world models can be efficiently optimized on synthetic data with perfect ground-truth, they lack an effective trainin…

Cited by 0SourceScholar
2026

ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval

CVPR 2026

Composed Image Retrieval (CIR) aims to retrieve target images based on a hybrid query comprising a reference image and a modification text. Early dual-tower Vision-Language Models (VLMs) struggle with cross-modality compositional reasoning required for this task. While adapting generative Multimodal

Cited by 0SourcecodeScholar
2026

WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval

CVPR 2026

Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve target images given a multimodal query (comprising a reference image and a modification text), without training on annotated triplets. Existing methods typically convert the multimodal query into a single modality--either as an edited capt

Cited by 0SourcecodeScholar
2026

WildActor: Unconstrained Identity-Preserving Video Generation

ICML 2026poster

Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. Prior methods often suffer from face-centric behavior that neglects body-level c…

Cited by 0SourceScholar
2025

Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis

ACL 2025finding

Due to the widespread use of LLMs and the rising critical ethical and safety concerns, LLM unlearning methods have been developed to remove harmful knowledge and undesirable capabilities. In this context, evaluations are mostly based on single-value metrics such as QA accuracy. However, these metric…

2025

CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP

ACL 2025long

Machine unlearning (MU) has gained significant attention as a means to remove the influence of specific data from a trained model without requiring full retraining. While progress has been made in unimodal domains like text and image classification, unlearning in multimodal models remains relatively…

Cited by 0SourcePDFScholar
2025

Physics: Benchmarking Foundation Models on University-Level Physics Problem Solving

ACL 2025finding

We introduce Physics, a comprehensive benchmark for university-level physics problem solving. It contains 1,297 expert-annotated problems covering six core areas: classical mechanics, quantum mechanics, thermodynamics and statistical mechanics, electromagnetism, atomic physics, and optics.Each probl…

2025

Robust Utility-Preserving Text Anonymization Based on Large Language Models

ACL 2025long

Anonymizing text that contains sensitive information is crucial for a wide range of applications. Existing techniques face the emerging challenges of the re-identification ability of large language models (LLMs), which have shown advanced capability in memorizing detailed information and reasoning o…

2025

Self-Improvement in Multimodal Large Language Models: A Survey

EMNLP 2025

Recent advancements in self-improvement for Large Language Models (LLMs) have efficiently enhanced model capabilities without significantly increasing costs, particularly in terms of human effort. While this area is still relatively young, its extension to the multimodal domain holds immense potenti

2025

StableDepth: Scene-Consistent and Scale-Invariant Monocular Depth

ICCV 2025poster

Recent advances in monocular depth estimation significantly improve robustness and accuracy. However, relative depth models exhibit flickering and 3D inconsistency in video data, limiting 3D reconstruction applications. We introduce StableDepth, a scene-consistent and scale-invariant depth estimatio…

Cited by 0SourcePDFScholar
2024

A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing

CVPR 2024poster

This paper presents a video inversion approach for zero-shot video editing which models the input video with low-rank representation during the inversion process. The existing video editing methods usually apply the typical 2D DDIM inversion or naive spatial-temporal DDIM inversion before editing wh…

2024

AddMe: Zero-shot Group-photo Synthesis by Inserting People into Scenes

ECCV 2024poster

"While large text-to-image diffusion models have made significant progress in high-quality image generation, challenges persist when users insert their portraits into existing photos, especially group photos. Concretely, existing customization methods struggle to insert facial identities at desired…

2024

GPAvatar: Generalizable and Precise Head Avatar from Image(s)

ICLR 2024poster

Head avatar reconstruction, crucial for applications in virtual reality, online meetings, gaming, and film industries, has garnered substantial attention within the computer vision community. The fundamental objective of this field is to faithfully recreate the head avatar and precisely control expr…

2024

OMG: Occlusion-friendly Personalized Multi-concept Generation in Diffusion Models

ECCV 2024poster

"Personalization is an important topic in text-to-image generation, especially the challenging multi-concept personalization. Current multi-concept methods are struggling with identity preservation, occlusion, and the harmony between foreground and background. In this work, we propose OMG, an occlus…

2024

Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic Prompts

ICLR 2024poster

Recent text-to-3D generation methods achieve impressive 3D content creation capacity thanks to the advances in image diffusion models and optimizing strategies. However, current methods struggle to generate correct 3D content for a complex prompt in semantics, i.e., a prompt describing multiple inte…

Cited by 42SourcePDFScholar
2024

SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering

EMNLP 2024finding

Audio-Visual Question Answering (AVQA) is a challenging task that involves answering questions based on both auditory and visual information in videos. A significant challenge is interpreting complex multi-modal scenes, which include both visual objects and sound sources, and connecting them to the…

Cited by 0SourcePDFScholar
2024

SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark

ACL 2024short

The paper introduces SceMQA, a novel benchmark for scientific multimodal question answering at the college entrance level. It addresses a critical educational phase often overlooked in existing benchmarks, spanning high school to pre-college levels. SceMQA focuses on core science subjects including…

Cited by 5SourcePDFScholar
2024

Symbol as Points: Panoptic Symbol Spotting via Point-based Representation

ICLR 2024poster

This work studies the problem of panoptic symbol spotting, which is to spot and parse both countable object instances (windows, doors, tables, etc.) and uncountable stuff (wall, railing, etc.) from computer-aided design (CAD) drawings. Existing methods typically involve either rasterizing the vector…

2024

TOSS: High-quality Text-guided Novel View Synthesis from a Single Image

ICLR 2024poster

In this paper, we present TOSS, which introduces text to the task of novel view synthesis (NVS) from just a single RGB image. While Zero123 has demonstrated impressive zero-shot open-set NVS capabilities, it treats NVS as a pure image-to-image translation problem. This approach suffers from the cha…

Cited by 18SourcePDFScholar
2023

Dior-CVAE: Pre-trained Language Models and Diffusion Priors for Variational Dialog Generation

EMNLP 2023long findings

Current variational dialog models have employed pre-trained language models (PLMs) to parameterize the likelihood and posterior distributions. However, the Gaussian assumption made on the prior distribution is incompatible with these distributions, thus restricting the diversity of generated respons…

Cited by 0SourcecodeScholar
2023

DropMAE: Masked Autoencoders With Spatial-Attention Dropout for Tracking Tasks

CVPR 2023poster

In this paper, we study masked autoencoder (MAE) pretraining on videos for matching-based downstream tasks, including visual object tracking (VOT) and video object segmentation (VOS). A simple extension of MAE is to randomly mask out frame patches in videos and reconstruct the frame pixels. However,…

2023

Learning Deep Hierarchical Features with Spatial Regularization for One-Class Facial Expression Recognition

AAAI 2023technical

Existing methods on facial expression recognition (FER) are mainly trained in the setting when multi-class data is available. However, to detect the alien expressions that are absent during training, this type of methods cannot work. To address this problem, we develop a Hierarchical Spatial One Cla…

2023

UniMath: A Foundational and Multimodal Mathematical Reasoner

EMNLP 2023short main

While significant progress has been made in natural language processing (NLP), existing methods exhibit limitations in effectively interpreting and processing diverse mathematical modalities. Therefore, we introduce UniMath, a versatile and unified system designed for multimodal mathematical reasoni…

Cited by 0SourceScholar
2022

Exploring Denoised Cross-Video Contrast for Weakly-Supervised Temporal Action Localization

CVPR 2022poster

Weakly-supervised temporal action localization aims to localize actions in untrimmed videos with only video-level labels. Most existing methods address this problem with a "localization-by-classification" pipeline that localizes action regions based on snippet-wise classification sequences. Snippet-…

Cited by 75PDFcodeScholar
2022

LocVTP: Video-Text Pre-training for Temporal Localization

ECCV 2022poster

"Video-Text Pre-training (VTP) aims to learn transferable representations for various downstream tasks from large-scale web videos. To date, almost all existing VTP methods are limited to retrieval-based downstream tasks, e.g., video retrieval, whereas their transfer potentials on localization-based…

2022

Motion-Aware Contrastive Video Representation Learning via Foreground-Background Merging

CVPR 2022poster

In light of the success of contrastive learning in the image domain, current self-supervised video representation learning methods usually employ contrastive loss to facilitate video representation learning. When naively pulling two augmented views of a video closer, the model however tends to learn…

Cited by 71PDFcodeScholar
2022

SWEM: Towards Real-Time Video Object Segmentation With Sequential Weighted Expectation-Maximization

CVPR 2022poster

Matching-based methods, especially those based on space-time memory, are significantly ahead of other solutions in semi-supervised video object segmentation (VOS). However, continuously growing and redundant template features lead to an inefficient inference. To alleviate this, we propose a novel Se…

Cited by 58PDFcodeScholar
2022

Unsupervised Pre-Training for Temporal Action Localization Tasks

CVPR 2022poster

Unsupervised video representation learning has made remarkable achievements in recent years. However, most existing methods are designed and optimized for video classification. These pre-trained models can be sub-optimal for temporal localization tasks due to the inherent discrepancy between video-l…

Cited by 64PDFcodeScholar
2021

VideoMoCo: Contrastive Video Representation Learning With Temporally Adversarial Examples

CVPR 2021poster

MoCo is effective for unsupervised image representation learning. In this paper, we propose VideoMoCo for unsupervised video representation learning. Given a video sequence as an input sample, we improve the temporal feature representations of MoCo from two perspectives. First, we introduce a genera…

Cited by 295PDFcodeScholar