← Search

Ning Yu

41 accepted papers

2026

Vista4D: Video Reshooting with 4D Point Clouds

CVPR 2026

We present **Vista4D**, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re-synthesizes the scene with the same dynamics from a different camera trajectory and viewpoint. Existing vide

Cited by 0SourcecodeScholar
2025

DREAM: Improving Video-Text Retrieval Through Relevance-Based Augmentation Using Large Foundation Models

NAACL 2025long

Recent progress in video-text retrieval has been driven largely by advancements in model architectures and training strategies. However, the representation learning capabilities of video-text retrieval models remain constrained by low-quality and limited training data annotations. To address this is…

Cited by 0SourcePDFScholar
2025

Detecting Adversarial Data Using Perturbation Forgery

CVPR 2025poster

As a defense strategy against adversarial attacks, adversarial detection aims to identify and filter out adversarial data from the data flow based on discrepancies in distribution and noise patterns between natural and adversarial data. Although previous detection methods achieve high performance in…

2025

FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution

ICCV 2025poster

A versatile video depth estimation model should be consistent and accurate across frames, produce high-resolution depth maps, and support real-time streaming. We propose a method, FlashDepth, that satisfies all three requirements, performing depth estimation for a 2044x1148 streaming video at 24 FPS…

2025

Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise

CVPR 2025poster

Generative modeling aims to transform random noise into structured outputs. In this work, we enhance video diffusion models by allowing motion control via structured latent noise sampling. This is achieved by just a change in data: we pre-process training videos to yield structured noise. Consequent…

2025

Infinite-Resolution Integral Noise Warping for Diffusion Models

ICLR 2025poster

Adapting pretrained image-based diffusion models to generate temporally consistent videos has become an impactful generative modeling research direction. Training-free noise-space manipulation has proven to be an effective technique, where the challenge is to preserve the Gaussian white noise distri…

Cited by 1SourcePDFScholar
2025

Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset

CVPR 2025poster

Video portrait relighting remains challenging because the results need to be both photorealistic and temporally stable.This typically requires a strong model design that can capture complex facial reflections as well as intensive training on a high-quality paired video dataset, such as dynamic one-l…

Cited by 1SourcePDFScholar
2025

Reference-Based 3D-Aware Image Editing with Triplanes

CVPR 2025highlight

Generative Adversarial Networks (GANs) have emerged as powerful tools for high-quality image generation and real image editing by manipulating their latent spaces. Recent advancements in GANs include 3D-aware models such as EG3D, which feature efficient triplane-based architectures capable of recons…

Cited by 5SourcePDFScholar
2025

Text2Data: Low-Resource Data Generation with Textual Control

AAAI 2025technical

Natural language serves as a common and straightforward control signal for humans to interact seamlessly with machines. Recognizing the importance of this interface, the machine learning community is investing considerable effort in generating data that is semantically coherent with textual instruct…

2024

"X-InstructBLIP: A Framework for Aligning Image, 3D, Audio, Video to LLMs and its Emergent Cross-modal Reasoning"

ECCV 2024poster

"Recent research has achieved significant advancements in visual reasoning tasks through learning image-to-language projections and leveraging the impressive reasoning abilities of Large Language Models (LLMs). This paper introduces an efficient and effective framework that integrates multiple modal…

2024

C-RAG: Certified Generation Risks for Retrieval-Augmented Language Models

ICML 2024poster

Despite the impressive capabilities of large language models (LLMs) across diverse applications, they still suffer from trustworthiness issues, such as hallucinations and misalignments. Retrieval-augmented language models (RAG) have been proposed to enhance the credibility of generations by groundin…

2024

FASTTRACK: Reliable Fact Tracing via Clustering and LLM-Powered Evidence Validation

EMNLP 2024finding

Fact tracing seeks to identify specific training examples that serve as the knowledge source for a given query. Existing approaches to fact tracing rely on assessing the similarity between each training sample and the query along a certain dimension, such as lexical similarity, gradient, or embeddin…

2024

HIVE: Harnessing Human Feedback for Instructional Visual Editing

CVPR 2024poster

Incorporating human feedback has been shown to be crucial to align text generated by large language models to human preferences. We hypothesize that state-of-the-art instructional image editing models where outputs are generated based on an input image and an editing instruction could similarly bene…

2024

Hierarchical Point Attention for Indoor 3D Object Detection

ICRA 2024poster

3D object detection is an essential vision technique for various robotic systems, such as augmented reality and domestic robots. Transformers as versatile network architectures have recently seen great success in 3D point cloud object detection. However, the lack of hierarchy in a plain transformer…

Cited by 1SourceScholar
2024

RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content

ICML 2024poster

Recent advancements in Large Language Models (LLMs) have showcased remarkable capabilities across various tasks in different domains. However, the emergence of biases and the potential for generating harmful content in LLMs, particularly under malicious inputs, pose significant challenges. Current m…

2024

Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language Models

NeurIPS 2024poster

Vision-Language Models (VLMs) excel in generating textual responses from visual inputs, but their versatility raises security concerns. This study takes the first step in exposing VLMs’ susceptibility to data poisoning attacks that can manipulate responses to innocuous, everyday prompts. We introduc…

2024

SimSCOOD: Systematic Analysis of Out-of-Distribution Generalization in Fine-tuned Source Code Models

NAACL 2024findings

Large code datasets have become increasingly accessible for pre-training source code models. However, for the fine-tuning phase, obtaining representative training data that fully covers the code distribution for specific downstream tasks remains challenging due to the task-specific nature and limite…

2024

T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition

NeurIPS 2024poster

While widespread access to the Internet and the rapid advancement of generative models boost people's creativity and productivity, the risk of encountering inappropriate or harmful content also increases. To address the aforementioned issue, researchers managed to incorporate several harmful content…

2024

ULIP-2: Towards Scalable Multimodal Pre-training for 3D Understanding

CVPR 2024poster

Recent advancements in multimodal pre-training have shown promising efficacy in 3D representation learning by aligning multimodal features across 3D shapes their 2D counterparts and language descriptions. However the methods used by existing frameworks to curate such multimodal data in particular la…

2023

Can't Steal? Cont-Steal! Contrastive Stealing Attacks Against Image Encoders

CVPR 2023poster

Self-supervised representation learning techniques have been developing rapidly to make full use of unlabeled images. They encode images into rich features that are oblivious to downstream tasks. Behind their revolutionary representation power, the requirements for dedicated model designs and a mass…

2023

Detecting Adversarial Faces Using Only Real Face Self-Perturbations

IJCAI 2023poster

Adversarial attacks aim to disturb the functionality of a target system by adding specific noise to the input samples, bringing potential threats to security and robustness when applied to facial recognition systems. Although existing defense techniques achieve high accuracy in detecting some specif…

2023

GlueGen: Plug and Play Multi-modal Encoders for X-to-image Generation

ICCV 2023poster

Text-to-image (T2I) models based on diffusion processes have achieved remarkable success in controllable image generation using user-provided captions. However, the tight coupling between the current text encoder and image decoder in T2I models makes it challenging to replace or upgrade. Such change…

Cited by 26PDFcodeScholar
2023

Learning Prototype Classifiers for Long-Tailed Recognition

IJCAI 2023poster

The problem of long-tailed recognition (LTR) has received attention in recent years due to the fundamental power-law distribution of objects in the real-world. Most recent works in LTR use softmax classifiers that are biased in that they correlate classifier norm with the amount of training data for…

2023

Mask-Free OVIS: Open-Vocabulary Instance Segmentation Without Manual Mask Annotations

CVPR 2023poster

Existing instance segmentation models learn task-specific information using manual mask annotations from base (training) categories. These mask annotations require tremendous human effort, limiting the scalability to annotate novel (new) categories. To alleviate this problem, Open-Vocabulary (OV) me…

2023

UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

NeurIPS 2023poster

Achieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when prompted with arbitrary languages. However, they often fall…

2022

RelaxLoss: Defending Membership Inference Attacks without Losing Utility

ICLR 2022spotlight

As a long-term threat to the privacy of training data, membership inference attacks (MIAs) emerge ubiquitously in machine learning models. Existing works evidence strong connection between the distinguishability of the training and testing loss distributions and the model's vulnerability to MIAs. Mo…

2022

RepMix: Representation Mixing for Robust Attribution of Synthesized Images

ECCV 2022poster

"Rapid advances in Generative Adversarial Networks (GANs) raise new challenges for image attribution; detecting whether an image is synthetic and, if so, determining which GAN architecture created it. Uniquely, we present a solution to this task capable of 1) matching images invariant to their seman…

2022

Responsible Disclosure of Generative Models Using Scalable Fingerprinting

ICLR 2022spotlight

Over the past years, deep generative models have achieved a new level of performance. Generated data has become difficult, if not impossible, to be distinguished from real data. While there are plenty of use cases that benefit from this technology, there are also strong concerns on how this new tech…

2021

Artificial Fingerprinting for Generative Models: Rooting Deepfake Attribution in Training Data

ICCV 2021poster

Photorealistic image generation has reached a new level of quality due to the breakthroughs of generative adversarial networks (GANs). Yet, the dark side of such deepfakes, the malicious use of generated media, raises concerns about visual misinformation. While existing research work on deepfake det…

Cited by 263PDFcodeScholar
2021

Dual Contrastive Loss and Attention for GANs

ICCV 2021poster

Generative Adversarial Networks (GANs) produce impressive results on unconditional image generation when powered with large-scale image datasets. Yet generated images are still easy to spot especially on datasets with high variance (e.g. bedroom, church). In this paper, we propose various improvemen…

Cited by 71PDFcodeScholar
2020

Inclusive GAN: Improving Data and Minority Coverage in Generative Models

ECCV 2020poster

Generative Adversarial Networks (GANs) have brought about rapid progress towards generating photorealistic images. Yet the equitable allocation of their modeling capacity among subgroups has received less attention, which could lead to potential biases against underrepresented minorities if left unc…

2019

Texture Mixer: A Network for Controllable Synthesis and Interpolation of Texture

CVPR 2019poster

This paper addresses the problem of interpolating visual textures. We formulate this problem by requiring (1) by-example controllability and (2) realistic and smooth interpolation among an arbitrary number of texture samples. To solve it we propose a neural network trained simultaneously on a recons…

Cited by 54PDFcodeScholar
2016

Automated pick-up of carbon nanotubes inside a scanning electron microscope

IROS 2016poster

It is of great importance to pick up a single carbon nanotube (CNT) from a bulk of CNTs for nanodevice fabrication. In this study, we have proposed a nanorobotic manipulation system allowing automated pick-up of CNTs based on visual feedback. We utilize histogram normalization for automatic binariza…

Cited by 2SourceScholar
2016

Microbubbles for High-Speed Assembly of Cell-Laden Vascular-Like Microtube

RA-L 2016

Vascular-like microtube takes an important role in delivering oxygen and nutrient to keep cell alive in the generated tissue. In this paper, we present an automated micromanipulation system to assemble 2-D gel micro-rings to vascular-like microtubes by means of generating and controlling microbubble

Cited by 1SourceScholar
2015

Automated bubble-based assembly of cell-laden microgels into vascular-like microtubes

IROS 2015poster

Fabrication of artificial blood vessels in micro scale significantly benefits the regeneration of functional human vascular networks. In this paper, we develop an efficient multi-microrobotic system with an innovative motorized sample holder (MSH) and two manipulators. Air is injected into the solut…

Cited by 3SourceScholar