← Search

Jing Shi

36 accepted papers

2026

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

ICML 2026poster

Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally incomplete. Such reliance often yields unnatural or implausible outcomes, especially by missing secondary causal consequences. To address this, we intro…

Cited by 0SourceScholar
2026

Plot’n Polish: Zero-Shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models

AAAI 2026technical

Text-to-image diffusion models have demonstrated significant capabilities to generate diverse and detailed visuals in various domains, and story visualization is emerging as a particularly promising application. However, as their use in real-world creative domains increases, the need for providing e

Cited by 0SourcePDFScholar
2026

RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward

CVPR 2026

Recent advances in multimodal large language models (MLLMs) have shown great potential for extending vision-language reasoning to professional tool-based image editing, enabling intuitive and creative editing. A promising direction is to use reinforcement learning (RL) to enable MLLMs to reason abou

Cited by 0SourceScholar
2026

Seeing Through Words: Controlling Visual Retrieval Quality with Language

ICLR 2026poster

Text-to-image retrieval is a fundamental task in vision--language learning, yet in real-world scenarios it is often challenged by short and underspecified user queries. Such queries are typically only one or two words long, making them semantically ambiguous, prone to collisions across diverse visua…

Cited by 0SourcecodeScholar
2026

Seeing is Solving: Unlocking Efficient Multimodal RL via View Alignment

ICML 2026poster

Although Reinforcement Learning Fine-Tuning (RLFT) applied to Vision-Language Models (VLMs) substantially enhances multimodal reasoning capabilities, their prohibitive training cost limits broad adoption. Surprisingly, most existing methods simply port Large Language Model (LLM) RLFT techniques to V…

Cited by 0SourceScholar
2026

Stepwise Credit Assignment for GRPO on Flow-Matching Models

CVPR 2026

Flow-GRPO successfully applies reinforcement learning to flow models, but uses uniform credit assignment across all steps. This ignores the temporal structure of diffusion generation: early steps determine composition and content (low-frequency structure), while late steps resolve details and textur

Cited by 0SourceScholar
2026

Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play

ICLR 2026poster

Although reinforcement learning (RL) can effectively enhance the reasoning capabilities of vision–language models (VLMs), current methods remain heavily dependent on labor-intensive datasets that require extensive manual construction and verification, leading to extremely high training costs and con…

Cited by 0SourcecodeScholar
2026

VividCam: Learning Unconventional Camera Motions from Virtual Synthetic Videos

ICML 2026poster

Although recent text-to-video generative models are getting more capable of following external camera controls, imposed by either text descriptions or camera trajectories, they still struggle to generalize to unconventional camera motions, which is crucial in creating truly original and artistic vid…

Cited by 0SourceScholar
2025

DiffTell: A High-Quality Dataset for Describing Image Manipulation Changes

ICCV 2025poster

The image difference captioning (IDC) task is to describe the distinctions between two images. However, existing datasets do not offer comprehensive coverage across all image-difference categories. In this work, we introduce a high-quality dataset, DiffTell with various types of image manipulations,…

Cited by 0SourcePDFScholar
2025

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

CVPR 2025poster

The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate integration of visual and textual information across various applications, including image and video captioning, visual question answering, and cross-modal retrieva…

Cited by 6SourcePDFScholar
2025

Improving Large Vision and Language Models by Learning from a Panel of Peers

ICCV 2025poster

Traditional alignment methods for Large Vision and Language Models (LVLMs) primarily rely on human-curated preference data. Human-generated preference data is costly; machine-generated preference data is limited in quality; and self-supervised preference data often introduces hallucinations. To over…

2025

MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities

ACL 2025long

While originally designed for unidirectional generative modeling, decoder-only large language models (LLMs) are increasingly being adapted for bidirectional modeling. However, unidirectional and bidirectional models are typically trained separately with distinct objectives (generation and representa…

Cited by 0SourcePDFScholar
2025

Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters

AAAI 2025technical

Scaling Deep Neural Networks (DNNs) requires significant computational resources in terms of GPU quantity and compute capacity. In practice, there usually exists a large number of heterogeneous GPU devices due to the rapid release cycle of GPU products. It is highly needed to efficiently and economi…

Cited by 0SourcePDFScholar
2025

The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographers

CVPR 2025poster

Photographer, curator, and former director of photography at the Museum of Modern Art (MoMA), John Szarkowski remarked in *William Eggleston's Guide*, "While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky." Szarkowski insightfull…

Cited by 0SourcePDFScholar
2025

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage

ICML 2025poster

Multimodal large language models (MLLMs) excel at generating highly detailed captions but often produce hallucinations. Our analysis reveals that existing hallucination detection methods struggle with detailed captions. We attribute this to the increasing reliance of MLLMs on their generated text, r…

Cited by 1SourcePDFScholar
2025

Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference

NeurIPS 2025oral

Large Language Models (LLMs) are now integral across various domains and have demonstrated impressive performance. Progress, however, rests on the premise that benchmark scores are both accurate and reproducible. We demonstrate that the reproducibility of LLM performance is fragile: changing system…

Cited by 0SourcecodeScholar
2025

Visual Persona: Foundation Model for Full-Body Human Customization

CVPR 2025poster

We introduce Visual Persona, a foundation model for text-to-image full-body human customization that, given a single in-the-wild human image, generates diverse images of the individual guided by text descriptions. Unlike prior methods that focus solely on preserving facial identity, our approach cap…

Cited by 0SourcePDFScholar
2025

Yo'Chameleon: Personalized Vision and Language Generation

CVPR 2025poster

Large Multimodal Models (e.g., GPT-4, Gemini, Chameleon) have evolved into powerful tools with millions of users. However, they remain generic models and lack personalized knowledge of specific user concepts. Previous work has explored personalization for text generation, yet it remains unclear how…

Cited by 1SourcePDFScholar
2024

Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models

ECCV 2024poster

"Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterpart, motion customization, has not yet been well investigated. To address the c…

2024

FineMatch: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction

ECCV 2024poster

"Recent progress in large-scale pre-training has led to the development of advanced vision-language models (VLMs) with remarkable proficiency in comprehending and generating multimodal content. Despite the impressive ability to perform complex reasoning for VLMs, current models often struggle to eff…

2024

InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning

CVPR 2024poster

Recent advances in personalized image generation have enabled pre-trained text-to-image models to learn new concepts from specific image sets. However these methods often necessitate extensive test-time finetuning for each new concept leading to inefficiencies in both time and scalability. To addres…

Cited by 272SourcePDFScholar
2024

VIXEN: Visual Text Comparison Network for Image Difference Captioning

AAAI 2024technical

We present VIXEN - a technique that succinctly summarizes in text the visual differences between a pair of images in order to highlight any content manipulation present. Our proposed network linearly maps image features in a pairwise manner, constructing a soft prompt for a pretrained large language…

2024

ViLaS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition

ICASSP 2024accepted

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived from human lip motions. In fact, context-dependent visual and…

Cited by 0SourceScholar
2023

Matching-Based Term Semantics Pre-Training for Spoken Patient Query Understanding

ICASSP 2023accepted

Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficient term semantics learning makes existing approaches hard to capture semantically identical but colloquial expressions of…

Cited by 0SourceScholar
2022

SpaceEdit: Learning a Unified Editing Space for Open-Domain Image Color Editing

CVPR 2022poster

Recently, large pretrained models (e.g., BERT, StyleGAN, CLIP) show great knowledge transfer and generalization capability on various downstream tasks within their domains. Inspired by these efforts, in this paper we propose a unified model for open-domain image editing focusing on color and tone ad…

Cited by 19PDFScholar
2021

A Simple Baseline for Weakly-Supervised Scene Graph Generation

ICCV 2021poster

We investigate the weakly-supervised scene graph generation, which is a challenging task since no correspondence of label and object is provided. The previous work regards such correspondence as a latent variable which is iteratively updated via nested optimization of the scene graph generation obje…

Cited by 36PDFcodeScholar
2021

Language-Guided Global Image Editing via Cross-Modal Cyclic Mechanism

ICCV 2021poster

Editing an image automatically via a linguistic request can significantly save laborious manual work and is friendly to photography novice. In this paper, we focus on the task of language-guided global image editing. Existing works suffer from imbalanced data distribution of real-world datasets and…

Cited by 29PDFScholar
2021

Learning To Generate Scene Graph From Natural Language Supervision

ICCV 2021poster

Learning from image-text data has demonstrated recent success for many recognition tasks, yet is currently limited to visual features or individual visual concepts such as objects. In this paper, we propose one of the first methods that learn from image-sentence pairs to extract a graphical represen…

Cited by 87PDFcodeScholar
2021

Learning by Planning: Language-Guided Global Image Editing

CVPR 2021poster

Recently, language-guided global image editing draws increasing attention with growing application potentials. However, previous GAN-based methods are not only confined to domain-specific, low-resolution data but also lacking in interpretability. To overcome the collective difficulties, we develop a…

Cited by 40PDFcodeScholar
2021

Recent Developments on Espnet Toolkit Boosted By Conformer

ICASSP 2021accepted

In this study, we present recent developments on ESPnet: End-to- End Speech Processing toolkit, which mainly involves a recently proposed architecture called Conformer, Convolution-augmented Transformer. This paper shows the results for a wide range of end- to-end speech processing applications, suc…

Cited by 0SourceScholar
2021

Training Noisy Single-Channel Speech Separation with Noisy Oracle Sources: A Large Gap and a Small Step

ICASSP 2021accepted

As the performance of single-channel speech separation systems has improved, there has been a desire to move to more challenging conditions than the clean, near-field speech that initial systems were developed on. When training deep learning separation models, a need for ground truth leads to traini…

Cited by 0SourceScholar
2020

Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture Signals

NeurIPS 2020poster

Neural sequence-to-sequence models are well established for applications which can be cast as mapping a single input sequence into a single output sequence. In this work, we focus on one-to-many sequence transduction problems, such as extracting multiple sequential sources from a mixture sequence.…

2019

Not All Frames Are Equal: Weakly-Supervised Video Grounding With Contextual Similarity and Visual Clustering Losses

CVPR 2019poster

We invest the problem of weakly-supervised video grounding, where only video-level sentences are provided. This is a challenging task, and previous Multi-Instance Learning (MIL) based image grounding methods turn to fail in the video domain. Recent work attempts to decompose the video-level MIL int…

Cited by 61PDFScholar
2018

Audio-Visual Event Localization in Unconstrained Videos

ECCV 2018poster

In this paper, we introduce a novel problem of audio-visual event localization in unconstrained videos. We define an audio-visual event as an event that is both visible and audible in a video segment. We collect an Audio-Visual Event (AVE) dataset to systemically investigate three temporal localizati…

Cited by 575SourcePDFScholar