← Search

Kaixiong Gong

12 accepted papers

2026

JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization

CVPR 2026

Agent-based editing models have substantially advanced interactive experiences, processing quality, and creative flexibility. However, two critical challenges persist: (1) instruction hallucination--text-only chain-of-thought (CoT) reasoning cannot fully prevent factual errors due to inherent inform

Cited by 0SourcecodeScholar
2026

Twins: Learn to Predict Unified Representations with Focal Loss

ICML 2026poster

Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codebook, whereas continuous pipelines often rely on two disparate representations—semantic features (e.g., ViT) for understa…

Cited by 0SourceScholar
2025

Video-R1: Reinforcing Video Reasoning in MLLMs

NeurIPS 2025poster

Inspired by DeepSeek-R1's success in eliciting reasoning abilities through rule-based reinforcement learning (RL), we introduce Video-R1 as the first attempt to systematically explore the R1 paradigm for incentivizing video reasoning within multimodal large language models (MLLMs). However, directly…

Cited by 0SourcecodeScholar
2024

$\textit{Bifr\"ost}$: 3D-Aware Image Compositing with Language Instructions

NeurIPS 2024poster

This paper introduces $\textit{Bifröst}$, a novel 3D-aware framework that is built upon diffusion models to perform instruction-based image composition. Previous methods concentrate on image compositing at the 2D level, which fall short in handling complex spatial relationships ($\textit{e.g.}$, occ…

Cited by 0SourcePDFScholar
2024

Better Regression Makes Better Test-time Adaptive 3D Object Detection

ECCV 2024poster

"Domain Adaptation (DA) has been widely explored and made significant progress on cross-domain 3D tasks recently. Despite being effective, existing works fail to deal with rapidly changing domains due to the unpredictable test time scenarios and meanwhile fast response time requirement. Thus, we exp…

2024

Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities

CVPR 2024poster

We propose to improve transformers of a specific modality with irrelevant data from other modalities e.g. improve an ImageNet model with audio or point cloud datasets. We would like to highlight that the data samples of the target modality are irrelevant to the other modalities which distinguishes o…

2024

OneLLM: One Framework to Align All Modalities with Language

CVPR 2024poster

Multimodal large language models (MLLMs) have gained significant attention due to their strong multimodal understanding capability. However existing works rely heavily on modality-specific encoders which usually differ in architecture and are limited to common modalities. In this paper we present On…

2024

Subjective Topic meets LLMs: Unleashing Comprehensive, Reflective and Creative Thinking through the Negation of Negation

EMNLP 2024main

Large language models (LLMs) exhibit powerful reasoning capacity, as evidenced by prior studies focusing on objective topics that with unique standard answers such as arithmetic and commonsense reasoning. However, the reasoning to definite answers emphasizes more on logical thinking, and falls short…

Cited by 1SourcePDFScholar
2024

Text-to-3D Generation with Bidirectional Diffusion using both 2D and 3D priors

CVPR 2024poster

Most 3D generation research focuses on up-projecting 2D foundation models into the 3D space either by minimizing 2D Score Distillation Sampling (SDS) loss or fine-tuning on multi-view datasets. Without explicit 3D priors these methods often lead to geometric anomalies and multi-view inconsistency. R…

2021

MetaSAug: Meta Semantic Augmentation for Long-Tailed Visual Recognition

CVPR 2021poster

Real-world training data usually exhibits long-tailed distribution, where several majority classes have a significantly larger number of samples than the remaining minority classes. This imbalance degrades the performance of typical supervised learning algorithms designed for balanced training sets.…

Cited by 201PDFcodeScholar
2021

Transferable Semantic Augmentation for Domain Adaptation

CVPR 2021poster

Domain adaptation has been widely explored by transferring the knowledge from a label-rich source domain to a related but unlabeled target domain. Most existing domain adaptation algorithms attend to adapting feature representations across two domains with the guidance of a shared source-supervised…

Cited by 165PDFcodeScholar