← Search

Jie Cao

28 accepted papers

2026

CoGrad3D: Spatially-Coupled Timestep Optimization with Orthogonal Gradient Fusion for 3D Generation

AAAI 2026technical

Score Distillation Sampling has driven recent advances in text-to-3D generation. However, current approaches often fail to produce 3D assets that are both rich in detail and consistent across viewpoints. These limitations primarily arise from imbalanced guidance on fine-grained details and an overde

Cited by 0SourcePDFScholar
2026

MARKSWEEP: A NO-BOX REMOVAL ATTACK ON AI-GENERATED IMAGE WATERMARKING VIA NOISE INTENSIFICATION AND FREQUENCY-AWARE DENOISING

ICASSP 2026poster

AI watermarking embeds invisible signals within images to provide provenance information and identify content as AI-generated. In this paper, we introduce MarkSweep, a novel watermark removal attack that effectively erases the embedded watermarks from AI-generated images without degrading visual qua…

Cited by 0SourcePDFScholar
2026

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis

ICML 2026poster

The advancement of Medical Vision-Language Models (VLMs) for 3D Computed Tomography (CT) analysis is hindered by a misalignment between optimization objectives and clinical rigor. Current Reinforcement Learning (RL) paradigms rely on lexical proxy signals that induce ``\textbf{evaluation hallucinati…

Cited by 0SourceScholar
2026

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

ICML 2026poster

Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diversity due to the over-incentivization of positive rewards. Although methods like Negative Sample Reinforcement (NSR) mitigate this issue by upweighting…

Cited by 0SourceScholar
2026

ResT: Reshaping Token-Level Policy Gradients for Tool-Use Large Language Models

ICLR 2026poster

Large language models (LLMs) transcend passive generation and act as goal-directed agents by invoking external tools. Reinforcement learning (RL) offers a principled framework for optimizing these emergent tool-use policies, yet the prevailing paradigm relies exclusively on sparse outcome rewards an…

Cited by 0SourcecodeScholar
2026

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

ICML 2026poster

The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. However, vlog assessment is highly subjective and remains challenging due to a lack of standardized criteria, dataset and benchmark, and effective rewar…

Cited by 0SourceScholar
2025

Breaking Mental Set to Improve Reasoning through Diverse Multi-Agent Debate

ICLR 2025poster

Large Language Models (LLMs) have seen significant progress but continue to struggle with persistent reasoning mistakes. Previous methods of *self-reflection* have been proven limited due to the models’ inherent fixed thinking patterns. While Multi-Agent Debate (MAD) attempts to mitigate this by in…

2025

Do LLMs Encode Frame Semantics? Evidence from Frame Identification

EMNLP 2025

We investigate whether large language models encode latent knowledge of frame semantics, focusing on frame identification, a core challenge in frame semantic parsing that involves selecting the appropriate semantic frame for a target word in context. Using the FrameNet lexical resource, we evaluate

2025

Dual-PST: Dual-Branch SpatioTemporal-Planar Network for Video Forgery Detection

ICASSP 2025accepted

With the advancement of generative AI, distinguishing real and AI-generated faces in videos has become increasingly challenging. However, traditional methods struggle to capture local details and temporal dynamics simultaneously, making it difficult to achieve high detection accuracy while maintaini…

Cited by 0SourceScholar
2025

Enhancing Talk Moves Analysis in Mathematics Tutoring through Classroom Teaching Discourse

COLING 2025main

Human tutoring interventions play a crucial role in supporting student learning, improving academic performance, and promoting personal growth. This paper focuses on analyzing mathematics tutoring discourse using talk moves—a framework of dialogue acts grounded in Accountable Talk theory. However, s…

Cited by 2SourcePDFScholar
2025

Simulating Classroom Education with LLM-Empowered Agents

NAACL 2025long

Large language models (LLMs) have been applied across various intelligent educational tasks to assist teaching. While preliminary studies have focused on task-specific, independent LLM-empowered agents, the potential of LLMs within a multi-agent collaborative framework for classroom simulation with…

2025

Towards Robust Defense against Customization via Protective Perturbation Resistant to Diffusion-based Purification

ICCV 2025poster

Diffusion models like Stable Diffusion have become prominent in visual synthesis tasks due to their powerful customization capabilities, which also introduce significant security risks, including deepfakes and copyright infringement. In response, a class of methods known as protective perturbation e…

Cited by 0SourcePDFScholar
2024

Effective Trajectory Generation for Robots on General 3D Curved Surface

RA-L 2024

More and more robots are required to adsorb or crawl on the 3D curved surface in order to assist humans in some dangerous or tedious tasks. Existing methods on curved surface are merely able to plan in 2.5D environment at best, limiting the applications of the robots. In this letter, we propose an e

Cited by 0SourceScholar
2024

Hallo3D: Multi-Modal Hallucination Detection and Mitigation for Consistent 3D Content Generation

NeurIPS 2024poster

Recent advancements in 3D content generation have been significant, primarily due to the visual priors provided by pretrained diffusion models. However, large 2D visual models exhibit spatial perception hallucinations, leading to multi-view inconsistency in 3D content generated through Score Distill…

Cited by 1SourcePDFScholar
2023

Bi-directional Feature Reconstruction Network for Fine-Grained Few-Shot Image Classification

AAAI 2023technical

The main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting --…

2023

Mind the Gap between the Application Track and the Real World

ACL 2023short

Recent advances in NLP have led to a rise in inter-disciplinary and application-oriented research. While this demonstrates the growing real-world impact of the field, research papers frequently feature experiments that do not account for the complexities of realistic data and environments. To explor…

Cited by 2SourcePDFScholar
2023

Semi-Offline Reinforcement Learning for Optimized Text Generation

ICML 2023poster

Existing reinforcement learning (RL) mainly utilize online or offline settings. The online methods explore the environment with expensive time cost, and the offline methods efficiently obtain reward signals by sacrificing the exploration capability. We propose semi-offline RL, a novel paradigm that…

2022

An Augmented Benchmark Dataset for Geometric Question Answering through Dual Parallel Text Encoding

COLING 2022main

Automatic math problem solving has attracted much attention of NLP researchers recently. However, most of the works focus on the solving of Math Word Problems (MWPs). In this paper, we study on the Geometric Problem Solving based on neural networks. Solving geometric problems requires the integratio…

2022

OTSeq2Set: An Optimal Transport Enhanced Sequence-to-Set Model for Extreme Multi-label Text Classification

EMNLP 2022main

Extreme multi-label text classification (XMTC) is the task of finding the most relevant subset labels from an extremely large-scale label collection. Recently, some deep learning models have achieved state-of-the-art results in XMTC tasks. These models commonly predict scores for all labels by a ful…

2021

FaceInpainter: High Fidelity Face Adaptation to Heterogeneous Domains

CVPR 2021poster

In this work, we propose a novel two-stage framework named FaceInpainter to implement controllable Identity-Guided Face Inpainting (IGFI) under heterogeneous domains. Concretely, by explicitly disentangling foreground and background of the target face, the first stage focuses on adaptive face fittin…

Cited by 45PDFScholar
2021

ReMix: Towards Image-to-Image Translation With Limited Data

CVPR 2021poster

Image-to-image (I2I) translation methods based on generative adversarial networks (GANs) typically suffer from overfitting when limited training data is available. In this work, we propose a data augmentation method (ReMix) to tackle this issue. We interpolate training samples at the feature level a…

Cited by 37PDFcodeScholar
2020

Informative Sample Mining Network for Multi-Domain Image-to-Image Translation

ECCV 2020poster

The performance of multi-domain image-to-image translation has been significantly improved by recent progress in deep generative models. Existing approaches can use a unified model to achieve translations between all the visual domains. However, their outcomes are far from satisfying when there are…

Cited by 9SourcePDFScholar
2020

PSGAN: Pose and Expression Robust Spatial-Aware GAN for Customizable Makeup Transfer

CVPR 2020oral

In this paper, we address the makeup transfer task, which aims to transfer the makeup from a reference image to a source image. Existing methods have achieved promising progress in constrained scenarios, but transferring between images with large pose and expression differences is still challenging.…

Cited by 181PDFcodeScholar
2018

Learning a High Fidelity Pose Invariant Model for High-resolution Face Frontalization

NeurIPS 2018poster

Face frontalization refers to the process of synthesizing the frontal view of a face from a given profile. Due to self-occlusion and appearance distortion in the wild, it is extremely challenging to recover faithful results and preserve texture details in a high-resolution. This paper proposes a Hi…

Cited by 113SourcePDFScholar