← Search

Yicong Li

29 accepted papers

2026

AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation

AAAI 2026technical

Optimization‐based text‑to‑3D methods distill guidance from 2D generative models via Score Distillation Sampling (SDS), but implicitly treat this guidance as static. This work shows that ignoring source dynamics yields inconsistent trajectories that suppress or merge semantic cues, leading to "seman

Cited by 0SourcePDFScholar
2026

DRSoRec: Dual-Rectification of Social Networks for Recommendation

AAAI 2026technical

Leveraging social homophily to enhance user preference modeling, social recommendation has become a cornerstone of modern recommender systems. However, the raw social network contains inherent unreliability as it teems with noise---misclicks, bot-generated and transient ties---while many meaningful

Cited by 0SourcePDFScholar
2026

EgoTwin: Dreaming Body and View in First Person

ICLR 2026poster

While exocentric video synthesis has achieved great progress, egocentric video generation remains largely underexplored, which requires modeling first-person view content along with camera motion patterns induced by the wearer's body movements. To bridge this gap, we introduce a novel task of joint…

Cited by 0SourceScholar
2026

FairTCD: Dual-Teacher Temporal Contrastive Distillation for Twofold Fair Dynamic Graph Embedding

IJCAI 2026

Fair dynamic graph embedding is crucial for real-world systems, such as recommendation and social networks. Prior studies impose a single-axis fairness formulation, treating attribute and structural bias as separable artifacts. This overlooks their coupling relationship, under which debiasing along

Cited by 0Scholar
2026

Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing

ICLR 2026poster

Textured 3D morphing seeks to generate smooth and plausible transitions between two 3D assets, preserving both structural coherence and fine-grained appearance. This ability is crucial not only for advancing 3D generation research but also for practical applications in animation, editing, and digita…

Cited by 0SourcecodeScholar
2026

SCOPE: Safety-Constrained Online Preview Enforcement for Efficient Encirclement in Multi-UAV Pursuit-Evasion

IJCAI 2026

Unmanned aerial vehicle swarms in pursuit-evasion requires encirclement efficiency while maintaining safety constraints, facing a critical safety-efficiency trade-off. Existing safe multi-agent reinforcement learning (MARL) methods often yield either unsafe task policies or conservative policies. Th

Cited by 0Scholar
2026

Shot-Conditioned Vision-Language Adaptation for Effective Harmful Content Detection from Online Short Videos

IJCAI 2026

Short video harmful content detection aims to automatically identify diverse anomalies from user-generated media. This task presents unique challenges due to frequent editing cuts and highly variable anomaly densities, limiting the effectiveness of traditional surveillance-based approaches. Moreover

Cited by 0Scholar
2026

VA-p: Variational Policy Alignment for Pixel-Aware Autoregressive Generation

CVPR 2026

Autoregressive (AR) visual generation relies on tokenizers to map images to and from discrete sequences. However, tokenizers are trained to reconstruct clean images from ground-truth tokens, while AR generators are optimized only for token likelihood. This misalignment leads to generated token seque

Cited by 0SourcecodeScholar
2026

VINCIE: Unlocking In-context Image Editing from Video

ICLR 2026poster

In-context image editing aims to modify images based on a contextual sequence comprising text and previously generated images. Existing methods typically depend on task-specific pipelines and expert models (e.g., segmentation and inpainting) to curate training data. In this work, we explore whether…

Cited by 0SourcecodeScholar
2026

reAR: Rethinking Visual Autoregressive Models via Token-wise Consistency Regularization

ICLR 2026poster

Visual autoregressive (AR) generation offers a promising path toward unifying vision and language models, yet its performance remains suboptimal against diffusion models. Prior work often attributes this gap to tokenizer limitations and rasterization ordering. In this work, we identify a core bottle…

Cited by 0SourceScholar
2025

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

CVPR 2025poster

We introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text aware questions that reflect real user needs in outdoor driving and indoor house-keeping activities. The questions are d…

2025

Factor Graph-based Interpretable Neural Networks

ICLR 2025poster

Comprehensible neural network explanations are foundations for a better understanding of decisions, especially when the input data are infused with malicious perturbations. Existing solutions generally mitigate the impact of perturbations through adversarial training, yet they fail to generate compr…

2025

Geometric Alignment and Prior Modulation for View-Guided Point Cloud Completion on Unseen Categories

ICCV 2025poster

View-Guided Point Cloud Completion (VG-PCC) aims to reconstruct complete point clouds from partial inputs by referencing single-view images. While existing VG-PCC models perform well on in-class predictions, they exhibit significant performance drops when generalizing to unseen categories. We identi…

Cited by 0SourcePDFScholar
2025

Intermediate Connectors and Geometric Priors for Language-Guided Affordance Segmentation on Unseen Object Categories

ICCV 2025poster

Language-guided Affordance Segmentation (LASO) aims to identify actionable object regions based on text instructions. At the core of its practicality is learning generalizable affordance knowledge that captures functional regions across diverse objects. However, current LASO solutions struggle to ex…

2025

MSCI: Addressing CLIP's Inherent Limitations for Compositional Zero-Shot Learning

IJCAI 2025

Compositional Zero-Shot Learning (CZSL) aims to recognize unseen state-object combinations by leveraging known combinations. Existing studies basically rely on the cross-modal alignment capabilities of CLIP but tend to overlook its limitations in capturing fine-grained local features, which arise fr

2025

Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language Models

NeurIPS 2025poster

Recent advances in Large Vision-Language Models (LVLMs) have showcased strong reasoning abilities across multiple modalities, achieving significant breakthroughs in various real-world applications. Despite this great success, the safety guardrail of LVLMs may not cover the unforeseen domains introdu…

Cited by 0SourcecodeScholar
2025

SynTag: Enhancing the Geometric Robustness of Inversion-based Generative Image Watermarking

ICCV 2025poster

Robustness is significant for generative image watermarking, typically achieved by injecting distortion-invariant watermark features. The leading paradigm, i.e., inversion-based framework, excels against non-geometric distortions but struggles with geometric ones. To address this, we propose SynTag,…

Cited by 0SourcePDFScholar
2025

Visual Intention Grounding for Egocentric Assistants

ICCV 2025poster

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are egocentric, and objects may be referred to implicitly through needs a…

2024

Can I Trust Your Answer? Visually Grounded Video Question Answering

CVPR 2024highlight

We study visually grounded VideoQA in response to the emerging trends of utilizing pretraining techniques for video- language understanding. Specifically by forcing vision- language models (VLMs) to answer questions and simultane- ously provide visual evidence we seek to ascertain the extent to whic…

2024

LASO: Language-guided Affordance Segmentation on 3D Object

CVPR 2024poster

Segmenting affordance in 3D data is key for bridging perception and action in robots. Existing efforts mostly focus on the visual side and overlook the affordance knowledge from a semantic aspect. This oversight not only limits their generalization to unseen objects but more importantly hinders thei…

2024

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

ACL 2024findings

Humans use multiple senses to comprehend the environment. Vision and language are two of the most vital senses since they allow us to easily communicate our thoughts and perceive the world around us. There has been a lot of interest in creating video-language understanding systems with human-like se…

2023

An Empirical Study Towards Prompt-Tuning for Graph Contrastive Pre-Training in Recommendations

NeurIPS 2023poster

Graph contrastive learning (GCL) has emerged as a potent technology for numerous graph learning tasks. It has been successfully applied to real-world recommender systems, where the contrastive loss and the downstream recommendation objectives are always combined to form the overall objective functio…

Cited by 10SourcePDFScholar
2023

Discovering Spatio-Temporal Rationales for Video Question Answering

ICCV 2023poster

This paper strives to solve complex video question answering (VideoQA) which features long videos containing multiple objects and events at different time. To tackle the challenge, we highlight the importance of identifying question-critical temporal moments and spatial objects from the vast amount…

Cited by 27PDFcodeScholar
2022

Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning

CVPR 2022oral

Vision Transformers (ViTs) and their multi-scale and hierarchical variations have been successful at capturing image representations but their use has been generally studied for low-resolution images (e.g. - 256x256, 384x384). For gigapixel whole-slide imaging (WSI) in computational pathology, WSIs…

Cited by 556PDFcodeScholar
2022

Video Question Answering: Datasets, Algorithms and Challenges

EMNLP 2022main

This survey aims to sort out the recent advances in video question answering (VideoQA) and point towards future directions. We firstly categorize the datasets into 1) normal VideoQA, multi-modal VideoQA and knowledge-based VideoQA, according to the modalities invoked in the question-answer pairs, or…

2022

Video as Conditional Graph Hierarchy for Multi-Granular Question Answering

AAAI 2022technical

Video question answering requires the models to understand and reason about both the complex video and language data to correctly derive the answers. Existing efforts have been focused on designing sophisticated cross-modal interactions to fuse the information from two modalities, while encoding the…