← Search

Hongxia Xie

9 accepted papers

2026

Beyond Success: Refining Elegant Robot Manipulation from Mixed-Quality Data via Just-in-Time Intervention

CVPR 2026

Vision-Language-Action (VLA) models have enabled notable progress in general-purpose robotic manipulation, yet their learned policies often exhibit variable execution quality. We attribute this variability to the mixed-quality nature of human demonstrations, where the implicit principles that govern

Cited by 0SourceScholar
2026

KPLM-STA: Physically-Accurate Shadow Synthesis for Human Relighting via Keypoint-Based Light Modeling

AAAI 2026technical

Image composition aims to seamlessly integrate a foreground object into a background, where generating realistic and geometrically accurate shadows remains a persistent challenge. While recent diffusion-based methods have outperformed GAN-based approaches, existing techniques, such as the diffusion-

Cited by 0SourcePDFScholar
2026

MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents

CVPR 2026

Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision-making, and existing benchmarks focus solely on human mental states while ignoring the agent's own perspective, hinderi

Cited by 0SourcecodeScholar
2026

TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection

CVPR 2026

Advances in generative modeling have made it increasingly easy to fabricate realistic portrayals of individuals, creating serious risks for security, communication, and public trust. Detecting such person-driven manipulations requires systems that not only distinguish altered content from authentic

Cited by 0SourceScholar
2025

Future Sight and Tough Fights: Revolutionizing Sequential Recommendation with FENRec

AAAI 2025technical

Sequential recommendation (SR) systems predict user preferences by analyzing time-ordered interaction sequences. A common challenge for SR is data sparsity, as users typically interact with only a limited number of items. While contrastive learning has been employed in previous approaches to addres…

2024

EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning

CVPR 2024poster

Visual Instruction Tuning represents a novel learning paradigm involving the fine-tuning of pre-trained language models using task-specific instructions. This paradigm shows promising zero-shot results in various natural language processing tasks but is still unexplored in vision emotion understandi…

2024

The Fabrication of Reality and Fantasy: Scene Generation with LLM-Assisted Prompt Interpretation

ECCV 2024poster

"In spite of recent advancements in text-to-image generation, limitations persist in handling complex and imaginative prompts due to the restricted diversity and complexity of training data. This work explores how diffusion models can generate images from prompts requiring artistic creativity or spe…

2023

Most Important Person-Guided Dual-Branch Cross-Patch Attention for Group Affect Recognition

ICCV 2023poster

Group affect refers to the subjective emotion that is evoked by an external stimulus in a group, which is an important factor that shapes group behavior and outcomes. Recognizing group affect involves identifying important individuals and salient objects among a crowd that can evoke emotions. Howeve…

Cited by 10PDFScholar