← Search

Li-Jia Li

20 accepted papers

2025

Symbolic Representation for Any-to-Any Generative Tasks

CVPR 2025poster

We propose a symbolic generative task description language and a corresponding inference engine that can represent arbitrary multimodal tasks as structured symbolic flows. Unlike conventional generative models, which rely on large-scale training and implicit neural representations to learn cross-mod…

2025

Towards AI-Assisted Psychotherapy: Emotion-Guided Generative Interventions

EMNLP 2025

Large language models (LLMs) hold promise for therapeutic interventions, yet most existing datasets rely solely on text, overlooking non-verbal emotional cues essential to real-world therapy. To address this, we introduce a multimodal dataset of 1,441 publicly sourced therapy session videos containi

Cited by 0SourcePDFScholar
2025

Video-Bench: Human-Aligned Video Generation Benchmark

CVPR 2025poster

Video generation assessment is essential for ensuring that generative models produce visually realistic, high-quality videos while aligning with human expectations. Current video generation benchmarks fall into two main categories: traditional benchmarks, which use metrics and embeddings to evaluate…

2024

Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations

ECCV 2024poster

"We introduce Affective Visual Dialog, an emotion explanation and reasoning task as a testbed for research on understanding constructed emotions in response to visually grounded conversations. The task involves three skills: (1) Dialog-based Question Answering (2) Dialog-based Emotion Prediction and…

Cited by 4SourcePDFScholar
2019

Composing Text and Image for Image Retrieval - an Empirical Odyssey

CVPR 2019oral

In this paper, we study the task of image retrieval, where the input query is specified in the form of an image plus some text that describes desired modifications to the input image. For example, we may present an image of the Eiffel tower, and ask the system to find images which are visually simil…

Cited by 442PDFScholar
2019

Eidetic 3D LSTM: A Model for Video Prediction and Beyond

ICLR 2019poster

Spatiotemporal predictive learning, though long considered to be a promising self-supervised feature learning method, seldom shows its effectiveness beyond future video prediction. The reason is that it is difficult to learn good representations for both short-term frame dependency and long-term hig…

Cited by 524SourcePDFScholar
2019

NOTE-RCNN: NOise Tolerant Ensemble RCNN for Semi-Supervised Object Detection

ICCV 2019poster

The labeling cost of large number of bounding boxes is one of the main challenges for training modern object detectors. To reduce the dependence on expensive bounding box annotations, we propose a new semi-supervised object detection formulation, in which a few seed box level annotations and a large…

Cited by 123PDFScholar
2018

AMC: AutoML for Model Compression and Acceleration on Mobile Devices

ECCV 2018poster

Model compression is an effective technique to efficiently deploy neural network models on mobile devices which have limited computation resources and tight power budgets. Conventional model compression techniques rely on hand-crafted features and require domain experts to explore the large design s…

2018

Distributed Asynchronous Optimization with Unbounded Delays: How Slow Can You Go?

ICML 2018oral

One of the most widely used optimization methods for large-scale machine learning problems is distributed asynchronous stochastic gradient descent (DASGD). However, a key issue that arises here is that of delayed gradients: when a “worker” node asynchronously contributes a gradient update to the “ma…

Cited by 72SourcePDFScholar
2018

Focal Visual-Text Attention for Visual Question Answering

CVPR 2018poster

Recent insights on language and vision with neural networks have been successfully applied to simple single-image visual question answering. However, to tackle real-life question answering problems on multimedia collections such as personal photos, we have to look at whole collections with sequences…

2018

MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels

ICML 2018oral

Recent deep networks are capable of memorizing the entire data even when the labels are completely random. To overcome the overfitting on corrupted labels, we propose a novel technique of learning another neural network, called MentorNet, to supervise the training of the base deep networks, namely,…

2018

Progressive Neural Architecture Search

ECCV 2018poster

We propose a new method for learning the structure of convolutional neural networks (CNNs) that is more efficient than recent state-of-the-art methods based on reinforcement learning and evolutionary algorithms. Our approach uses a sequential model-based optimization (SMBO) strategy, in which we sea…

2018

Thoracic Disease Identification and Localization With Limited Supervision

CVPR 2018poster

Accurate identification and localization of abnormalities from radiology images play an integral part in clinical diagnosis and treatment planning. Building a highly accurate prediction model for these tasks usually requires a large number of images manually annotated with labels and finding sites o…

Cited by 455SourcePDFScholar
2017

Deep Reinforcement Learning-Based Image Captioning With Embedding Reward

CVPR 2017oral

Image captioning is a challenging problem owing to the complexity in understanding the image content and diverse ways of describing it in natural language. Recent advances in deep neural networks have substantially improved the performance of this task. Most state-of-the-art approaches follow an enc…

Cited by 426PDFScholar
2015

Image Retrieval Using Scene Graphs

CVPR 2015poster

This paper develops a novel framework for semantic image retrieval based on the notion of a scene graph. Our scene graphs represent objects ("man", "boat"), attributes of objects ("boat is white") and relationships between objects ("man standing on boat"). We use these scene graphs as queries to ret…

Cited by 1399SourcePDFScholar