← Search

Jun Guo

28 accepted papers

2026

FlowDreamer: A RGB-D World Model With Flow-Based Motion Representations for Robot Manipulation

RA-L 2026

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions. Specifically, we consider world models that operate on RGB-D frames (RGB-D world models). As opposed to canon

Cited by 12SourcecodeScholar
2026

FlowDreamer: A RGB-D World Model with Flow-Based Motion Representations for Robot Manipulation

ICRA 2026poster

This paper investigates training better visual world models for robot manipulation, i.e., models that can predict future visual observations by conditioning on past frames and robot actions. Specifically, we consider world models that operate on RGB-D frames (RGB-D world models). As opposed to canon…

2025

ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer

CVPR 2025poster

The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos, failing to transfer specific subject motion; 2) struggle to…

2025

Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding

ICCV 2025poster

This paper investigates the problem of understanding dynamic 3D scenes from egocentric observations, a key challenge in robotics and embodied AI. Unlike prior studies that explored this as long-form video understanding and utilized egocentric video only, we instead propose an LLM-based agent, Embodi…

Cited by 0SourcePDFScholar
2025

Empirical Study on Robustness and Resilience in Cooperative Multi-Agent Reinforcement Learning

NeurIPS 2025poster

In cooperative Multi-Agent Reinforcement Learning (MARL), it is a common practice to tune hyperparameters in ideal simulated environments to maximize cooperative performance. However, policies tuned for cooperation often fail to maintain robustness and resilience under real-world uncertainties. Buil…

Cited by 0SourceScholar
2025

Large Language Model-Empowered Adversarial Fusion for Typhoon Track Prediction

ICASSP 2025accepted

Accurate prediction of typhoon tracks is essential for effective disaster prevention strategies. Given that typhoon tracks can be conceptualized as a special class of time series, promising results have been achieved via the learning transferability of large language models (LLMs) in time series for…

Cited by 0SourceScholar
2025

MFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and Forecasting

CVPR 2025poster

Deep learning approaches for marine fog detection and forecasting have outperformed traditional methods, demonstrating significant scientific and practical importance. However, the limited availability of open-source datasets remains a major challenge. Existing datasets, often focused on a single re…

2024

Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video Understanding

NeurIPS 2024poster

With the emergence of large pre-trained multimodal video models, multiple benchmarks have been proposed to evaluate model capabilities. However, most of the benchmarks are human-centric, with evaluation data and tasks centered around human applications. Animals are an integral part of the natural wo…

2024

Benchmarking Segmentation Models with Mask-Preserved Attribute Editing

CVPR 2024poster

When deploying segmentation models in practice it is critical to evaluate their behaviors in varied and complex scenes. Different from the previous evaluation paradigms only in consideration of global attribute variations (e.g. adverse weather) we investigate both local and global attribute variatio…

2024

Byzantine Robust Cooperative Multi-Agent Reinforcement Learning as a Bayesian Game

ICLR 2024poster

In this study, we explore the robustness of cooperative multi-agent reinforcement learning (c-MARL) against Byzantine failures, where any agent can enact arbitrary, worst-case actions due to malfunction or adversarial attack. To address the uncertainty that any agent can be adversarial, we propose a…

2024

Dual-Prior Augmented Decoding Network for Long Tail Distribution in HOI Detection

AAAI 2024technical

Human object interaction detection aims at localizing human-object pairs and recognizing their interactions. Trapped by the long-tailed distribution of the data, existing HOI detection methods often have difficulty recognizing the tail categories. Many approaches try to improve the recognition of HO…

2023

Bi-directional Feature Reconstruction Network for Fine-Grained Few-Shot Image Classification

AAAI 2023technical

The main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting --…

2023

Semantic Memory Guided Image Representation for Polyp Segmentation

ICASSP 2023accepted

Polyp segmentation is important in the early diagnosis and treatment of colorectal cancer. Since polyps vary in shape, size, color, and texture, accurate polyp segmentation is very challenging. One promising solution is to model the contextual relation for each pixel. However, previous methods only…

Cited by 0SourceScholar
2022

Learning Invariant Visual Representations for Compositional Zero-Shot Learning

ECCV 2022poster

"Compositional Zero-Shot Learning (CZSL) aims to recognize novel compositions using knowledge learned from seen attribute-object compositions in the training set. Previous works mainly project an image and a composition into a common embedding space to measure their compatibility score. However, bot…

2021

A Joint Model for Dropped Pronoun Recovery and Conversational Discourse Parsing in Chinese Conversational Speech

ACL 2021long

In this paper, we present a neural model for joint dropped pronoun recovery (DPR) and conversational discourse parsing (CDP) in Chinese conversational speech. We show that DPR and CDP are closely related, and a joint model benefits both tasks. We refer to our model as DiscProReco, and it first encod…

2021

Give the Truth: Incorporate Semantic Slot into Abstractive Dialogue Summarization

EMNLP 2021finding

Abstractive dialogue summarization suffers from a lots of factual errors, which are due to scattered salient elements in the multi-speaker information interaction process. In this work, we design a heterogeneous semantic slot graph with a slot-level mask cross-attention to enhance the slot features…

Cited by 12SourcePDFScholar
2021

Joint Topology-Preserving and Feature-Refinement Network for Curvilinear Structure Segmentation

ICCV 2021poster

Curvilinear structure segmentation (CSS) is under semantic segmentation, whose applications include crack detection, aerial road extraction, and biomedical image segmentation. In general, geometric topology and pixel-wise features are two critical aspects of CSS. However, most semantic segmentation…

Cited by 54PDFScholar
2021

Your "Flamingo" is My "Bird": Fine-Grained, or Not

CVPR 2021poster

Whether what you see in Figure 1 is a "flamingo" or a "bird", is the question we ask in this paper. While fine-grained visual classification (FGVC) strives to arrive at the former, for the majority of us non-experts just "bird" would probably suffice. The real question is therefore -- how can we tai…

Cited by 145PDFcodeScholar
2020

Fine-Grained Visual Classification via Progressive Multi-Granularity Training of Jigsaw Patches

ECCV 2020poster

Fine-grained visual classification (FGVC) is much more challenging than traditional classification tasks due to the inherently subtle intra-class object variations. Recent works mainly tackle this problem by focusing on how to locate the most discriminative parts, more complementary parts, and parts o…

2020

Improving Abstractive Dialogue Summarization with Graph Structures and Topic Words

COLING 2020main

Recently, people have been beginning paying more attention to the abstractive dialogue summarization task. Since the information flows are exchanged between at least two interlocutors and key elements about a certain event are often spanned across multiple utterances, it is necessary for researchers…

Cited by 64SourcePDFScholar
2019

Multi-Order Information for Working Set Selection of Sequential Minimal Optimization

AISTATS 2019poster

A new working set selection method for sequential minimal optimization (SMO) is proposed in this paper. Instead of the method adopted in the current version of LIBSVM, which uses the second order information of the objective function to choose the violating pairs, we suggest a new method where a hig…

Cited by 3SourcePDFScholar
2019

Self-Supervised Convolutional Subspace Clustering Network

CVPR 2019poster

Subspace clustering methods based on data self-expression have become very popular for learning from data that lie in a union of low-dimensional linear subspaces. However, the applicability of subspace clustering has been limited because practical visual data in raw form do not necessarily lie in su…

Cited by 199PDFScholar
2018

SketchMate: Deep Hashing for Million-Scale Human Sketch Retrieval

CVPR 2018poster

We propose a deep hashing framework for sketch retrieval that, for the first time, works on a multi-million scale human sketch dataset.Leveraging on this large dataset, we explore a few sketch-specific traits that were otherwise under-studied in prior literature. Instead of following the conventiona…

Cited by 149SourcePDFScholar
2015

Learning Semi-Supervised Representation Towards a Unified Optimization Framework for Semi-Supervised Learning

ICCV 2015poster

State of the art approaches for Semi-Supervised Learning (SSL) usually follow a two-stage framework -- constructing an affinity matrix from the data and then propagating the partial labels on this affinity matrix to infer those unknown labels. While such a two-stage framework has been successful in…

Cited by 46PDFScholar
2015

Making Better Use of Edges via Perceptual Grouping

CVPR 2015poster

We propose a perceptual grouping framework that organizes image edges into meaningful structures and demonstrate its usefulness on various computer vision tasks. Our grouper formulates edge grouping as a graph partition problem, where a learning to rank method is developed to encode probabilities of…

Cited by 105SourcePDFScholar