← Search

Xiaolong Wang

181 accepted papers

2026

3D Aware Region Prompted Vision Language Model

ICLR 2026poster

We present Spatial Region 3D (SR-3D) aware vision-language model that connects single-view 2D images and multi-view 3D data through a shared visual token space. SR-3D supports flexible region prompting, allowing users to annotate regions with bounding boxes, segmentation masks on any frame, or direc…

Cited by 0SourcecodeScholar
2026

Cross-Embodied Co-Design for Dexterous Hands

ICLR 2026poster

Dexterous manipulation is limited by both control and design, without consensus as to what makes manipulators best for performing dexterous tasks. This raises a fundamental challenge: how should we design and control robot manipulators that are optimized for dexterity? We present a co-design framewo…

Cited by 0SourcecodeScholar
2026

Cross-Hand Latent Representation for Vision-Language-Action Models

CVPR 2026

Dexterous manipulation is essential for real-world robot autonomy, mirroring the central role of human hand coordination in daily activity. Humans rely on rich multimodal perception--vision, sound, and language-guided intent--to perform dexterous actions, motivating vision-based, language-conditione

Cited by 0SourceScholar
2026

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

RSS 2026poster

Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday environments. However, existing human datasets are often limi…

Cited by 0SourceScholar
2026

ExBody2: Advanced Expressive Humanoid Whole-Body Control

ICRA 2026poster

This paper tackles the challenge of enabling real-world humanoid robots to perform expressive and dynamic whole-body motions while maintaining stability. We propose ExBody2, a whole-body tracking framework trained in simulation with Reinforcement Learning and then transferred to the real world. The …

2026

GSWorld: Closed-Loop Photo-Realistic Simulation Suite for Robotic Manipulation

ICRA 2026poster

This paper presents GSWorld, a robust, photo-realistic simulator for robotics manipulation that combines 3D Gaussian Splatting with physics engines. Our framework advocates ‘closing the loop’ of developing manipulation policies with reproducible evaluation of policies learned from real-robot data an…

2026

Learning to Design Soft Hands Using Reward Models

ICRA 2026poster

Soft robotic hands promise to provide compliant and safe interaction with objects and environments. However, designing soft hands to be both compliant and functional across diverse use cases remains challenging. Although co-design of hardware and control better couples morphology to behavior, the re…

2026

Learning to Discover at Test Time

ICML 2026spotlight

How can we use AI to discover a new state of the art for a scientific problem? Prior work in test-time scaling, such as AlphaEvolve, performs search by prompting a frozen LLM. We perform reinforcement learning at test time, so the LLM can continue to train, but now with experience specific to the te…

Cited by 0SourceScholar
2026

MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks

AAAI 2026technical

As automated attack techniques rapidly advance, CAPTCHAs remain a critical defense mechanism against malicious bots. However, existing CAPTCHA schemes encompass a diverse range of modalities—from static distorted text and obfuscated images to interactive clicks, sliding puzzles, and logic-based ques

Cited by 0SourcePDFScholar
2026

TIPS: Turn-level Information-Potential Reward Shaping for Search-Augmented LLMs

ICLR 2026poster

Search-augmented large language models (LLMs) trained with reinforcement learning (RL) have achieved strong results on open-domain question answering (QA), but training still remains a significant challenge. The optimization is often unstable due to sparse rewards and difficult credit assignments ac…

Cited by 0SourcecodeScholar
2025

AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control

RSS 2025poster

Humanoid robots derive much of their dexterity from hyper-dexterous whole-body movements, enabling tasks that require a large operational workspace—such as picking objects off the ground. However, achieving these capabilities on real humanoids remains challenging due to their high degrees of freedom…

Cited by 0PDFScholar
2025

ARGenSeg: Image Segmentation with Autoregressive Image Generation Model

NeurIPS 2025poster

We propose a novel AutoRegressive Generation-based paradigm for image Segmentation (ARGenSeg), achieving multimodal understanding and pixel-level perception within a unified framework. Prior works integrating image segmentation into multimodal large language models (MLLMs) typically employ either b…

Cited by 0SourceScholar
2025

ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models

ACL 2025long

Active perception, a crucial human capability, involves setting a goal based on the current understanding of the environment and performing actions to achieve that goal. Despite significant efforts in evaluating Multimodal Large Language Models (MLLMs), active perception has been largely overlooked.…

2025

Bunny-VisionPro: Real-Time Bimanual Dexterous Teleoperation for Imitation Learning

IROS 2025

Teleoperation is a crucial tool for collecting human demonstrations, but controlling robots with bimanual dexterous hands remains a challenge. Existing teleoperation systems struggle to handle the complexity of coordinating two hands for intricate manipulations. We introduce Bunny-VisionPro, a real-

Cited by 129SourcecodeScholar
2025

Co-Design of Soft Gripper with Neural Physics

CoRL 2025poster

For robot manipulation, both the controller and end-effector design are crucial. Compared with rigid grippers, soft grippers are more generalizable by deforming to different geometries, but designing such a gripper and finding its grasp pose remains challenging. In this paper, we propose a co-design…

Cited by 0SourceScholar
2025

Dex1B: Learning with 1B Demonstrations for Dexterous Manipulation

RSS 2025poster

Generating large-scale demonstrations for dexterous manipulation remains a challenging problem, and various approaches have been proposed in recent years to address it. Among these, generative models have emerged as a promising paradigm, enabling the efficient generation of diverse and plausible dem…

Cited by 0PDFScholar
2025

EmoCharacter: Evaluating the Emotional Fidelity of Role-Playing Agents in Dialogues

NAACL 2025long

Role-playing agents (RPAs) powered by large language models (LLMs) have been widely utilized in dialogue systems for their capability to deliver personalized interactions. Current evaluations of RPAs mainly focus on personality fidelity, tone imitation, and knowledge consistency, while overlooking e…

Cited by 0SourcePDFScholar
2025

Gaussian-Augmented Physics Simulation and System Identification with Complex Colliders

NeurIPS 2025poster

System identification involving the geometry, appearance, and physical properties from video observations is a challenging task with applications in robotics and graphics. Recent approaches have relied on fully differentiable Material Point Method (MPM) and rendering for simultaneous optimization of…

Cited by 0SourceScholar
2025

HOVER: Versatile Neural Whole-Body Controller for Humanoid Robots

ICRA 2025

Humanoid whole-body control requires adapting to diverse tasks such as navigation, loco-manipulation, and tabletop manipulation, each demanding a different mode of control. For example, navigation relies on root velocity or position tracking, while tabletop manipulation prioritizes upper-body joint

Cited by 126SourceScholar
2025

Hallucination Detection in Structured Query Generation via LLM Self-Debating

EMNLP 2025

Hallucination remains a key challenge in applying large language models (LLMs) to structured query generation, especially for semi-private or domain-specific languages underrepresented in public training data. In this work, we focus on hallucination detection in these low-resource structured languag

Cited by 0SourcePDFScholar
2025

Helpful DoggyBot: Open-World Object Fetching using Legged Robots and Vision-Language Models

IROS 2025

Learning-Based methods have achieved strong performance for quadrupedal locomotion. However, several challenges prevent quadrupeds from learning helpful indoor skills that require interaction with environments and humans: lack of end-effectors for manipulation, limited semantic under-standing using

Cited by 15SourceScholar
2025

Hierarchical World Models as Visual Whole-Body Humanoid Controllers

ICLR 2025poster

Whole-body control for humanoids is challenging due to the high-dimensional nature of the problem, coupled with the inherent instability of a bipedal morphology. Learning from visual observations further exacerbates this difficulty. In this work, we explore highly data-driven approaches to visual wh…

Cited by 5SourcePDFScholar
2025

HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography Estimation

AAAI 2025technical

Feature matching between image pairs is a fundamental problem in computer vision that drives many applications, such as SLAM. Recently, semi-dense matching approaches have achieved substantial performance enhancements and established a widely-accepted coarse-to-fine paradigm. However, the majority…

Cited by 0SourcePDFScholar
2025

Learning Generalizable Feature Fields for Mobile Manipulation

IROS 2025

An open problem in mobile manipulation is how to represent objects and scenes in a unified manner so that robots can use both for navigation and manipulation. The latter requires capturing intricate geometry while understanding fine-grained semantics, whereas the former involves capturing the comple

Cited by 49SourceScholar
2025

Learning to (Learn at Test Time): RNNs with Expressive Hidden States

ICML 2025spotlight

Self-attention performs well in long context but has quadratic complexity. Existing RNN layers have linear complexity, but their performance in long context is limited by the expressive power of their hidden states. We present a practical framework for instantiating sequence modeling layers with lin…

2025

Lucid-XR: An Extended-Reality Data Engine for Robotic Manipulation

CoRL 2025poster

We introduce Lucid-XR, a generative data engine for creating diverse and realistic-looking data to train real-world robot systems. At the core of Lucid-XR is vuer, a web-based physics simulation environment that runs directly on the XR headset, enabling internet-scale access to immersive, latency-fr…

Cited by 0SourceScholar
2025

MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models

EMNLP 2025

Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. Due to their strong performance in image-text alignment, MLLMs can effectively understand image-text pairs with clear meanings. However, effectively resolving the inherent ambiguiti

2025

ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training

CoRL 2025poster

Generative models based on flow matching offer significant potential for learning robot policies, particularly in generating high-dimensional, dexterous behaviors that are conditioned on diverse observations. In this work, we introduce ManiFlow, an advanced flow matching model specifically designed…

Cited by 0SourceScholar
2025

Mobile-TeleVision: Predictive Motion Priors for Humanoid Whole-Body Control

ICRA 2025

Humanoid robots require both robust lower-body locomotion and precise upper-body manipulation. While recent Reinforcement Learning (RL) approaches provide whole-body loco-manipulation policies, they lack precise manipulation with high DoF arms. In this paper, we propose decoupling upper-body control

Cited by 85SourceScholar
2025

NaVILA: Legged Robot Vision-Language-Action Model for Navigation

RSS 2025poster

This paper proposes to solve the problem of Vision-and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes. However, it is non-trivial to translate human language inst…

Cited by 13PDFScholar
2025

One-Minute Video Generation with Test-Time Training

CVPR 2025poster

Transformers today still struggle to generate one-minute videos because self-attention layers are inefficient for long context. Alternatives such as Mamba layers struggle to produce coherent scenes because their hidden states are small and less expressive. We experiment with Test-Time Training (TTT)…

2025

Parallel Sequence Modeling via Generalized Spatial Propagation Network

CVPR 2025poster

We present the Generalized Spatial Propagation Network (GSPN), a new attention mechanism optimized for vision tasks that inherently captures 2D spatial structures. Existing attention models, including transformers, linear attention, and state-space models like Mamba, process multi-dimensional data a…

Cited by 0SourcePDFScholar
2025

Perspective Transition of Large Language Models for Solving Subjective Tasks

ACL 2025finding

Large language models (LLMs) have revolutionized the field of natural language processing, enabling remarkable progress in various tasks. Different from objective tasks such as commonsense reasoning and arithmetic question-answering, the performance of LLMs on subjective tasks is still limited, wher…

2025

SPOT: SE(3) Pose Trajectory Diffusion for Object-Centric Manipulation

ICRA 2025

We introduce SPOT, an object-centric imitation learning framework. The key idea is to capture each task by an object-centric representation, specifically the SE(3) object pose trajectory relative to the target. This approach decouples embodiment actions from sensory inputs, facilitating learning fro

Cited by 34SourcecodeScholar
2025

VT-Refine: Learning Bimanual Assembly with Visuo-Tactile Feedback via Simulation Fine-Tuning

CoRL 2025poster

Humans excel at bimanual assembly tasks by adapting to rich tactile feedback—a capability that remains difficult to replicate in robots through behavioral cloning alone, due to the suboptimality and limited diversity of human demonstrations. In this work, we present VT-Refine, a visuo-tactile policy…

Cited by 0SourcecodeScholar
2025

WildLMa: Long Horizon Loco-Manipulation in the Wild

ICRA 2025

‘In-the-wild’ mobile manipulation aims to deploy robots in diverse real-world environments, which requires the robot to (1) have skills that generalize across object configurations; (2) be capable of long-horizon task execution in diverse environments; and (3) perform complex manipulation beyond pic

Cited by 16SourceScholar
2025

WorldModelBench: Judging Video Generation Models As World Models

NeurIPS 2025poster

Video generation models have rapidly progressed, positioning themselves as video world models capable of supporting decision-making applications like robotics and autonomous driving. However, current benchmarks fail to rigorously evaluate these claims, focusing only on general video quality, ignorin…

Cited by 0SourcecodeScholar
2024

3D Reconstruction with Generalizable Neural Fields using Scene Priors

ICLR 2024poster

High-fidelity 3D scene reconstruction has been substantially advanced by recent progress in neural fields. However, most existing methods train a separate network from scratch for each individual scene. This is not scalable, inefficient, and unable to yield good results given limited views. While le…

2024

A Simulation Benchmark for Autonomous Racing with Large-Scale Human Data

NeurIPS 2024poster

Despite the availability of international prize-money competitions, scaled vehicles, and simulation environments, research on autonomous racing and the control of sports cars operating close to the limit of handling has been limited by the high costs of vehicle acquisition and management, as well as…

2024

ACE: A Cross-platform and visual-Exoskeletons System for Low-Cost Dexterous Teleoperation

CoRL 2024poster

Bimanual robotic manipulation with dexterous hands has a large potential workability and a wide workspace as it follows the most natural human workflow. Learning from human demonstrations has proven highly effective for learning a dexterous manipulation policy. To collect such data, teleoperation se…

Cited by 37SourceScholar
2024

CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language Models

ACL 2024long

Multimodal large language models (MLLMs) have demonstrated promising results in a variety of tasks that combine vision and language. As these models become more integral to research and applications, conducting comprehensive evaluations of their capabilities has grown increasingly important. However…

Cited by 8SourcePDFScholar
2024

COLMAP-Free 3D Gaussian Splatting

CVPR 2024highlight

While neural rendering has led to impressive advances in scene reconstruction and novel view synthesis it relies heavily on accurately pre-computed camera poses. To relax this constraint multiple efforts have been made to train Neural Radiance Fields (NeRFs) without pre-processed camera poses. Howev…

2024

CyberDemo: Augmenting Simulated Human Demonstration for Real-World Dexterous Manipulation

CVPR 2024poster

We introduce CyberDemo a novel approach to robotic imitation learning that leverages simulated human demonstrations for real-world tasks. By incorporating extensive data augmentation in a simulated environment CyberDemo outperforms traditional in-domain real-world demonstrations when transferred to…

2024

DEEM: Dynamic Experienced Expert Modeling for Stance Detection

COLING 2024main

Recent work has made a preliminary attempt to use large language models (LLMs) to solve the stance detection task, showing promising results. However, considering that stance detection usually requires detailed background knowledge, the vanilla reasoning method may neglect the domain knowledge to ma…

2024

Editable Image Elements for Controllable Synthesis

ECCV 2024poster

"Diffusion models have made significant advances in text-guided synthesis tasks. However, editing user-provided images remains challenging, as the high dimensional noise input space of diffusion models is not naturally suited for image inversion or spatial editing. In this work, we propose an image…

Cited by 8SourcePDFScholar
2024

Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages

ACL 2024long

While large language models (LLMs) have been pre-trained on multilingual corpora, their performance still lags behind in most languages compared to a few resource-rich languages. One common approach to mitigate this issue is to translate training data from resource-rich languages into other language…

2024

Expressive Whole-Body Control for Humanoid Robots

RSS 2024poster

Can we enable humanoid robots to generate rich, diverse, and expressive motions in the real world? We propose to learn a whole-body control policy on a human-sized robot to mimic human motions as realistic as possible. To train such a policy, we leverage the large-scale human motion capture data fro…

Cited by 94SourcePDFScholar
2024

GenSim: Generating Robotic Simulation Tasks via Large Language Models

ICLR 2024spotlight

Collecting large amounts of real-world interaction data to train general robotic policies is often prohibitively expensive, thus motivating the use of simulation data. However, existing methods for data generation have generally focused on scene-level diversity (e.g., object instances and poses) rat…

2024

Generalized Animal Imitator: Agile Locomotion with Versatile Motion Prior

CoRL 2024poster

The agility of animals, particularly in complex activities such as running, turning, jumping, and backflipping, stands as an exemplar for robotic system design. Transferring this suite of behaviors to legged robotic systems introduces essential inquiries: How can a robot be trained to learn multiple…

Cited by 19SourceScholar
2024

HOIDiffusion: Generating Realistic 3D Hand-Object Interaction Data

CVPR 2024poster

3D hand-object interaction data is scarce due to the hardware constraints in scaling up the data collection process. In this paper we propose HOIDiffusion for generating realistic and diverse 3D hand-object interaction data. Our model is a conditional diffusion model that takes both the 3D hand-obje…

2024

Image Neural Field Diffusion Models

CVPR 2024highlight

Diffusion models have shown an impressive ability to model complex data distributions with several key advantages over GANs such as stable training better coverage of the training distribution's modes and the ability to solve inverse problems without extra training. However most diffusion models lea…

Cited by 6SourcePDFScholar
2024

Investigating and Mitigating the Side Effects of Noisy Views for Self-Supervised Clustering Algorithms in Practical Multi-View Scenarios

CVPR 2024poster

Multi-view clustering (MVC) aims at exploring category structures among multi-view data in self-supervised manners. Multiple views provide more information than single views and thus existing MVC methods can achieve satisfactory performance. However their performance might seriously degenerate when…

2024

Language-Driven Physics-Based Scene Synthesis and Editing via Feature Splatting

ECCV 2024poster

"Scene representations using 3D Gaussian primitives have produced excellent results in modeling the appearance of static and dynamic 3D scenes. Many graphics applications, however, demand the ability to manipulate both the appearance and the physical properties of objects. We introduce Feature Splat…

Cited by 12SourcePDFScholar
2024

Lessons from Learning to Spin “Pens”

CoRL 2024poster

In-hand manipulation of pen-like objects is a most basic and important skill in our daily lives, as many tools such as hammers and screwdrivers are similarly shaped. However, current learning-based methods struggle with this task due to a lack of high-quality demonstrations and the significant gap b…

Cited by 16SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

Open-TeleVision: Teleoperation with Immersive Active Visual Feedback

CoRL 2024poster

Teleoperation serves as a powerful method for collecting on-robot data essential for robot learning from demonstrations. The intuitiveness and ease of use of the teleoperation system are crucial for ensuring high-quality, diverse, and scalable data. To achieve this, we propose an immersive teleopera…

Cited by 99SourceScholar
2024

Pluggable Neural Machine Translation Models via Memory-augmented Adapters

COLING 2024main

Although neural machine translation (NMT) models perform well in the general domain, it remains rather challenging to control their generation behavior to satisfy the requirement of different users. Given the expensive training cost and the data scarcity challenge of learning a new model from scratc…

2024

PointLLM: Empowering Large Language Models to Understand Point Clouds

ECCV 2024oral

"The unprecedented advancements in Large Language Models (LLMs) have shown a profound impact on natural language processing but are yet to fully embrace the realm of 3D understanding. This paper introduces PointLLM, a preliminary effort to fill this gap, empowering LLMs to understand point clouds an…

2024

RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos

CVPR 2024poster

We introduce a new RGB-D object dataset captured in the wild called WildRGB-D. Unlike most existing real-world object-centric datasets which only come with RGB capturing the direct capture of the depth channel allows better 3D annotations and broader downstream applications. WildRGB-D comprises larg…

2024

Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models

ACL 2024long

Large Language Models (LLMs) have achieved remarkable performance in objective tasks such as open-domain question answering and mathematical reasoning, which can often be solved through recalling learned factual knowledge or chain-of-thought style reasoning. However, we find that the performance of…

2024

Robot Synesthesia: In-Hand Manipulation with Visuotactile Sensing

ICRA 2024poster

Executing contact-rich manipulation tasks necessitates the fusion of tactile and visual feedback. However, the distinct nature of these modalities poses significant challenges. In this paper, we introduce a system that leverages visual and tactile sensory inputs to enable dexterous in-hand manipulat…

Cited by 47SourcecodeScholar
2024

Sim2Real Manipulation on Unknown Objects with Tactile-based Reinforcement Learning

ICRA 2024poster

Using tactile sensors for manipulation remains one of the most challenging problems in robotics. At the heart of these challenges is generalization: How can we train a tactile-based policy that can manipulate unseen and diverse objects? In this paper, we propose to perform Reinforcement Learning wit…

Cited by 6SourcecodeScholar
2024

SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models

NeurIPS 2024poster

Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Region GPT (SpatialRGPT) to enhance VLMs’ spatial perception and reasoning capabilit…

Cited by 61SourcePDFScholar
2024

Visual Whole-Body Control for Legged Loco-Manipulation

CoRL 2024poster

We study the problem of mobile manipulation using legged robots equipped with an arm, namely legged loco-manipulation. The robot legs, while usually utilized for mobility, offer an opportunity to amplify the manipulation capabilities by conducting whole-body control. That is, the robot can control t…

Cited by 47SourceScholar
2023

ActorsNeRF: Animatable Few-shot Human Rendering with Generalizable NeRFs

ICCV 2023poster

While NeRF-based human representations have shown impressive novel view synthesis results, most methods still rely on a large number of images / views for training. In this work, we propose a novel animatable NeRF called ActorsNeRF. It is first pre-trained on diverse human subjects, and then adapted…

Cited by 25PDFScholar
2023

AnyTeleop: A General Vision-Based Dexterous Robot Arm-Hand Teleoperation System

RSS 2023poster

Vision-based teleoperation offers the possibility to endow robots with human-level intelligence to physically interact with the environment, while only requiring low-cost camera sensors. However, current vision-based teleoperation systems are designed and engineered towards a particular robot model…

Cited by 114SourcePDFScholar
2023

Cross-Modality Person Re-identification with Memory-Based Contrastive Embedding

AAAI 2023technical

Visible-infrared person re-identification (VI-ReID) aims to retrieve the person images of the same identity from the RGB to infrared image space, which is very important for real-world surveillance system. In practice, VI-ReID is more challenging due to the heterogeneous modality discrepancy, which…

Cited by 14SourcePDFScholar
2023

DexArt: Benchmarking Generalizable Dexterous Manipulation With Articulated Objects

CVPR 2023poster

To enable general-purpose robots, we will require the robot to operate daily articulated objects as humans do. Current robot manipulation has heavily relied on using a parallel gripper, which restricts the robot to a limited set of objects. On the other hand, operating with a multi-finger robot hand…

2023

Dynamic Handover: Throw and Catch with Bimanual Hands

CoRL 2023poster

Humans throw and catch objects all the time. However, such a seemingly common skill introduces a lot of challenges for robots to achieve: The robots need to operate such dynamic actions at high-speed, collaborate precisely, and interact with diverse objects. In this paper, we design a system with tw…

Cited by 50SourcecodeScholar
2023

Dynamic Inference With Grounding Based Vision and Language Models

CVPR 2023poster

Transformers have been recently utilized for vision and language tasks successfully. For example, recent image and language models with more than 200M parameters have been proposed to learn visual grounding in the pre-training step and show impressive results on downstream vision and language tasks.…

2023

Efficient Bimanual Handover and Rearrangement via Symmetry-Aware Actor-Critic Learning

ICRA 2023poster

Bimanual manipulation is important for building intelligent robots that unlock richer skills than single arms. We consider a multi-object bimanual rearrangement task, where a reinforcement learning (RL) agent aims to jointly control two arms to rearrange these objects as fast as possible. Solving th…

Cited by 15SourceScholar
2023

FeatureNeRF: Learning Generalizable NeRFs by Distilling Foundation Models

ICCV 2023poster

Recent works on generalizable NeRFs have shown promising results on novel view synthesis from single or few images. However, such models have rarely been applied on other downstream tasks beyond synthesis such as semantic understanding and parsing. In this paper, we propose a novel framework named F…

Cited by 46PDFcodeScholar
2023

Fine-Grained Cross-View Geo-Localization Using a Correlation-Aware Homography Estimator

NeurIPS 2023poster

In this paper, we introduce a novel approach to fine-grained cross-view geo-localization. Our method aligns a warped ground image with a corresponding GPS-tagged satellite image covering the same area using homography estimation. We first employ a differentiable spherical transform, adhering to geom…

2023

Finetuning Offline World Models in the Real World

CoRL 2023oral

Reinforcement Learning (RL) is notoriously data-inefficient, which makes training on a real robot difficult. While model-based RL algorithms (world models) improve data-efficiency to some extent, they still require hours or days of interaction to learn skills. Recently, offline RL has been proposed…

Cited by 23SourceScholar
2023

GNFactor: Multi-Task Real Robot Learning with Generalizable Neural Feature Fields

CoRL 2023oral

It is a long-standing problem in robotics to develop agents capable of executing diverse manipulation tasks from visual observations in unstructured real-world environments. To achieve this goal, the robot will need to have a comprehensive understanding of the 3D structure and semantics of the scen…

Cited by 88SourcecodeScholar
2023

GPViT: A High Resolution Non-Hierarchical Vision Transformer with Group Propagation

ICLR 2023top-25%

We present the Group Propagation Vision Transformer (GPViT): a novel non- hierarchical (i.e. non-pyramidal) transformer model designed for general visual recognition with high-resolution features. High-resolution features (or tokens) are a natural fit for tasks that involve perceiving fine-grained d…

2023

Learning Continuous Grasping Function With a Dexterous Hand From Human Demonstrations

RA-L 2023

We propose to learn to generate grasping motion for manipulation with a dexterous hand using implicit functions. With continuous time inputs, the model can generate a continuous and smooth grasping plan. We name the proposed model Continuous Grasping Function (CGF). CGF is learned via generative mod

Cited by 75SourcecodeScholar
2023

MoDem: Accelerating Visual Model-Based Reinforcement Learning with Demonstrations

ICLR 2023poster

Poor sample efficiency continues to be the primary challenge for deployment of deep Reinforcement Learning (RL) algorithms for real-world applications, and in particular for visuo-motor control. Model-based RL has the potential to be highly sample efficient by concurrently learning a world model and…

2023

MonoNeRF: Learning Generalizable NeRFs from Monocular Videos without Camera Poses

ICML 2023poster

We propose a generalizable neural radiance fields - MonoNeRF, that can be trained on large-scale monocular videos of moving in static scenes without any ground-truth annotations of depth and camera poses. MonoNeRF follows an Autoencoder-based architecture, where the encoder estimates the monocular d…

2023

On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch Baseline

ICML 2023poster

In this paper, we examine the effectiveness of pre-training for visuo-motor control tasks. We revisit a simple Learning-from-Scratch (LfS) baseline that incorporates data augmentation and a shallow ConvNet, and find that this baseline is surprisingly competitive with recent approaches (PVR, MVP, R3M…

2023

Open-Vocabulary Panoptic Segmentation With Text-to-Image Diffusion Models

CVPR 2023highlight

We present ODISE: Open-vocabulary DIffusion-based panoptic SEgmentation, which unifies pre-trained text-image diffusion and discriminative models to perform open-vocabulary panoptic segmentation. Text-to-image diffusion models have the remarkable ability to generate high-quality images with diverse…

2023

Policy Adaptation From Foundation Model Feedback

CVPR 2023poster

Recent progress on vision-language foundation models have brought significant advancement to building general-purpose robots. By using the pre-trained models to encode the scene and instructions as inputs for decision making, the instruction-conditioned policy can generalize across different objects…

Cited by 13SourcePDFScholar
2023

Rotating without Seeing: Towards In-hand Dexterity through Touch

RSS 2023

Tactile information plays a critical role in human dexterity. It reveals useful contact information that may not be inferred directly from vision. In fact, humans can even perform in-hand dexterous manipulation without using vision. Can we enable the same ability for the multi-finger robot hand? In

Cited by 70SourceScholar
2023

Self-Supervised Geometric Correspondence for Category-Level 6D Object Pose Estimation in the Wild

ICLR 2023poster

While 6D object pose estimation has wide applications across computer vision and robotics, it remains far from being solved due to the lack of annotations. The problem becomes even more challenging when moving to category-level 6D pose, which requires generalization to unseen instances. Current appr…

2023

Visual Reinforcement Learning With Self-Supervised 3D Representations

RA-L 2023

A prominent approach to visual Reinforcement Learning (RL) is to learn an internal state representation using self-supervised methods, which has the potential benefit of improved sample-efficiency and generalization through additional learning signal and inductive biases. However, while the real wor

Cited by 74SourcecodeScholar
2023

Zero-Shot Pose Transfer for Unrigged Stylized 3D Characters

CVPR 2023poster

Transferring the pose of a reference avatar to stylized 3D characters of various shapes is a fundamental task in computer graphics. Existing methods either require the stylized characters to be rigged, or they use the stylized character in the desired pose as ground truth at training. We present a z…

2022

Category-Level 6D Object Pose Estimation in the Wild: A Semi-Supervised Learning Approach and A New Dataset

NeurIPS 2022accept

6D object pose estimation is one of the fundamental problems in computer vision and robotics research. While a lot of recent efforts have been made on generalizing pose estimation to novel object instances within the same category, namely category-level 6D pose estimation, it is still restricted in…

2022

CoordGAN: Self-Supervised Dense Correspondences Emerge From GANs

CVPR 2022poster

Recent advances show that Generative Adversarial Networks (GANs) can synthesize images with smooth variations along semantically meaningful latent directions, such as pose, expression, layout, etc. While this indicates that GANs implicitly learn pixel-level correspondences across images, few studies…

Cited by 22PDFcodeScholar
2022

DexMV: Imitation Learning for Dexterous Manipulation from Human Videos

ECCV 2022poster

"While in computer vision we have made significant progress on understanding hand-object interactions, it is still very challenging for robots to perform complex dexterous manipulation. In this paper, we propose a new platform and pipeline, DexMV (Dexterous Manipulation from Videos), for imitation l…

2022

DexPoint: Generalizable Point Cloud Reinforcement Learning for Sim-to-Real Dexterous Manipulation

CoRL 2022poster

We propose a sim-to-real framework for dexterous manipulation which can generalize to new objects of the same category in the real world. The key of our framework is to train the manipulation policy with point cloud inputs and dexterous hands. We propose two new techniques to enable joint learning o…

Cited by 79SourcecodeScholar
2022

From One Hand to Multiple Hands: Imitation Learning for Dexterous Manipulation From Single-Camera Teleoperation

RA-L 2022

We propose to perform imitation learning for dexterous manipulation with multi-finger robot hand from human demonstrations, and transfer the policy to the real robot hand. We introduce a novel single-camera teleoperation system to collect the 3D demonstrations efficiently with only an iPad and a com

Cited by 145SourceScholar
2022

GIFS: Neural Implicit Function for General Shape Representation

CVPR 2022poster

Recent development of neural implicit function has shown tremendous success on high-quality 3D shape reconstruction. However, most works divide the space into inside and outside of the shape, which limits their representing power to single-layer and watertight shapes. This limitation leads to tediou…

Cited by 75PDFcodeScholar
2022

Graph Inverse Reinforcement Learning from Diverse Videos

CoRL 2022oral

Research on Inverse Reinforcement Learning (IRL) from third-person videos has shown encouraging results on removing the need for manual reward design for robotic tasks. However, most prior works are still limited by training from a relatively restricted domain of videos. In this paper, we argue that…

Cited by 57SourcecodeScholar
2022

GroupViT: Semantic Segmentation Emerges From Text Supervision

CVPR 2022poster

Grouping and recognition are important components of visual scene understanding, e.g., for object detection and semantic segmentation. With end-to-end deep learning systems, grouping of image regions usually happens implicitly via top-down supervision from pixel-level recognition labels. Instead, in…

Cited by 612PDFcodeScholar
2022

Joint Hand Motion and Interaction Hotspots Prediction From Egocentric Videos

CVPR 2022poster

We propose to forecast future hand-object interactions given an egocentric video. Instead of predicting action labels or pixels, we directly predict the hand motion trajectory and the future contact points on the next active object (i.e., interaction hotspots). This relatively low-dimensional repres…

Cited by 105PDFcodeScholar
2022

Learning Continuous Environment Fields via Implicit Functions

ICLR 2022poster

We propose a novel scene representation that encodes reaching distance -- the distance between any position in the scene to a goal along a feasible trajectory. We demonstrate that this environment field representation can directly guide the dynamic behaviors of agents in 2D mazes or 3D indoor scenes…

Cited by 12SourcePDFScholar
2022

Learning Generalizable Dexterous Manipulation from Human Grasp Affordance

CoRL 2022poster

Dexterous manipulation with a multi-finger hand is one of the most challenging problems in robotics. While recent progress in imitation learning has largely improved the sample efficiency compared to Reinforcement Learning, the learned policy can hardly generalize to manipulate novel objects, given…

Cited by 69SourcecodeScholar
2022

Learning Implicit Feature Alignment Function for Semantic Segmentation

ECCV 2022poster

"Integrating high-level context information with low-level details is of central importance in semantic segmentation. Towards this end, most existing segmentation models apply bilinear up-sampling and convolutions to feature maps of different scales, and then align them at the same resolution. Howev…

2022

Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers

ICLR 2022spotlight

We propose to address quadrupedal locomotion tasks using Reinforcement Learning (RL) with a Transformer-based model that learns to combine proprioceptive information and high-dimensional depth sensor inputs. While learning-based locomotion has made great advances using RL, most methods still rely on…

2022

Look Closer: Bridging Egocentric and Third-Person Views With Transformers for Robotic Manipulation

RA-L 2022

Learning to solve precision-based manipulation tasks from visual feedback using Reinforcement Learning (RL) could drastically reduce the engineering efforts required by traditional robot systems. However, performing fine-grained motor control from visual inputs alone is challenging, especially with

Cited by 86SourcecodeScholar
2022

Look Outside the Room: Synthesizing a Consistent Long-Term 3D Scene Video From a Single Image

CVPR 2022poster

Novel view synthesis from a single image has recently attracted a lot of attention, and it has been primarily advanced by 3D deep learning and rendering techniques. However, most work is still limited by synthesizing new views within relatively small camera motions. In this paper, we propose a novel…

Cited by 59PDFcodeScholar
2022

Multi-Role Event Argument Extraction as Machine Reading Comprehension with Argument Match Optimization

ICASSP 2022accepted

Extracting arguments for the pre-defined roles is a crucial step for event extraction. Recently, there are some insightful works that view it as a machine reading comprehension problem and achieve significant progress. However, most of them need multi-turns to extract the arguments of each role inde…

Cited by 0SourceScholar
2022

Online Adaptation for Implicit Object Tracking and Shape Reconstruction in the Wild

RA-L 2022

Tracking and reconstructing 3D objects from cluttered scenes are the key components for computer vision, robotics and autonomous driving systems. While recent progress in implicit function has shown encouraging results on high-quality 3D shape reconstruction, it is still very challenging to generali

Cited by 9SourcecodeScholar
2022

Scraping Textures from Natural Images for Synthesis and Editing

ECCV 2022poster

"Existing texture synthesis methods focus on generating large texture images given a small texture sample. But such samples are typically assumed to be highly curated: rectangular, clean, and stationary. This paper aims to scrape textures directly from natural images of everyday objects and scenes,…

Cited by 4SourcePDFScholar
2022

VideoINR: Learning Video Implicit Neural Representation for Continuous Space-Time Super-Resolution

CVPR 2022poster

Videos typically record the streaming and continuous visual data as discrete consecutive frames. Since the storage cost is expensive for videos of high fidelity, most of them are stored in a relatively low resolution and frame rate. Recent works of Space-Time Video Super-Resolution (STVSR) are devel…

Cited by 120PDFcodeScholar
2022

Vision-Guided Quadrupedal Locomotion in the Wild with Multi-Modal Delay Randomization

IROS 2022poster

Developing robust vision-guided controllers for quadrupedal robots in complex environments with various obstacles, dynamical surroundings and uneven terrains is very challenging. While Reinforcement Learning (RL) provides a promising paradigm for agile locomotion skills with vision inputs in simulat…

Cited by 29SourcecodeScholar
2021

A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape Representation

ICCV 2021poster

Recent work has made significant progress on using implicit functions, as a continuous representation for 3D rigid object shape reconstruction. However, much less effort has been devoted to modeling general articulated objects. Compared to rigid objects, articulated objects have higher degrees of fr…

Cited by 117PDFScholar
2021

Compositional Video Synthesis with Action Graphs

ICML 2021spotlight

Videos of actions are complex signals containing rich compositional structure in space and time. Current video generation methods lack the ability to condition the generation on multiple coordinated and potentially simultaneous timed actions. To address this challenge, we propose to represent the ac…

2021

Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization

ICLR 2021poster

We propose a simple, general and effective technique, Reward Randomization for discovering diverse strategic policies in complex multi-agent games. Combining reward randomization and policy gradient, we derive a new algorithm, Reward-Randomized Policy Gradient (RPG). RPG is able to discover a set of…

Cited by 63SourcePDFScholar
2021

Hand-Object Contact Consistency Reasoning for Human Grasps Generation

ICCV 2021poster

While predicting robot grasps with parallel jaw grippers have been well studied and widely applied in robot manipulation tasks, the study on natural human grasp generation with a multi-finger hand remains a very challenging problem. In this paper, we propose to generate human grasps given a 3D objec…

Cited by 189PDFcodeScholar
2021

Learning Cross-Domain Correspondence for Control with Dynamics Cycle-Consistency

ICLR 2021oral

At the heart of many robotics problems is the challenge of learning correspondences across domains. For instance, imitation learning requires obtaining correspondence between humans and robots; sim-to-real requires correspondence between physics simulators and real hardware; transfer learning requir…

Cited by 73SourcePDFScholar
2021

Learning Long-term Visual Dynamics with Region Proposal Interaction Networks

ICLR 2021poster

Learning long-term dynamics models is the key to understanding physical common sense. Most existing approaches on learning dynamics from visual input sidestep long-term predictions by resorting to rapid re-planning with short-term models. This not only requires such models to be super accurate but a…

2021

Meta-Baseline: Exploring Simple Meta-Learning for Few-Shot Learning

ICCV 2021poster

Meta-learning has been the most common framework for few-shot learning in recent years. It learns the model from collections of few-shot classification tasks, which is believed to have a key advantage of making the training objective consistent with the testing objective. However, some recent works…

Cited by 520PDFScholar
2021

Multi-Person 3D Motion Prediction with Multi-Range Transformers

NeurIPS 2021poster

We propose a novel framework for multi-person 3D motion trajectory prediction. Our key observation is that a human's action and behaviors may highly depend on the other persons around. Thus, instead of predicting each human pose trajectory in isolation, we introduce a Multi-Range Transformers model…

2021

NovelD: A Simple yet Effective Exploration Criterion

NeurIPS 2021poster

Efficient exploration under sparse rewards remains a key challenge in deep reinforcement learning. Previous exploration methods (e.g., RND) have achieved strong results in multiple hard tasks. However, if there are multiple novel areas to explore, these methods often focus quickly on one without suf…

2021

Region Similarity Representation Learning

ICCV 2021poster

We present Region Similarity Representation Learning (ReSim), a new approach to self-supervised representation learning for localization-based tasks such as object detection and segmentation. While existing work has largely focused on learning global representations for an entire image, ReSim learns…

Cited by 139PDFcodeScholar
2021

Rethinking Preventing Class-Collapsing in Metric Learning With Margin-Based Losses

ICCV 2021poster

Metric learning seeks perceptual embeddings where visually similar instances are close and dissimilar instances are apart, but learned representations can be sub-optimal when the distribution of intra-class samples is diverse and distinct sub-clusters are present. Although theoretically with optimal…

Cited by 16PDFScholar
2021

Rethinking Self-Supervised Correspondence Learning: A Video Frame-Level Similarity Perspective

ICCV 2021poster

Learning a good representation for space-time correspondence is the key for various computer vision tasks, including tracking object bounding boxes and performing video object pixel segmentation. To learn generalizable representation for correspondence in large-scale, a variety of self-supervised pr…

Cited by 114PDFcodeScholar
2021

Robust Object Detection via Instance-Level Temporal Cycle Confusion

ICCV 2021poster

Building reliable object detectors that are robust to domain shifts, such as various changes in context, viewpoint, and object appearances, is critical for real-world applications. In this work, we study the effectiveness of auxiliary self-supervised tasks to improve the out-of-distribution generali…

Cited by 34PDFcodeScholar
2021

Self-Supervised Policy Adaptation during Deployment

ICLR 2021spotlight

In most real world scenarios, a policy trained by reinforcement learning in one environment needs to be deployed in another, potentially quite different environment. However, generalization across different environments is known to be hard. A natural solution would be to keep training after deployme…

2021

Semi-Supervised 3D Hand-Object Poses Estimation With Interactions in Time

CVPR 2021poster

Estimating 3D hand and object pose from a single image is an extremely challenging problem: hands and objects are often self-occluded during interactions, and the 3D annotations are scarce as even humans cannot directly label the ground-truths from a single image perfectly. To tackle these challenge…

Cited by 195PDFcodeScholar
2021

Solving Compositional Reinforcement Learning Problems via Task Reduction

ICLR 2021poster

We propose a novel learning paradigm, Self-Imitation via Reduction (SIR), for solving compositional reinforcement learning problems. SIR is based on two core ideas: task reduction and self-imitation. Task reduction tackles a hard-to-solve task by actively reducing it to an easier task whose solution…

2021

Stabilizing Deep Q-Learning with ConvNets and Vision Transformers under Data Augmentation

NeurIPS 2021poster

While agents trained by Reinforcement Learning (RL) can solve increasingly challenging tasks directly from visual observations, generalizing learned skills to novel environments remains very challenging. Extensive use of data augmentation is a promising technique for improving generalization in RL,…

2021

State-Only Imitation Learning for Dexterous Manipulation

IROS 2021poster

Modern model-free reinforcement learning methods have recently demonstrated impressive results on a number of problems. However, complex domains like dexterous manipulation remain a challenge due to the high sample complexity. To address this, current approaches employ expert demonstrations in the f…

Cited by 135SourceScholar
2021

Synthesizing Long-Term 3D Human Motion and Interaction in 3D Scenes

CVPR 2021poster

Synthesizing 3D human motion plays an important role in many graphics applications as well as understanding human activity. While many efforts have been made on generating realistic and natural human motion, most approaches neglect the importance of modeling human-scene interactions and affordances.…

Cited by 145PDFScholar
2021

Test-Time Personalization with a Transformer for Human Pose Estimation

NeurIPS 2021poster

We propose to personalize a 2D human pose estimator given a set of test images of a person without using any manual annotations. While there is a significant advancement in human pose estimation, it is still very challenging for a model to generalize to different unknown environments and unseen pers…

2021

Video Autoencoder: Self-Supervised Disentanglement of Static 3D Structure and Motion

ICCV 2021poster

We present Video Autoencoder for learning disentangled representations of 3D structure and camera pose from videos in a self-supervised manner. Relying on temporal continuity in videos, our work assumes that the 3D scene structure in nearby video frames remains static. Given a sequence of video fram…

Cited by 39PDFScholar
2020

Deep Isometric Learning for Visual Recognition

ICML 2020poster

Initialization, normalization, and skip connections are believed to be three indispensable techniques for training very deep convolutional neural networks and obtaining state-of-the-art performance. This paper shows that deep vanilla ConvNets without normalization nor skip connections can also be tr…

2020

Hierarchical Style-based Networks for Motion Synthesis

ECCV 2020poster

Generating diverse and natural behaviors is one of the long-standing goals for creating intelligent characters in the animated world. In this paper, we propose an unsupervised method for generating long-range, diverse and plausible behaviors to achieve a specific goal location. Our proposed method l…

Cited by 35SourcePDFScholar
2020

Learning to Generate Diverse Questions from Keywords

ICASSP 2020accepted

Diverse text generation has been emerging as an important topic of natural language generation. Traditional studies on question generation mainly investigate how to generate one question based on a given input (one-to-one). In this paper, we focus on a more complex question generation task, i.e., ge…

Cited by 0SourceScholar
2020

MedWriter: Knowledge-Aware Medical Text Generation

COLING 2020main

To exploit the domain knowledge to guarantee the correctness of generated text has been a hot topic in recent years, especially for high professional domains such as medical. However, most of recent works only consider the information of unstructured text rather than structured information of the kn…

Cited by 7SourcePDFScholar
2020

Online Adaptation for Consistent Mesh Reconstruction in the Wild

NeurIPS 2020poster

This paper presents an algorithm to reconstruct temporally consistent 3D meshes of deformable object instances from videos in the wild. Without requiring annotations of 3D mesh, 2D keypoints, or camera pose for each video frame, we pose video-based reconstruction as a self-supervised online adaptati…

Cited by 61SourcePDFScholar
2020

Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction Networks

CVPR 2020poster

Human action is naturally compositional: humans can easily recognize and perform actions with objects that are different from those used in training demonstrations. In this paper, we study the compositionality of action by looking into the dynamics of subject-object interactions. We propose a novel…

Cited by 218PDFScholar
2020

Test-Time Training with Self-Supervision for Generalization under Distribution Shifts

ICML 2020poster

In this paper, we propose Test-Time Training, a general approach for improving the performance of predictive models when training and test data come from different distributions. We turn a single unlabeled test sample into a self-supervised learning problem, on which we update the model parameters b…

Cited by 945SourcePDFScholar
2019

Joint-task Self-supervised Learning for Temporal Correspondence

NeurIPS 2019poster

This paper proposes to learn reliable dense correspondence from videos in a self-supervised manner. Our learning process integrates two highly related tasks: tracking large image regions and establishing fine-grained pixel-level associations between consecutive video frames. We exploit the synergy b…

2019

Putting Humans in a Scene: Learning Affordance in 3D Indoor Environments

CVPR 2019poster

Affordance modeling plays an important role in visual understanding. In this paper, we aim to predict affordances of 3D indoor scenes, specifically what human poses are afforded by a given indoor environment, such as sitting on a chair or standing on the floor. In order to predict valid affordances…

Cited by 125PDFScholar
2019

Visual Semantic Navigation using Scene Priors

ICLR 2019poster

How do humans navigate to target objects in novel scenes? Do we use the semantic/functional priors we have built over years to efficiently search and navigate? For example, to search for mugs, we search cabinets near the coffee machine and for fruits we try the fridge. In this work, we focus on inco…

Cited by 391SourcePDFScholar
2018

3D Human Pose Estimation in the Wild by Adversarial Learning

CVPR 2018poster

Recently, remarkable advances have been achieved in 3D human pose estimation from monocular images because of the powerful Deep Convolutional Neural Networks (DCNNs). Despite their success on large-scale datasets collected in the constrained lab environment, it is difficult to obtain the 3D pose ann…

Cited by 493SourcePDFScholar
2018

A Topological Approach to Workspace and Motion Planning for a Cable-Controlled Robot in Cluttered Environments

RA-L 2018

There is a rising demand for multiple-cable controlled robots in stadiums or warehouses due to its low cost, longer operation time, and higher safety standards. In a cluttered environment, the cables can wrap around obstacles, but careful choice needs to be made for the initial cable configurations

Cited by 14SourceScholar
2017

A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection

CVPR 2017poster

How do we learn an object detector that is invariant to occlusions and deformations? Our current solution is to use a data-driven strategy -- collect large-scale datasets which have object instances under different conditions. The hope is that the final classifier can use these examples to learn inv…

Cited by 802PDFcodeScholar
2017

Temporal Dynamic Graph LSTM for Action-Driven Video Object Detection

ICCV 2017poster

In this paper, we investigate a weakly-supervised object detection framework. Most existing frameworks focus on using static images to learn object detectors. However, these detectors often fail to generalize to videos because of the existing domain shift. Therefore, we investigate learning these de…

Cited by 108PDFcodeScholar
2017

Transitive Invariance for Self-Supervised Visual Representation Learning

ICCV 2017poster

Learning visual representations with self-supervised learning has become popular in computer vision. The idea is to design auxiliary tasks where labels are free to obtain. Most of these tasks end up providing data to learn specific kinds of invariance useful for recognition. In this paper, we propos…

Cited by 182PDFcodeScholar