← Search

Liang Lin

190 accepted papers

2026

A Causal Marriage between VLM and IRM from Understanding to Reasoning

CVPR 2026

Vision-Language Models (VLMs) like CLIP exhibit extraordinary out-of-distribution (OOD) generalization, while the theoretical foundations underlying this robustness remain largely unexplored. This work establishes a connection between CLIP and Invariant Risk Minimization (IRM), the principled paradi

Cited by 0SourcecodeScholar
2026

AlphaAgentEvo: Evolution-Oriented Alpha Mining via Self-Evolving Agentic Reinforcement Learning

ICLR 2026poster

Alpha mining seeks to identify predictive alpha factors that generate excess returns beyond the market from a vast and noisy search space; however, existing approaches struggle to facilitate the systematic evolution of alphas. Traditional methods, such as genetic programming, are unable to interpret…

Cited by 0SourceScholar
2026

AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots

CVPR 2026

Recent advances in Visual-Language-Action (VLA) models have shown promising potential for robotic manipulation tasks.However, real-world robotic tasks often involve long-horizon, multi-step problem-solving and require generalization for continual skill acquisition, extending beyond single actions or

Cited by 0SourceScholar
2026

DDP-WM: Disentangled Dynamics Prediction for Efficient World Models

ICML 2026poster

World models are essential for autonomous robotic planning. However, the substantial computational overhead of existing dense Transformer-based models significantly hinders real-time deployment. To address this efficiency-performance bottleneck, we introduce DDP-WM, a novel world model centered on t…

Cited by 0SourceScholar
2026

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior

CVPR 2026

Storyboard synthesis plays a crucial role in visual storytelling, aiming to generate coherent shot sequences that visually narrate cinematic events with consistent characters, scenes, and transitions. However, existing approaches are mostly adapted from text-to-image diffusion models, which struggle

Cited by 0SourceScholar
2026

Failure-Driven Workflow Refinement

ICML 2026spotlight

Workflow optimization for tool-using LLM agents is often cast as global search over candidate graphs, scored by a scalar metric. This collapses rich, multi-step failure traces into binary outcomes, obscuring recurring failure structure and making refinement inefficient. We reframe optimization as \e…

Cited by 0SourceScholar
2026

Great Minds Think Alike: Contextual Tacit Communication for Decentralized LLM-Agent Cooperation

ICML 2026poster

Large language models (LLMs) are increasingly used as planners for cooperative embodied agents, but multi-agent settings amplify inconsistency under partial observability and make explicit communication costly or even unavailable. Many existing approaches rely on online message passing; when communi…

Cited by 0SourceScholar
2026

Hidden in the Noise: Unveiling Backdoors in Audio LLMs Alignment Through Latent Acoustic Pattern Triggers

AAAI 2026technical

As Audio Large Language Models (ALLMs) emerge as powerful tools for speech processing, their safety implications demand urgent attention. While considerable research has explored textual and vision safety, audio’s distinct characteristics present significant challenges. This paper first investigates

Cited by 0SourcePDFScholar
2026

Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search

AAAI 2026technical

Recent progress in robotics and embodied AI is largely driven by Large Multimodal Models (LMMs). However, a key challenge remains underexplored: how can we advance LMMs to discover tasks that assist humans in open-future scenarios, where human intentions are highly concurrent and dynamic. In this wo

Cited by 0SourcePDFScholar
2026

Large Vision–Language Models Get Lost in Attention

ICML 2026poster

Despite the rapid evolution of training paradigms, the decoder backbone of large vision--language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer architecture. Therefore, deciphering the distinct roles of internal modules is critical for understanding model mechani…

Cited by 0SourceScholar
2026

Learning Heterogeneous Degradation Representation for Real-World Super-Resolution

ICLR 2026poster

Real-World Super-Resolution (RWSR) aims to reconstruct high-resolution images from low-resolution inputs captured under complex, real-life conditions, where diverse distortions result in significant degradation heterogeneity. Many methods rely on degradation representations, yet they struggle with t…

Cited by 0SourceScholar
2026

Learning to See and Act: Task-Aware Virtual View Exploration for Robotic Manipulation

CVPR 2026

Recent vision-language-action (VLA) models for multi-task robot manipulation often rely on fixed camera setups and shared visual encoders, which limit their performance under occlusions and during cross-task transfer. To address these challenges, we propose Task-aware Virtual View Exploration (TVVE)

Cited by 0SourcecodeScholar
2026

LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

CVPR 2026

Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex urban environments.While recent advances have achieved progress through large-scale memory graphs and lookahead path planning, they remain limited b

Cited by 0SourceScholar
2026

Massive Editing for Large Language Models Based on Dynamic Weight Generation

ICLR 2026poster

Knowledge Editing (KE) is a field that studies how to modify some knowledge in Large Language Models (LLMs) at a low cost (compared to pre-training). Currently, performing large-scale edits on LLMs while ensuring the Reliability, Generality, and Locality metrics of the edits remain a challenge. This…

Cited by 0SourceScholar
2026

Memoria-Bench: A Comprehensive Benchmark for Evaluating Memory in Long-Horizon Autonomous Agents

ICML 2026poster

Memory is a core capability of autonomous agents, yet existing benchmarks evaluate it primarily in constrained settings such as short dialogues or synthetic tasks, failing to reflect realistic agent deployments. We present \textbf{Memoria-Bench}, a benchmark for evaluating agent memory grounded in c…

Cited by 0SourceScholar
2026

OptiMVMap: Offline Vectorized Map Construction via Optimal Multi-vehicle Perspectives

CVPR 2026

Offline vectorized maps constitute critical infrastructure for high-precision autonomous driving and mapping services. Existing approaches rely predominantly on single ego-vehicle trajectories, which fundamentally suffer from viewpoint insufficiency: while memory-based methods extend observation tim

Cited by 0SourcecodeScholar
2026

OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment

ICLR 2026poster

Large language model (LLM) alignment faces a critical dilemma when addressing multiple human preferences: improvements in one dimension frequently come at the expense of others, creating unavoidable trade-offs between competing objectives like helpfulness and harmlessness. While prior work mainly fo…

Cited by 0SourcecodeScholar
2026

PhyScene3D: Physically Consistent 3D Interactive Tabletop Scene Generation

ICML 2026poster

Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense object hierarchies and irregular affordances. Existing methods, ranging from decoupled symbolic solvers to end-to-end regress…

Cited by 0SourceScholar
2026

SOLAR for Offline MARL: Plateau-Triggered Potential Shaping under World-Model Uncertainty

ICML 2026poster

Reward shaping can accelerate reinforcement learning, but in sparse-reward \emph{offline} multi-agent RL it is often brittle: dense intrinsic rewards may alter the underlying Markov game, while world-model guidance can amplify model bias. We find that shaping becomes reliable when it is (i) activate…

Cited by 0SourceScholar
2026

Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection

ICLR 2026poster

Current AI-Generated Image (AIGI) detection approaches predominantly rely on binary classification to distinguish real from synthetic images, often lacking interpretable or convincing evidence to substantiate their decisions. This limitation stems from existing AIGI detection benchmarks, which, desp…

Cited by 0SourcecodeScholar
2026

VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling

CVPR 2026

Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittleness primarily arises from misalignment in Spatial Modeling, rather than Physical Modeling. To address this, we propose

Cited by 0SourceScholar
2026

When Local Rules Create Global Order: Self-Organized Representation Learning for Latent Diffusion Models

CVPR 2026

This work studies how latent space structure impacts the performance of Latent Diffusion Models (LDMs). We show that effective generation requires a latent space that is simultaneously locally smooth, enabling stable and reliable reconstruction, and globally dispersive, allowing the model to draw di

Cited by 0SourceScholar
2026

When Preference Labels Fall Short: Aligning Diffusion Models from Real Data

ICML 2026poster

Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existing approaches rely on preference pairs constructed from model-generated images. Such supervision is inherently relative and can be ambiguous when bot…

Cited by 0SourceScholar
2026

Where Detectors Fail: Probing Generative Space for Generalizable AI-Generated Image Detection

ICML 2026poster

Detecting AI-generated images (AIGI) remains challenging because detectors often fail to generalize to unseen generators. Although existing methods are trained on large datasets, their performance still degrades when generation settings change, indicating that data scale alone is insufficient and th…

Cited by 0SourceScholar
2025

Anima2: Cross-Species Animal Animation through Image-to-Video Synthesis with Subject Alignment

ICASSP 2025accepted

Recent video editing advancements rely on accurate pose sequences to animate human actors. However, these efforts are not suitable for cross-species animation due to pose misalignment between species (for example, the poses of a cat differ greatly from that of a pig due to their distinct body struct…

Cited by 0SourceScholar
2025

Are High-Quality AI-Generated Images More Difficult for Models to Detect?

ICML 2025poster

The remarkable evolution of generative models has enabled the generation of high-quality, visually attractive images, often perceptually indistinguishable from real photographs to human eyes. This has spurred significant attention on AI-generated image (AIGI) detection. Intuitively, higher image qua…

2025

Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering

ICCV 2025poster

Embodied Question Answering (EQA) is a challenging task in embodied intelligence that requires agents to dynamically explore 3D environments, actively gather visual information, and perform multi-step reasoning to answer questions. However, current EQA approaches suffer from critical limitations in…

2025

Boosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty Estimation

CVPR 2025poster

Over the last decade, significant efforts have been dedicated to designing efficient models for the challenge of ultra-high resolution (UHR) semantic segmentation. These models mainly follow the dual-stream architecture and generally fall into three subcategories according to the improvement objecti…

2025

Can We Achieve Efficient Diffusion Without Self-Attention? Distilling Self-Attention into Convolutions

ICCV 2025poster

Contemporary diffusion models built upon U-Net or Diffusion Transformer (DiT) architectures have revolutionized image generation through transformer-based attention mechanisms. The prevailing paradigm has commonly employed self-attention with quadratic computational complexity to handle global spati…

Cited by 0SourcePDFScholar
2025

Chain of Methodologies: Scaling Test Time Computation without Training

ACL 2025finding

Large Language Models (LLMs) often struggle with complex reasoning tasks due to insufficient in-depth insights in their training data, which are frequently absent in publicly available documents. This paper introduces the Chain of Methodologies (CoM), a simple and innovative iterative prompting fram…

Cited by 0SourcePDFScholar
2025

Cool-Fusion: Fuse Large Language Models without Training

ACL 2025long

We focus on the problem of fusing two or more heterogeneous large language models (LLMs) to leverage their complementary strengths. One of the challenges of model fusion is high computational load, specifically in fine-tuning or aligning vocabularies. To address this, we propose Cool-Fusion, a simpl…

2025

Cross-modal Causal Relation Alignment for Video Question Grounding

CVPR 2025highlight

Video question grounding (VideoQG) requires models to answer the questions and simultaneously infer the relevant video segments to support the answers. However, existing VideoQG methods usually suffer from spurious cross-modal correlations, leading to a failure to identify the dominant visual scenes…

2025

DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh

CVPR 2025poster

Text-driven avatar generation has gained significant attention owing to its convenience. However, existing methods typically model the human body with all garments as a single 3D model, limiting its usability, such as clothing replacement, and reducing user control over the generation process. To ov…

Cited by 2SourcePDFScholar
2025

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

CVPR 2025poster

3D Question Answering (3D QA) requires the model to comprehensively understand its situated 3D scene described by the text, then reason about its surrounding environment and answer a question under that situation. However, existing methods usually rely on global scene perception from pure 3D point c…

2025

Delving into Cascaded Instability: A Lipschitz Continuity View on Image Restoration and Object Detection Synergy

NeurIPS 2025poster

To improve detection robustness in adverse conditions (e.g., haze and low light), image restoration is commonly applied as a pre-processing step to enhance image quality for the detector. However, the functional mismatch between restoration and detection networks can introduce instability and hinder…

Cited by 0SourceScholar
2025

DreamFuse: Adaptive Image Fusion with Diffusion Transformer

ICCV 2025poster

Image fusion seeks to seamlessly integrate foreground objects with background scenes, producing realistic and harmonious fused images. Unlike existing methods that directly insert objects into the background, adaptive and interactive fusion remains a challenging yet appealing task. It requires the f…

Cited by 0SourcePDFScholar
2025

Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference

ICCV 2025poster

Video Multimodal Large Language Models (Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the underlying LLMs, existing Video-MLLMs typically exhibit suboptimal performance on long video scenarios. To understand e…

2025

Hybrid Re-matching for Continual Learning with Parameter-Efficient Tuning

NeurIPS 2025poster

Continual learning seeks to enable a model to assimilate knowledge from non-stationary data streams without catastrophic forgetting. Recently, methods based on Parameter-Efficient Tuning (PET) have achieved superior performance without even storing any historical exemplars, which train much fewer sp…

Cited by 0SourcecodeScholar
2025

HyperCRS: Hypergraph-Aware Multi-Grained Preference Learning to Burst Filter Bubbles in Conversational Recommendation System

ACL 2025finding

The filter bubble is a notorious issue in Recommender Systems (RSs), characterized by users being confined to a limited corpus of information or content that strengthens and amplifies their pre-established preferences and beliefs. Most existing methods primarily aim to analyze filter bubbles in the…

2025

IntelliCockpitBench: A Comprehensive Benchmark to Evaluate VLMs for Intelligent Cockpit

ACL 2025finding

The integration of sophisticated Vision-Language Models (VLMs) in vehicular systems is revolutionizing vehicle interaction and safety, performing tasks such as Visual Question Answering (VQA). However, a critical gap persists due to the lack of a comprehensive benchmark for multimodal VQA models in…

2025

Language Models as Implicit Tree Search

ICML 2025poster

Despite advancing language model (LM) alignment, direct preference optimization (DPO) falls short in LM reasoning with the free lunch from reinforcement learning (RL). As the breakthrough, this work proposes a new RL-free preference optimization method aiming to achieve DPO along with learning anoth…

Cited by 0SourcePDFScholar
2025

MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models

ACL 2025long

Long Context Understanding (LCU) is a critical area for exploration in current large language models (LLMs). However, due to the inherently lengthy nature of long-text data, existing LCU benchmarks for LLMs often result in prohibitively high evaluation costs, like testing time and inference expenses…

2025

Monitoring Primitive Interactions During the Training of DNNs

AAAI 2025technical

This paper focuses on the newly emerged research topic, i.e., whether the complex decision-making logic of a DNN can be mathematically summarized into a few simple logics. Beyond the explanation of a static DNN, in this paper, we hope to show that the seemingly complex learning dynamics of a DNN can…

Cited by 0SourcePDFScholar
2025

No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image Recognition

CVPR 2025highlight

Over the last decade, many notable methods have emerged to tackle the computational resource challenge of the high resolution image recognition (HRIR). They typically focus on identifying and aggregating a few salient regions for classification, discarding sub-salient areas for low training consumpt…

2025

PS-Diffusion: Photorealistic Subject-Driven Image Editing with Disentangled Control and Attention

CVPR 2025poster

Diffusion models pre-trained on large-scale paired image-text data achieve significant success in image editing. To convey more fine-grained visual details, subject-driven editing integrates subjects in user-provided reference images into existing scenes. However, it is challenging to obtain photore…

2025

Quadratic Coreset Selection: Certifying and Reconciling Sequence and Token Mining for Efficient Instruction Tuning

NeurIPS 2025poster

Instruction-Tuning (IT) was recently found the impressive data efficiency in post-training large language models (LLMs). While the pursuit of efficiency predominantly focuses on sequence-level curation, often overlooking the nuanced impact of critical tokens and the inherent risks of token noise and…

Cited by 0SourceScholar
2025

Reproducible Vision-Language Models Meet Concepts Out of Pre-Training

CVPR 2025poster

Contrastive Language-Image Pre-training (CLIP) models as a milestone of modern multimodal intelligence, its generalization mechanism grasped massive research interests in the community. While existing studies limited in the scope of pre-training knowledge, hardly underpinned its generalization to co…

Cited by 0SourcePDFScholar
2025

RoboPearls: Editable Video Simulation for Robot Manipulation

ICCV 2025poster

The development of generalist robot manipulation policies has seen significant progress, driven by large-scale demonstration data across diverse environments. However, the high cost and inefficiency of collecting real-world demonstrations hinder the scalability of data acquisition. While existing si…

Cited by 0SourcePDFScholar
2025

Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention

NeurIPS 2025poster

Egocentric Referring Video Object Segmentation (Ego-RVOS) aims to segment the specific object actively involved in a human action, as described by a language query, within first-person videos. This task is critical for understanding egocentric human behavior. However, achieving such segmentation rob…

Cited by 0SourceScholar
2025

RouterEval: A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in LLMs

EMNLP 2025

Routing large language models (LLMs) is a new paradigm that uses a router to recommend the best LLM from a pool of candidates for a given input. In this paper, our comprehensive analysis with more than 8,500 LLMs reveals a novel model-level scaling up phenomenon in Routing LLMs, i.e., a capable rout

2025

SR-FoT: A Syllogistic-Reasoning Framework of Thought for Large Language Models Tackling Knowledge-based Reasoning Tasks

AAAI 2025technical

Deductive reasoning is a crucial logical capability that assists us in solving complex problems based on existing knowledge. Although augmented by Chain-of-Thought prompts, Large Language Models (LLMs) might not follow the correct reasoning paths. Enhancing the deductive reasoning abilities of LLMs,…

2025

Sim-DETR: Unlock DETR for Temporal Sentence Grounding

ICCV 2025poster

Temporal sentence grounding aims to identify exact moments in a video that correspond to a given textual query, typically addressed with detection transformer (DETR) solutions. However, we find that typical strategies designed to enhance DETR do not improve, and may even degrade, its performance in…

Cited by 0SourcePDFScholar
2025

Thinking Before You Speak: A Proactive Test-time Scaling Approach

EMNLP 2025

Large Language Models (LLMs) often exhibit deficiencies with complex reasoning tasks, such as maths, which we attribute to the discrepancy between human reasoning patterns and those presented in the LLMs’ training data. When dealing with complex problems, humans tend to think carefully before expres

Cited by 0SourcePDFScholar
2025

Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method

CVPR 2025poster

Existing Vision-Language Navigation (VLN) methods primarily focus on single-stage navigation, limiting their effectiveness in multi-stage and long-horizon tasks within complex and dynamic environments. To address these limitations, we propose a novel VLN task, named Long-Horizon Vision-Language Navi…

Cited by 5SourcePDFScholar
2025

Towards Understanding the Robustness of Diffusion-Based Purification: A Stochastic Perspective

ICLR 2025poster

Diffusion-Based Purification (DBP) has emerged as an effective defense mechanism against adversarial attacks. The success of DBP is often attributed to the forward diffusion process, which reduces the distribution gap between clean and adversarial images by adding Gaussian noise. Although this expla…

Cited by 0SourcePDFScholar
2025

VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction

CVPR 2025poster

Virtual Try-On (VTON) is a transformative technology in e-commerce and fashion design, enabling realistic digital visualization of clothing on individuals. In this work, we propose VTON 360, a novel 3D VTON method that addresses the open challenge of achieving high-fidelity VTON that supports any-vi…

Cited by 1SourcePDFScholar
2025

Why Multi-Interest Fairness Matters: Hypergraph Contrastive Multi-Interest Learning for Fair Conversational Recommender System

ACL 2025finding

Unfairness is a well-known challenge in Recommender Systems (RSs), often resulting in biased outcomes that disadvantage users or items based on attributes such as gender, race, age, or popularity. Although some approaches have started to improve fairness recommendation in offline or static contexts,…

2024

AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis

CVPR 2024highlight

Neural implicit fields have been a de facto standard in novel view synthesis. Recently there exist some methods exploring fusing multiple modalities within a single field aiming to share implicit features from different modalities to enhance reconstruction performance. However these modalities often…

2024

AttNS: Attention-Inspired Numerical Solving For Limited Data Scenarios

ICML 2024poster

We propose the attention-inspired numerical solver (AttNS), a concise method that helps the generalization and robustness issues faced by the AI-Hybrid numerical solver in solving differential equations due to limited data. AttNS is inspired by the effectiveness of attention modules in Residual Neur…

Cited by 5SourcePDFScholar
2024

Diagnosing and Rectifying Fake OOD Invariance: A Restructured Causal Approach

AAAI 2024technical

Invariant representation learning (IRL) encourages the prediction from invariant causal features to labels deconfounded from the environments, advancing the technical roadmap of out-of-distribution (OOD) generalization. Despite spotlights around, recent theoretical result verified that some causal f…

Cited by 1SourcePDFScholar
2024

EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE

AAAI 2024technical

Building scalable vision-language models to learn from diverse, multimodal data remains an open challenge. In this paper, we introduce an Efficient Vision-languagE foundation model, namely EVE, which is one unified multimodal Transformer pre-trained solely by one unified pre-training task. Specifica…

Cited by 11SourcePDFScholar
2024

FacetCRS: Multi-Faceted Preference Learning for Pricking Filter Bubbles in Conversational Recommender System

AAAI 2024technical

The filter bubble is a notorious issue in Recommender Systems (RSs), which describes the phenomenon whereby users are exposed to a limited and narrow range of information or content that reinforces their existing dominant preferences and beliefs. This results in a lack of exposure to diverse and var…

Cited by 0SourcePDFScholar
2024

HyCoRec: Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation

ACL 2024long

The Matthew effect is a notorious issue in Recommender Systems (RSs), i.e., the rich get richer and the poor get poorer, wherein popular items are overexposed while less popular ones are regularly ignored. Most methods examine Matthew effect in static or nearly-static recommendation scenarios. Howev…

2024

Learning Adaptive Spatial Coherent Correlations for Speech-Preserving Facial Expression Manipulation

CVPR 2024highlight

Speech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current works depend on inaccessible paired training samples for the person where two aligned frames exhibit the same speech cont…

2024

Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection

CVPR 2024poster

Open vocabulary object detection (OVD) aims at seeking an optimal object detector capable of recognizing objects from both base and novel categories. Recent advances leverage knowledge distillation to transfer insightful knowledge from pre-trained large-scale vision-language models to the task of ob…

Cited by 16SourcePDFScholar
2024

Let's Think Outside the Box: Exploring Leap-of-Thought in Large Language Models with Creative Humor Generation

CVPR 2024poster

Chain-of-Thought (CoT) guides large language models (LLMs) to reason step-by-step and can motivate their logical reasoning ability. While effective for logical tasks CoT is not conducive to creative problem-solving which often requires out-of-box thoughts and is crucial for innovation advancements.…

2024

MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection

ECCV 2024poster

"Learning from pseudo-labels that generated with VLMs (Vision Language Models) has been shown as a promising solution to assist open vocabulary detection (OVD) in recent studies. However, due to the domain gap between VLM and vision-detection tasks, pseudo-labels produced by the VLMs are prone to be…

2024

Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation

EMNLP 2024main

The Matthew effect is a big challenge in Recommender Systems (RSs), where popular items tend to receive increasing attention, while less popular ones are often overlooked, perpetuating existing disparities. Although many existing methods attempt to mitigate Matthew effect in the static or quasi-stat…

2024

VisDiaHalBench: A Visual Dialogue Benchmark For Diagnosing Hallucination in Large Vision-Language Models

ACL 2024long

Despite the significant success of large vision-language models (LVLMs), some studies have revealed that LVLMs suffer from the hallucination problem, where the LVLMs’ response contains descriptions of non-existent objects. Although various benchmarks have been proposed to investigate this problem, t…

2024

WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models

ECCV 2024poster

"Video virtual try-on aims to generate realistic sequences that maintain garment identity and adapt to a person’s pose and body shape in source videos. Traditional image-based methods, relying on warping and blending, struggle with complex human movements and occlusions, limiting their effectiveness…

Cited by 11SourcePDFScholar
2023

A Retrospect to Multi-prompt Learning across Vision and Language

ICCV 2023poster

The vision community is undergoing the unprecedented progress with the emergence of Vision-Language Pretraining Models (VLMs). Prompt learning plays as the holy grail of accessing VLMs since it enables their fast adaptation to downstream tasks with limited resources. Whereas existing research millin…

Cited by 7PDFcodeScholar
2023

Actional Atomic-Concept Learning for Demystifying Vision-Language Navigation

AAAI 2023technical

Vision-Language Navigation (VLN) is a challenging task which requires an agent to align complex visual observations to language instructions to reach the goal position. Most existing VLN agents directly learn to align the raw directional features and visual features trained using one-hot labels to l…

Cited by 5SourcePDFScholar
2023

Adapting Object Size Variance and Class Imbalance for Semi-supervised Object Detection

AAAI 2023technical

Semi-supervised object detection (SSOD) attracts extensive research interest due to its great significance in reducing the data annotation effort. Collecting high-quality and category-balanced pseudo labels for unlabeled images is critical to addressing the SSOD problem. However, most of the existin…

Cited by 13SourcePDFScholar
2023

Being Comes From Not-Being: Open-Vocabulary Text-to-Motion Generation With Wordless Training

CVPR 2023highlight

Text-to-motion generation is an emerging and challenging problem, which aims to synthesize motion with the same semantics as the input text. However, due to the lack of diverse labeled training data, most approaches either limit to specific types of text annotations or require online optimizations t…

2023

Coordinate Transformer: Achieving Single-stage Multi-person Mesh Recovery from Videos

ICCV 2023poster

Multi-person 3D mesh recovery from videos is a critical first step towards automatic perception of group behavior in virtual reality, physical therapy and beyond. However, existing approaches rely on multi-stage paradigms, where the person detection and tracking stages are performed in a multi-perso…

Cited by 5PDFcodeScholar
2023

De-biased Teacher: Rethinking IoU Matching for Semi-supervised Object Detection

AAAI 2023technical

Most of the recent research in semi-supervised object detection follows the pseudo-labeling paradigm evolved from the semi-supervised image classification task. However, the training paradigm of the two-stage object detector inevitably makes the pseudo-label learning process for unlabeled images ful…

2023

DenseLight: Efficient Control for Large-scale Traffic Signals with Dense Feedback

IJCAI 2023poster

Traffic Signal Control (TSC) aims to reduce the average travel time of vehicles in a road network, which in turn enhances fuel utilization efficiency, air quality, and road safety, benefiting society as a whole. Due to the complexity of long-horizon control and coordination, most prior TSC methods l…

2023

DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment

ICCV 2023poster

Cross-modal garment synthesis and manipulation will significantly benefit the way fashion designers generate garments and modify their designs via flexible linguistic interfaces. However, despite the significant progress that has been made in generic image synthesis using diffusion models, producing…

Cited by 17PDFScholar
2023

HutCRS: Hierarchical User-Interest Tracking for Conversational Recommender System

EMNLP 2023long main

Conversational Recommender System (CRS) aims to explicitly acquire user preferences towards items and attributes through natural language conversations. However, existing CRS methods ask users to provide explicit answers (yes/no) for each attribute they require, regardless of users' knowledge or int…

Cited by 0SourcecodeScholar
2023

Identity-Preserving Talking Face Generation With Landmark and Appearance Priors

CVPR 2023poster

Generating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for training or fine-tuning. Existing person-generic methods have difficulty in generating realistic and lip-synced videos whi…

2023

LAW-Diffusion: Complex Scene Generation by Diffusion with Layouts

ICCV 2023poster

Thanks to the rapid development of diffusion models, unprecedented progress has been witnessed in image synthesis. Prior works mostly rely on pre-trained linguistic models, but a text is often too abstract to properly specify all the spatial properties of an image, e.g., the layout configuration of…

Cited by 14PDFScholar
2023

Long-term Wind Power Forecasting with Hierarchical Spatial-Temporal Transformer

IJCAI 2023poster

Wind power is attracting increasing attention around the world due to its renewable, pollution-free, and other advantages. However, safely and stably integrating the high permeability intermittent power energy into electric power systems remains challenging. Accurate wind power forecasting (WPF) can…

2023

Masked Images Are Counterfactual Samples for Robust Fine-Tuning

CVPR 2023poster

Deep learning models are challenged by the distribution shift between the training data and test data. Recently, the large models pre-trained on diverse data have demonstrated unprecedented robustness to various distribution shifts. However, fine-tuning these models can lead to a trade-off between i…

2023

RankMatch: Fostering Confidence and Consistency in Learning with Noisy Labels

ICCV 2023poster

Learning with noisy labels (LNL) is one of the most important and challenging problems in weakly-supervised learning. Recent advances adopt the sample selection strategy to mitigate the interference of noisy labels and use small-loss criteria to select clean samples. However, the one-dimensional los…

Cited by 14PDFScholar
2023

ScaleLong: Towards More Stable Training of Diffusion Model via Scaling Network Long Skip Connection

NeurIPS 2023poster

In diffusion models, UNet is the most popular network backbone, since its long skip connects (LSCs) to connect distant network blocks can aggregate long-distant information and alleviate vanishing gradient. Unfortunately, UNet often suffers from unstable training in diffusion models which can be all…

2023

SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-training

ICCV 2023poster

Skeleton sequence representation learning has shown great advantages for action recognition due to its promising ability to model human joints and topology. However, the current methods usually require sufficient labeled data for training computationally expensive models. Moreover, these methods ign…

Cited by 59PDFcodeScholar
2023

Towards Real-World Burst Image Super-Resolution: Benchmark and Method

ICCV 2023poster

Despite substantial advances, single-image super-resolution (SISR) is always in a dilemma to reconstruct high-quality images with limited information from one input image, especially in realistic scenarios. In this paper, we establish a large-scale real-world burst super-resolution dataset, i.e., Re…

Cited by 16PDFcodeScholar
2023

Understanding Self-attention Mechanism via Dynamical System Perspective

ICCV 2023poster

The self-attention mechanism (SAM) is widely used in various fields of artificial intelligence and has successfully boosted the performance of different models. However, current explanations of this mechanism are mainly based on intuitions and experiences, while there still lacks direct modeling for…

Cited by 24PDFScholar
2022

Continual Object Detection via Prototypical Task Correlation Guided Gating Mechanism

CVPR 2022poster

Continual learning is a challenging real-world problem for constructing a mature AI system when data are provided in a streaming fashion. Despite recent progress in continual classification, the researches of continual object detection are impeded by the diverse sizes and numbers of objects in each…

Cited by 44PDFcodeScholar
2022

Divide and Contrast: Source-free Domain Adaptation via Adaptive Contrastive Learning

NeurIPS 2022accept

We investigate a practical domain adaptation task, called source-free domain adaptation (SFUDA), where the source pretrained model is adapted to the target domain without access to the source data. Existing techniques mainly leverage self-supervised pseudo-labeling to achieve class-wise global align…

2022

Double-Check Soft Teacher for Semi-Supervised Object Detection

IJCAI 2022poster

In the semi-supervised object detection task, due to the scarcity of labeled data and the diversity and complexity of objects to be detected, the quality of pseudo-labels generated by existing methods for unlabeled data is relatively low, which severely restricts the performance of semi-supervised o…

2022

Dual Adversarial Adaptation for Cross-Device Real-World Image Super-Resolution

CVPR 2022oral

Due to the sophisticated imaging process, an identical scene captured by different cameras could exhibit distinct imaging patterns, introducing distinct proficiency among the super-resolution (SR) models trained on images from different devices. In this paper, we investigate a novel and practical ta…

Cited by 21PDFcodeScholar
2022

Enhancing Prototypical Few-Shot Learning By Leveraging The Local-Level Strategy

ICASSP 2022accepted

Aiming at recognizing the samples from novel categories with few reference samples, few-shot learning (FSL) is a challenging problem. We found that the existing works often build their few-shot model based on the image-level feature by mixing all local-level features, which leads to the discriminati…

Cited by 0SourceScholar
2022

LogicSolver: Towards Interpretable Math Word Problem Solving with Logical Prompt-enhanced Learning

EMNLP 2022finding

Recently, deep learning models have made great progress in MWP solving on answer accuracy. However, they are uninterpretable since they mainly rely on shallow heuristics to achieve high performance without understanding and reasoning the grounded math logic. To address this issue and make a step tow…

2022

Semantic-Aware Auto-Encoders for Self-Supervised Representation Learning

CVPR 2022poster

The resurgence of unsupervised learning can be attributed to the remarkable progress of self-supervised learning, which includes generative (G) and discriminative (D) models. In computer vision, the mainstream self-supervised learning algorithms are D models. However, designing a D model could be ov…

Cited by 11PDFcodeScholar
2022

Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels

AAAI 2022technical

Training the multi-label image recognition models with partial labels, in which merely some labels are known while others are unknown for each image, is a considerably challenging and practical task. To address this task, current algorithms mainly depend on pre-training classification or similarity…

2022

Structure-Preserving 3D Garment Modeling with Neural Sewing Machines

NeurIPS 2022accept

3D Garment modeling is a critical and challenging topic in the area of computer vision and graphics, with increasing attention focused on garment representation learning, garment reconstruction, and controllable garment manipulation, whereas existing methods were constrained to model garments under…

Cited by 16SourcePDFScholar
2022

Structured Semantic Transfer for Multi-Label Recognition with Partial Labels

AAAI 2022technical

Multi-label image recognition is a fundamental yet practical task because real-world images inherently possess multiple semantic labels. However, it is difficult to collect large-scale multi-label annotations due to the complexity of both the input images and output label spaces. To reduce the annot…

2022

UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression

EMNLP 2022main

Geometry problem solving is a well-recognized testbed for evaluating the high-level multi-modal reasoning capability of deep models. In most existing works, two main geometry problems: calculation and proving, are usually treated as two specific tasks, hindering a deep model to unify its reasoning c…

2022

Unsupervised Domain Adaptive Salient Object Detection through Uncertainty-Aware Pseudo-Label Learning

AAAI 2022technical

Recent advances in deep learning significantly boost the performance of salient object detection (SOD) at the expense of labeling larger-scale per-pixel annotations. To relieve the burden of labor-intensive labeling, deep unsupervised SOD methods have been proposed to exploit noisy labels generated…

2021

AU-Expression Knowledge Constrained Representation Learning for Facial Expression Recognition

ICRA 2021poster

Recognizing human emotion/expressions automatically is quite an expected ability for intelligent robotics, as it can promote better communication and cooperation with humans. Current deep-learning-based algorithms may achieve impressive performance in some lab-controlled environments, but they alway…

Cited by 28SourcecodeScholar
2021

Adversarial Meta Sampling for Multilingual Low-Resource Speech Recognition

AAAI 2021technical

Low-resource automatic speech recognition (ASR) is challenging, as the low-resource target language data cannot well train an ASR model. To solve this issue, meta-learning formulates ASR for each source language into many small ASR tasks and meta-learns a model initialization on all tasks from diffe…

Cited by 35SourcePDFScholar
2021

Continuous Transition: Improving Sample Efficiency for Continuous Control Problems via MixUp

ICRA 2021poster

Although deep reinforcement learning (RL) has been successfully applied to a variety of robotic control tasks, it’s still challenging to apply it to real-world tasks, due to the poor sample efficiency. Attempting to overcome this shortcoming, several works focus on reusing the collected trajectory d…

Cited by 17SourcecodeScholar
2021

Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting

CVPR 2021poster

Crowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited information of RGB images and cannot well discover potential pedestrians in unconstrained scenarios. In this work, we f…

Cited by 165PDFcodeScholar
2021

Deductive Learning for Weakly-Supervised 3D Human Pose Estimation via Uncalibrated Cameras

AAAI 2021technical

Without prohibitive and laborious 3D annotations, weakly-supervised 3D human pose methods mainly employ the model regularization with geometric projection consistency or geometry estimation from multi-view images. Nevertheless, those approaches explicitly need known parameters of calibrated cameras,…

2021

Graph-Evolving Meta-Learning for Low-Resource Medical Dialogue Generation

AAAI 2021technical

Human doctors with well-structured medical knowledge can diagnose a disease merely via a few conversations with patients about symptoms. In contrast, existing knowledge-grounded dialogue systems often require a large number of dialogue instances to learn as they fail to capture the correlations betw…

2021

Linguistically Routing Capsule Network for Out-of-Distribution Visual Question Answering

ICCV 2021poster

Generalization on out-of-distribution (OOD) test data is an essential but underexplored topic in visual question answering. Current state-of-the-art VQA models often exploit the biased correlation between data and labels, which results in a large performance drop when the test and training data have…

Cited by 16PDFScholar
2021

Neural-Symbolic Solver for Math Word Problems with Auxiliary Tasks

ACL 2021long

Previous math word problem solvers following the encoder-decoder paradigm fail to explicitly incorporate essential math symbolic constraints, leading to unexplainable and unreasonable predictions. Herein, we propose Neural-Symbolic Solver (NS-Solver) to explicitly and seamlessly incorporate differen…

2021

Pi-NAS: Improving Neural Architecture Search by Reducing Supernet Training Consistency Shift

ICCV 2021poster

Recently proposed neural architecture search (NAS) methods co-train billions of architectures in a supernet and estimate their potential accuracy using the network weights detached from the supernet. However, the ranking correlation between the architectures' predicted accuracy and their actual capa…

Cited by 22PDFcodeScholar
2021

Rethinking the Pruning Criteria for Convolutional Neural Network

NeurIPS 2021poster

Channel pruning is a popular technique for compressing convolutional neural networks (CNNs), where various pruning criteria have been proposed to remove the redundant filters. From our comprehensive experiments, we found two blind spots of pruning criteria: (1) Similarity: There are some strong simi…

Cited by 67SourcePDFScholar
2021

Solving Inefficiency of Self-Supervised Representation Learning

ICCV 2021poster

Self-supervised learning (especially contrastive learning) has attracted great interest due to its huge potential in learning discriminative representations in an unsupervised manner. Despite the acknowledged successes, existing contrastive learning methods suffer from very low learning efficiency,…

Cited by 66PDFcodeScholar
2021

Towards Quantifiable Dialogue Coherence Evaluation

ACL 2021long

Automatic dialogue coherence evaluation has attracted increasing attention and is crucial for developing promising dialogue systems. However, existing metrics have two major limitations: (a) they are mostly trained in a simplified two-level setting (coherent vs. incoherent), while humans give Likert…

2021

Trash To Treasure: Harvesting OOD Data With Cross-Modal Matching for Open-Set Semi-Supervised Learning

ICCV 2021poster

Open-set semi-supervised learning (open-set SSL) investigates a challenging but practical scenario where out-of-distribution (OOD) samples are contained in the unlabeled data. While the mainstream technique seeks to completely filter out the OOD samples for semi-supervised learning (SSL), we propose…

Cited by 76PDFScholar
2021

Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition

EMNLP 2021finding

Unifying acoustic and linguistic representation learning has become increasingly crucial to transfer the knowledge learned on the abundance of high-resource language data for low-resource speech recognition. Existing approaches simply cascade pre-trained acoustic and language models to learn the tra…

2021

Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

IJCAI 2021poster

In this paper, we introduce a novel task, referred to as Weakly-Supervised Spatio-Temporal Anomaly Detection (WSSTAD) in surveillance video. Specifically, given an untrimmed video, WSSTAD aims to localize a spatio-temporal tube (i.e., a sequence of bounding boxes at consecutive times) that encloses…

Cited by 75SourcePDFScholar
2020

Auto-Panoptic: Cooperative Multi-Component Architecture Search for Panoptic Segmentation

NeurIPS 2020poster

Panoptic segmentation is posed as a new popular test-bed for the state-of-the-art holistic scene understanding methods with the requirement of simultaneously segmenting both foreground things and background stuff. The state-of-the-art panoptic segmentation network exhibits high structural complexity…

2020

Bidirectional Graph Reasoning Network for Panoptic Segmentation

CVPR 2020poster

Recent researches on panoptic segmentation resort to a single end-to-end network to combine the tasks of instance segmentation and semantic segmentation. However, prior models only unified the two related tasks at the architectural level via a multi-branch scheme or revealed the underlying correlati…

Cited by 79PDFScholar
2020

Block-Wisely Supervised Neural Architecture Search With Knowledge Distillation

CVPR 2020poster

Neural Architecture Search (NAS), aiming at automatically designing network architectures by machines, is expected to bring about a new revolution in machine learning. Despite these high expectation, the effectiveness and efficiency of existing NAS solutions are unclear, with some recent works going…

Cited by 244PDFcodeScholar
2020

Collaborative Training between Region Proposal Localization and Classification for Domain Adaptive Object Detection

ECCV 2020poster

Object detectors are usually trained with large amount of labeled data, which is expensive and labor-intensive. Pre-trained detectors applied to unlabeled dataset always suffer from the difference of dataset distribution, also called domain shift. Domain adaptation for object detection tries to adap…

2020

Component Divide-and-Conquer for Real-World Image Super-Resolution

ECCV 2020poster

In this paper, we present a large-scale Diverse Real-world image Super-Resolution dataset, i.e., DRealSR, as well as a divide-and-conquer Super-Resolution (SR) network, exploring the utility of guiding SR model with low-level image components. DRealSR establishes a new SR benchmark with diverse real…

2020

Transferable, Controllable, and Inconspicuous Adversarial Attacks on Person Re-identification With Deep Mis-Ranking

CVPR 2020oral

The success of DNNs has driven the extensive applications of person re-identification (ReID) into a new era. However, whether ReID inherits the vulnerability of DNNs remains unexplored. To examine the robustness of ReID systems is rather important because the insecurity of ReID systems may cause sev…

Cited by 106PDFcodeScholar
2019

Blending-Target Domain Adaptation by Adversarial Meta-Adaptation Networks

CVPR 2019oral

(Unsupervised) Domain Adaptation (DA) seeks for classifying target instances when solely provided with source labeled and target unlabeled examples for training. Learning domain-invariant features helps to achieve this goal, whereas it underpins unlabeled samples drawn from a single or multiple expl…

Cited by 120PDFcodeScholar
2019

ClusterNet: Deep Hierarchical Cluster Network With Rigorously Rotation-Invariant Representation for Point Cloud Analysis

CVPR 2019poster

Current neural networks for 3D object recognition are vulnerable to 3D rotation. Existing works mostly rely on massive amounts of rotation-augmented data to alleviate the problem, which lacks solid guarantee of the 3D rotation invariance. In this paper, we address the issue by introducing a novel po…

Cited by 217PDFScholar
2019

Crowd Counting With Deep Structured Scale Integration Network

ICCV 2019poster

Automatic estimation of the number of people in unconstrained crowded scenes is a challenging task and one major difficulty stems from the huge scale variation of people. In this paper, we propose a novel Deep Structured Scale Integration Network (DSSINet) for crowd counting, which addresses the sca…

Cited by 305PDFScholar
2019

Fashion Retrieval via Graph Reasoning Networks on a Similarity Pyramid

ICCV 2019oral

Matching clothing images from customers and online shopping stores has rich applications in E-commerce. Existing algorithms encoded an image as a global feature vector and performed retrieval with the global representation. However, discriminative local information on clothes are submerged in this g…

Cited by 118PDFScholar
2019

Graphonomy: Universal Human Parsing via Graph Transfer Learning

CVPR 2019poster

Prior highly-tuned human parsing models tend to fit towards each dataset in a specific domain or with discrepant label granularity, and can hardly be adapted to other human parsing tasks without extensive re-training. In this paper, we aim to learn a single universal human parsing model that can tac…

Cited by 228PDFcodeScholar
2019

Larger Norm More Transferable: An Adaptive Feature Norm Approach for Unsupervised Domain Adaptation

ICCV 2019oral

Domain adaptation enables the learner to safely generalize into novel environments by mitigating domain shifts across distributions. Previous works may not effectively uncover the underlying reasons that would lead to the drastic model degradation on the target task. In this paper, we empirically re…

Cited by 656PDFcodeScholar
2019

Layout-Graph Reasoning for Fashion Landmark Detection

CVPR 2019poster

Detecting dense landmarks for diverse clothes, as a fundamental technique for clothes analysis, has attracted increasing research attention due to its huge application potential. However, due to the lack of modeling underlying semantic layout constraints among landmarks, prior works often detect amb…

Cited by 52PDFScholar
2019

Learning Semantic-Specific Graph Representation for Multi-Label Image Recognition

ICCV 2019poster

Recognizing multiple labels of images is a practical and challenging task, and significant progress has been made by searching semantic-aware regions and modeling label dependency. However, current methods cannot locate the semantic regions accurately due to the lack of part-level supervision or sem…

Cited by 384PDFcodeScholar
2019

Meta R-CNN: Towards General Solver for Instance-Level Low-Shot Learning

ICCV 2019poster

Resembling the rapid learning capability of human, low-shot learning empowers vision systems to understand new concepts by training with few samples. Leading approaches derived from meta-learning on images with a single visual object. Obfuscated by a complex background and multiple objects in one im…

Cited by 651PDFcodeScholar
2019

Multivariate-Information Adversarial Ensemble for Scalable Joint Distribution Matching

ICML 2019oral

A broad range of cross-$m$-domain generation researches boil down to matching a joint distribution by deep generative models (DGMs). Hitherto algorithms excel in pairwise domains while as $m$ increases, remain struggling to scale themselves to fit a joint distribution. In this paper, we propose a dom…

2019

NADPEx: An on-policy temporally consistent exploration method for deep reinforcement learning

ICLR 2019poster

Reinforcement learning agents need exploratory behaviors to escape from local optima. These behaviors may include both immediate dithering perturbation and temporally consistent exploration. To achieve these, a stochastic policy model that is inherently consistent through a period of time is in desi…

Cited by 9SourcePDFScholar
2019

Reasoning-RCNN: Unifying Adaptive Global Reasoning Into Large-Scale Object Detection

CVPR 2019oral

In this paper, we address the large-scale object detection problem with thousands of categories, which poses severe challenges due to long-tail data distributions, heavy occlusions, and class ambiguities. However, the dominant object detection paradigm is limited by treating each object region separ…

Cited by 117PDFcodeScholar
2019

Semi-Supervised Video Salient Object Detection Using Pseudo-Labels

ICCV 2019poster

Deep learning-based video salient object detection has recently achieved great success with its performance significantly outperforming any other unsupervised methods. However, existing data-driven approaches heavily rely on a large quantity of pixel-wise annotated video frames to deliver such promi…

Cited by 154PDFScholar
2019

Spatially Variant Linear Representation Models for Joint Filtering

CVPR 2019poster

Joint filtering mainly uses an additional guidance image as a prior and transfers its structures to the target image in the filtering process. Different from existing algorithms that rely on locally linear models or hand-designed objective functions to extract the structural information from the gui…

Cited by 53PDFScholar
2019

Weakly-Supervised Discovery of Geometry-Aware Representation for 3D Human Pose Estimation

CVPR 2019oral

Recent studies have shown remarkable advances in 3D human pose estimation from monocular images, with the help of large-scale in-door 3D datasets and sophisticated network architectures. However, the generalizability to different environments remains an elusive goal. In this work, we propose a geome…

Cited by 139PDFScholar
2018

Avoidance of High-Speed Obstacles Based on Velocity Obstacles

ICRA 2018poster

For obstacles moving with high speeds, existing motion planning methods can rarely guarantee collision avoidance. This paper proposes a viable two-period velocity obstacle algorithm where one period predicts potential collisions within a limited time horizon, and the second period foresees collision…

Cited by 20SourceScholar
2018

Crafting a Toolchain for Image Restoration by Deep Reinforcement Learning

CVPR 2018poster

We investigate a novel approach for image restoration by reinforcement learning. Unlike existing studies that mostly train a single large network for a specialized task, we prepare a toolbox consisting of small-scale convolutional networks of different complexities and specialized in different tasks…

Cited by 235SourcePDFScholar
2018

Deep Cocktail Network: Multi-Source Unsupervised Domain Adaptation With Category Shift

CVPR 2018poster

Most existing unsupervised domain adaptation (UDA) methods are based upon the assumption that source labeled data come from an identical underlying distribution. Whereas in practical scenario, labeled instances are typically collected from diverse sources. Moreover, those sources may not completely…

2018

Embedding Temporally Consistent Depth Recovery for Real-time Dense Mapping in Visual-inertial Odometry

IROS 2018poster

Dense mapping is always the desire of simultaneous localization and mapping (SLAM), especially for the applications that require fast and dense scene information. Visual-inertial odometry (VIO) is a light-weight and effective solution to fast self-localization. However, VIO-based SLAM systems have d…

Cited by 3SourceScholar
2018

Flow Guided Recurrent Neural Encoder for Video Salient Object Detection

CVPR 2018poster

Image saliency detection has recently witnessed significant progress due to deep convolutional neural networks. However, extending state-of-the-art saliency detectors from image to video is challenging. The performance of salient object detection suffers from object or camera motion and the dramatic…

Cited by 205SourcePDFScholar
2018

Fusing Object Context to Detect Functional Area for Cognitive Robots

ICRA 2018poster

A cognitive robot usually needs to perform multiple tasks in practice and needs to locate the desired area for each task. Since deep learning has achieved substantial progress in image recognition, to solve this area detection problem, it is straightforward to label a functional area (affordance) im…

Cited by 0SourceScholar
2018

Hybrid Knowledge Routed Modules for Large-scale Object Detection

NeurIPS 2018poster

Abstract The dominant object detection approaches treat the recognition of each region separately and overlook crucial semantic correlations between objects in one scene. This paradigm leads to substantial performance drop when facing heavy long-tail problems, where very few samples are available fo…

2018

Instance-level Human Parsing via Part Grouping Network

ECCV 2018poster

Instance-level human parsing towards real-world human analysis scenarios is still under-explored due to the absence of sufficient data resources and technical difficulty in parsing multiple instances in a single pass. Several related works all follow the ``parsing-by-detection" pipeline that heavily…

2018

Interpretable Video Captioning via Trajectory Structured Localization

CVPR 2018poster

Automatically describing open-domain videos with natural language are attracting increasing interest in the field of artificial intelligence. Most existing methods simply borrow ideas from image captioning and obtain a compact video representation from an ensemble of global image feature before feed…

Cited by 70SourcePDFScholar
2018

Kalman Normalization: Normalizing Internal Representations Across Network Layers

NeurIPS 2018poster

As an indispensable component, Batch Normalization (BN) has successfully improved the training of deep neural networks (DNNs) with mini-batches, by normalizing the distribution of the internal representation for each hidden layer. However, the effectiveness of BN would diminish with the scenario of…

Cited by 31SourcePDFScholar
2018

Learning Warped Guidance for Blind Face Restoration

ECCV 2018poster

This paper studies the problem of blind face restoration from an unconstrained blurry, noisy, low-resolution, or compressed image (i.e., degraded observation). For better recovery of fine facial details, we modify the problem setting by taking both the degraded observation and a high-quality guided…

2018

Monocular Depth Estimation with Affinity, Vertical Pooling, and Label Enhancement

ECCV 2018poster

While significant progress has been made in monocular depth estimation with Convolutional Neural Networks (CNNs) extracting absolute features, such as edges and textures, the depth constraint of neighboring pixels, namely relative features, has been mostly ignored by recent methods. To overcome this…

Cited by 148SourcePDFScholar
2018

Robust Object-Aware Sample Consensus with Application to Lidar Odometry

ICASSP 2018accepted

Random sample consensus (RANSAC) is a popular paradigm for parameter estimation with outlier detection, which plays an essential role in 3D robot vision, especially for LiDAR odometry. The success of RANSAC strongly depends on the probability of selecting a subset of pure inliers, which sets barrier…

Cited by 0SourceScholar
2018

Toward Characteristic-Preserving Image-based Virtual Try-On Network

ECCV 2018poster

Image-based virtual try-on systems for fitting new in-shop clothes into a person image have attracted increasing research attention, yet is still challenging. A desirable pipeline should not only transform the target clothes into the most fitting shape seamlessly but also preserve well the clothes i…

2018

Towards Human-Machine Cooperation: Self-Supervised Sample Mining for Object Detection

CVPR 2018poster

Though quite challenging, leveraging large-scale unlabeled or partially labeled images in a cost-effective way has increasingly attracted interests for its great importance to computer vision. To tackle this problem, many Active Learning (AL) methods have been developed. However, these methods mainl…

Cited by 136SourcePDFScholar
2018

Visual Question Reasoning on General Dependency Tree

CVPR 2018poster

The collaborative reasoning for understanding each image-question pair is very critical but under-explored for an interpretable Visual Question Answering (VQA) system. Although very recent works also tried the explicit compositional processes to assemble multiple sub-tasks embedded in the questions…

Cited by 42SourcePDFScholar
2018

Zoom and Learn: Generalizing Deep Stereo Matching to Novel Domains

CVPR 2018poster

Despite the recent success of stereo matching with convolutional neural networks (CNNs), it remains arduous to generalize a pre-trained deep stereo model to a novel domain. A major difficulty is to collect accurate ground-truth disparities for stereo pairs in the target domain. In this work, we prop…

2017

Attention-Aware Face Hallucination via Deep Reinforcement Learning

CVPR 2017poster

Face hallucination is a domain-specific super-resolution problem with the goal to generate high-resolution (HR) faces from low-resolution (LR) input images. In contrast to existing methods that often learn a single patch-to-patch mapping from LR to HR images and are regardless of the contextual inte…

Cited by 249PDFScholar
2017

Decentralized navigation of multiple agents based on ORCA and model predictive control

IROS 2017poster

This paper presents a decentralized strategy for collision-free navigation of multiple agents. This strategy combines the Optimal Reciprocal Collision Avoidance (ORCA) algorithm and Model Predictive Control (MPC). Concretely, each agent applies the decentralized ORCA algorithm to compute the collisi…

Cited by 62SourceScholar
2017

Interpretable Structure-Evolving LSTM

CVPR 2017spotlight

This paper develops a general framework for learning interpretable data representation via Long Short-Term Memory (LSTM) recurrent neural networks over hierarchal graph structures. Instead of learning LSTM models over the pre-fixed structures, we propose to further learn the intermediate interpretab…

Cited by 121PDFScholar
2017

Joint Detection and Identification Feature Learning for Person Search

CVPR 2017spotlight

Existing person re-identification benchmarks and methods mainly focus on matching cropped pedestrian images between queries and candidates. However, it is different from real-world scenarios where the annotations of pedestrian bounding boxes are unavailable and the target person needs to be searched…

Cited by 1086PDFcodeScholar
2017

Learning Object Interactions and Descriptions for Semantic Image Segmentation

CVPR 2017poster

Recent advanced deep convolutional networks (CNNs) achieved great successes in many computer vision tasks, because of their compelling learning complexity and the presences of large-scale labeled data. However, as obtaining per-pixel annotations is expensive, performances of CNNs in semantic image s…

Cited by 58PDFScholar
2017

Look Into Person: Self-Supervised Structure-Sensitive Learning and a New Benchmark for Human Parsing

CVPR 2017poster

Human parsing has recently attracted a lot of research interests due to its huge application potentials. However existing datasets have limited number of images and annotations, and lack the variety of human appearances and the coverage of challenging cases in unconstrained environment. In this pape…

Cited by 621PDFcodeScholar
2017

Multi-Label Image Recognition by Recurrently Discovering Attentional Regions

ICCV 2017poster

This paper proposes a novel deep architecture to address multi-label image recognition, a fundamental and practical task towards general visual understanding. Current solutions for this task usually rely on an extra step of extracting hypothesis regions (i.e., region proposals), resulting in redunda…

Cited by 394PDFScholar
2016

Deep Structured Scene Parsing by Learning With Image Descriptions

CVPR 2016oral

This paper addresses the problem of structured scene parsing, i.e., parsing the input scene into a configuration including hierarchical semantic objects with their interaction relations. We propose a deep architecture consisting of two networks: i) a convolutional neural network (CNN) extracting the…

Cited by 40PDFScholar
2016

Dictionary Pair Classifier Driven Convolutional Neural Networks for Object Detection

CVPR 2016poster

Feature representation and object category classification are two key components of most object detection methods. While significant improvements have been achieved for deep feature representation learning, traditional SVM/softmax classifiers remain the dominant methods for final object category cla…

Cited by 53PDFScholar
2016

Joint Learning of Single-Image and Cross-Image Representations for Person Re-Identification

CVPR 2016poster

Person re-identification has been usually solved as either the matching of single-image representation (SIR) or the classification of cross-image representation (CIR). In this work, we exploit the connection between these two categories of methods, and propose a joint learning framework to unify SIR…

Cited by 518PDFScholar
2016

Reversible Recursive Instance-Level Object Segmentation

CVPR 2016poster

In this work, we propose a novel Reversible Recursive Instance-level Object Segmentation (R2-IOS) framework to address the challenging instance-level object segmentation task. R2-IOS consists of a reversible proposal refinement sub-network that predicts bounding box offsets for refining the object p…

Cited by 65PDFScholar
2016

Semantic Object Parsing With Local-Global Long Short-Term Memory

CVPR 2016spotlight

Semantic object parsing is a fundamental task for understanding objects in detail in computer vision community, where incorporating multi-level contextual information is critical for achieving such fine-grained pixel-level recognition. Prior methods often leverage the contextual information through…

Cited by 215PDFScholar
2015

Discriminative Learning of Iteration-Wise Priors for Blind Deconvolution

CVPR 2015poster

The maximum a posterior (MAP)-based blind deconvolution framework generally involves two stages: blur kernel estimation and non-blind restoration. For blur kernel estimation, sharp edge prediction and carefully designed image priors are vital to the success of MAP. In this paper, we propose a blind…

Cited by 49SourcePDFScholar
2015

Human Parsing With Contextualized Convolutional Neural Network

ICCV 2015oral

In this work, we address the human parsing task with a novel Contextualized Convolutional Neural Network (Co-CNN) architecture, which well integrates the cross-layer context, global image-level context, within-super-pixel context and cross-super-pixel neighborhood context into a unified network. Giv…

Cited by 356PDFScholar
2015

Matching-CNN Meets KNN: Quasi-Parametric Human Parsing

CVPR 2015poster

Both parametric and non-parametric approaches have demonstrated encouraging performances in the human parsing task, namely segmenting a human image into several semantic regions (e.g., hat, bag, left arm, face). In this work, we aim to develop a new solution with the advantages of both methodologie…

Cited by 203SourcePDFScholar
2015

SOLD: Sub-Optimal Low-rank Decomposition for Efficient Video Segmentation

CVPR 2015poster

This paper investigates how to perform robust and efficient unsupervised video segmentation while suppressing the effects of data noises and/or corruptions. We propose a general algorithm, called Sub-Optimal Low-rank Decomposition (SOLD), which pursues the low-rank representation for video segmentat…

Cited by 54SourcePDFScholar
2015

Towards Computational Baby Learning: A Weakly-Supervised Approach for Object Detection

ICCV 2015poster

Intuitive observations show that a baby may inherently possess the capability of recognizing a new visual concept (e.g., chair, dog) by learning from only very few positive instances taught by parent(s) or others, and this recognition capability can be gradually further improved by exploring and/or…

Cited by 114PDFScholar