← Search

Lin Ma

104 accepted papers

2026

Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation

ICLR 2026poster

While reinforcement learning (RL) has proven highly effective for general reasoning in vision-language models, its application to tasks requiring deep understanding of information-rich images and structured output generation remains underexplored. Chart-to-code generation exemplifies this challenge,…

Cited by 0SourcecodeScholar
2026

Cross-Modal Attention Calibration for LVLM Hallucination Mitigation

CVPR 2026

Large vision-language models (LVLMs) have shown remarkable capabilities in visual-language understanding. Despite their success, LVLMs still suffer from generating hallucinations in complex generation tasks, leading to inconsistencies between visual inputs and generated content. To address this issu

Cited by 0SourceScholar
2026

DBGroup: Dual-Branch Point Grouping for Weakly Supervised 3D Semantic Instance Segmentation

AAAI 2026technical

Weakly supervised 3D instance segmentation is essential for 3D scene understanding, especially as the growing scale of data and high annotation costs associated with fully supervised approaches. Existing methods primarily rely on two forms of weak supervision: one-thing-one-click annotations and bou

Cited by 0SourcePDFScholar
2026

Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction

CVPR 2026

We present Gen3R, a method that bridges the strong priors of foundational reconstruction models and video diffusion models for scene-level 3D generation. We repurpose the VGGT reconstruction model to produce geometric latents by training an adapter on its tokens, which are regularized to align with

Cited by 0SourcecodeScholar
2026

Leveraging Visual Blur Perception Characteristics for EEG Decoding

AAAI 2026technical

In recent years, electroencephalography (EEG)-based visual decoding research has become a key direction for revealing brain processing mechanisms and realizing brain-computer interfaces. This emerging field has attracted extensive attention in the fields of brain science, cognitive neuroscience, and

Cited by 0SourcePDFScholar
2026

M4V: Multimodal Mamba for Efficient Text-to-Video Generation

CVPR 2026

Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. However, modeling the vast spatiotemporal space remains computationally demanding, particularly when employing Transformers, which incur quadratic complexity in sequ

Cited by 0SourceScholar
2026

OddGridBench: Exposing the Lack of Fine-Grained Visual Discrepancy Sensitivity in Multimodal Large Language Models

CVPR 2026

Multimodal large language models (MLLMs) have achieved remarkable performance across a wide range of vision-language tasks. However, their ability in low-level visual perception, particularly in detecting fine-grained visual discrepancies, remains underexplored and lacks systematic analysis.In this

Cited by 0SourcecodeScholar
2026

OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds

ICLR 2026poster

Multimodal large language models are progressively advancing toward multimodal agents that can proactively execute tasks. Existing research on multimodal agents primarily targets either GUI or embodied scenarios, corresponding to interactions within 2D virtual world and 3D physical world, respective…

Cited by 0SourceScholar
2026

Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning

CVPR 2026

Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Large Language Models (LLMs) and is now being applied to Vision-Language Models (VLMs). However, vanilla RLVR for VLMs verifies only the final textual output, critically neglecting the foun

Cited by 0SourcecodeScholar
2026

Reading or Reasoning? Format Decoupled Reinforcement Learning for Document OCR

CVPR 2026

Reading text from images or scanned documents via OCR models has been a longstanding focus of researchers. Intuitively, text reading is perceived as a straightforward perceptual task, and existing work primarily focuses on constructing enriched data engineering to enhance SFT capabilities. In this w

Cited by 0SourcecodeScholar
2026

SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth Estimation

CVPR 2026

Panoramic depth estimation enables a complete 360^\circ understanding of 3D environments but faces significant challenges in generalizing to real-world scenes. While recent zero-shot depth models like Depth Anything achieve remarkable generalization on perspective images, their performance sharply d

Cited by 0SourceScholar
2026

SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start

ICLR 2026poster

Reinforcement learning (RL) with verifiable rewards has recently catalyzed a wave of “MLLM-r1” approaches that bring RL to vision language models. Most representative paradigms begin with a cold start, typically employing supervised fine-tuning (SFT), to initialize the policy before RL. However, SFT…

Cited by 0SourcecodeScholar
2026

Scalable Training for Vector-Quantized Networks with 100% Codebook Utilization

ICLR 2026poster

Vector quantization (VQ) is a key component in discrete tokenizers for image generation, but its training is often unstable due to straight-through estimation bias, one-step-behind updates, and sparse codebook gradients, which lead to suboptimal reconstruction performance and low codebook usage. In…

Cited by 0SourceScholar
2026

Stereo World Model: Camera-Guided Stereo Video Generation

CVPR 2026

We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within the RGB modality, while simultaneously grounding geometry dire

Cited by 0SourcecodeScholar
2026

UniComp: Rethinking Video Compression Through Informational Uniqueness

CVPR 2026

Distinct from attention-based compression methods, this paper presents an information uniqueness driven video compression framework, termed UniComp, which aims to maximize the information fidelity of video representations under constrained computational budgets. Starting from the information-theoret

Cited by 0SourcecodeScholar
2026

X-SAM: From Segment Anything to Any Segmentation

AAAI 2026technical

Large Language Models (LLMs) demonstrate strong capabilities in broad knowledge representation, yet they are inherently deficient in pixel-level perceptual understanding. Although the Segment Anything Model (SAM) represents a significant advancement in visual-prompt-driven image segmentation, it exh

Cited by 0SourcePDFScholar
2025

Affordances-Oriented Planning Using Foundation Models for Continuous Vision-Language Navigation

AAAI 2025technical

LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) task. However, existing LLM-based methods often focus only on solving high-level task planning by selecting nodes in predefined navigation graphs for movements, overlooking low-level control in na…

Cited by 8SourcePDFScholar
2025

CO-MOT: Boosting End-to-end Transformer-based Multi-Object Tracking via Coopetition Label Assignment and Shadow Sets

ICLR 2025poster

Existing end-to-end Multi-Object Tracking (e2e-MOT) methods have not surpassed non-end-to-end tracking-by-detection methods. One possible reason lies in the training label assignment strategy that consistently binds the tracked objects with tracking queries and assigns few newborns to detection quer…

2025

Climate Downscaling Using Neural Operator: Spatiotemporal Multimodal Fusion Operator with State-Query Coupled Kernel

ICASSP 2025accepted

Climate downscaling is crucial for detailed small- scale analysis and for acquiring climate data in regions without weather stations. Operator learning has proven potential for this task. However, several challenges remain in operator learning, such as multimodal fusion, spatiotemporal fusion and in…

Cited by 0SourceScholar
2025

DINOv2-Based UAV Visual Self-Localization in Low-Altitude Urban Environments

RA-L 2025

Visual self-localization technology is essential for unmanned aerial vehicles(UAVs) to achieve autonomous navigation and mission execution in environments where global navigation satellite system (GNSS) signals are unavailable. This technology estimates the UAV's geographic location by performing cr

Cited by 15SourceScholar
2025

DisTime: Distribution-based Time Representation for Video Large Language Models

ICCV 2025poster

Despite advances in general video understanding, Video Large Language Models (Video-LLMs) face challenges in precise temporal localization due to discrete time representations and limited temporally aware datasets. Existing methods for temporal expression either conflate time with text-based numeric…

2025

FlexVAR: Flexible Visual Autoregressive Modeling without Residual Prediction

NeurIPS 2025poster

This work challenges the residual prediction paradigm in visual autoregressive modeling and presents FlexVAR, a new Flexible Visual AutoRegressive image generation paradigm. FlexVAR facilitates autoregressive learning with ground-truth prediction, enabling each step to independently produce plausibl…

Cited by 0SourcecodeScholar
2025

GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection

NeurIPS 2025poster

Fine-grained open-vocabulary object detection (FG-OVD) aims to detect novel object categories described by attribute-rich texts. While existing open-vocabulary detectors show promise at the base-category level, they underperform in fine-grained settings due to the semantic entanglement of subjects a…

Cited by 0SourceScholar
2025

Learning Dynamical Coupled Operator For High-dimensional Black-box Partial Differential Equations

IJCAI 2025

The deep operator networks (DON), a class of neural operators that learn mappings between function spaces, have recently emerged as surrogate models for parametric partial differential equations (PDEs). However, their full potential for accurately approximating general black-box PDEs remains underex

2025

MCF-Spouse: A Multi-Label Causal Feature Selection Method with Optimal Spouses Discovery

IJCAI 2025

Multi-label causal feature selection has garnered considerable attention for its ability to identify the most informative features while accounting for the causal dependencies between labels and features. However, previous work often overlooks the unique contributions of labels to the target variabl

2025

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning

EMNLP 2025

Instruction tuning has underscored the significant potential of large language models (LLMs) in producing more human controllable and effective outputs in various domains. In this work, we focus on the data selection problem for task-specific instruction tuning of LLMs. Prevailing methods primarily

2025

RoboTrom-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and Prediction

ICCV 2025poster

In language-guided visual navigation, agents locate target objects in unseen environments using natural language instructions. For reliable navigation in unfamiliar scenes, agents should possess strong perception, planning, and prediction capabilities. Additionally, when agents revisit previously ex…

Cited by 0SourcePDFScholar
2025

RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

ICCV 2025poster

Large Multimodal Models (LMMs) have demonstrated exceptional comprehension and interpretation capabilities in Autonomous Driving (AD) by incorporating large language models. Despite the advancements, current data-driven AD approaches tend to concentrate on a single dataset and specific tasks, neglec…

Cited by 0SourcePDFScholar
2025

RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation

ICCV 2025poster

Recently, robotics has advanced significantly through the integration of larger models and large-scale datasets. However, challenges remain in applying these models to 3D spatial interactions and managing data collection costs. To address these issues, we propose the multimodal robotic manipulation…

2025

RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case

ICCV 2025poster

Collecting real-world data for rare high-risk scenarios, long-tailed driving events, and complex interactions remains challenging, leading to poor performance of existing autonomous driving systems in these critical situations. In this paper, we propose RoboTron-Sim that improves real-world driving…

Cited by 0SourcePDFScholar
2025

SSPNet: Leveraging Robust Medication Recommendation with History and Knowledge

IJCAI 2025

Automated medication recommendation is a crucial task within the domain of artificial intelligence in healthcare, where recommender systems are supposed to deliver precise, personalized drug combinations tailored to the evolving health states of patients. Existing approaches often treat clinical rec

2025

TimeStacker: A Novel Framework with Multilevel Observation for Capturing Nonstationary Patterns in Time Series Forecasting

ICML 2025poster

Real-world time series inherently exhibit significant non-stationarity, posing substantial challenges for forecasting. To address this issue, this paper proposes a novel prediction framework, TimeStacker, designed to overcome the limitations of existing models in capturing the characteristics of non…

Cited by 0SourcePDFScholar
2025

Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy

NeurIPS 2025poster

In this work, we first revisit the sampling issues in current autoregressive (AR) image generation models and identify that image tokens, unlike text tokens, exhibit lower information density and non-uniform spatial distribution. Accordingly, we present an entropy-informed decoding strategy that fac…

Cited by 0SourceScholar
2025

Towards Efficient Foundation Model for Zero-shot Amodal Segmentation

CVPR 2025poster

Aiming to predict the complete shape of partially occluded objects, amodal segmentation is an important capacity towards visual intelligence. In order to promote the practicability, zero-shot foundation model competent for the open world gains growing attention in this field. Nevertheless, prior mod…

Cited by 0SourcePDFScholar
2025

VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction-Editing Data and Long Captions

NeurIPS 2025poster

Despite the success of Vision-Language Models (VLMs) like CLIP in aligning vision and language, their proficiency in detailed, fine-grained visual comprehension remains a key challenge. We present CLIP-IN, a novel framework that bolsters CLIP's fine-grained perception through two core innovations. F…

Cited by 0SourcecodeScholar
2025

VITRIX-UniViTAR: Unified Vision Transformer with Native Resolution

NeurIPS 2025poster

Conventional Vision Transformer streamlines visual modeling by employing a uniform input resolution, which underestimates the inherent variability of natural visual data and incurs a cost in spatial-contextual fidelity. While preliminary explorations have superficially investigated native resolution…

Cited by 0SourceScholar
2024

3D Weakly Supervised Semantic Segmentation with 2D Vision-Language Guidance

ECCV 2024poster

"In this paper, we propose 3DSS-VLG, a weakly supervised approach for 3D Semantic Segmentation with 2D Vision-Language Guidance, an alternative approach that a 3D model predicts dense-embedding for each point which is co-embedded with both the aligned image and text spaces from the 2D vision-languag…

2024

A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation

COLING 2024main

In this paper, we propose a new setting for generating product descriptions from images, augmented by marketing keywords. It leverages the combined power of visual and textual information to create descriptions that are more tailored to the unique features of products. For this setting, previous met…

2024

AlignSAM: Aligning Segment Anything Model to Open Context via Reinforcement Learning

CVPR 2024poster

Powered by massive curated training data Segment Anything Model (SAM) has demonstrated its impressive generalization capabilities in open-world scenarios with the guidance of prompts. However the vanilla SAM is class-agnostic and heavily relies on user-provided prompts to segment objects of interest…

2024

Aux-NAS: Exploiting Auxiliary Labels with Negligibly Extra Inference Cost

ICLR 2024poster

We aim at exploiting additional auxiliary labels from an independent (auxiliary) task to boost the primary task performance which we focus on, while preserving a single task inference cost of the primary task. While most existing auxiliary learning methods are optimization-based relying on loss weig…

2024

ColNeRF: Collaboration for Generalizable Sparse Input Neural Radiance Field

AAAI 2024technical

Neural Radiance Fields (NeRF) have demonstrated impressive potential in synthesizing novel views from dense input, however, their effectiveness is challenged when dealing with sparse input. Existing approaches that incorporate additional depth or semantic supervision can alleviate this issue to an e…

2024

EEGPT: Pretrained Transformer for Universal and Reliable Representation of EEG Signals

NeurIPS 2024poster

Electroencephalography (EEG) is crucial for recording brain activity, with applications in medicine, neuroscience, and brain-computer interfaces (BCI). However, challenges such as low signal-to-noise ratio (SNR), high inter-subject variability, and channel mismatch complicate the extraction of…

2024

InstaGen: Enhancing Object Detection by Training on Synthetic Dataset

CVPR 2024poster

In this paper we present a novel paradigm to enhance the ability of object detector e.g. expanding categories or improving detection performance by training on syn- thetic dataset generated from diffusion models. Specifically we integrate an instance-level grounding head into a pre- trained generati…

Cited by 13SourcePDFScholar
2024

Instance-Aware Multi-Camera 3D Object Detection with Structural Priors Mining and Self-Boosting Learning

AAAI 2024technical

Camera-based bird-eye-view (BEV) perception paradigm has made significant progress in the autonomous driving field. Under such a paradigm, accurate BEV representation construction relies on reliable depth estimation for multi-camera images. However, existing approaches exhaustively predict depths fo…

2024

LESS: Label-Efficient and Single-Stage Referring 3D Segmentation

NeurIPS 2024poster

Referring 3D Segmentation is a visual-language task that segments all points of the specified object from a 3D point cloud described by a sentence of query. Previous works perform a two-stage paradigm, first conducting language-agnostic instance segmentation then matching with given text query. Howe…

2024

Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models

NeurIPS 2024poster

Large Multimodal Model (LMM) is a hot research topic in the computer vision area and has also demonstrated remarkable potential across multiple disciplinary fields. A recent trend is to further extend and enhance the perception capabilities of LMMs. The current methods follow the paradigm of adaptin…

2024

Making Large Language Models Better Planners with Reasoning-Decision Alignment

ECCV 2024oral

"Data-driven approaches for autonomous driving (AD) have been widely adopted in the past decade but are confronted with dataset bias and uninterpretability. Inspired by the knowledge-driven nature of human driving, recent approaches explore the potential of large language models (LLMs) to improve un…

Cited by 12SourcePDFScholar
2024

Misalignment-Robust Frequency Distribution Loss for Image Transformation

CVPR 2024poster

This paper aims to address a common challenge in deep learning-based image transformation methods such as image enhancement and super-resolution which heavily rely on precisely aligned paired datasets with pixel-level alignments. However creating precisely aligned paired images presents significant…

2024

Splatter a Video: Video Gaussian Representation for Versatile Processing

NeurIPS 2024poster

Video representation is a long-standing problem that is crucial for various downstream tasks, such as tracking, depth prediction, segmentation, view synthesis, and editing. However, current methods either struggle to model complex motions due to the absence of 3D structure or rely on implicit 3D rep…

Cited by 7SourcePDFScholar
2024

Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs

ACL 2024findings

Tables contrast with unstructured text data by its structure to organize the information.In this paper, we investigate the efficiency of various LLMs in interpreting tabular data through different prompting strategies and data formats. Our analysis extends across six benchmarks for table-related tas…

Cited by 10SourcePDFScholar
2024

UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection

ECCV 2024poster

"Temporal Action Detection (TAD) focuses on detecting pre-defined actions, while Moment Retrieval (MR) aims to identify the events described by open-ended natural language within untrimmed videos. Despite that they focus on different events, we observe they have a significant connection. For instanc…

2023

A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual Clues

ACL 2023long

Conditional inference on joint textual and visual clues is a multi-modal reasoning task that textual clues provide prior permutation or external knowledge, which are complementary with visual content and pivotal to deducing the correct option. Previous methods utilizing pretrained vision-language mo…

2023

A Neural Divide-and-Conquer Reasoning Framework for Image Retrieval from Linguistically Complex Text

ACL 2023long

Pretrained Vision-Language Models (VLMs) have achieved remarkable performance in image retrieval from text. However, their performance drops drastically when confronted with linguistically complex texts that they struggle to comprehend. Inspired by the Divide-and-Conquer algorithm and dual-process t…

2023

Adaptive Sparse Pairwise Loss for Object Re-Identification

CVPR 2023poster

Object re-identification (ReID) aims to find instances with the same identity as the given probe from a large gallery. Pairwise losses play an important role in training a strong ReID network. Existing pairwise losses densely exploit each instance as an anchor and sample its triplets in a mini-batch…

2023

AeDet: Azimuth-Invariant Multi-View 3D Object Detection

CVPR 2023poster

Recent LSS-based multi-view 3D object detection has made tremendous progress, by processing the features in Brid-Eye-View (BEV) via the convolutional detector. However, the typical convolution ignores the radial symmetry of the BEV features and increases the difficulty of the detector optimization.…

2023

Curriculum Multi-Negative Augmentation for Debiased Video Grounding

AAAI 2023technical

Video Grounding (VG) aims to locate the desired segment from a video given a sentence query. Recent studies have found that current VG models are prone to over-rely the groundtruth moment annotation distribution biases in the training set. To discourage the standard VG model's behavior of exploiting…

2023

MSMDFusion: Fusing LiDAR and Camera at Multiple Scales With Multi-Depth Seeds for 3D Object Detection

CVPR 2023poster

Fusing LiDAR and camera information is essential for accurate and reliable 3D object detection in autonomous driving systems. This is challenging due to the difficulty of combining multi-granularity geometric and semantic features from two drastically different modalities. Recent approaches aim at e…

2023

Open-Vocabulary Semantic Segmentation with Decoupled One-Pass Network

ICCV 2023poster

Recently, the open-vocabulary semantic segmentation problem has attracted increasing attention and the best performing methods are based on two-stream networks: one stream for proposal mask generation and the other for segment classification using a pre-trained visual-language model. However, existi…

Cited by 48PDFcodeScholar
2023

Planning Assembly Sequence with Graph Transformer

ICRA 2023poster

Assembly Sequence Planning (ASP) is the essential process for modern manufacturing, proven to be NP-complete thus its effective and efficient solution has been a challenge for researchers in the field. In this paper, we present a graph-transformer based framework for the ASP problem which is trained…

Cited by 23SourcecodeScholar
2023

Punctuation-level Attack: Single-shot and Single Punctuation Can Fool Text Models

NeurIPS 2023poster

The adversarial attacks have attracted increasing attention in various fields including natural language processing. The current textual attacking models primarily focus on fooling models by adding character-/word-/sentence-level perturbations, ignoring their influence on human perception. In this p…

Cited by 3SourcePDFScholar
2023

Tri-MipRF: Tri-Mip Representation for Efficient Anti-Aliasing Neural Radiance Fields

ICCV 2023oral

Despite the tremendous progress in neural radiance fields (NeRF), we still face a dilemma of the trade-off between quality and efficiency, e.g., MipNeRF presents fine-detailed and anti-aliased renderings but takes days for training, while Instant-ngp can accomplish the reconstruction in a few minute…

Cited by 142PDFcodeScholar
2023

TriDet: Temporal Action Detection With Relative Boundary Modeling

CVPR 2023poster

In this paper, we present a one-stage framework TriDet for temporal action detection. Existing methods often suffer from imprecise boundary predictions due to the ambiguous action boundaries in videos. To alleviate this problem, we propose a novel Trident-head to model the action boundary via an est…

2022

Expansion and Shrinkage of Localization for Weakly-Supervised Semantic Segmentation

NeurIPS 2022accept

Generating precise class-aware pseudo ground-truths, a.k.a, class activation maps (CAMs), is essential for Weakly-Supervised Semantic Segmentation. The original CAM method usually produces incomplete and inaccurate localization maps. To tackle with this issue, this paper proposes an Expansion and Sh…

2022

Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence Grounding

AAAI 2022technical

Weakly supervised temporal sentence grounding aims to temporally localize the target segment corresponding to a given natural language query, where it provides video-query pairs without temporal annotations during training. Most existing methods use the fused visual-linguistic feature to reconstruct…

2022

MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes

ECCV 2022poster

"3D dense captioning is a recently-proposed novel task, where point clouds contain more geometric information than the 2D counterpart. However, it is also more challenging due to the higher complexity and wider variety of inter-object relations contained in point clouds. Existing methods only treat…

2022

PromptDet: Towards Open-Vocabulary Detection Using Uncurated Images

ECCV 2022poster

"The goal of this work is to establish a scalable pipeline for expanding an object detector towards novel/unseen categories, using zero manual annotations. To achieve that, we make the following four contributions: (i) in pursuit of generalisation, we propose a two-stage open-vocabulary object detec…

2022

ReAct: Temporal Action Detection with Relational Queries

ECCV 2022poster

"This work aims at advancing temporal action detection (TAD) using an encoder-decoder framework with action queries, similar to DETR, which has shown great success in object detection. However, the framework suffers from several problems if directly applied to TAD: the insufficient exploration of in…

2021

Similarity Reasoning and Filtration for Image-Text Matching

AAAI 2021technical

Image-text matching plays a critical role in bridging the vision and language, and great progress has been made by exploiting the global alignment between image and sentence, or local alignments between regions and words. However, how to make the most of these alignments to infer more accurate match…

2020

Beyond Monocular Deraining: Stereo Image Deraining via Semantic Understanding

ECCV 2020poster

Rain is a common natural phenomenon. Taking images in the rain however often results in degraded quality of images, thus compromises the performance of many computer vision systems. Most existing de-rain algorithms use only one single input image and aim to recover a clean image. Few work has exploi…

Cited by 54SourcePDFScholar
2020

Consensus-Aware Visual-Semantic Embedding for Image-Text Matching

ECCV 2020poster

Image-text matching plays a central role in bridging vision and language. Most existing approaches only rely on the image-text instance pair to learn their representations, thereby exploiting their matching relationships and making the corresponding alignments. Such approaches only exploit the super…

2020

Cops-Ref: A New Dataset and Task on Compositional Referring Expression Comprehension

CVPR 2020poster

Referring expression comprehension (REF) aims at identifying a particular object in a scene by a natural language expression. It requires joint reasoning over the textual and visual domains to solve the problem. Some popular referring expression datasets, however, fail to provide an ideal test bed f…

Cited by 75PDFScholar
2020

Fine-Grained Image-to-Image Transformation Towards Visual Recognition

CVPR 2020poster

Existing image-to-image transformation approaches primarily focus on synthesizing visually pleasing data. Generating images with correct identity labels is challenging yet much less explored. It is even more challenging to deal with image transformation tasks with large deformation in poses, viewpoi…

Cited by 35PDFScholar
2019

Controllable Video Captioning With POS Sequence Guidance Based on Gated Fusion Network

ICCV 2019poster

In this paper, we propose to guide the video caption generation with Part-of-Speech (POS) information, based on a gated fusion of multiple representations of input videos. We construct a novel gated fusion network, with one particularly designed cross-gating (CG) block, to effectively encode and fus…

Cited by 232PDFcodeScholar
2019

Exploiting Local and Global Structure for Point Cloud Semantic Segmentation with Contextual Point Representations

NeurIPS 2019poster

In this paper, we propose one novel model for point cloud semantic segmentation,which exploits both the local and global structures within the point cloud based onthe contextual point representations. Specifically, we enrich each point represen-tation by performing one novel gated fusion on the poin…

2019

Image Deformation Meta-Networks for One-Shot Learning

CVPR 2019oral

Humans can robustly learn novel visual concepts even when images undergo various deformations and loose certain information. Mimicking the same behavior and synthesizing deformed instances of new concepts may help visual recognition systems perform better one-shot learning, i.e., learning concepts f…

Cited by 303PDFcodeScholar
2019

Liquid Warping GAN: A Unified Framework for Human Motion Imitation, Appearance Transfer and Novel View Synthesis

ICCV 2019poster

We tackle the human motion imitation, appearance transfer, and novel view synthesis within a unified framework, which means that the model once being trained can be used to handle all these tasks. The existing task-specific methods mainly use 2D keypoints (pose) to estimate the human body structure.…

Cited by 333PDFcodeScholar
2019

Semantic Conditioned Dynamic Modulation for Temporal Sentence Grounding in Videos

NeurIPS 2019poster

Temporal sentence grounding in videos aims to detect and localize one target video segment, which semantically corresponds to a given sentence. Existing methods mainly tackle this task via matching and aligning semantics between a sentence and candidate video segments, while neglect the fact that th…

2018

Bidirectional Attentive Fusion With Context Gating for Dense Video Captioning

CVPR 2018poster

Dense video captioning is a newly emerging task that aims at both localizing and describing all events in a video. We identify and tackle two challenges on this task, namely, (1) how to utilize both past and future contexts for accurate event proposal predictions, and (2) how to construct informativ…

Cited by 272SourcePDFScholar
2018

Deep Non-Blind Deconvolution via Generalized Low-Rank Approximation

NeurIPS 2018poster

In this paper, we present a deep convolutional neural network to capture the inherent properties of image degradation, which can handle different kernels and saturated pixels in a unified framework. The proposed neural network is motivated by the low-rank property of pseudo-inverse kernels. We firs…

Cited by 99SourcePDFScholar
2018

Gated Fusion Network for Single Image Dehazing

CVPR 2018poster

In this paper, we propose an efficient algorithm to directly restore a clear image from a hazy input. The proposed algorithm hinges on an end-to-end trainable neural network that consists of an encoder and a decoder. The encoder is exploited to capture the context of the derived input images, while…

Cited by 1032SourcePDFScholar
2018

Learning to Generate Time-Lapse Videos Using Multi-Stage Dynamic Generative Adversarial Networks

CVPR 2018poster

Taking a photo outside, can we predict the immediate future, e.g., how would the cloud move in the sky? We address this problem by presenting a generative adversarial network (GAN) based two-stage approach to generating realistic time-lapse videos of high resolution. Given the first frame, our model…

Cited by 209SourcePDFScholar
2018

Parsimonious Quantile Regression of Financial Asset Tail Dynamics via Sequential Learning

NeurIPS 2018poster

We propose a parsimonious quantile regression framework to learn the dynamic tail behaviors of financial asset returns. Our model captures well both the time-varying characteristic and the asymmetrical heavy-tail property of financial time series. It combines the merits of a popular sequential neura…

Cited by 31SourcePDFScholar
2018

Regularizing RNNs for Caption Generation by Reconstructing the Past With the Present

CVPR 2018poster

Recently, caption generation with an encoder-decoder framework has been extensively studied and applied in different domains, such as image captioning, code captioning, and so on. In this paper, we propose a novel architecture, namely Auto-Reconstructor Network (ARNet), which, coupling with the conv…

2018

Unsupervised Image-to-Image Translation with Stacked Cycle-Consistent Adversarial Networks

ECCV 2018poster

Recent studies on unsupervised image-to-image translation have made remarkable progress by training a pair of generative adversarial networks with a cycle-consistent loss. However, such unsupervised methods may generate inferior results when the image resolution is high or the two image domains are…

Cited by 124SourcePDFScholar
2015

Multi-task rank learning for image quality assessment

ICASSP 2015accepted

In practice, multiple types of distortions are associated with an image quality degradation process. The existing machine learning (ML) based image quality assessment (IQA) approaches generally established a unified model for all distortion types, or each model is trained independently for each dist…

Cited by 0SourceScholar