← Search

Lu Zhang

93 accepted papers

2026

A Diagnostic Study of Multi-Agent LLMs for Real-World Debates

ICML 2026poster

Multi-agent LLM debates are increasingly deployed in domains such as policy analysis and city planning, where no objective ground truth exists. Despite this, debate quality is typically evaluated using outcome-based proxies such as LLM-as-judge scores that provide little insight into whether meaning…

Cited by 0SourceScholar
2026

Agent as Student: Learning From Informative Cues for Active Open-Vocabulary Recognition

RA-L 2026

Active recognition, a fundamental task in embodied vision, aims to improve recognition performance by dynamically adjusting viewpoints and poses to mitigate the negative impacts of occlusion and blind spots. Although existing active recognition methods possess basic viewpoint adaptation capabilities

Cited by 0SourceScholar
2026

DePO: Demonstration-guided Policy Optimization for Molecular Optimization

ICLR 2026poster

Large language models (LLMs) exhibit remarkable mathematical reasoning abilities through supervised fine-tuning (SFT) or reinforcement learning with verifiable rewards (RLVR). However, adapting them to scientific domains like molecular optimization is challenging: its datasets provide only reference…

Cited by 0SourceScholar
2026

Disentangled Representation Learning for Parametric Partial Differential Equations

ICLR 2026poster

Neural operators (NOs) excel at learning mappings between function spaces, serving as efficient forward solution approximators for PDE-governed systems. However, as black-box solvers, they offer limited insight into the underlying physical mechanism, due to the lack of interpretable representations…

Cited by 0SourcecodeScholar
2026

GPR-GSLAM: Gaussian Process Regression-Enhanced Real-Time RGB-D SLAM Using Gaussian Splatting

RA-L 2026

3D Gaussian Splatting (3DGS) has recently revolutionized novel view synthesis and provided a new paradigm for photorealistic Simultaneous Localization and Mapping (SLAM). However, current 3DGS-based RGB-D SLAM systems still face three key challenges: incomplete depth observations due to sensor noise

Cited by 0SourceScholar
2026

MSAT-LDM: Toward Transferable High-Fidelity Watermarking for Latent Diffusion Model via Modular Self-Augmented Training

AAAI 2026technical

The rapid proliferation of AI-generated images necessitates effective watermarking techniques to protect intellectual property and detect fraudulent content. While existing training-based watermarking methods show promise, they often struggle with generalization across diverse prompts, introduce vis

Cited by 0SourcePDFScholar
2026

Reinforcing Video Object Segmentation to Think before it Segments

CVPR 2026

Video reasoning segmentation (VRS) endeavors to delineate referred objects in videos guided by implicit instructions that encapsulate human intent and temporal logic. Previous approaches leverage large vision language models (LVLMs) to encode object semantics into \SEG tokens for mask prediction. Ho

Cited by 0SourceScholar
2026

SeesawNet: Towards Non-stationary Time Series Forecasting with Balanced Modeling of Common and Specific Dependencies

IJCAI 2026

Instance normalization (IN) is widely used in non-stationary multivariate time series forecasting to reduce distribution shifts and highlight common patterns across samples. However, IN can over-smooth instance-specific structural information that is essential for modeling temporal and cross-channel

Cited by 0Scholar
2026

Structure-to-Intensity Diffusion for Adverse-Weather LiDAR Generation

CVPR 2026

Adverse-weather LiDAR point cloud generation is challenged by complex weather-induced degradations. These degradations affect geometry and reflectance in fundamentally different ways, making joint modeling difficult and ambiguous, especially when diverse real-world training data is limited. To addre

Cited by 0SourceScholar
2026

UETrack: A Unified and Efficient Framework for Single Object Tracking

CVPR 2026

With growing real-world demands, efficient tracking has received increasing attention. However, most existing methods are limited to RGB inputs and struggle in multi-modal scenarios. Moreover, current multi-modal tracking approaches typically use complex designs, making them too heavy and slow for r

Cited by 0SourcecodeScholar
2025

Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception

CVPR 2025poster

Large Vision-Language Models (LVLMs) have achieved impressive results across various cross-modal tasks. However, hallucinations, i.e., the models generating counterfactual responses, remain a challenge. Though recent studies have attempted to alleviate object perception hallucinations, they focus on…

2025

Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding

AAAI 2025technical

Injecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view segmentation and semantic understanding, their heavy reliance o…

2025

FedDiffRec: A Module-wise Training Approach for Diffusion-Based Recommendation in Federated Learning

ICASSP 2025accepted

Federated Learning (FL) has become a prominent framework for maintaining privacy in recommender systems by enabling decentralized model training. Despite its benefits, traditional Federated Recommender Systems (FRSs)—often relying on collaborative filtering or generative models such as Variational A…

Cited by 0SourceScholar
2025

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

NeurIPS 2025poster

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images---particul…

Cited by 0SourceScholar
2025

GARD: A Geometry-Informed and Uncertainty-Aware Baseline Method for Zero-Shot Roadside Monocular Object Detection

RA-L 2025

Roadside camera-based perception methods are in high demand for developing efficient vehicle-infrastructure collaborative perception systems. By focusing on object-level depth prediction, we explore the potential benefits of integrating environmental priors into such systems and propose a geometry-b

Cited by 1SourceScholar
2025

GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory Prediction

ICML 2025poster

Trajectory prediction for surrounding agents is a challenging task in autonomous driving due to its inherent uncertainty and underlying multimodality. Unlike prevailing data-driven methods that primarily rely on supervised learning, in this paper, we introduce a novel **G**raph-**o**riented **I**nve…

Cited by 0SourcePDFScholar
2025

Grammar-Based Code Representation: Is It a Worthy Pursuit for LLMs?

ACL 2025finding

Grammar serves as a cornerstone in programming languages and software engineering, providing frameworks to define the syntactic space and program structure. Existing research demonstrates the effectiveness of grammar-based code representations in small-scale models, showing their ability to reduce s…

Cited by 0SourcePDFScholar
2025

MT-PCR: Leveraging Modality Transformation for Large-Scale Point Cloud Registration with Limited Overlap

ICRA 2025

Large-scale scene point cloud registration with limited overlap is a challenging task due to computational load and constrained data acquisition. To tackle these issues, we propose a point cloud registration method, MT-PCR, based on Modality Transformation. MT-PCR leverages a Bird's Eye View (BEV) c

Cited by 0SourceScholar
2025

Multimodal Integrated Prediction and Decision-making with Adaptive Interaction Modality Explorations

IROS 2025

Navigating dense and dynamic environments poses a significant challenge for autonomous driving systems, owing to the intricate nature of multimodal interaction, wherein the actions of various traffic participants and the autonomous vehicle are complex and implicitly coupled. In this paper, we propos

Cited by 7SourcecodeScholar
2025

Optimal Information Retention for Time-Series Explanations

ICML 2025poster

Explaining deep models for time-series data is crucial for identifying key patterns in sensitive domains, such as healthcare and finance. However, due to the lack of unified optimization criterion, existing explanation methods often suffer from redundancy and incompleteness, where irrelevant pattern…

2025

Root Cause Analysis of Anomalies in Multivariate Time Series through Granger Causal Discovery

ICLR 2025oral

Identifying the root causes of anomalies in multivariate time series is challenging due to the complex dependencies among the series. In this paper, we propose a comprehensive approach called AERCA that inherently integrates Granger causal discovery with root cause analysis. By defining anomalies as…

Cited by 0SourcePDFScholar
2025

Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge

ICLR 2025poster

Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long video sequences, supporting multi-turn dialogues, and adapti…

2025

The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

CVPR 2025poster

Existing methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To overcome these challenges, we propose VRS-HQ, an end-to-end video reasoning segme…

2025

Towards Explicit Geometry-Reflectance Collaboration for Generalized LiDAR Segmentation in Adverse Weather

CVPR 2025poster

Existing LiDAR semantic segmentation models often suffer from decreased accuracy when exposed to adverse weather conditions. Recent methods addressing this issue focus on enhancing training data through weather simulation or universal augmentation techniques. However, few works have studied the nega…

Cited by 0SourcePDFScholar
2025

Zero-Shot Adaptation at Task-Level via Coarse-to-Fine Policy Refinement and Holistic-Local Contrastive Representation

RA-L 2025

Meta-reinforcement learning offers a mechanism for zero-shot adaptation, enabling agents to handle new tasks with parametric variation in real-world environments. However, existing methods still struggle with task-level adaptation, which demands generalization beyond simple variations within tasks,

Cited by 0SourceScholar
2025

posteriordb: Testing, Benchmarking and Developing Bayesian Inference Algorithms

AISTATS 2025oral

The general applicability and robustness of posterior inference algorithms is critical to widely used probabilistic programming languages such as Stan, PyMC, Pyro, and Turing.jl. When designing a new inference algorithm, whether it involves Monte Carlo sampling or variational approximation, the fund…

Cited by 0SourcecodeScholar
2024

BEVoxSeg: BEV-Voxel Representation for Fast and Accurate Camera-Based 3D Segmentation

ICASSP 2024accepted

Recent research has demonstrated the advantages of Bird’s-eye-view (BEV) representation in the field of 3D perception. However, due to the lack of height information, BEV representation alone is insufficient to accurately reconstruct the complete surrounding 3D scene. On the other hand, voxel repres…

Cited by 0SourceScholar
2024

Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters

CVPR 2024poster

Continual learning can empower vision-language models to continuously acquire new knowledge without the need for access to the entire historical dataset. However mitigating the performance degradation in large-scale models is non-trivial due to (i) parameter shifts throughout lifelong learning and (…

2024

Harnessing the Power of Neural Operators with Automatically Encoded Conservation Laws

ICML 2024spotlight

Neural operators (NOs) have emerged as effective tools for modeling complex physical systems in scientific machine learning. In NOs, a central characteristic is to learn the governing physical laws directly from data. In contrast to other machine learning applications, partial knowledge is often kno…

2024

LLMs Can Evolve Continually on Modality for $\mathbb{X}$-Modal Reasoning

NeurIPS 2024poster

Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily on extensive modal-specific pretraining and joint-modal tuning, leading to significant computational burdens when expand…

2024

Neural Atoms: Propagating Long-range Interaction in Molecular Graphs through Efficient Communication Channel

ICLR 2024poster

Graph Neural Networks (GNNs) have been widely adopted for drug discovery with molecular graphs. Nevertheless, current GNNs mainly excel in leveraging short-range interactions (SRI) but struggle to capture long-range interactions (LRI), both of which are crucial for determining molecular properties.…

2024

OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer

EMNLP 2024main

Recent advancements in Large Language Models (LLMs) have expanded their capabilities to multimodal contexts, including comprehensive video understanding. However, processing extensive videos such as 24-hour CCTV footage or full-length films presents significant challenges due to the vast data and pr…

2024

PathRL: An End-to-End Path Generation Method for Collision Avoidance via Deep Reinforcement Learning

ICRA 2024poster

Robot navigation using deep reinforcement learning (DRL) has shown great potential in improving the performance of mobile robots. Nevertheless, most existing DRL-based navigation methods primarily focus on training a policy that directly commands the robot with low-level controls, like linear and an…

Cited by 8SourceScholar
2024

SIMPL: A Simple and Efficient Multi-Agent Motion Prediction Baseline for Autonomous Driving

RA-L 2024

This letter presents a <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">S</u> imple and eff <underline xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">I</u> cient <underline xmlns:mml="http://www.w3.org/199

Cited by 66SourcecodeScholar
2024

Tail Classes Matter: Long-Tailed Object Detection Revisited

ICASSP 2024accepted

Real-world data ubiquitously exhibit long-tailed distribution, which sparks the increasing interest in long-tailed object detection (LTOD). However, existing methods neglect that a lack of diverse data in tail classes will cause underrepresented tail class features, making their efforts for balancin…

Cited by 0SourceScholar
2024

Towards Resource-Efficient and Secure Federated Multimedia Recommendation

ICASSP 2024accepted

Federated multimedia recommendation remains unexplored due to the high dimensionality of multimedia context, which limits the federated optimization on resource-constrained user devices. To address this issue, we propose a resource-efficient and secure federated learning framework for multimedia rec…

Cited by 0SourceScholar
2024

Weakly Supervised Few-Shot Object Detection with DETR

AAAI 2024technical

In recent years, Few-shot Object Detection (FSOD) has become an increasingly important research topic in computer vision. However, existing FSOD methods require strong annotations including category labels and bounding boxes, and their performance is heavily dependent on the quality of box annotatio…

Cited by 3SourcePDFScholar
2023

DrugOOD: Out-of-Distribution Dataset Curator and Benchmark for AI-Aided Drug Discovery – a Focus on Affinity Prediction Problems with Noise Annotations

AAAI 2023technical

AI-aided drug discovery (AIDD) is gaining popularity due to its potential to make the search for new pharmaceuticals faster, less expensive, and more effective. Despite its extensive use in numerous fields (e.g., ADMET prediction, virtual screening), little research has been conducted on the out-of-…

Cited by 122SourcePDFScholar
2023

Unseen Object Instance Segmentation with Fully Test-time RGB-D Embeddings Adaptation

ICRA 2023poster

Segmenting unseen objects is a crucial ability for the robot since it may encounter new environments during the operation. Recently, a popular solution is leveraging RGB-D features of large-scale synthetic data and directly applying the model to unseen real-world scenarios. However, the domain shift…

Cited by 11SourceScholar
2023

Video Diffusion Models with Local-Global Context Guidance

IJCAI 2023poster

Diffusion models have emerged as a powerful paradigm in video synthesis tasks including prediction, generation, and interpolation. Due to the limitation of the computational budget, existing methods usually implement conditional diffusion models with an autoregressive inference pipeline, in which th…

2022

Adversarial Learning in Transformer Based Neural Network in Radio Signal Classification

ICASSP 2022accepted

Deep Learning has attracted significant interests in wireless communication design problems. However, recent studies discovered that the deep neural network is vulnerable to adversarial attacks in the sense that a carefully designed and imperceptible perturbation to the input of the neural network c…

Cited by 0SourceScholar
2022

Design of Real-Time System Based on Machine Learning for Snoring and OSA Detection

ICASSP 2022accepted

Obstructive sleep apnea (OSA) is a common sleep disorder. The diagnosis of OSA based on snoring is low-cost, convenient and non-invasive. In this study, we place a microphone under the patient’s bed and combined with full-night polysomnography to record audio signals. Five machine learning models an…

Cited by 0SourceScholar
2022

EPiDA: An Easy Plug-in Data Augmentation Framework for High Performance Text Classification

NAACL 2022long

Recent works have empirically shown the effectiveness of data augmentation (DA) in NLP tasks, especially for those suffering from data scarcity. Intuitively, given the size of generated data, their diversity and quality are crucial to the performance of targeted tasks. However, to the best of our kn…

2022

FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional Network

ICASSP 2022accepted

In recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time enhancement of full-band (48 kHz) speech signals is still very challenging. Becaus…

Cited by 0SourceScholar
2022

Generalized Equivariance and Preferential Labeling for GNN Node Classification

AAAI 2022technical

Existing graph neural networks (GNNs) largely rely on node embeddings, which represent a node as a vector by its identity, type, or content. However, graphs with unattributed nodes widely exist in real-world applications (e.g., anonymized social networks). Previous GNNs either assign random labels t…

2022

Lyra: A Benchmark for Turducken-Style Code Generation

IJCAI 2022poster

Recently, neural techniques have been used to generate source code automatically. While promising for declarative languages, these approaches achieve much poorer performance on datasets for imperative languages. Since a declarative language is typically embedded in an imperative language (i.e., the…

2022

Next Point-of-Interest Recommendation with Inferring Multi-step Future Preferences

IJCAI 2022poster

Existing studies on next point-of-interest (POI) recommendation mainly attempt to learn user preference from the past and current sequential behaviors. They, however, completely ignore the impact of future behaviors on the decision-making, thus hindering the quality of user preference learning. Intu…

2022

Towards edible drones for rescue missions: design and flight of nutritional wings

IROS 2022poster

Drones have shown to be useful aerial vehicles for unmanned transport missions such as food and medical supply delivery. This can be leveraged to deliver life-saving nutrition and medicine for people in emergency situations. However, commercial drones can generally only carry 10 %–30 % of their own…

Cited by 13SourceScholar
2022

You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object Segmentation

AAAI 2022technical

We present YOFO (You Only inFer Once), a new paradigm for referring video object segmentation (RVOS) that operates in an one-stage manner. Our key insight is that the language descriptor should serve as target-specific guidance to identify the target object, while a direct feature fusion of image an…

Cited by 59SourcePDFScholar
2021

A Generative Adversarial Framework for Bounding Confounded Causal Effects

AAAI 2021technical

Causal inference from observational data is receiving wide applications in many fields. However, unidentifiable situations, where causal effects cannot be uniquely computed from observational data, pose critical barriers to applying causal inference to complicated real applications. In this paper, w…

2021

Accurate Few-Shot Object Detection With Support-Query Mutual Guidance and Hybrid Loss

CVPR 2021poster

Most object detection methods require huge amounts of annotated data and can detect only the categories that appear in the training set. However, in reality acquiring massive annotated training data is both expensive and time-consuming. In this paper, we propose a novel two-stage detector for accura…

Cited by 75PDFScholar
2021

DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled Samples

NeurIPS 2021poster

The scarcity of labeled data is a critical obstacle to deep learning. Semi-supervised learning (SSL) provides a promising way to leverage unlabeled data by pseudo labels. However, when the size of labeled data is very small (say a few labeled samples per class), SSL performs poorly and unstably, pos…

Cited by 36SourcePDFScholar
2021

HSAN: A Hierarchical Self-Attention Network for Multi-Turn Dialogue Generation

ICASSP 2021accepted

In the multi-turn dialogue system, response generation is not only related to the sentences in context but also relies on the words in each utterance. Although there are lots of methods that pay attention to model words and utterances, there still exist problems such as tending to generate common re…

Cited by 0SourceScholar
2021

Learning Motion-Appearance Co-Attention for Zero-Shot Video Object Segmentation

ICCV 2021poster

How to make the appearance and motion information interact effectively to accommodate complex scenarios is a fundamental issue in flow-based zero-shot video object segmentation. In this paper, we propose an Attentive Multi-Modality Collaboration Network (AMC-Net) to utilize appearance and motion inf…

Cited by 78PDFcodeScholar
2021

Weakly-supervised Text Classification Based on Keyword Graph

EMNLP 2021main

Weakly-supervised text classification has received much attention in recent years for it can alleviate the heavy burden of annotating massive data. Among them, keyword-driven methods are the mainstream where user-provided keywords are exploited to generate pseudo-labels for unlabeled texts. However,…

2020

An Interactive Multi-Task Learning Framework for Next POI Recommendation with Uncertain Check-ins

IJCAI 2020poster

Studies on next point-of-interest (POI) recommendation mainly seek to learn users' transition patterns with certain historical check-ins. However, in reality, users' movements are typically uncertain (i.e., fuzzy and incomplete) where most existing methods suffer from the transition pattern vanishin…

2020

Binary Probability Model for Learning Based Image Compression

ICASSP 2020accepted

In this paper, we propose to enhance learned image compression systems with a richer probability model for the latent variables. Previous works model the latents with a Gaussian or a Laplace distribution. Inspired by binary arithmetic coding, we propose to signal the latents with three binary values…

Cited by 0SourceScholar
2020

Efficient Uncertainty-aware Decision-making for Automated Driving Using Guided Branching

ICRA 2020poster

Decision-making in dense traffic scenarios is challenging for automated vehicles (AVs) due to potentially stochastic behaviors of other traffic participants and perception uncertainties (e.g., tracking noise and prediction errors, etc.). Although the partially observable Markov decision process (POM…

Cited by 60SourcecodeScholar
2020

Manet: Multi-Scale Aggregated Network For Light Field Depth Estimation

ICASSP 2020accepted

We present a novel end-to-end network, MANet, for light field depth estimation. MANet is a parameter-effective and effi-cient multi-scale aggregated network, which is about 3 times smaller and 3 times faster than the current top-performing method Epinet. The MANet architecture is performed for estim…

Cited by 0SourceScholar
2020

NLocalSAT: Boosting Local Search with Solution Prediction

IJCAI 2020poster

The Boolean satisfiability problem (SAT) is a famous NP-complete problem in computer science. An effective way for solving a satisfiable SAT problem is the stochastic local search (SLS). However, in this method, the initialization is assigned in a random manner, which impacts the effectiveness of SL…

2020

Unsupervised Video Object Segmentation with Joint Hotspot Tracking

ECCV 2020poster

Object tracking is a well-studied problem in computer vision while identifying salient spots of objects in a video is a less explored direction in the literature. Video eye gaze estimation methods aim to tackle a related task but salient spots in those methods are not bounded by objects and tend to…

2019

CapSal: Leveraging Captioning to Boost Semantics for Salient Object Detection

CVPR 2019poster

Detecting salient objects in cluttered scenes is a big challenge. To address this problem, we argue that the model needs to learn discriminative semantic features for salient objects. To this end, we propose to leverage captioning as an auxiliary semantic task to boost salient object detection in c…

Cited by 136PDFScholar
2019

PC-Fairness: A Unified Framework for Measuring Causality-based Fairness

NeurIPS 2019poster

A recent trend of fair machine learning is to define fairness as causality-based notions which concern the causal connection between protected attributes and decisions. However, one common challenge of all causality-based fairness notions is identifiability, i.e., whether they can be uniquely measur…

Cited by 151SourcePDFScholar
2019

Safe Trajectory Generation for Complex Urban Environments Using Spatio-Temporal Semantic Corridor

RA-L 2019

Planning safe trajectories for autonomous vehicles in complex urban environments is challenging since there are numerous semantic elements (such as dynamic agents, traffic lights, and speed limits) to consider. These semantic elements may have different mathematical descriptions, such as obstacle, c

Cited by 134SourcecodeScholar
2019

Weakly Aligned Cross-Modal Learning for Multispectral Pedestrian Detection

ICCV 2019poster

Multispectral pedestrian detection has shown great advantages under poor illumination conditions, since the thermal modality provides complementary information for the color image. However, real multispectral data suffers from the position shift problem, i.e. the color-thermal image pairs are not st…

Cited by 241PDFcodeScholar
2018

A Bi-Directional Message Passing Model for Salient Object Detection

CVPR 2018poster

Recent progress on salient object detection is beneficial from Fully Convolutional Neural Network (FCN). The saliency cues contained in multi-level convolutional features are complementary for detecting salient objects. How to integrate multi-level features becomes an open problem in saliency detect…

Cited by 579SourcePDFScholar
2018

Low Complexity Joint RDO of Prediction Units Couples for HEVC Intra Coding

ICASSP 2018accepted

HEVC is the latest block-based video compression standard, outperforming H.264/AVC by 50% bitrate savings for the same perceptual quality. An HEVC encoder provides Rate-Distortion optimization coding tools for block-wise compression. Because of complexity limitations, Rate-Distortion Optimization (R…

Cited by 0SourceScholar
2017

Inter-block dependencies consideration for intra coding in H.264/AVC and HEVC standards

ICASSP 2017accepted

Recent MPEG video compression standards are still block-based: blocks of pixels are sequentially coded using spatial or temporal prediction schemes. For each block, a vector of coding parameters has to be selected. In order to limit the complexity of this decision, independence between blocks is ass…

Cited by 0SourceScholar
2017

NIQSV: A no reference image quality assessment metric for 3D synthesized views

ICASSP 2017accepted

The popularity of 3D applications, such as Free View-point TV (FTV) and Multi-view Video plus Depth (MVD), induces a heavy requirement of synthesized views. However, the quality assessment of synthesized views is very challenging because the corresponding original views (reference views) are usually…

Cited by 0SourceScholar
2016

Beyond F-Formations: Determining Social Involvement in Free Standing Conversing Groups From Static Images

CVPR 2016poster

In this paper, we present the first attempt to analyse differing levels of social involvement in free standing conversing groups (or the so-called F-formations) from static images. In addition, we enrich state-of-the-art F-formation modelling by learning a frustum of attention that accounts for the…

Cited by 50PDFScholar
2016

Tag recommendation via robust probabilistic discriminative matrix factorization

ICASSP 2016accepted

Low-rank matrix factorization serves as a key technique in learning latent factor models for many applications in machine learning. However, in many applications, observed data often exhibits different levels of noise. To address this issue, we propose a Robust Probabilistic Discriminative Matrix Fa…

Cited by 0SourceScholar
2015

A multi-slice model observer for medical image quality assessment

ICASSP 2015accepted

Model observers (MOs) have been developed for the medical image quality assessment. Nowadays, numerous modern medical instruments are capable of producing 3D images, while few researchers have conducted MO studies on 3D data. In this paper, we propose a multi-slice MO when considering a relatively m…

Cited by 0SourceScholar