← Search

Yi ZHOU

158 accepted papers

2026

Bidirectional Channel-selective Semantic Interaction for Semi-Supervised Medical Segmentation

AAAI 2026technical

Semi-supervised medical image segmentation is an effective method for addressing scenarios with limited labeled data. Existing methods mainly rely on frameworks such as mean teacher and dual-stream consistency learning. These approaches often face issues like error accumulation and model structural

Cited by 0SourcePDFScholar
2026

Enhancing Molecular Property Predictions by Learning from Bond Modelling and Interactions

ICLR 2026poster

Molecule representation learning is crucial for understanding and predicting molecular properties. However, conventional atom-centric models, which treat chemical bonds merely as pairwise interactions, often overlook complex bond-level phenomena like resonance and stereoselectivity. This oversight l…

Cited by 0SourceScholar
2026

Entropy-Aware On-Policy Distillation of Language Models

ICML 2026poster

On-policy distillation is a promising approach for transferring knowledge between language models, where a student learns from dense token-level signals along its own trajectories. This framework typically uses reverse KL divergence, encouraging the student to match the teacher's high-confidence pre…

Cited by 0SourceScholar
2026

GneissWeb: Preparing High Quality Data for LLMs at Scale

ICLR 2026poster

Data quantity and quality play a vital role in determining the performance of Large Language Models (LLMs). High-quality data, in particular, can significantly boost the LLM's ability to generalize on a wide range of downstream tasks. In this paper, we introduce **GneissWeb**, a large dataset of aro…

Cited by 0SourceScholar
2026

InvariantCloud: A Globally Invariant, Uniquely Indexed Point Cloud Framework for Robust 6-DoF Tactile Pose Tracking

ICRA 2026poster

Recent advances in imitation learning and vision–language models highlight the need for high-fidelity tactile perception, with 6-DoF tactile object pose estimation providing a crucial foundation for precise robotic manipulation. We introduce InvariantCloud, a 6-DoF pose estimation framework that lev…

2026

LEAR: Learning Edge-Aware Representations for Event-To-LiDAR Localization

ICRA 2026poster

Event cameras offer high-temporal-resolution sensing that remains reliable under high-speed motion and challenging lighting, making them promising for localization from LiDAR point clouds in GPS-denied and visually degraded environments. However, aligning sparse, asynchronous events with dense LiDAR…

2026

MedKCO: Medical Vision-Language Pretraining via Knowledge-Driven Cognitive Orchestration

CVPR 2026

Medical vision-language pretraining (VLP) models have recently been investigated for their generalization to diverse downstream tasks. However, current medical VLP methods typically force the model to learn simple and complex concepts simultaneously. This anti-cognitive process leads to suboptimal f

Cited by 0SourcecodeScholar
2026

MeshSplatting: Differentiable Rendering with Opaque Meshes

CVPR 2026

Primitive-based splatting methods like 3D Gaussian Splatting (3DGS) have revolutionized novel view synthesis with real-time rendering. However, their point-based representations remain incompatible with mesh-based pipelines that power AR/VR and game engines. We present Mesh Splatting, a mesh-based r

Cited by 0SourcecodeScholar
2026

Multi-Agent Reinforcement Learning with Submodular Reward

ICML 2026poster

In this paper, we study cooperative multi-agent reinforcement learning (MARL) where the joint reward exhibits submodularity, which is a natural property capturing diminishing marginal returns when adding agents to a team. Unlike standard MARL with additive rewards, submodular rewards model realistic…

Cited by 0SourceScholar
2026

Neural Predictor-Corrector: Solving Homotopy Problems with Reinforcement Learning

ICLR 2026poster

The Homotopy paradigm, a general principle for solving challenging problems, appears across diverse domains such as robust optimization, global optimization, polynomial root-finding, and sampling. Practical solvers for these problems typically follow a predictor-corrector (PC) structure, but rely on…

Cited by 0SourceScholar
2026

Real-Time Motion Segmentation with Event-Based Normal Flow

ICRA 2026poster

Event-based cameras are bio-inspired sensors with pixels that independently and asynchronously respond to brightness changes at microsecond resolution, offering the potential to handle visual tasks in challenging scenarios. However, due to the sparse information content in individual events, directl…

2026

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation

CVPR 2026

Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generators with human preferences. However, existing GRPO pipelines rely on a single scalar reward per sample, treating each im

Cited by 0SourceScholar
2026

SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentation

CVPR 2026

Semi-supervised learning addresses label scarcity and high annotation costs in medical image segmentation by exploiting the latent information in unlabeled data to enhance model performance. Traditional discriminative segmentation relies on segmentation masks, neglecting feature-level distribution c

Cited by 0SourcecodeScholar
2026

StrokeFusion: Vector Sketch Generation via Joint Stroke-UDF Encoding and Latent Sequence Diffusion

AAAI 2026technical

In the field of sketch generation, raster-format trained models often produce non-stroke artifacts, while vector-format trained models typically lack a holistic understanding of sketches, leading to compromised recognizability. Moreover, existing methods struggle to extract common features from simi

Cited by 0SourcePDFScholar
2026

TTAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Events

CVPR 2026

Tracking any point (TAP) is a fundamental yet challenging task in computer vision, requiring high precision and long-term motion reasoning. Recent attempts to combine RGB frames and event streams have shown promise, yet they typically rely on synchronous or non-adaptive fusion, leading to temporal m

Cited by 0SourcecodeScholar
2025

A Novel Local Search Algorithm for the Vertex Bisection Minimization Problem

IJCAI 2025

The vertex bisection minimization problem (VBMP) is a fundamental graph partitioning problem with numerous real-world applications. In this study, we propose a (k, l, S)-cluster guided local search algorithm to address this challenge. First, we propose a novel (k,l,S)-cluster enumeration procedure,

2025

Annotating Training Data for Conditional Semantic Textual Similarity Measurement using Large Language Models

EMNLP 2025

Semantic similarity between two sentences depends on the aspects considered between those sentences. To study this phenomenon, Deshpande et al. (2023) proposed the Conditional Semantic Textual Similarity (C-STS) task and annotated a human-rated similarity dataset containing pairs of sentences compar

2025

BRIGHTER: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages

ACL 2025long

People worldwide use language in subtle and complex ways to express emotions. Although emotion recognition–an umbrella term for several NLP tasks–impacts various applications within NLP and beyond, most work in this area has focused on high-resource languages. This has led to significant disparities…

2025

Convex Relaxation for Robust Vanishing Point Estimation in Manhattan World

CVPR 2025award

Determining the vanishing points (VPs) in a Manhattan world, as a fundamental task in many 3D vision applications, consists of jointly inferring the line-VP association and locating each VP. Existing methods are, however, either sub-optimal solvers or pursuing global optimality at a significant cost…

2025

DMesh++: An Efficient Differentiable Mesh for Complex Shapes

ICCV 2025poster

Recent probabilistic methods for 3D triangular meshes capture diverse shapes by differentiable mesh connectivity, but face high computational costs with increased shape details. We introduce a new differentiable mesh processing method that addresses this challenge and efficiently handles meshes with…

2025

DocMMIR: A Framework for Document Multi-modal Information Retrieval

EMNLP 2025

The rapid advancement of unsupervised representation learning and large-scale pre-trained vision-language models has significantly improved cross-modal retrieval tasks. However, existing multi-modal information retrieval (MMIR) studies lack a comprehensive exploration of document-level retrieval and

Cited by 0SourcePDFScholar
2025

E-MoFlow: Learning Egomotion and Optical Flow from Event Data via Implicit Regularization

NeurIPS 2025poster

The estimation of optical flow and 6-DoF ego-motion—two fundamental tasks in 3-D vision—has typically been addressed independently. For neuromorphic vision (e.g., event cameras), however, the lack of robust data association makes solving the two problems separately an ill-posed challenge, especiall…

Cited by 0SourceScholar
2025

EvTTC: An Event Camera Dataset for Time-to-Collision Estimation

RA-L 2025

Time-to-Collision (TTC) estimation lies in the core of the forward collision warning (FCW) functionality, which is key to all Automatic Emergency Braking (AEB) systems. Although the success of solutions using frame-based cameras (e.g., <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xli

Cited by 4SourceScholar
2025

FaceLift: Learning Generalizable Single Image 3D Face Reconstruction from Synthetic Heads

ICCV 2025poster

We present FaceLift, a novel feed-forward approach for generalizable high-quality 360-degree 3D head reconstruction from a single image. Our pipeline first employs a multi-view latent diffusion model to generate consistent side and back views from a single facial input, which then feed into a transf…

Cited by 0SourcePDFScholar
2025

HUMOTO: A 4D Dataset of Mocap Human Object Interactions

ICCV 2025poster

We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 735 sequences (7,875 seconds at 30 fps), HUMOTO captures interactions with 63 precisely modeled objects and 72 articulated…

2025

KARLM: Enhancing LLM-based Recommendation Systems with Knowledge Bases

ICASSP 2025accepted

Large language models signify a pivotal advancement in general artificial intelligence, exhibiting capabilities that exceed human performance in diverse tasks. Nevertheless, these models often lack expertise in specialized knowledge areas. To augment the performance of LLMs in downstream application…

Cited by 0SourceScholar
2025

Large-Scale Trade-Off Curve Computation for Incentive Allocation with Cardinality and Matroid Constraints

IJCAI 2025

We consider a large-scale incentive allocation problem where the entire trade-off curve between budget and profit has to be maintained approximately at all time. The application originally comes from assigning coupons to users of the ride-sharing apps, where each user can have a limit on the number

Cited by 0SourcePDFScholar
2025

MAP: Multi-Human-Value Alignment Palette

ICLR 2025oral

Ensuring that generative AI systems align with human values is essential but challenging, especially when considering multiple human values and their potential trade-offs. Since human values can be personalized and dynamically change over time, the desirable levels of value alignment vary across dif…

Cited by 3SourcePDFScholar
2025

Make Domain Shift a Catastrophic Forgetting Alleviator in Class-Incremental Learning

AAAI 2025technical

In the realm of class-incremental learning (CIL), alleviating the catastrophic forgetting problem is a pivotal challenge. This paper discovers a counter-intuitive observation: by incorporating domain shift into CIL tasks, the forgetting rate is significantly reduced. Our comprehensive studies demons…

2025

MemGS: Memory-Efficient Gaussian Splatting for Real-Time SLAM

IROS 2025

Recent advancements in 3D Gaussian Splatting (3DGS) have made a significant impact on rendering and reconstruction techniques. Current research predominantly focuses on improving rendering performance and reconstruction quality using high-performance desktop GPUs, largely overlooking applications fo

Cited by 3SourceScholar
2025

Perm: A Parametric Representation for Multi-Style 3D Hair Modeling

ICLR 2025spotlight

We present Perm, a learned parametric representation of human 3D hair designed to facilitate various hair-related applications. Unlike previous work that jointly models the global hair structure and local curl patterns, we propose to disentangle them using a PCA-based strand representation in the fr…

2025

Probabilistic Person-in-Bed Detection Using Accelerometer Signals

ICASSP 2025accepted

Using accelerometer data in smart bed systems offers a cost-effective solution for person-in-bed detection. In this work, we propose a lightweight probabilistic model for this task. The accelerometer time series is first divided into multiple patches, with high-frequency noise filtered through a com…

Cited by 0SourceScholar
2025

Revisiting Large-Scale Non-convex Distributionally Robust Optimization

ICLR 2025poster

Distributionally robust optimization (DRO) is a powerful technique to train robust machine learning models that perform well under distribution shifts. Compared with empirical risk minimization (ERM), DRO optimizes the expected loss under the worst-case distribution in an uncertainty set of distribu…

Cited by 0SourcePDFScholar
2025

SegmentDreamer: Towards High-fidelity Text-to-3D Synthesis with Segmented Consistency Trajectory Distillation

ICCV 2025poster

Recent advancements in text-to-3D generation improve the visual quality of Score Distillation Sampling (SDS) and its variants by directly connecting Consistency Distillation (CD) to score distillation.However, due to the imbalance between self-consistency and cross-consistency, these CD-based method…

Cited by 0SourcePDFScholar
2025

Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions

CVPR 2025poster

We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and t…

Cited by 1SourcePDFScholar
2025

WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines

NAACL 2025long

Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate their understanding of such knowledge, we introduce WorldCuisines, a massive-scale benchmark for multilingual and multicul…

2024

A Fast Exact Solver with Theoretical Analysis for the Maximum Edge-Weighted Clique Problem

AAAI 2024technical

The maximum vertex-weighted clique problem (MVWCP) and the maximum edge-weighted clique problem (MEWCP) are two natural extensions of the fundamental maximum clique problem. In this paper, we systematically study MEWCP and make the following major contributions: (1) We show that MEWCP is NP-hard ev…

2024

Approximate Kernel Density Estimation under Metric-based Local Differential Privacy

UAI 2024poster

Kernel Density Estimation (KDE) is a fundamental problem with broad machine learning applications. In this paper, we investigate the KDE problem under Local Differential Privacy (LDP), a setting in which users privatize data on their own devices before sending them to an untrusted server for analyti…

Cited by 0SourcePDFScholar
2024

BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages

NeurIPS 2024poster

Large language models (LLMs) often lack culture-specific everyday knowledge, especially across diverse regions and non-English languages. Existing benchmarks for evaluating LLMs' cultural sensitivities are usually limited to a single language or online sources like Wikipedia, which may not reflect t…

2024

BeNeRF:Neural Radiance Fields from a Single Blurry Image and Event Stream

ECCV 2024poster

"Implicit scene representation has attracted a lot of attention in recent research of computer vision and graphics. Most prior methods focus on how to reconstruct 3D scene representation from a set of images. In this work, we demonstrate the possibility to recover the neural radiance fields (NeRF) f…

2024

Carve3D: Improving Multi-view Reconstruction Consistency for Diffusion Models with RL Finetuning

CVPR 2024poster

Multi-view diffusion models obtained by applying Supervised Finetuning (SFT) to text-to-image diffusion models have driven recent breakthroughs in text-to-3D research. However due to the limited size and quality of existing 3D datasets they still suffer from multi-view inconsistencies and Neural Rad…

2024

DMesh: A Differentiable Mesh Representation

NeurIPS 2024poster

We present a differentiable representation, DMesh, for general 3D triangular meshes. DMesh considers both the geometry and connectivity information of a mesh. In our design, we first get a set of convex tetrahedra that compactly tessellates the domain based on Weighted Delaunay Triangulation (WDT),…

2024

Effective Data Distillation for Tabular Datasets (Student Abstract)

AAAI 2024technical

Data distillation is a technique of reducing a large dataset into a smaller dataset. The smaller dataset can then be used to train a model which can perform comparably to a model trained on the full dataset. Past works have examined this approach for image datasets, focusing on neural networks as ta…

Cited by 3SourcePDFScholar
2024

Enhancing In-context Learning via Linear Probe Calibration

AISTATS 2024poster

In-context learning (ICL) is a new paradigm for natural language processing that utilizes Generative Pre-trained Transformer (GPT)-like models. This approach uses prompts that include in-context demonstrations to generate the corresponding output for a new query input. However, applying ICL in real…

2024

Evaluating Short-Term Temporal Fluctuations of Social Biases in Social Media Data and Masked Language Models

EMNLP 2024main

Social biases such as gender or racial biases have been reported in language models (LMs), including Masked Language Models (MLMs). Given that MLMs are continuously trained with increasing amounts of additional data collected over time, an important yet unanswered question is how the social biases e…

2024

Evaluating Unsupervised Dimensionality Reduction Methods for Pretrained Sentence Embeddings

COLING 2024main

Sentence embeddings produced by Pretrained Language Models (PLMs) have received wide attention from the NLP community due to their superior performance when representing texts in numerous downstream applications. However, the high dimensionality of the sentence embeddings produced by PLMs is problem…

Cited by 4SourcePDFScholar
2024

HIMap: HybrId Representation Learning for End-to-end Vectorized HD Map Construction

CVPR 2024poster

Vectorized High-Definition (HD) map construction requires predictions of the category and point coordinates of map elements (e.g. road boundary lane divider pedestrian crossing etc.). State-of-the-art methods are mainly based on point-level representation learning for regressing accurate point coord…

Cited by 24SourcePDFScholar
2024

Is Your HD Map Constructor Reliable under Sensor Corruptions?

NeurIPS 2024poster

Driving systems often rely on high-definition (HD) maps for precise environmental information, which is crucial for planning and navigation. While current HD map constructors perform well under ideal conditions, their resilience to real-world challenges, \eg, adverse weather and sensor failures, is…

Cited by 17SourcePDFScholar
2024

LRM-Zero: Training Large Reconstruction Models with Synthesized Data

NeurIPS 2024poster

We present LRM-Zero, a Large Reconstruction Model (LRM) trained entirely on synthesized 3D data, achieving high-quality sparse-view 3D reconstruction. The core of LRM-Zero is our procedural 3D dataset, Zeroverse, which is automatically synthesized from simple primitive shapes with random texturing a…

2024

Large-Scale Non-convex Stochastic Constrained Distributionally Robust Optimization

AAAI 2024technical

Distributionally robust optimization (DRO) is a powerful framework for training robust models against data distribution shifts. This paper focuses on constrained DRO, which has an explicit characterization of the robustness level. Existing studies on constrained DRO mostly focus on convex loss func…

Cited by 5SourcePDFScholar
2024

MBFusion: A New Multi-modal BEV Feature Fusion Method for HD Map Construction

ICRA 2024poster

HD map construction is a fundamental and challenging task in autonomous driving to understand the surrounding environment. Recently, Camera-LiDAR BEV feature fusion methods have attracted increasing attention in HD map construction task, which can significantly boost the benchmark. However, existing…

Cited by 11SourceScholar
2024

Memory-Assisted Sub-Prototype Mining for Universal Domain Adaptation

ICLR 2024poster

Universal domain adaptation aims to align the classes and reduce the feature gap between the same category of the source and target domains. The target private category is set as the unknown class during the adaptation process, as it is not included in the source domain. However, most existing metho…

Cited by 2SourcePDFScholar
2024

Non-Asymptotic Analysis for Single-Loop (Natural) Actor-Critic with Compatible Function Approximation

ICML 2024poster

Actor-critic (AC) is a powerful method for learning an optimal policy in reinforcement learning, where the critic uses algorithms, e.g., temporal difference (TD) learning with function approximation, to evaluate the current policy and the actor updates the policy along an approximate gradient direct…

Cited by 12SourcePDFScholar
2024

Spatio-Temporal Calibration for Omni-Directional Vehicle-Mounted Event Cameras

RA-L 2024

We present a solution to the problem of spatio-temporal calibration for event cameras mounted on an onmi-directional vehicle. Different from traditional methods that typically determine the camera's pose with respect to the vehicle's body frame using alignment of trajectories, our approach leverages

Cited by 7SourcecodeScholar
2023

A Critical Analysis of Document Out-of-Distribution Detection

EMNLP 2023long findings

Large-scale pre-training is widely used in recent document understanding tasks. During deployment, one may expect that models should trigger a conservative fallback policy when encountering out-of-distribution (OOD) samples, which highlights the importance of OOD detection. However, most existing OO…

Cited by 0SourceScholar
2023

A Fast Maximum k-Plex Algorithm Parameterized by the Degeneracy Gap

IJCAI 2023poster

Given a graph, the k-plex is a vertex set in which each vertex is not adjacent to at most k-1 other vertices in the set. The maximum k-plex problem, which asks for the largest k-plex from a given graph, is an important but computationally challenging problem in applications like graph search and com…

2023

A Predictive Factor Analysis of Social Biases and Task-Performance in Pretrained Masked Language Models

EMNLP 2023long main

Various types of social biases have been reported with pretrained Masked Language Models (MLMs) in prior work. However, multiple underlying factors are associated with an MLM such as its model size, size of the training data, training objectives, the domain from which pretraining data is sampled, to…

Cited by 0SourceScholar
2023

A Word Sense Distribution-based approach for Semantic Change Prediction

EMNLP 2023long findings

Semantic Change Detection of words is an important task for various NLP applications that must make time-sensitive predictions. Some words are used over time in novel ways to express new meanings, and these new meanings establish themselves as novel senses of existing words. On the other hand, Word…

Cited by 0SourceScholar
2023

Boosting Multi-modal Model Performance with Adaptive Gradient Modulation

ICCV 2023poster

While the field of multi-modal learning keeps growing fast, the deficiency of the standard joint training paradigm has become clear through recent studies. They attribute the sub-optimal performance of the jointly trained model to the modality competition phenomenon. Existing works attempt to improv…

Cited by 29PDFcodeScholar
2023

Conic10K: A Challenging Math Problem Understanding and Reasoning Dataset

EMNLP 2023long findings

Mathematical understanding and reasoning are crucial tasks for assessing the capabilities of artificial intelligence (AI). However, existing benchmarks either require just a few steps of reasoning, or only contain a small amount of data in one specific topic, making it hard to analyse AI's behaviour…

Cited by 0SourcecodeScholar
2023

Cross-View Geo-Localization via Learning Disentangled Geometric Layout Correspondence

AAAI 2023technical

Cross-view geo-localization aims to estimate the location of a query ground image by matching it to a reference geo-tagged aerial images database. As an extremely challenging task, its difficulties root in the drastic view changes and different capturing time between two views. Despite these difficu…

2023

Efficient Multilingual Language Model Compression through Vocabulary Trimming

EMNLP 2023long findings

Multilingual language models (LMs) have become a powerful tool in NLP, especially for non-English languages. Nevertheless, model parameters of multilingual LMs remain large due to the larger embedding matrix of the vocabulary covering tokens in different languages. Instead, monolingual LMs can be tr…

Cited by 0SourceScholar
2023

Generalized-Smooth Nonconvex Optimization is As Efficient As Smooth Nonconvex Optimization

ICML 2023poster

Various optimal gradient-based algorithms have been developed for smooth nonconvex optimization. However, many nonconvex machine learning problems do not belong to the class of smooth functions and therefore the existing algorithms are sub-optimal. Instead, these problems have been shown to satisfy…

2023

LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning

ICML 2023poster

We propose LESS-VFL, a communication-efficient feature selection method for distributed systems with vertically partitioned data. We consider a system of a server and several parties with local datasets that share a sample ID space but have different feature sets. The parties wish to collaboratively…

Cited by 30SourcePDFScholar
2023

Language Model is Suitable for Correction of Handwritten Mathematical Expressions Recognition

EMNLP 2023long main

Handwritten mathematical expression recognition (HMER) is a multidisciplinary task that generates LaTeX sequences from images. Existing approaches, employing tree decoders within attention-based encoder-decoder architectures, aim to capture the hierarchical tree structure, but are limited by CFGs an…

Cited by 0SourceScholar
2023

Learning Dynamic Contextualised Word Embeddings via Template-based Temporal Adaptation

ACL 2023long

Dynamic contextualised word embeddings (DCWEs) represent the temporal semantic variations of words. We propose a method for learning DCWEs by time-adapting a pretrained Masked Language Model (MLM) using time-sensitive templates. Given two snapshots C1 and C2 of a corpus taken respectively at two dis…

2023

Normal-Guided Garment UV Prediction for Human Re-Texturing

CVPR 2023highlight

Clothes undergo complex geometric deformations, which lead to appearance changes. To edit human videos in a physically plausible way, a texture map must take into account not only the garment transformation induced by the body movements and clothes fitting, but also its 3D fine-grained surface geome…

Cited by 15SourcePDFScholar
2023

Single-shot General Hyper-parameter Optimization for Federated Learning

ICLR 2023top-25%

We address the problem of hyper-parameter optimization (HPO) for federated learning (FL-HPO). We introduce Federated Loss SuRface Aggregation (FLoRA), a general FL-HPO solution framework that can address use cases of tabular data and any Machine Learning (ML) model including gradient boosting traini…

Cited by 16SourcePDFScholar
2023

Solving Cosine Similarity Underestimation between High Frequency Words by ℓ2 Norm Discounting

ACL 2023findings

Cosine similarity between two words, computed using their contextualised token embeddings obtained from masked language models (MLMs) such as BERT has shown to underestimate the actual similarity between those words CITATION.This similarity underestimation problem is particularly severe for high fre…

Cited by 2SourcePDFScholar
2023

Structure-informed Language Models Are Protein Designers

ICML 2023oral

This paper demonstrates that language models are strong structure-based protein designers. We present LM-Design, a generic approach to reprogramming sequence-based protein language models (pLMs), that have learned massive sequential evolutionary knowledge from the universe of natural protein sequenc…

2023

Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQA

ACL 2023long

Natural language is ambiguous. Resolving ambiguous questions is key to successfully answering them. Focusing on questions about images, we create a dataset of ambiguous examples. We annotate these, grouping answers by the underlying question they address and rephrasing the question for each group to…

2022

A Repulsive Force Unit for Garment Collision Handling in Neural Networks

ECCV 2022poster

"Despite recent success, deep learning-based methods for predicting 3D garment deformation under body motion suffer from interpenetration problems between the garment and the body. To address this problem, we propose a novel collision handling neural network layer called Repulsive Force Unit (ReFU).…

Cited by 15SourcePDFScholar
2022

DDDM: A Brain-Inspired Framework for Robust Classification

IJCAI 2022poster

Despite their outstanding performance in a broad spectrum of real-world tasks, deep artificial neural networks are sensitive to input noises, particularly adversarial perturbations. On the contrary, human and animal brains are much less vulnerable. In contrast to the one-shot inference performed by…

2022

DarkVisionNet: Low-Light Imaging via RGB-NIR Fusion with Deep Inconsistency Prior

AAAI 2022technical

RGB-NIR fusion is a promising method for low-light imaging. However, high-intensity noise in low-light images amplifies the effect of structure inconsistency between RGB-NIR images, which fails existing algorithms. To handle this, we propose a new RGB-NIR fusion algorithm called Dark Vision Net (DVN…

2022

Data sampling affects the complexity of online SGD over dependent data

UAI 2022poster

Conventional machine learning applications typically assume that data samples are independently and identically distributed (i.i.d.). However, practical scenarios often involve a data-generating process that produces highly dependent data samples, which are known to heavily bias the stochastic optim…

Cited by 4SourcePDFScholar
2022

Delving Into the Estimation Shift of Batch Normalization in a Network

CVPR 2022poster

Batch normalization (BN) is a milestone technique in deep learning. It normalizes the activation using mini-batch statistics during training but the estimated population statistics during inference. This paper focuses on investigating the estimation of population statistics. We define the estimation…

Cited by 28PDFcodeScholar
2022

Finding Correlated Equilibrium of Constrained Markov Game: A Primal-Dual Approach

NeurIPS 2022accept

Constrained Markov game is a fundamental problem that covers many applications, where multiple players compete with each other under behavioral constraints. The existing literature has proved the existence of Nash equilibrium for constrained Markov games, which turns out to be PPAD-complete and cann…

Cited by 12SourcePDFScholar
2022

ICASSP 2022 L3DAS22 Challenge: Ensemble of Resnet-Conformers with Ambisonics Data Augmentation for Sound Event Localization and Detection

ICASSP 2022accepted

It remains a tough challenge to tackle sound event localization and detection (SELD) problem, especially when sound scene complexity increases and overlapping acoustic sources appear. To improve the SELD performance, we propose an ensemble system, which consists of a ResNet and Conformer backbone ne…

Cited by 0SourceScholar
2022

Learning Visibility for Robust Dense Human Body Estimation

ECCV 2022poster

"Estimating 3D human pose and shape from 2D images is a crucial yet challenging task. While prior methods with model-based representations can perform reasonably well on whole-body images, they often fail when parts of the body are occluded or outside the frame. Moreover, these results usually do no…

2022

NeMF: Neural Motion Fields for Kinematic Animation

NeurIPS 2022accept

We present an implicit neural representation to learn the spatio-temporal space of kinematic motions. Unlike previous work that represents motion as discrete sequential samples, we propose to express the vast motion space as a continuous function over time, hence the name Neural Motion Fields (NeMF)…

Cited by 58SourcePDFScholar
2022

Regularized Molecular Conformation Fields

NeurIPS 2022accept

Predicting energetically favorable 3-dimensional conformations of organic molecules from molecular graph plays a fundamental role in computer-aided drug discovery research. However, effectively exploring the high-dimensional conformation space to identify (meta) stable conformers is anything but tri…

Cited by 7SourcePDFScholar
2022

SHAPE: An Unified Approach to Evaluate the Contribution and Cooperation of Individual Modalities

IJCAI 2022poster

As deep learning advances, there is an ever-growing demand for models capable of synthesizing information from multi-modal resources to address the complex tasks raised from real-life applications. Recently, many large multi-modal datasets have been collected, on which researchers actively explore d…

2022

Sample Efficient Stochastic Policy Extragradient Algorithm for Zero-Sum Markov Game

ICLR 2022poster

Two-player zero-sum Markov game is a fundamental problem in reinforcement learning and game theory. Although many algorithms have been proposed for solving zero-sum Markov games in the existing literature, many of them either require a full knowledge of the environment or are not sample-efficient. I…

Cited by 21SourcePDFScholar
2022

Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time Analysis

ICML 2022spotlight

Actor-critic (AC) algorithms have been widely used in decentralized multi-agent systems to learn the optimal joint control policy. However, existing decentralized AC algorithms either need to share agents’ sensitive information or lack communication-efficiency. In this work, we develop decentralized…

Cited by 36SourcePDFScholar
2022

Sense Embeddings are also Biased – Evaluating Social Biases in Static and Contextualised Sense Embeddings

ACL 2022long

Sense embedding learning methods learn different embeddings for the different senses of an ambiguous word. One sense of an ambiguous word might be socially biased while its other senses remain unbiased. In comparison to the numerous prior work evaluating the social biases in pretrained word embeddin…

2022

Slot-VPS: Object-Centric Representation Learning for Video Panoptic Segmentation

CVPR 2022poster

Video Panoptic Segmentation (VPS) aims at assigning a class label to each pixel, uniquely segmenting and identifying all object instances consistently across all frames. Classic solutions usually decompose the VPS task into several sub-tasks and utilize multiple surrogates (e.g. boxes and masks, cen…

Cited by 30PDFcodeScholar
2022

UNISON: Unpaired Cross-Lingual Image Captioning

AAAI 2022technical

Image captioning has emerged as an interesting research field in recent years due to its broad application scenarios. The traditional paradigm of image captioning relies on paired image-caption datasets to train the model in a supervised manner. However, creating such paired datasets for every targe…

2021

CCT-Net: Category-Invariant Cross-Domain Transfer for Medical Single-to-Multiple Disease Diagnosis

ICCV 2021poster

A medical imaging model is usually explored for the diagnosis of a single disease. However, with the expanding demand for multi-disease diagnosis in clinical applications, multi-function solutions need to be investigated. Previous works proposed to either exploit different disease labels to conduct…

Cited by 11PDFScholar
2021

Curse or Redemption? How Data Heterogeneity Affects the Robustness of Federated Learning

AAAI 2021technical

Data heterogeneity has been identified as one of the key features in federated learning but often overlooked in the lens of robustness to adversarial attacks. This paper focuses on characterizing and understanding its impact on backdooring attacks in federated learning through comprehensive experime…

2021

Defense against Synonym Substitution-based Adversarial Attacks via Dirichlet Neighborhood Ensemble

ACL 2021long

Although deep neural networks have achieved prominent performance on many NLP tasks, they are vulnerable to adversarial examples. We propose Dirichlet Neighborhood Ensemble (DNE), a randomized method for training a robust model to defense synonym substitution-based attacks. During training, DNE form…

2021

Enhancing Balanced Graph Edge Partition with Effective Local Search

AAAI 2021technical

Graph partition is a key component to achieve workload balance and reduce job completion time in parallel graph processing systems. Among the various partition strategies, edge partition has demonstrated more promising performance in power-law graphs than vertex partition and thereby has been more w…

2021

Exploiting Semantic Embedding and Visual Feature for Facial Action Unit Detection

CVPR 2021poster

Recent study on detecting facial action units (AU) has utilized auxiliary information (i.e., facial landmarks, relationship among AUs and expressions, web facial images, etc.), in order to improve the AU detection performance. As of now, no semantic information of AUs has yet been explored for such…

Cited by 78PDFcodeScholar
2021

Greedy-GQ with Variance Reduction: Finite-time Analysis and Improved Complexity

ICLR 2021poster

Greedy-GQ is a value-based reinforcement learning (RL) algorithm for optimal control. Recently, the finite-time analysis of Greedy-GQ has been developed under linear function approximation and Markovian sampling, and the algorithm is shown to achieve an $\epsilon$-stationary point with a sample comp…

Cited by 19SourcePDFScholar
2021

Group Whitening: Balancing Learning Efficiency and Representational Capacity

CVPR 2021poster

Batch normalization (BN) is an important technique commonly incorporated into deep learning models to perform standardization within mini-batches. The merits of BN in improving a model's learning efficiency can be further amplified by applying whitening, while its drawbacks in estimating population…

Cited by 24PDFcodeScholar
2021

Group-Wise Semantic Mining for Weakly Supervised Semantic Segmentation

AAAI 2021technical

Acquiring sufficient ground-truth supervision to train deep vi- sual models has been a bottleneck over the years due to the data-hungry nature of deep learning. This is exacerbated in some structured prediction tasks, such as semantic segmen- tation, which requires pixel-level annotations. This work…

2021

Improving Maximum k-plex Solver via Second-Order Reduction and Graph Color Bounding

AAAI 2021technical

In a graph, a k-plex is a vertex set in which every vertex is not adjacent to at most k vertices of this set. The maximum k-plex problem, which asks for the largest k-plex from the given graph, is a key primitive in a variety of real-world applications like community detection and so on. In the pape…

2021

Many-to-One Distribution Learning and K-Nearest Neighbor Smoothing for Thoracic Disease Identification

AAAI 2021technical

Chest X-rays are an important and accessible clinical imaging tool for the detection of many thoracic diseases. Over the past decade, deep learning, with a focus on the convolutional neural network (CNN), has become the most powerful computer-aided diagnosis technology for improving disease identifi…

Cited by 14SourcePDFScholar
2021

Non-Asymptotic Analysis for Two Time-scale TDC with General Smooth Function Approximation

NeurIPS 2021poster

Temporal-difference learning with gradient correction (TDC) is a two time-scale algorithm for policy evaluation in reinforcement learning. This algorithm was initially proposed with linear function approximation, and was later extended to the one with general smooth function approximation. The asymp…

Cited by 14SourcePDFScholar
2021

On the Transferability of Adversarial Attacks against Neural Text Classifier

EMNLP 2021main

Deep neural networks are vulnerable to adversarial attacks, where a small perturbation to an input alters the model prediction. In many cases, malicious inputs intentionally crafted for one model can fool another model. In this paper, we present the first study to systematically investigate the tran…

Cited by 28SourcePDFScholar
2021

Proximal Gradient Descent-Ascent: Variable Convergence under KŁ Geometry

ICLR 2021poster

The gradient descent-ascent (GDA) algorithm has been widely applied to solve minimax optimization problems. In order to achieve convergent policy parameters for minimax optimization, it is important that GDA generates convergent variable sequences rather than convergent sequences of function value o…

Cited by 37SourcePDFScholar
2021

Specificity-Preserving RGB-D Saliency Detection

ICCV 2021poster

RGB-D saliency detection has attracted increasing attention, due to its effectiveness and the fact that depth cues can now be conveniently captured. Existing works often focus on learning a shared representation through various fusion strategies, with few methods explicitly considering how to preser…

Cited by 256PDFcodeScholar
2021

Visual-Textual Attentive Semantic Consistency for Medical Report Generation

ICCV 2021poster

Diagnosing diseases from medical radiographs and writing reports requires professional knowledge and is time-consuming. To address this, automatic medical report generation approaches have recently gained interest. However, identifying diseases as well as correctly predicting their corresponding siz…

Cited by 24PDFScholar
2020

A Statistical Mechanics Framework for Task-Agnostic Sample Design in Machine Learning

NeurIPS 2020poster

In this paper, we present a statistical mechanics framework to understand the effect of sampling properties of training data on the generalization gap of machine learning (ML) algorithms. We connect the generalization gap to the spatial properties of a sample design characterized by the pair correla…

2020

Fully Convolutional Mesh Autoencoder using Efficient Spatially Varying Kernels

NeurIPS 2020poster

Learning latent representations of registered meshes is useful for many 3D tasks. Techniques have recently shifted to neural mesh autoencoders. Although they demonstrate higher precision than traditional methods, they remain unable to capture fine-grained deformations. Furthermore, these methods can…

Cited by 97SourcePDFScholar
2020

History-Gradient Aided Batch Size Adaptation for Variance Reduced Algorithms

ICML 2020poster

Variance-reduced algorithms, although achieve great theoretical performance, can run slowly in practice due to the periodic gradient estimation with a large batch of data. Batch-size adaptation thus arises as a promising approach to accelerate such algorithms. However, existing schemes either apply…

Cited by 20SourcePDFScholar
2020

Perception-Distortion Trade-Off with Restricted Boltzmann Machines

ICASSP 2020accepted

In this work, we introduce a new procedure for applying Restricted Boltzmann Machines (RBMs) to missing data inference tasks, based on linearization of the effective energy function governing the distribution of observations. We compare the performance of our proposed procedure with those obtained u…

Cited by 0SourceScholar
2020

Proximal Gradient Algorithm with Momentum and Flexible Parameter Restart for Nonconvex Optimization

IJCAI 2020poster

Various types of parameter restart schemes have been proposed for proximal gradient algorithm with momentum to facilitate their convergence in convex optimization. However, under parameter restart, the convergence of proximal gradient algorithm with momentum remains obscure in nonconvex optimization…

Cited by 0SourcePDFScholar
2020

Understanding the Impact of Model Incoherence on Convergence of Incremental SGD with Random Reshuffle

ICML 2020poster

Although SGD with random reshuffle has been widely-used in machine learning applications, there is a limited understanding of how model characteristics affect the convergence of the algorithm. In this work, we introduce model incoherence to characterize the diversity of model characteristics and stu…

Cited by 6SourcePDFScholar
2020

Variance-Reduced Off-Policy TDC Learning: Non-Asymptotic Convergence Analysis

NeurIPS 2020poster

Variance reduction techniques have been successfully applied to temporal-difference (TD) learning and help to improve the sample complexity in policy evaluation. However, the existing work applied variance reduction to either the less popular one time-scale TD algorithm or the two time-scale GTD alg…

Cited by 21SourcePDFScholar
2019

A unified variance-reduced accelerated gradient method for convex optimization

NeurIPS 2019poster

We propose a novel randomized incremental gradient algorithm, namely, VAriance-Reduced Accelerated Gradient (Varag), for finite-sum optimization. Equipped with a unified step-size policy that adjusts itself to the value of the conditional number, Varag exhibits the unified optimal rates of convergen…

Cited by 74SourcePDFScholar
2019

Building Detail-Sensitive Semantic Segmentation Networks With Polynomial Pooling

CVPR 2019poster

Semantic segmentation is an important computer vision task, which aims to allocate a semantic label to each pixel in an image. When training a segmentation model, it is common to fine-tune a classification network pre-trained on a large-scale dataset. However, as an intrinsic property of the classif…

Cited by 34PDFScholar
2019

Collaborative Learning of Semi-Supervised Segmentation and Classification for Medical Images

CVPR 2019poster

Medical image analysis has two important research areas: disease grading and fine-grained lesion segmentation. Although the former problem often relies on the latter, the two are usually studied separately. Disease severity grading can be treated as a classification problem, which only requires imag…

Cited by 327PDFScholar
2019

Cross-lingual Voice Conversion with Bilingual Phonetic Posteriorgram and Average Modeling

ICASSP 2019accepted

This paper presents a cross-lingual voice conversion approach using bilingual Phonetic PosteriorGram (PPG) and average modeling. The proposed approach makes use of bilingual PPGs to represent speaker-independent features of speech signals from different languages in the same feature space. In partic…

Cited by 0SourceScholar
2019

Improved Zeroth-Order Variance Reduced Algorithms and Analysis for Nonconvex Optimization

ICML 2019oral

Two types of zeroth-order stochastic algorithms have recently been designed for nonconvex optimization respectively based on the first-order techniques SVRG and SARAH/SPIDER. This paper addresses several important issues that are still open in these methods. First, all existing SVRG-type zeroth-orde…

2019

Iterative Normalization: Beyond Standardization Towards Efficient Whitening

CVPR 2019poster

Batch Normalization (BN) is ubiquitously employed for accelerating neural network training and improving the generalization capability by performing standardization within mini-batches. Decorrelated Batch Normalization (DBN) further boosts the above effectiveness by whitening. However, DBN relies…

Cited by 183PDFcodeScholar
2019

On the Continuity of Rotation Representations in Neural Networks

CVPR 2019poster

In neural networks, it is often desirable to work with various representations of the same space. For example, 3D rotations can be represented with quaternions or Euler angles. In this paper, we advance a definition of a continuous representation, which can be helpful for training deep neural netwo…

Cited by 1552PDFScholar
2019

SGD Converges to Global Minimum in Deep Learning via Star-convex Path

ICLR 2019poster

Stochastic gradient descent (SGD) has been found to be surprisingly effective in training a variety of deep neural networks. However, there is still a lack of understanding on how and why SGD can train these complex networks towards a global minimum. In this study, we establish the convergence of SG…

Cited by 86SourcePDFScholar
2019

SpiderBoost and Momentum: Faster Variance Reduction Algorithms

NeurIPS 2019poster

SARAH and SPIDER are two recently developed stochastic variance-reduced algorithms, and SPIDER has been shown to achieve a near-optimal first-order oracle complexity in smooth nonconvex optimization. However, SPIDER uses an accuracy-dependent stepsize that slows down the convergence in practice, and…

Cited by 213SourcePDFScholar
2019

Stochastic Variance-Reduced Cubic Regularization for Nonconvex Optimization

AISTATS 2019poster

Cubic regularization (CR) is an optimization method with emerging popularity due to its capability to escape saddle points and converge to second-order stationary solutions for nonconvex optimization. However, CR encounters a high sample complexity issue for finite-sum problems with a large data siz…

Cited by 67SourcePDFScholar
2019

Toward Understanding the Impact of Staleness in Distributed Machine Learning

ICLR 2019poster

Most distributed machine learning (ML) systems store a copy of the model parameters locally on each machine to minimize network communication. In practice, in order to reduce synchronization waiting time, these copies of the model are not necessarily updated in lock-step, and can become stale. Despi…

Cited by 100SourcePDFScholar
2018

Auto-Conditioned Recurrent Networks for Extended Complex Human Motion Synthesis

ICLR 2018poster

We present a real-time method for synthesizing highly complex human motions using a novel training regime we call the auto-conditioned Recurrent Neural Network (acRNN). Recently, researchers have attempted to synthesize new motion by using autoregressive techniques, but existing methods tend to free…

Cited by 264SourcePDFScholar
2018

Convergence of Cubic Regularization for Nonconvex Optimization under KL Property

NeurIPS 2018spotlight

Cubic-regularized Newton's method (CR) is a popular algorithm that guarantees to produce a second-order stationary solution for solving nonconvex optimization problems. However, existing understandings of convergence rate of CR are conditioned on special types of geometrical properties of the object…

Cited by 28SourcePDFScholar
2018

HairNet: Single-View Hair Reconstruction using Convolutional Neural Networks

ECCV 2018poster

We introduce a deep learning-based method to generate full 3D hair geometry from an unconstrained image. Our method can recover local strand details and has real-time performance. State-of-the-art hair modeling techniques rely on large hairstyle collections for nearest neighbor retrieval and then pe…

Cited by 86SourcePDFScholar
2018

Semi-Dense 3D Reconstruction with a Stereo Event Camera

ECCV 2018poster

Event cameras are bio-inspired sensors that offer several advantages, such as low latency, high-speed and high dynamic range, to tackle challenging scenarios in computer vision. This paper presents a solution to the problem of 3D reconstruction from data captured by a stereo event-camera rig moving…

Cited by 196SourcePDFScholar
2017

Convergence Analysis of Proximal Gradient with Momentum for Nonconvex Optimization

ICML 2017poster

In this work, we investigate the accelerated proximal gradient method for nonconvex programming (APGnc). The method compares between a usual proximal gradient step and a linear extrapolation step, and accepts the one that has a lower function value to achieve a monotonic decrease. In specific, under…

Cited by 106SourcePDFScholar
2017

Learning Latent Space Models with Angular Constraints

ICML 2017poster

The large model capacity of latent space models (LSMs) enables them to achieve great performance on various applications, but meanwhile renders LSMs to be prone to overfitting. Several recent studies investigate a new type of regularization approach, which encourages components in LSMs to be diverse…

Cited by 27SourcePDFScholar
2017

Realistic Dynamic Facial Textures From a Single Image Using GANs

ICCV 2017poster

We present a novel method to realistically puppeteer and animate a face from a single RGB image using a source video sequence. We begin by fitting a multilinear PCA model to obtain the 3D geometry and a single texture of the target face. In order for the animation to be realistic, however, we need d…

Cited by 116PDFScholar
2017

Semi-dense visual odometry for RGB-D cameras using approximate nearest neighbour fields

ICRA 2017poster

This paper presents a robust and efficient semidense visual odometry solution for RGB-D cameras. The core of our method is a 2D-3D ICP pipeline which estimates the pose of the sensor by registering the projection of a 3D semidense map of a reference frame with the 2D semi-dense region extracted in t…

Cited by 17SourceScholar
2016

On Convergence of Model Parallel Proximal Gradient Algorithm for Stale Synchronous Parallel System

AISTATS 2016poster

With ever growing data volume and model size, an error-tolerant, communication efficient, yet versatile parallel algorithm has become a vital part for the success of many large-scale applications. In this work we propose mspg, an extension of the flexible proximal gradient algorithm to the model par…

Cited by 40SourcePDFScholar
2016

Real-time rotation estimation for dense depth sensors in piece-wise planar environments

IROS 2016poster

Low-drift rotation estimation is a crucial part of any accurate odometry system. In this paper, we focus on the problem of 3D rotation estimation with dense depth sensors in environments that consist of piece-wise planar structures, such as corridors and office rooms. An efficient mean-shift paradig…

Cited by 12SourceScholar