← Search

Tianyi Zhang

96 accepted papers

2026

Backjump-on-Graph: Empowering LLMs with Reinforced Retrospective Exploration for Agentic KG Reasoning

ICML 2026poster

Grounding Large Language Models (LLMs) in Knowledge Graphs (KGs) has shown significant promise for complex Question Answering (QA) tasks. Since LLMs' limited context window cannot accommodate the sheer volume of large-scale KGs, existing work usually utilizes agents to reason on real-world KGs, whic…

Cited by 0SourceScholar
2026

Bi-Manual Joint Camera Calibration and Scene Representation

ICRA 2026poster

Robot manipulation, especially bimanual manipulation, often requires setting up multiple cameras on multiple robot manipulators. Before robot manipulators can generate motion or even build representations of their environments, the cameras rigidly mounted to the robot need to be calibrated. Camera c…

2026

CPQS-Tuning: A Model Self-Perception-Based Data Filtering Algorithm for Efficient Instruction Fine-Tuning

ICLR 2026poster

Instruction fine-tuning is a key technique for enhancing the performance of large language models (LLMs), but low-quality and redundant data often hinder its effectiveness. Recent studies suggest that filtering a small amount of high-quality data for instruction fine-tuning can achieve faster and mo…

Cited by 0SourcecodeScholar
2026

Caracal: Causal Architecture via Spectral Mixing

ICML 2026poster

The scalability of Large Language Models to long sequences is hindered by the quadratic cost of self-attention and the limitations of positional encodings. To address these, we introduce **Caracal**, a novel architecture that replaces self-attention with a parameter-efficient, $\mathcal{O}(L \log L)…

Cited by 0SourceScholar
2026

Cross-Chirality Generalization by Axial Vectors for Hetero-Chiral Protein-Peptide Interaction Design

ICML 2026poster

D-peptide binders targeting L-proteins have promising therapeutic potential. Despite rapid advances in machine learning-based target-conditioned peptide design, generating D-peptide binders remains largely unexplored. In this work, we show that by injecting axial features to E(3)-equivariant (polar)…

Cited by 0SourceScholar
2026

DOSE3: Diffusion-Based Unified Out-Of-Distribution Detection on SE(3) Trajectories

ICRA 2026poster

Out-Of-Distribution (OOD) detection, the task of identifying when an input falls outside the distribution seen at training time, is critical for deploying safe and reliable systems. Traditional OOD methods require retraining models whenever the in‐distribution has changed. Recent work introduces uni…

Cited by 0Scholar
2026

DOSE3: Diffusion-Based Unified Out-of-Distribution Detection on $\mathbb{SE}(3)$ Trajectories

RA-L 2026

Out-of-Distribution (OOD) detection, the task of identifying when an input falls outside the distribution seen at training time, is critical for deploying safe and reliable systems. Traditional OOD methods require retraining models whenever the in-distribution has changed. Recent work introduces <it

Cited by 1SourceScholar
2026

DreamSea: Photorealistic 3D Underwater Terrain Generation by Latent Fractal Diffusion Models

ICRA 2026poster

This paper tackles the problem of generating representations of underwater 3D terrain. Off-the-shelf generative models, trained on Internet-scale data but not on specialized underwater images, exhibit downgraded realism, as images of the seafloor are relatively uncommon. To this end, we introduce Dr…

Cited by 0Scholar
2026

Efficient Construction of Implicit Surface Models from a Single Image for Motion Generation

ICRA 2026poster

Implicit representations have been widely applied in robotics for obstacle avoidance and path planning. In this paper, we explore the problem of constructing an implicit distance representation from a single image. Past methods for implicit surface reconstruction, such as NeuS and its variants gener…

2026

Enhancing Vision Transformers for Object Detection via Context-Aware Token Selection and Packing

ICLR 2026poster

In recent years, the long-range attention mechanism of vision transformers has driven significant performance breakthroughs across various computer vision tasks. However, these advancements come at the cost of inefficiency and substantial computational expense, especially when dealing with sparse da…

Cited by 0SourceScholar
2026

Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling

ICLR 2026poster

The reasoning process of Large Language Models (LLMs) is often plagued by hallucinations and missing facts in question-answering tasks. A promising solution is to ground LLMs' answers in verifiable knowledge sources, such as Knowledge Graphs (KGs). Prevailing KG-enhanced methods typically constrain…

Cited by 0SourceScholar
2026

FedGLoRA: Grassmann-Manifold Federated Learning via Dual LoRA for Large EEG Models

IJCAI 2026

Large EEG Models (LEMs) are drawing increasing attention in EEG, as large-scale pretraining yields transferable representations that improve generalization. As EEG research moves to real-world deployment, objectives and paradigms diversify, yielding increasingly heterogeneous and unevenly scaled dat

Cited by 0Scholar
2026

FossilWriter: Learning Hypergraph World Models with Latent Narratives for Creative Story Generation

IJCAI 2026

Creative story generation has achieved notable progress with large language models. Current methods construct narratives through hierarchical planning or incremental expansion. These approaches produce structurally complete stories but offer limited support for organic narrative development. Many fi

Cited by 0Scholar
2026

Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

AAAI 2026technical

Vision-Language-Action (VLA) models frequently encounter challenges in generalizing to real-world environments due to inherent discrepancies between observation and action spaces. Although training data are collected from diverse camera perspectives, the models typically predict end-effector poses w

Cited by 0SourcePDFScholar
2026

Langevin Rollout Optimization for Modelic Reinforcement Learning

ICML 2026poster

Planning-driven model-based (modelic) reinforcement learning has achieved impressive success in continuous control tasks but predominantly relies on zero-order optimizers like Model Predictive Path Integral (MPPI). While robust for global exploration, MPPI updates actions solely through sampling and…

Cited by 0SourceScholar
2026

Time-Aware One Step Diffusion Network for Real-World Image Super-Resolution

CVPR 2026

Diffusion-based real-world image super-resolution (Real-ISR) methods have demonstrated impressive performance. To achieve efficient Real-ISR, many works employ Variational Score Distillation (VSD) to distill a pre-trained stable-diffusion (SD) model for one-step SR with a fixed timestep. However, si

Cited by 0SourcecodeScholar
2026

To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration

ICLR 2026poster

The scaling of Generative AI (GenAI) models into the hundreds of billions of parameters makes low-precision computation indispensable for efficient deployment. We argue that the fundamental solution lies in developing low-precision \emph{floating-point} formats, which inherently provide numerical st…

Cited by 0SourcecodeScholar
2026

Toward More Reliable Agent Evaluation: A Component-Based Benchmark Auditing Pipeline

ICML 2026poster

Reliable evaluation of large language model (LLM) agents depends critically on benchmark validity. However, agent benchmarks are increasingly complex and often contain hidden flaws arising from interactions among user instructions, environments, tools, ground-truth trajectories, and evaluation proto…

Cited by 0SourceScholar
2026

VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection

ICML 2026poster

Privacy protection has become a critical requirement in the era of ubiquitous visual data sharing, imposing higher demands on efficient and robust privacy detection algorithms. However, current robust detection models are severely hindered by the lack of comprehensive datasets. Existing privacy-orie…

Cited by 0SourceScholar
2026

Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning

ICLR 2026poster

While significant research has focused on developing embodied reasoning capabilities using Vision-Language Models (VLMs) or integrating advanced VLMs into Vision-Language-Action (VLA) models for end-to-end robot control, few studies directly address the critical gap between upstream VLM-based reason…

Cited by 0SourcecodeScholar
2025

70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)

NeurIPS 2025poster

Large-scale AI models, such as Large Language Models (LLMs) and Diffusion Models (DMs), have grown rapidly in size, creating significant challenges for efficient deployment on resource-constrained hardware. In this paper, we introduce Dynamic-Length Float (DFloat11), a lossless compression framework…

Cited by 0SourceScholar
2025

A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation

EMNLP 2025

Transformer-based Large Language Models (LLMs) struggle with inputs exceeding their training context window due to positional out-of-distribution (O.O.D.) issues that disrupt attention. Existing solutions, including fine-tuning and training-free methods, face challenges like inefficiency, redundant

2025

A Variable Stiffness Supernumerary Robotic Limb with Pneumatic-Tendon Coupled Actuation *

IROS 2025

Supernumerary robotic limbs (SRLs) can assist humans in achieving efficient and comfortable work in daily life or industrial assembly scenarios, requiring SRLs to switch between rigidity and flexibility to perform compliant movements while also providing stable support for humans to reduce fatigue f

Cited by 0SourceScholar
2025

Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model Generation

EMNLP 2025

Recent decoding methods improve the factuality of large language models (LLMs) by refining how the next token is selected during generation. These methods typically operate at the token level, leveraging internal representations to suppress superficial patterns. Nevertheless, LLMs remain prone to ha

2025

Adaptive Merchant-Centric Risk Control via Unbiased Decision-Making and Dynamic Optimization in E-Commerce

AAAI 2025technical

In the domain of merchant-oriented risk control decisions within e-commerce, balancing the effectiveness of risk management with merchant satisfaction remains a critical challenge. Strict risk control strategies, while effectively mitigating risks, often lead to increased merchant dissatisfaction. C…

Cited by 0SourcePDFScholar
2025

Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM Pretraining

NeurIPS 2025poster

Low-rank optimization has emerged as a promising approach to enabling memory-efficient training of large language models (LLMs). Existing low-rank optimization methods typically project gradients onto a low-rank subspace, reducing the memory cost of storing optimizer states. A key challenge in these…

Cited by 0SourceScholar
2025

Building 3D Representations and Generating Motions From a Single Image via Video-Generation

NeurIPS 2025poster

Autonomous robots typically need to construct representations of their surroundings and adapt their motions to the geometry of their environment. Here, we tackle the problem of constructing a policy model for collision-free motion generation, consistent with the environment, from a single input RGB…

Cited by 0SourceScholar
2025

CasiaHand: Design and Evaluation of a 15-DoF Tendon-Driven Anthropomorphic Robotic Hand

RA-L 2025

Anthropomorphic dexterous hands significantly enhance the manipulation capabilities of robots; however, balancing structural complexity with functional dexterity remains a major challenge. In this work, we propose the CasiaHand, a 15-DoF tendon-driven anthropomorphic dexterous hand featuring human-l

Cited by 7SourceScholar
2025

CoVE: Compressed Vocabulary Expansion Makes Better LLM-based Recommender Systems

ACL 2025finding

Recommender systems play a pivotal role in providing relevant content to users. With the rapid development of large language models (LLMs), researchers have begun utilizing LLMs to build more powerful recommender systems. However, existing approaches that focus on aligning LLMs with recommendation t…

2025

Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

ICCV 2025poster

While recent vision-language-action models trained on diverse robot datasets exhibit promising generalization capabilities with limited in-domain data, their reliance on compact action heads to predict discretized or continuous actions constrains adaptability to heterogeneous action spaces. We prese…

Cited by 0SourcePDFScholar
2025

From Poses to Identity: Training-Free Person Re-Identification via Feature Centralization

CVPR 2025poster

Person re-identification (ReID) aims to extract accurate identity representation features. However, during feature extraction, individual samples are inevitably affected by noise (background, occlusions, and model limitations). Considering that features from the same identity follow a normal distrib…

2025

LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid

ICLR 2025poster

Large language models (LLMs) have shown immense potential across various domains, but their high memory requirements and inference costs remain critical challenges for deployment. Post-training quantization (PTQ) has emerged as a promising technique to reduce memory requirements and decoding latency…

2025

MASTER: A Multi-Agent System with LLM Specialized MCTS

NAACL 2025long

Large Language Models (LLM) are increasingly being explored for problem-solving tasks. However, their strategic planning capability is often viewed with skepticism. Recent studies have incorporated the Monte Carlo Tree Search (MCTS) algorithm to augment the planning capacity of LLM. Despite its pote…

Cited by 0SourcePDFScholar
2025

MambaXCTrack: Mamba-Based Tracker With SSM Cross-Correlation and Motion Prompt for Ultrasound Needle Tracking

RA-L 2025

Ultrasound (US)-guided needle insertion is widely employed in percutaneous interventions. However, providing feedback on the needle tip position via US imaging presents challenges due to noise, artifacts, and the thin imaging plane of US, which degrades needle features and leads to intermittent tip

Cited by 5SourceScholar
2025

One Filters All: A Generalist Filter For State Estimation

NeurIPS 2025poster

Estimating hidden states in dynamical systems, also known as optimal filtering, is a long-standing problem in various fields of science and engineering. In this paper, we introduce a general filtering framework, $\textbf{LLM-Filter}$, which leverages large language models (LLMs) for state estimation…

Cited by 0SourceScholar
2025

Probing Political Ideology in Large Language Models: How Latent Political Representations Generalize Across Tasks

EMNLP 2025

Large language models (LLMs) encode rich internal representations of political ideology, but it remains unclear how these representations contribute to model decision-making, and how these latent dimensions interact with one another. In this work, we investigate whether ideological directions identi

2025

RS-ModCubes: Self-Reconfigurable, Scalable, Modular Cubic Robots for Underwater Operations

RA-L 2025

This paper introduces a reconfigurable underwater robot system, RS-ModCubes, which allows scalable multi-robot configurations. An RS-ModCubes system comprises multiple ModCube modules, that can travel underwater with 6 DoFs and assemble with each other into a larger structure with onboard electromag

Cited by 9SourceScholar
2025

RecGS: Removing Water Caustic With Recurrent Gaussian Splatting

RA-L 2025

Water caustics are commonly observed in seafloor imaging data from shallow-water areas. Traditional methods that remove caustic patterns from images often rely on 2D filtering or pre-training on an annotated dataset, hindering the performance when generalizing to real-world seafloor data with 3D str

Cited by 17SourceScholar
2025

Robust State Estimation for Legged Robots With Dual Beta Kalman Filter

RA-L 2025

Existing state estimation algorithms for legged robots that rely on proprioceptive sensors often overlook foot slippage and leg deformation in the physical world, leading to large estimation errors. To address this limitation, we propose a comprehensive measurement model that accounts for both foot

Cited by 6SourceScholar
2025

Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video

ICLR 2025poster

We aim to redefine robust ego-motion estimation and photorealistic 3D reconstruction by addressing a critical limitation: the reliance on noise-free data in existing models. While such sanitized conditions simplify evaluation, they fail to capture the unpredictable, noisy complexities of real-world…

2025

Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM Adaptation

ICML 2025poster

Adapting pre-trained large language models (LLMs) is crucial but challenging due to their enormous size. Parameter-efficient fine-tuning (PEFT) techniques typically employ additive adapters applied to frozen model weights. To further reduce memory usage, model weights are often compressed through qu…

Cited by 0SourcePDFScholar
2025

Uncertainty-Based Extensible Codebook for Discrete Federated Learning in Heterogeneous Data Silos

ICML 2025poster

Federated learning (FL), aimed at leveraging vast distributed datasets, confronts a crucial challenge: the heterogeneity of data across different silos. While previous studies have explored discrete representations to enhance model generalization across minor distributional shifts, these approaches…

2025

Underwater Target Tracking with Unknown Maneuver by Remotely Operated Vehicles: A Digital Twin-Driven Strategy

IROS 2025

Underwater target tracking is a critical challenge in marine exploration and defense applications due to the unknown maneuvers of target and the complex marine environment. To overcome the above challenge, this paper develops a digital twin (DT)-driven unknown maneuver target tracking strategy via r

Cited by 0SourceScholar
2024

A new approach for fine-tuning sentence transformers for intent classification and out-of-scope detection tasks

EMNLP 2024industry

In virtual assistant (VA) systems it is important to reject or redirect user queries that fall outside the scope of the system. One of the most accurate approaches for out-of-scope (OOS) rejection is to combine it with the task of intent classification on in-scope queries, and to use methods based o…

2024

Compositional Inversion for Stable Diffusion Models

AAAI 2024technical

Inversion methods, such as Textual Inversion, generate personalized images by incorporating concepts of interest provided by user images. However, existing methods often suffer from overfitting issues, where the dominant presence of inverted concepts leads to the absence of other desired concepts. I…

2024

DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision

CVPR 2024poster

We have witnessed significant progress in deep learning-based 3D vision ranging from neural radiance field (NeRF) based 3D representation learning to applications in novel view synthesis (NVS). However existing scene-level datasets for deep learning-based 3D vision limited to either synthetic enviro…

Cited by 85SourcePDFScholar
2024

DarkGS: Learning Neural Illumination and 3D Gaussians Relighting for Robotic Exploration in the Dark

IROS 2024poster

Humans have the remarkable ability to construct consistent mental models of an environment, even under limited or varying levels of illumination. We wish to endow robots with this same capability. In this paper, we tackle the challenge of constructing a photorealistic scene representation under poor…

Cited by 21SourcecodeScholar
2024

Design and Control of a Soft Supernumerary Robotic Limb Based on Fiber-Reinforced Actuator

IROS 2024poster

Supernumerary robotic limbs (SRLs) provide additional wearable limbs to enhance the user’s physical abilities. Most SRLs employ rigid structures, resulting in uncomfortable wearing experience and insufficient flexible manipulation. As a new type of SRL, soft SRLs offer operational flexibility, light…

Cited by 0SourceScholar
2024

Dynamic-Range Focal Sweep: Seamless Continuous Autofocus Based on High-Speed Vision for Magnified Object Tracking

IROS 2024poster

This paper presents an innovative continuous autofocus (C-AF) approach based on high-speed vision. It consistently provides focused images with stable and sufficiently high frame rates, aiming to improve the ability to track small, fast-moving objects in a highly magnified scene. To achieve this, we…

Cited by 1SourceScholar
2024

GroupCover: A Secure, Efficient and Scalable Inference Framework for On-device Model Protection based on TEEs

ICML 2024poster

Due to the high cost of training DNN models, how to protect the intellectual property of DNN models, especially when the models are deployed to users' devices, is becoming an important topic. One practical solution is to use Trusted Execution Environments (TEEs) and researchers have proposed various…

Cited by 2SourcePDFScholar
2024

Human-Robot Interaction Control for Multi-Mode Exosuit with Reinforcement Learning

IROS 2024poster

Soft exoskeleton robots have promising potential in walking assistance with comfortable wearing experience. In this study, an exosuit equipped with a twisted string actuator (TSA) is developed to provide powerful driving force and diverse operating modes for hemiplegic patients in daily life. It is…

Cited by 0SourceScholar
2024

Instructing Robots by Sketching: Learning from Demonstration via Probabilistic Diagrammatic Teaching

ICRA 2024poster

Learning from Demonstration (LfD) enables robots to acquire new skills by imitating expert demonstrations, allowing users to communicate their instructions intuitively. Recent progress in LfD often relies on kinesthetic teaching or teleoperation as the medium for users to specify the demonstrations.…

Cited by 10SourceScholar
2024

KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization

NeurIPS 2024poster

Efficient deployment of Large Language Models (LLMs) requires batching multiple requests together to improve throughput. As batch size, context length, or model size increases, the size of key and value (KV) cache quickly becomes the main contributor to GPU memory usage and the bottleneck of inferen…

Cited by 23SourcePDFScholar
2024

Multi-Modality Affinity Inference for Weakly Supervised 3D Semantic Segmentation

AAAI 2024technical

3D point cloud semantic segmentation has a wide range of applications. Recently, weakly supervised point cloud segmentation methods have been proposed, aiming to alleviate the expensive and laborious manual annotation process by leveraging scene-level labels. However, these methods have not effectiv…

2024

NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention

NeurIPS 2024poster

Large Language Model (LLM) inference on Central Processing Units (CPU) is challenging due to the vast quantities of Multiply-Add (MAD) matrix operations in the attention computations. This paper highlights a rare gem in modern CPUs, Single-Instruction-Multiple-Data (SIMD) registers, which allows fo…

2024

Robust Incremental Structure-from-Motion with Hybrid Features

ECCV 2024poster

"Structure-from-Motion (SfM) has become a ubiquitous tool for camera calibration and scene reconstruction with many downstream applications in computer vision and beyond. While the state-of-the-art SfM pipelines have reached a high level of maturity in well-textured and well-configured scenes over t…

2024

Simultaneous Geometry and Pose Estimation of Held Objects Via 3D Foundation Models

RA-L 2024

Humans have the remarkable ability to use held objects as tools to interact with their environment. Humans internally estimate how hand movements affect the object's movement. We wish to endow robots with this capability. We contribute methodology to jointly estimate the geometry and pose of objects

Cited by 9SourceScholar
2024

Stronger Than You Think: Benchmarking Weak Supervision on Realistic Tasks

NeurIPS 2024poster

Weak supervision (WS) is a popular approach for label-efficient learning, leveraging diverse sources of noisy but inexpensive *weak labels* to automatically annotate training data. Despite its wide usage, WS and its practical value are challenging to benchmark due to the many knobs in its setup, inc…

2024

Teaching Robots Where To Go And How To Act With Human Sketches via Spatial Diagrammatic Instructions

IROS 2024poster

This paper introduces Spatial Diagrammatic Instructions (SDIs), an approach for human operators to specify objectives and constraints that are related to spatial regions in the working environment. Human operators are enabled to sketch out regions directly on camera images that correspond to the obj…

Cited by 0SourceScholar
2024

Transformer-Based Selective Super-resolution for Efficient Image Refinement

AAAI 2024technical

Conventional super-resolution methods suffer from two drawbacks: substantial computational cost in upscaling an entire large image, and the introduction of extraneous or potentially detrimental information for downstream computer vision tasks during the refinement of the background. To solve these i…

2024

Unifying Representation and Calibration With 3D Foundation Models

RA-L 2024

Representing the environment is a central challenge in robotics, and is essential for effective decision-making. Traditionally, before capturing images with a manipulator-mounted camera, users need to calibrate the camera using a specific external marker, such as a checkerboard or AprilTag. However,

Cited by 12SourceScholar
2023

AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

NeurIPS 2023spotlight

Large language models (LLMs) such as ChatGPT have seen widespread adoption due to their ability to follow user instructions well. Developing these LLMs involves a complex yet poorly understood workflow requiring training with human feedback. Replicating and understanding this instruction-following p…

Cited by 523SourcePDFScholar
2023

Beyond NeRF Underwater: Learning Neural Reflectance Fields for True Color Correction of Marine Imagery

RA-L 2023

Underwater imagery often exhibits distorted coloration as a result of light-water interactions, which complicates the study of benthic environments in marine biology and geography. In this research, we propose an algorithm to restore the true color (albedo) in underwater imagery by jointly learning

Cited by 38SourcecodeScholar
2023

CDFI: Cross Domain Feature Interaction for Robust Bronchi Lumen Detection

ICRA 2023poster

Endobronchial intervention is increasingly used as a minimally invasive means for the treatment of pulmonary diseases. In order to reduce the difficulty of manipulation in complex airway networks, robust lumen detection is essential for intraoperative guidance. However, these methods are sensitive t…

Cited by 1SourceScholar
2023

Coder Reviewer Reranking for Code Generation

ICML 2023poster

Sampling diverse programs from a code language model and reranking with model likelihood is a popular method for code generation but it is prone to preferring degenerate solutions. Inspired by collaborative programming, we propose Coder-Reviewer reranking. We augment Coder language models from past…

2023

DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

ICML 2023poster

We introduce DS-1000, a code generation benchmark with a thousand data science problems spanning seven Python libraries, such as Numpy and Pandas. Compared to prior works, DS-1000 incorporates three core features. First, our problems reflect diverse, realistic, and practical use cases since we colle…

2023

Efficient Graph Field Integrators Meet Point Clouds

ICML 2023poster

We present two new classes of algorithms for efficient field integration on graphs encoding point cloud data. The first class, $\mathrm{SeparatorFactorization}$ (SF), leverages the bounded genus of point cloud mesh graphs, while the second class, $\mathrm{RFDiffusion}$ (RFD), uses popular $\epsilon$…

2023

Graph Self-supervised Learning via Proximity Distribution Minimization

UAI 2023poster

Self-supervised learning (SSL) for graphs is an essential problem since graph data are ubiquitous and labeling can be costly. We argue that existing SSL approaches for graphs have two limitations. First, they rely on corruption techniques such as node attribute perturbation and edge dropping to gene…

Cited by 0SourcePDFScholar
2023

Neural Frailty Machine: Beyond proportional hazard assumption in neural survival regressions

NeurIPS 2023poster

We present neural frailty machine (NFM), a powerful and flexible neural modeling framework for survival regressions. The NFM framework utilizes the classical idea of multiplicative frailty in survival analysis as a principled way of extending the proportional hazard assumption, at the same time bein…

2023

SAD: Semi-Supervised Anomaly Detection on Dynamic Graphs

IJCAI 2023poster

Anomaly detection aims to distinguish abnormal instances that deviate significantly from the majority of benign ones. As instances that appear in the real world are naturally connected and can be represented with graphs, graph neural networks become increasingly popular in tackling the anomaly detec…

2023

TempLM: Distilling Language Models into Template-Based Generators

ACL 2023findings

While pretrained language models (PLMs) have greatly improved text generation, they have also been known to produce unfaithful or inappropriate content. In contrast, classic template-based systems provide strong guarantees of faithfulness at the cost of fluency. We propose TempLM, which achieves the…

2022

Decentralized Training of Foundation Models in Heterogeneous Environments

NeurIPS 2022accept

Training foundation models, such as GPT-3 and PaLM, can be extremely expensive, often involving tens of thousands of GPUs running continuously for months. These models are typically trained in specialized clusters featuring fast, homogeneous interconnects and using carefully designed software system…

2022

Dual-camera High Magnification Surveillance System with Non-delay Gaze Control and Always-in-focus Function in Indoor Scenes

IROS 2022poster

This study proposes a dual-camera system for indoor high magnification surveillance which is capable of achieving always-in-focus and non-delay gaze control based on high-speed vision. The users are enabled to move the mouse freely on the wide-view screen while observing its in-focal zoom-in monitor…

Cited by 6SourceScholar
2022

From block-Toeplitz matrices to differential equations on graphs: towards a general theory for scalable masked Transformers

ICML 2022spotlight

In this paper we provide, to the best of our knowledge, the first comprehensive approach for incorporating various masking mechanisms into Transformers architectures in a scalable way. We show that recent results on linear causal attention (Choromanski et al., 2021) and log-linear RPE-attention (Luo…

2022

LADIS: Language Disentanglement for 3D Shape Editing

EMNLP 2022finding

Natural language interaction is a promising direction for democratizing 3D shape design. However, existing methods for text-driven 3D shape editing face challenges in producing decoupled, local edits to 3D shapes. We address this problem by learning disentangled latent representations that ground la…

2022

Retaining Knowledge for Learning with Dynamic Definition

NeurIPS 2022accept

Machine learning models are often deployed in settings where they must be constantly updated in response to the changes in class definitions while retaining high accuracy on previously learned definitions. A classical use case is fraud detection, where new fraud schemes come one after another. While…

Cited by 2SourcePDFScholar
2022

Structural Contrastive Representation Learning for Zero-shot Multi-label Text Classification

EMNLP 2022finding

Zero-shot multi-label text classification (ZMTC) is a fundamental task in natural language processing with applications in the cold start problem of recommendation systems. Ideally, one would learn an expressive representation of both input text and label features so that ZMTC is transformed into a…

Cited by 13SourcePDFScholar
2021

Design Paradigms Based on Spring Agonists for Underactuated Robot Hands: Concepts and Application

ICRA 2021poster

In this paper, we focus on a rarely used paradigm in the design of underactuated robot hands: the use of springs as agonists and tendons as antagonists. We formalize this approach in a design matrix also considering its interplay with the underactuation method used (one tendon for multiple joints vs…

Cited by 4SourceScholar
2021

On the Inductive Bias of Masked Language Modeling: From Statistical to Syntactic Dependencies

NAACL 2021long

We study how masking and predicting tokens in an unsupervised fashion can give rise to linguistic structures and downstream performance gains. Recent theories have suggested that pretrained language models acquire useful inductive biases through masks that implicitly act as cloze reductions for down…

2021

PSRR-MaxpoolNMS: Pyramid Shifted MaxpoolNMS With Relationship Recovery

CVPR 2021poster

Non-maximum Suppression (NMS) is an essential post-processing step in modern convolutional neural networks for object detection. Unlike convolutions which are inherently parallel, the de-facto standard for NMS, namely GreedyNMS, cannot be easily parallelized and thus could be the performance bottlen…

Cited by 12PDFcodeScholar
2021

Revisiting Few-sample BERT Fine-tuning

ICLR 2021poster

This paper is a study of fine-tuning of BERT contextual representations, with focus on commonly observed instabilities in few-sample scenarios. We identify several factors that cause this instability: the common use of a non-standard optimization method with biased gradient estimation; the limited a…

2020

Demystifying Orthogonal Monte Carlo and Beyond

NeurIPS 2020poster

Orthogonal Monte Carlo (OMC) is a very effective sampling algorithm imposing structural geometric conditions (orthogonality) on samples for variance reduction. Due to its simplicity and superior performance as compared to its Quasi Monte Carlo counterparts, OMC is used in a wide spectrum of challeng…

2020

Identifying Mislabeled Data using the Area Under the Margin Ranking

NeurIPS 2020poster

Not all data in a typical training set help with generalization; some samples can be overly ambiguous or outrightly mislabeled. This paper introduces a new method to identify such samples and mitigate their impact when training neural networks. At the heart of our algorithm is the Area Under the Mar…

2020

Mitigating Overfitting in Supervised Classification from Two Unlabeled Datasets: A Consistent Risk Correction Approach

AISTATS 2020poster

The recently proposed unlabeled-unlabeled (UU) classification method allows us to train a binary classifier only from two unlabeled datasets with different class priors. Since this method is based on the empirical risk minimization, it works as if it is a supervised classification method, compatible…

Cited by 71SourcePDFScholar
2020

Splitting vs. Merging: Mining Object Regions with Discrepancy and Intersection Loss for Weakly Supervised Semantic Segmentation

ECCV 2020poster

In this paper we focus on the task of weakly-supervised semantic segmentation supervised with image-level labels. Since the pixel-level annotation is not available in the training process, we rely on region mining models to estimate the pseudo-masks from the image-level labels. Thus, in order to imp…

Cited by 81SourcePDFScholar
2019

SWALP : Stochastic Weight Averaging in Low Precision Training

ICML 2019oral

Low precision operations can provide scalability, memory savings, portability, and energy efficiency. This paper proposes SWALP, an approach to low precision training that averages low-precision SGD iterates with a modified learning rate schedule. SWALP is easy to implement and can match the perform…

2019

Simplifying Graph Convolutional Networks

ICML 2019oral

Graph Convolutional Networks (GCNs) and their variants have experienced significant attention and have become the de facto methods for learning graph representations. GCNs derive inspiration primarily from recent deep learning approaches, and as a result, may inherit unnecessary complexity and redun…