← Search

Xinyu Zhang

123 accepted papers

2026

Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language Models

CVPR 2026

While model merging has demonstrated remarkable success for large language models (LLMs), its application to vision-language models (VLMs) remains largely underexplored. Recent methods attempt to enhance VLM reasoning capabilities by integrating specialized LLM parameters through layer-wise merging.

Cited by 0SourceScholar
2026

Constrained Bayesian Experimental Design via Online Planning

ICML 2026poster

Bayesian experimental design (BED) is a principled framework for data-efficient design of sequential experiments. However, existing BED methods are unable to adapt to dynamic constraints inherent in real-world tasks due to budget limitations, varying costs, or physical constraints that restrict how …

Cited by 0SourceScholar
2026

Correspondence Coverage Matters for Multi-Modal Dataset Distillation

AAAI 2026technical

Multi-modal dataset distillation (DD) condenses large datasets into compact ones that retain task efficacy by capturing correspondence patterns, i.e., shared semantics between paired modalities. However, such patterns rely on cross-modal similarity and cannot be faithfully captured by intra-modal si

Cited by 0SourcePDFScholar
2026

Cutting the Skip: Training Residual-Free Transformers

ICLR 2026poster

Transformers have achieved remarkable success across a wide range of applications, a feat often attributed to their scalability. Yet training them without residual (skip) connections remains notoriously difficult. While skips stabilize optimization, they also disrupt the hierarchical structure of re…

Cited by 0SourceScholar
2026

Encode Geometric Diagram as Geo-Graph in Geometry Problem Solving

AAAI 2026technical

Geometry Problem Solving has become a hot topic these years due to its complexity of enabling the machine with geometric abstraction, multi-modal reasoning and mathematical capabilities. Majority of research works place their attention on the fusion of multi-modal data or the synergistic combination

Cited by 0SourcePDFScholar
2026

Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves

CVPR 2026

Understanding hand-object interaction (HOI) is fundamental to computer vision, robotics, and AR/VR. However, conventional hand videos often lack essential physical information, such as contact forces and motion dynamics, and are prone to frequent occlusions. To address these challenges, we present G

Cited by 0SourceScholar
2026

IntentMotion: Learning Intent-Aware Human Motion from Language in 3D Scenes

AAAI 2026technical

Generating human motion in complex 3D scenes from text is a challenging task with broad applications. However, existing methods often overlook realistic physical contact, resulting in visually plausible but physically unrealistic motion, e.g., penetration. To alleviate this, we propose IntentMotion,

Cited by 0SourcePDFScholar
2026

Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer

CVPR 2026

Motion transfer has emerged as a promising direction for controllable video generation, yet existing methods largely focus on single-object scenarios and struggle when multiple objects require distinct motion patterns. In this work, we present FlexiMMT, the first implicit image-to-video (I2V) motion

Cited by 0SourcecodeScholar
2026

MICE-Bench: A Challenging and Comprehensive Benchmark for Multi-Reference Image Creation and Editing

ICML 2026poster

The paradigm of visual generation is rapidly shifting from single-image conditioning toward multi-image conditioning, making the ability to synthesize and edit images based on multiple visual references a critical capability. Despite this trend, existing benchmarks remain largely limited to single-r…

Cited by 0SourceScholar
2026

OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

CVPR 2026

Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties, yet dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annotations, insufficient modality-specific modeling, and challenges in reliability. I

Cited by 0SourceScholar
2026

Satellite-Text-Prompted Large Language Model for Photovoltaic Power Forecasting

AAAI 2026technical

Photovoltaic (PV) power forecasting is critical for the operation of solar power plants and the coordination of energy within power grids. This work aims to predict future PV power time series by leveraging multimodal data. While recent studies have incorporated numerical modalities such as satellit

Cited by 0SourcePDFScholar
2026

Stochastic Neural Ray Tracing for Radio Frequency Channel Modeling

ICML 2026poster

Wireless channel modeling is essential for the design, analysis, and optimization of modern wireless sensing and communication systems. However, accurately modeling wireless channels in electrically large and complex environments remains a long-standing challenge, owing to the intricate interactions…

Cited by 0SourceScholar
2026

Task-Agnostic Amortized Multi-Objective Optimization

ICLR 2026poster

Balancing competing objectives is omnipresent across disciplines, from drug design to autonomous systems. Multi-objective Bayesian optimization is a promising solution for such expensive, black-box problems: it fits probabilistic surrogates and selects new designs via an acquisition function that ba…

Cited by 0SourceScholar
2026

Temporal Equilibrium MeanFlow: Bridging the Scale Gap for One-Step Generation

CVPR 2026

MeanFlow is a powerful few-step generative framework that can be trained from scratch, but its performance degrades significantly when the one-step loss uses a large portion of training data. This stems from a temporal scale imbalance: gradients from different stages of generation contribute unevenl

Cited by 0SourceScholar
2025

Adapting to Observation Length of Trajectory Prediction via Contrastive Learning

CVPR 2025poster

The ability to adapt to varying observation lengths is crucial for human trajectory prediction tasks, particularly in scenarios with limited observation lengths or missing data. Existing approaches mainly focus on introducing novel architectures or additional structural components, which substantial…

2025

Are Images Indistinguishable to Humans Also Indistinguishable to Classifiers?

CVPR 2025poster

The ultimate goal of generative models is to perfectly capture the data distribution. For image generation, common metrics of visual quality (e.g., FID) and the perceived truthfulness of generated images seem to suggest that we are nearing this goal. However, through distribution classification task…

Cited by 2SourcePDFScholar
2025

AudioCache: Accelerate Audio Generation With Training-Free Layer Caching

ICASSP 2025accepted

Diffusion models have become the primary choice in audio generation. However, their slow generation speed necessitates acceleration techniques. While current audio generation methods primarily target U-Net-based models, the Diffusion Transformer (DiT) is emerging as the trend in audio generation. As…

Cited by 0SourceScholar
2025

Autoregressive Action Sequence Learning for Robotic Manipulation

RA-L 2025

Designing a universal policy architecture that performs well across diverse robots and task configurations remains a key challenge. In this work, we address this by representing robot actions as sequential data and generating actions through autoregressive sequence modeling. Existing autoregressive

Cited by 37SourcecodeScholar
2025

Beyond the Surface: Measuring Self-Preference in LLM Judgments

EMNLP 2025

Recent studies show that large language models (LLMs) exhibit self-preference bias when serving as judges, meaning they tend to favor their own responses over those generated by other models. Existing methods typically measure this bias by calculating the difference between the scores a judge model

2025

Bi-Tuning with Collaborative Information for Controllable LLM-based Sequential Recommendation

ACL 2025long

Sequential recommender systems, which leverage historical interactions to deliver targeted recommendations, have been significantly advanced by large language models (LLMs). However, LLM-based generative sequential recommendation often faces two key challenges: the lack of collaborative knowledge an…

Cited by 0SourcePDFScholar
2025

Co-MTP: A Cooperative Trajectory Prediction Framework with Multi-Temporal Fusion for Autonomous Driving

ICRA 2025

Vehicle-to-everything technologies (V2X) have become an ideal paradigm to extend the perception range and see through the occlusion. Exiting efforts focus on single-frame cooperative perception, however, how to capture the temporal cue between frames with V2X to facilitate the prediction task even t

Cited by 16SourcecodeScholar
2025

CoFFT: Chain of Foresight-Focus Thought for Visual Language Models

NeurIPS 2025poster

Despite significant advances in Vision Language Models (VLMs), they remain constrained by the complexity and redundancy of visual input. When images contain large amounts of irrelevant information, VLMs are susceptible to interference, thus generating excessive task-irrelevant reasoning processes or…

Cited by 0SourceScholar
2025

CogReact: A Reinforced Framework to Model Human Cognitive Reaction Modulated by Dynamic Intervention

ICML 2025poster

Using deep neural networks as computational models to simulate cognitive processes can provide key insights into human behavioral dynamics. Challenges arise when environments are highly dynamic, obscuring stimulus-behavior relationships. However, the majority of current research focuses on simulatin…

Cited by 0SourcePDFScholar
2025

CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis

EMNLP 2025

Cultural competence, defined as the ability to understand and adapt to multicultural contexts, is increasingly vital for large language models (LLMs) in global environments. While several cultural benchmarks exist to assess LLMs’ cultural competence, current evaluations suffer from fragmented taxono

2025

Dually Self-Improved Counterfactual Data Augmentation Using Large Language Model

ACL 2025long

Counterfactual data augmentation, which generates minimally edited tokens to alter labels, has become a key approach to improving model robustness in natural language processing (NLP). It is usually implemented by first identifying the causal terms and then modifying these terms to create counterfac…

Cited by 0SourcePDFScholar
2025

Enhancing Image Editing with Chain-of-Thought Reasoning and Multimodal Large Language Models

ICASSP 2025accepted

Image editing in our daily lives often requires models to first understand user’s intention and then proceed with the editing. Despite significant advancements in image editing technology, understanding and executing complex instructions remains a substantial challenge. Existing image editing models…

Cited by 0SourceScholar
2025

EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding

AAAI 2025technical

Chart understanding enables automated data analysis for humans, which requires models to achieve highly accurate visual comprehension. While existing Visual Language Models (VLMs) have shown progress in chart understanding, the lack of high-quality training data and comprehensive evaluation benchmar…

2025

Exploring Multimodal Relation Extraction of Hierarchical Tabular Data with Multi-task Learning

ACL 2025long

Relation Extraction (RE) is a key task in table understanding, aiming to extract semantic relations between columns. However, complex tables with hierarchical headers are hard to obtain high-quality textual formats (e.g., Markdown) for input under practical scenarios like webpage screenshots and sca…

2025

Failure Forecasting Boosts Robustness of Sim2Real Rhythmic Insertion Policies

IROS 2025

This paper addresses the challenges of Rhythmic Insertion Tasks (RIT), where a robot must repeatedly perform high-precision insertions, such as screwing a nut into a bolt with a wrench. The inherent difficulty of RIT lies in achieving millimeter-level accuracy and maintaining consistent performance

Cited by 0SourcecodeScholar
2025

FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts

NeurIPS 2025poster

In this paper, we introduce FedMGP, a new paradigm for personalized federated prompt learning in vision-language models (VLMs). Existing federated prompt learning (FPL) methods often rely on a single, text-only prompt representation, which leads to client-specific overfitting and unstable aggregatio…

Cited by 0SourcecodeScholar
2025

Generalizing Experience for Language Agents with Hierarchical MetaFlows

NeurIPS 2025poster

Recent efforts to employ large language models (LLMs) as agents have demonstrated promising results in a wide range of multi-step agent tasks. However, existing agents lack an effective experience reuse approach to leverage historical completed tasks. In this paper, we propose a novel experience reu…

Cited by 0SourceScholar
2025

HAC-LOCO: Learning Hierarchical Active Compliance Control for Quadruped Locomotion under Continuous External Disturbances

IROS 2025

Despite recent remarkable achievements in quadruped control, it remains challenging to ensure robust and compliant locomotion in the presence of unforeseen external disturbances. Existing methods prioritize locomotion robustness over compliance, often leading to stiff, high-frequency motions, and en

Cited by 4SourceScholar
2025

Learning Upright and Forward-Facing Object Poses using Category-level Canonical Representations

IROS 2025

Constructing a unified canonical pose representation for 3D object categories is crucial for pose estimation and robotic scene understanding. Previous unified pose representations often relied on manual alignment, such as in ShapeNet and ModelNet. Recently, self-supervised canonicalization methods h

Cited by 0SourcecodeScholar
2025

PABBO: Preferential Amortized Black-Box Optimization

ICLR 2025spotlight

Preferential Bayesian Optimization (PBO) is a sample-efficient method to learn latent user utilities from preferential feedback over a pair of designs. It relies on a statistical surrogate model for the latent function, usually a Gaussian process, and an acquisition strategy to select the next candi…

2025

PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

ACL 2025long

Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physics theorems and constraints. We present PhysReason, a 1,200-problem benchmark co…

Cited by 0SourcePDFScholar
2025

Radio Frequency Ray Tracing with Neural Object Representation for Enhanced RF Modeling

CVPR 2025poster

Radio frequency (RF) propagation modeling poses unique electromagnetic simulation challenges. While recent neural representations have shown success in visible spectrum rendering, the fundamentally different scales and physics of RF signals require novel modeling paradigms. In this paper, we introdu…

Cited by 0SourcePDFScholar
2025

Rectifying Soft-Label Entangled Bias in Long-Tailed Dataset Distillation

NeurIPS 2025poster

Dataset distillation compresses large-scale datasets into compact, highly informative synthetic data, significantly reducing storage and training costs. However, existing research primarily focuses on balanced datasets and struggles to perform under real-world long-tailed distributions. In this work…

Cited by 0SourceScholar
2025

Revisit Self-Debugging with Self-Generated Tests for Code Generation

ACL 2025long

Large language models (LLMs) have demonstrated significant advancements in code generation, yet they still face challenges when tackling tasks that extend beyond their basic capabilities. Recently, the concept of self-debugging has been proposed as a way to enhance code generation performance by lev…

Cited by 0SourcePDFScholar
2025

Revisiting Chain-of-Thought Prompting: Zero-shot Can Be Stronger than Few-shot

EMNLP 2025

In-Context Learning (ICL) is an essential emergent ability of Large Language Models (LLMs), and recent studies introduce CoT to exemplars of ICL to enhance the reasoning capability, especially in mathematics tasks. However, given the continuous advancement of model capabilities, it remains unclear w

Cited by 0SourcePDFScholar
2025

Reward Mixology: Crafting Hybrid Signals for Reinforcement Learning Driven In-Context Learning

EMNLP 2025

In-context learning (ICL) performance heavily relies on the quality and ordering of demonstrations. Iterative selection (IS) is a promising approach to address this issue, but existing IS methods face two key challenges: the oversimplification of process reward signals that guide intermediate steps

Cited by 0SourcePDFScholar
2025

SCoder: Progressive Self-Distillation for Bootstrapping Small-Scale Data Synthesizers to Empower Code LLMs

EMNLP 2025

Existing code large language models (LLMs) often rely on large-scale instruction data distilled from proprietary LLMs for fine-tuning, which typically incurs high costs. In this paper, we explore the potential of small-scale open-source LLMs (e.g., 7B) as synthesizers for high-quality code instructi

2025

Scaling Diffusion Transformers Efficiently via $\mu$P

NeurIPS 2025poster

Diffusion Transformers have emerged as the foundation for vision generative models, but their scalability is limited by the high cost of hyperparameter (HP) tuning at large scales. Recently, Maximal Update Parametrization ($\mu$P) was proposed for vanilla Transformers, which enables stable HP transf…

Cited by 0SourceScholar
2025

TAGMO: Temporal Control Audio Generation for Multiple Visual Objects Without Training

ICASSP 2025accepted

With the great popularity of Sora, video-based audio generation has become indispensable. While numerous video-to-audio generation models have emerged, they frequently face difficulties including semantic incompatibilities and synchronization problems, especially in situations with multiple objects.…

Cited by 0SourceScholar
2025

TTA-FedDG: Leveraging Test-Time Adaptation to Address Federated Domain Generalization

AAAI 2025technical

In recent years, Federated Domain Generalization (FedDG) has succeeded in generalizing to unknown clients (domains). However, current methods only utilize training data, and when there is a significant difference between the unknown client and source client domains (domain shift), these methods cann…

Cited by 0SourcePDFScholar
2025

Towards Neurorobotic Interface for Finger Joint Angle Estimation: A Multi-Stage CNN-LSTM Network with Transfer Learning

ICRA 2025

To maximize the autonomy of individuals with upper limb amputations in daily activities, leveraging forearm muscle information to infer movement intent is a promising research direction. While current prosthetic hand technologies can utilize forearm muscle data to achieve basic movements such as gra

Cited by 3SourceScholar
2025

Unified Graph and Hypergraph Neural Network for Next-item Recommendation

ICASSP 2025accepted

The task of next-item recommendation is a crucial component in recommendation systems. The challenge of this task lies in extracting complex interaction information from users’ historical interactions with items. While prior research has transformed users’ interaction histories into graphs and hyper…

Cited by 0SourceScholar
2025

V2X-Radar: A Multi-modal Dataset with 4D Radar for Cooperative Perception

NeurIPS 2025spotlight

Modern autonomous vehicle perception systems often struggle with occlusions and limited perception range. Previous studies have demonstrated the effectiveness of cooperative perception in extending the perception range and overcoming occlusions, thereby enhancing the safety of autonomous driving. In…

Cited by 0SourceScholar
2025

VProChart: Answering Chart Question Through Visual Perception Alignment Agent and Programmatic Solution Reasoning

AAAI 2025technical

Charts are widely used for data visualization across various fields, including education, research, and business. Chart Question Answering (CQA) is an emerging task focused on the automatic interpretation and reasoning of data presented in charts. However, chart images are inherently difficult to in…

2024

Backdoor Contrastive Learning via Bi-level Trigger Optimization

ICLR 2024poster

Contrastive Learning (CL) has attracted enormous attention due to its remarkable capability in unsupervised representation learning. However, recent works have revealed the vulnerability of CL to backdoor attacks: the feature extractor could be misled to embed backdoored data close to an attack targ…

2024

CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question Answering

CVPR 2024poster

Diagram Question Answering (DQA) is a challenging task requiring models to answer natural language questions based on visual diagram contexts. It serves as a crucial basis for academic tutoring technical support and more practical applications. DQA poses significant challenges such as the demand for…

Cited by 6SourcePDFScholar
2024

CompetEvo: Towards Morphological Evolution from Competition

IJCAI 2024poster

Training an agent to adapt to specific tasks through co-optimization of morphology and control has widely attracted attention. However, whether there exists an optimal configuration and tactics for agents in a multiagent competition scenario is still an issue that is challenging to definitively conc…

2024

DAP: Diffusion-based Affordance Prediction for Multi-modality Storage

IROS 2024poster

Solving storage problems—where objects must be accurately placed into containers with precise orientations and positions—presents a distinct challenge that extends beyond traditional rearrangement tasks. These challenges are primarily due to the need for fine-grained 6D manipulation and the inherent…

Cited by 1SourcecodeScholar
2024

Debatrix: Multi-dimensional Debate Judge with Iterative Chronological Analysis Based on LLM

ACL 2024findings

How can we construct an automated debate judge to evaluate an extensive, vibrant, multi-turn debate? This task is challenging, as judging a debate involves grappling with lengthy texts, intricate argument relationships, and multi-dimensional assessments.At the same time, current research mainly focu…

2024

Debiased Offline Representation Learning for Fast Online Adaptation in Non-stationary Dynamics

ICML 2024poster

Developing policies that can adapt to non-stationary environments is essential for real-world reinforcement learning applications. Nevertheless, learning such adaptable policies in offline settings, with only a limited set of pre-collected trajectories, presents significant challenges. A key difficu…

2024

Efficient Multi-scale Network with Learnable Discrete Wavelet Transform for Blind Motion Deblurring

CVPR 2024poster

Coarse-to-fine schemes are widely used in traditional single-image motion deblur; however in the context of deep learning existing multi-scale algorithms not only require the use of complex modules for feature fusion of low-scale RGB images and deep semantics but also manually generate low-resolutio…

2024

Estimating Conditional Average Treatment Effects via Sufficient Representation Learning

IJCAI 2024poster

Estimating the conditional average treatment effects (CATE) is very important in causal inference and has a wide range of applications across many fields. In the estimation process of CATE, the unconfoundedness assumption is typically required to ensure the identifiability of the regression problems…

Cited by 1SourcePDFScholar
2024

Evaluation of Text-to-Video Generation Models: A Dynamics Perspective

NeurIPS 2024poster

Comprehensive and constructive evaluation protocols play an important role when developing sophisticated text-to-video (T2V) generation models. Existing evaluation protocols primarily focus on temporal consistency and content continuity, yet largely ignore dynamics of video content. Such dynamics is…

2024

HOIAnimator: Generating Text-prompt Human-object Animations using Novel Perceptive Diffusion Models

CVPR 2024poster

To date the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehensive physics-centric…

Cited by 11SourcePDFScholar
2024

LiDAR-PTQ: Post-Training Quantization for Point Cloud 3D Object Detection

ICLR 2024poster

Due to highly constrained computing power and memory, deploying 3D lidar-based detectors on edge devices equipped in autonomous vehicles and robots poses a crucial challenge. Being a convenient and straightforward model compression approach, Post-Training Quantization (PTQ) has been widely adopted i…

2024

Low Category Uncertainty and High Training Potential Instance Learning for Unsupervised Domain Adaptation

AAAI 2024technical

Recently, instance contrastive learning achieves good results in unsupervised domain adaptation. It reduces the distances between positive samples and the anchor, increases the distances between negative samples and the anchor, and learns discriminative feature representations for target samples. Ho…

2024

Modeling Route Representation With Mixed-Scale Hierarchical Transformer

ICASSP 2024accepted

Modeling route representation aims to obtain contextual representations of an entire route for various traffic-related tasks. In reality, spatial-temporal data often exhibits multi-scale characteristics, which are utilized by many studies to enhance their performance. However, there is still a lack…

Cited by 0SourceScholar
2024

Multi-Objective Forward Reasoning and Multi-Reward Backward Refinement for Product Review Summarization

COLING 2024main

Product review summarization aims to generate a concise summary based on product reviews to facilitate purchasing decisions. This intricate task gives rise to three challenges in existing work: factual accuracy, aspect comprehensiveness, and content relevance. In this paper, we first propose an FB-T…

Cited by 1SourcePDFScholar
2024

One-Shot Imitation Learning with Invariance Matching for Robotic Manipulation

RSS 2024poster

Learning a single universal policy that can perform a diverse set of manipulation tasks is a promising new direction in robotics. However, existing techniques are limited to learning policies that can only perform tasks that are encountered during training, and require a large number of demonstratio…

2024

PA-LOCO: Learning Perturbation-Adaptive Locomotion for Quadruped Robots

IROS 2024poster

Locomotion control is still a challenging task for quadruped robots traversing diverse terrains amidst unforeseen disturbances. Recently, privileged learning has been employed to learn reliable and robust quadrupedal locomotion over various terrains based on a teacher-student architecture. However,…

Cited by 3SourceScholar
2024

Scaling Manipulation Learning with Visual Kinematic Chain Prediction

CoRL 2024poster

Learning general-purpose models from diverse datasets has achieved great success in machine learning. In robotics, however, existing methods in multi-task learning are typically constrained to a single robot and workspace, while recent work such as RT-X requires a non-trivial action normalization pr…

Cited by 1SourcecodeScholar
2024

Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement Learning

NeurIPS 2024spotlight

As a marriage between offline RL and meta-RL, the advent of offline meta-reinforcement learning (OMRL) has shown great promise in enabling RL agents to multi-task and quickly adapt while acquiring knowledge safely. Among which, context-based OMRL (COMRL) as a popular paradigm, aims to learn a univer…

2024

Trajectory set Empowered Hypergraph Transformer for Mobile Sensor Based Traffic Prediction

ICASSP 2024accepted

Traffic speed prediction is vital for intelligent transportation systems. However, most existing methods focus on costly static sensors. In contrast, utilizing GPS devices from vehicles as mobile sensors offers a cost-effective means to gather dynamic traffic data. Despite the presence of historical…

Cited by 0SourceScholar
2024

VRP-SAM: SAM with Visual Reference Prompt

CVPR 2024poster

In this paper we propose a novel Visual Reference Prompt (VRP) encoder that empowers the Segment Anything Model (SAM) to utilize annotated reference images as prompts for segmentation creating the VRP-SAM model. In essence VRP-SAM can utilize annotated reference images to comprehend specific objects…

2023

Argue with Me Tersely: Towards Sentence-Level Counter-Argument Generation

EMNLP 2023long main

Counter-argument generation—a captivating area in computational linguistics—seeks to craft statements that offer opposing views. While most research has ventured into paragraph-level generation, sentence-level counter-argument generation beckons with its unique constraints and brevity-focused challe…

Cited by 0SourcecodeScholar
2023

BEVHeight: A Robust Framework for Vision-Based Roadside 3D Object Detection

CVPR 2023poster

While most recent autonomous driving system focuses on developing perception methods on ego-vehicle sensors, people tend to overlook an alternative approach to leverage intelligent roadside cameras to extend the perception ability beyond the visual range. We discover that the state-of-the-art vision…

2023

CMG-Net: An End-to-End Contact-based Multi-Finger Dexterous Grasping Network

ICRA 2023poster

In this paper, we propose a novel representation for grasping using contacts between multi-finger robotic hands and objects to be manipulated. This representation significantly reduces the prediction dimensions and accelerates the learning process. We present an effective end-to-end network, CMG-Net…

Cited by 4SourceScholar
2023

CO-Net: Learning Multiple Point Cloud Tasks at Once with A Cohesive Network

ICCV 2023poster

We present CO-Net, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains. CO-Net maintains the characteristics of high storage efficiency since models with the preponderance of shared parameters can be assembled into a single model. Specific…

Cited by 7PDFScholar
2023

Diagram Visual Grounding: Learning to See with Gestalt-Perceptual Attention

IJCAI 2023poster

Diagram visual grounding aims to capture the correlation between language expression and local objects in the diagram, and plays an important role in the applications like textbook question answering and cross-modal retrieval. Most diagrams consist of several colors and simple geometries. This resul…

2023

Dual Memory Aggregation Network for Event-Based Object Detection with Learnable Representation

AAAI 2023technical

Event-based cameras are bio-inspired sensors that capture brightness change of every pixel in an asynchronous manner. Compared with frame-based sensors, event cameras have microsecond-level latency and high dynamic range, hence showing great potential for object detection under high-speed motion and…

2023

Evaluating Embedding APIs for Information Retrieval

ACL 2023industry

The ever-increasing size of language models curtails their widespread access to the community, thereby galvanizing many companies and startups into offering access to large language models through APIs. One particular API, suitable for dense retrieval, is the semantic embedding API that builds vecto…

Cited by 24SourcePDFScholar
2023

General Category Network: Handwritten Mathematical Expression Recognition with Coarse-Grained Recognition Task

ICASSP 2023accepted

Handwritten Mathematical Expression Recognition (HMER) is an important task in pattern recognition. It is a challenging task due to symbols resembling each other in appearance("z/2", "B/β") and the complex mathematical syntax. The encoder-decoder architecture has been widely used in recent HMER meth…

Cited by 0SourceScholar
2023

HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception

NeurIPS 2023poster

Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight…

2023

Hence, Socrates is mortal: A Benchmark for Natural Language Syllogistic Reasoning

ACL 2023findings

Syllogistic reasoning, a typical form of deductive reasoning, is a critical capability widely required in natural language understanding tasks, such as text entailment and question answering. To better facilitate research on syllogistic reasoning, we develop a benchmark called SylloBase that differs…

2023

Hi-ArG: Exploring the Integration of Hierarchical Argumentation Graphs in Language Pretraining

EMNLP 2023long main

The knowledge graph is a structure to store and represent knowledge, and recent studies have discussed its capability to assist language models for various applications. Some variations of knowledge graphs aim to record arguments and their relations for computational argumentation tasks. However, ma…

Cited by 0SourcecodeScholar
2023

IAG: Induction-Augmented Generation Framework for Answering Reasoning Questions

EMNLP 2023long main

Retrieval-Augmented Generation (RAG), by incorporating external knowledge with parametric memory of language models, has become the state-of-the-art architecture for open-domain QA tasks. However, common knowledge bases are inherently constrained by limited coverage and noisy information, making ret…

Cited by 0SourceScholar
2023

Reinforced Potential Field for Multi-Robot Motion Planning in Cluttered Environments

IROS 2023poster

Motion planning is challenging for multiple robots in cluttered environments without communication, especially in view of real-time efficiency, motion safety, distributed computation, and trajectory optimality, etc. In this paper, a reinforced potential field method is developed for distributed mult…

Cited by 6SourceScholar
2023

Sequential Texts Driven Cohesive Motions Synthesis with Natural Transitions

ICCV 2023poster

The intelligent synthesis/generation of daily-life motion sequences is fundamental and urgently needed for many VR/metaverse-related applications. However, existing approaches commonly focus on monotonic motion generation (e.g., walking, jumping, etc.) based on single instruction-like text, which is…

Cited by 15PDFcodeScholar
2023

Unified Pre-Training with Pseudo Texts for Text-To-Image Person Re-Identification

ICCV 2023poster

The pre-training task is indispensable for the text-to-image person re-identification (T2I-ReID) task. However, there are two underlying inconsistencies between these two tasks that may impact the performance: i) Data inconsistency. A large domain gap exists between the generic images/texts used in…

Cited by 45PDFcodeScholar
2022

AfriCLIRMatrix: Enabling Cross-Lingual Information Retrieval for African Languages

EMNLP 2022main

Language diversity in NLP is critical in enabling the development of tools for a wide range of users.However, there are limited resources for building such tools for many languages, particularly those spoken in Africa.For search, most existing datasets feature few or no African languages, directly i…

2022

Certified Error Control of Candidate Set Pruning for Two-Stage Relevance Ranking

EMNLP 2022main

In information retrieval (IR), candidate set pruning has been commonly used to speed up two-stage relevance ranking. However, such an approach lacks accurate error control and often trades accuracy against computational efficiency in an empirical fashion, missing theoretical guarantees. In this pape…

2022

Coarse-to-Fine: Hierarchical Multi-task Learning for Natural Language Understanding

COLING 2022main

Generalized text representations are the foundation of many natural language understanding tasks. To fully utilize the different corpus, it is inevitable that models need to understand the relevance among them. However, many methods ignore the relevance and adopt a single-channel model (a coarse par…

Cited by 4SourcePDFScholar
2022

Hyperlink-induced Pre-training for Passage Retrieval in Open-domain Question Answering

ACL 2022long

To alleviate the data scarcity problem in training question answering systems, recent works propose additional intermediate pre-training for dense passage retrieval (DPR). However, there still remains a large discrepancy between the provided upstream signals and the downstream question-passage relev…

2022

IPS300+: a Challenging multi-modal data sets for Intersection Perception System

ICRA 2022poster

Due to high complexity and occlusion, insufficient perception in the crowded urban intersection can be a serious safety risk for both human drivers and autonomous algorithms, whereas CVIS (Cooperative Vehicle Infrastructure System) is a proposed solution for full-participants perception under this s…

Cited by 36SourceScholar
2022

Implicit Sample Extension for Unsupervised Person Re-Identification

CVPR 2022poster

Most existing unsupervised person re-identification (Re-ID) methods use clustering to generate pseudo labels for model training. Unfortunately, clustering sometimes mixes different true identities together or splits the same identity into two or more sub clusters. Training on these noisy clusters su…

Cited by 135PDFcodeScholar
2022

InterFusion: Interaction-based 4D Radar and LiDAR Fusion for 3D Object Detection

IROS 2022poster

Many recent works detect 3D objects by several sensor modalities for autonomous driving, where high-resolution cameras and high-line LiDARs are mostly used but relatively expensive. To achieve a balance between overall cost and detection accuracy, many multi-modal fusion techniques have been suggest…

Cited by 26SourceScholar
2022

Replacing Labeled Real-Image Datasets With Auto-Generated Contours

CVPR 2022poster

In the present work, we show that the performance of formula-driven supervised learning (FDSL) can match or even exceed that of ImageNet-21k without the use of real images, human-, and self-supervision during the pre-training of Vision Transformers (ViTs). For example, ViT-Base pre-trained on ImageN…

Cited by 44PDFScholar
2022

Self-Guided Hard Negative Generation for Unsupervised Person Re-Identification

IJCAI 2022poster

Recent unsupervised person re-identification (reID) methods mostly apply pseudo labels from clustering algorithms as supervision signals. Despite great success, this fashion is very likely to aggregate different identities with similar appearances into the same cluster. In result, the hard negative…

Cited by 12SourcePDFScholar
2022

Towards Efficient NLP: A Standard Evaluation and A Strong Baseline

NAACL 2022long

Supersized pre-trained language models have pushed the accuracy of various natural language processing (NLP) tasks to a new state-of-the-art (SOTA). Rather than pursuing the reachless SOTA accuracy, more and more researchers start paying attention to model efficiency and usability. Different from ac…

2021

Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box Estimation

CVPR 2021poster

Visual object tracking aims to precisely estimate the bounding box for the given target, which is a challenging problem due to factors such as deformation and occlusion. Many recent trackers adopt the multiple-stage tracking strategy to improve the quality of bounding box estimation. These methods f…

Cited by 268PDFcodeScholar
2021

Diverse Knowledge Distillation for End-to-End Person Search

AAAI 2021technical

Person search aims to localize and identify a specific person from a gallery of images. Recent methods can be categorized into two groups, i.e., two-step and end-to-end approaches. The former views person search as two independent tasks and achieves dominant results using separately trained person d…

Cited by 47SourcePDFScholar
2021

Line-based Automatic Extrinsic Calibration of LiDAR and Camera

ICRA 2021poster

Reliable real-time extrinsic parameters of 3D Light Detection and Ranging (LiDAR) and camera are a key component of multi-modal perception systems. However, extrinsic transformation may drift gradually during operation, which can result in decreased accuracy of perception system. To solve this probl…

Cited by 54SourceScholar
2021

Robust Multimodal Vehicle Detection in Foggy Weather Using Complementary Lidar and Radar Signals

CVPR 2021poster

Vehicle detection with visual sensors like lidar and camera is one of the critical functions enabling autonomous driving. While they generate fine-grained point clouds or high-resolution images with rich information in good weather conditions, they fail in adverse weather (e.g., fog) where opaque pa…

Cited by 193PDFcodeScholar
2021

Window Loss for Bone Fracture Detection and Localization in X-ray Images with Point-based Annotation

AAAI 2021technical

Object detection methods are widely adopted for computer-aided diagnosis using medical images. Anomalous findings are usually treated as objects that are described by bounding boxes. Yet, many pathological findings, e.g., bone fractures, cannot be clearly defined by bounding boxes, owing to consider…

Cited by 2SourcePDFScholar
2020

Trajectory Similarity Learning with Auxiliary Supervision and Optimal Matching

IJCAI 2020poster

Trajectory similarity computation is a core problem in the field of trajectory data queries. However, the high time complexity of calculating the trajectory similarity has always been a bottleneck in real-world applications. Learning-based methods can map trajectories into a uniform embedding space…

2019

Self-Training With Progressive Augmentation for Unsupervised Cross-Domain Person Re-Identification

ICCV 2019poster

Person re-identification (Re-ID) has achieved great improvement with deep learning and a large amount of labelled training data. However, it remains a challenging task for adapting a model trained in a source domain of labelled data to a target domain of only unlabelled data available. In this work,…

Cited by 305PDFScholar
2019

Sell-corpus: an Open Source Multiple Accented Chinese-english Speech Corpus for L2 English Learning Assessment

ICASSP 2019accepted

We present SELL-CORPUS, a multiple accented speech corpus for L2 English learning in China, aiming at the potential research of multiple accented acoustic model, mispronunciation detection and pronunciation assessment for future nationwide oral English tests. Our corpus contains 31.6 hour speech rec…

Cited by 0SourceScholar
2019

Transferring Grasp Configurations using Active Learning and Local Replanning

ICRA 2019poster

We present a new approach to transfer grasp configurations from prior example objects to novel objects. We assume the novel and example objects have the same topology and similar shapes. We perform 3D segmentation on these objects using geometric and semantic shape characteristics. We compute a gras…

Cited by 22SourceScholar