← Search

ke Wang

126 accepted papers

2026

Activations as Features: Probing LLMs for Generalizable Essay Scoring Representations

AAAI 2026technical

Automated essay scoring (AES) is a challenging task in cross-prompt settings due to the diversity of scoring criteria. While previous studies have focused on the output of large language models (LLMs) to improve scoring accuracy, we believe activations from intermediate layers may also provide valua

Cited by 0SourcePDFScholar
2026

Agentic Reinforcement Learning with Implicit Step Rewards

ICLR 2026poster

Large language models (LLMs) are increasingly developed as autonomous agents using reinforcement learning (agentic RL) that reason and act in interactive environments. However, sparse and sometimes unverifiable rewards make it extremely challenging to assign credit when training LLM agents that serv…

Cited by 0SourceScholar
2026

BEV-CAR: Enhancing Monocular Bird's Eye View Segmentation with Context-Aware Rasterization

CVPR 2026

Bird's Eye View (BEV) semantic segmentation is essential for autonomous driving and mobile robotics, yet it still faces significant challenges on accurate segmentation of foreground object and efficient estimating of layout categories obscured by objects. To address these issues, we propose BEV-CAR,

Cited by 0SourcecodeScholar
2026

Beyond Client Clustering: Fine-Grained Preference Alignment in Federated RLHF via Self-Evolving Routing

IJCAI 2026

Federated Reinforcement Learning from Human Feedback (RLHF) enables the collaborative alignment of Large Language Models (LLMs) while preserving privacy, yet it faces critical bottlenecks arising from data heterogeneity. Existing approaches typically rely on rigid client-level clustering, which over

Cited by 0Scholar
2026

Cache Coherent Resampling for Efficient Test Time Scaling in LLM Reasoning via Adaptive Sequential Monte Carlo

ICML 2026poster

Recent work shows that chain based sampling for power shaped trajectory distributions can deliver large test time gains from a fixed base LLM and can approach RL trained reasoners such as GRPO. Deployment is the bottleneck. Autoregressive Metropolis Hastings is inherently serial, limits GPU utilizat…

Cited by 0SourceScholar
2026

CausalPlanner: A Causality-Enhanced Planning Framework for Generalizable Autonomous Driving

RA-L 2026

Imitation learning (IL) has been widely adopted for autonomous driving planning because of its data efficiency and stable optimization. Yet IL-based planners often suffer from causal confusion, fitting spurious correlations instead of genuine causal mechanisms, which leads to unreliable planning beh

Cited by 0SourceScholar
2026

Edit-Based Refinement for Parallel Masked Diffusion Language Models

ICML 2026poster

Masked diffusion language models enable parallel token generation and offer improved decoding efficiency over autoregressive models. However, their performance degrades significantly when generating multiple tokens simultaneously, due to a mismatch between token-level training objectives and the nee…

Cited by 0SourceScholar
2026

EmWorld: Emotion World Model with Latent State Evolution for Scenario-Incremental Dynamic Facial Expression Recognition

ICML 2026poster

Dynamic Facial Expression Recognition (DFER) models the temporal evolution of facial expressions in videos. In real-world deployments, changing scenarios distort expression trajectories over time, making it difficult for existing methods to maintain performance. While most current approaches address…

Cited by 0SourceScholar
2026

From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench

AAAI 2026technical

Large Language Models (LLMs) show significant potential in AI mathematical tutoring, yet current evaluations often rely on simplistic metrics or narrow pedagogical scenarios, failing to assess comprehensive, multi-turn teaching effectiveness. In this paper, we introduce KMP-Bench, a comprehensive K-

Cited by 0SourcePDFScholar
2026

FullStack-Agent: Enhancing Agentic Full-Stack Web Coding via Development-Oriented Testing and Repository Back-Translation

ICML 2026poster

Assisting non-expert users to develop complex interactive websites has become a popular task for LLM-powered code agents. However, existing code agents tend to only generate frontend web pages, masking the lack of real full-stack data processing and storage with fancy visual effects. Notably, constr…

Cited by 0SourceScholar
2026

HDGS: Hierarchical Dynamic Gaussian Splatting for Urban Driving Scenes

AAAI 2026technical

This paper tackles the challenging task of achieving storage-efficient yet high-fidelity motion representation in large-scale dynamic 3D Gaussian Splatting. Our motivation stems from the truth that existing urban-scale methods, which rely on massive and unstructured individual Gaussians for scene mo

Cited by 0SourcePDFScholar
2026

Hist2Style: Histogram-Guided Stylization with Bilateral Grids

CVPR 2026

Photorealistic style transfer aims to match the color and tone of an input image to that of a style target while preserving the content and details of the original scene. Although existing large image models can facilitate these kinds of appearance edits, their high computational demands, potential

Cited by 0SourcecodeScholar
2026

MetaDAT: Generalizable Trajectory Prediction Via Meta Pre-Training and Data-Adaptive Test-Time Updating

ICRA 2026poster

Existing trajectory prediction methods exhibit significant performance degradation under distribution shifts during test time. Although test-time training techniques have been explored to enable adaptation, current approaches rely on an offline pre-trained predictor that lacks online learning flexib…

2026

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

ICLR 2026poster

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modaliti…

Cited by 0SourcecodeScholar
2026

SG-Reg: Generalizable and Efficient Scene Graph Registration

ICRA 2026poster

This paper addresses the challenge of registering two rigid semantic scene graphs, an essential capability for autonomous agents to align with remote agents or prior maps. Traditional methods rely on hand-crafted descriptors or ground-truth annotations, limiting their applicability in real-world sce…

2026

Scalable Bayesian Semi-supervised Clustering with Feature Selection and Adaptive Constraint Weighting

ICML 2026poster

Constrained clustering incorporates prior knowledge in the form of pairwise constraints to guide data partitioning. While effective, existing Bayesian approaches are often limited in scalability to large datasets and provide weak interpretability due to the lack of explicit feature relevance modelin…

Cited by 0SourceScholar
2026

Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling

AAAI 2026technical

Multimedia documents such as slide presentations and posters are designed to be interactive and easy to modify. Yet, they are often distributed in a static raster format, which limits editing and customization. Restoring their editability requires converting these raster images back into structured

Cited by 0SourcePDFScholar
2026

Sparse Annotation, Dense Supervision: Unleashing Self-Training Power for Occupancy Prediction With 2D Labels

RA-L 2026

Serving as a fundamental task in robotic navigation and autonomous driving, occupancy prediction is gaining increasing attention for its fine-grained perception of the 3D environment. Most existing methods rely on dense 3D annotations, which are expensive, labor-intensive, and difficult to scale in

Cited by 1SourceScholar
2026

Structured Labeling Enables Faster Vision-Language Models for End-To-End Autonomous Driving

ICRA 2026poster

Vision-Language Models (VLMs) offer a promising approach to end-to-end autonomous driving due to their human-like reasoning capabilities. However, troublesome gaps remains between current VLMs and real-world autonomous driving applications. One major limitation is that existing datasets with loosely…

2026

Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning

ICML 2026poster

Traditional whole slide image (WSI) analysis methods typically rely on the multiple instance learning (MIL) paradigm, which extracts patch-level features at high magnification and aggregates them for slide-level prediction. However, such exhaustive patch-level processing is computationally expensive…

Cited by 0SourceScholar
2026

WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning

ICLR 2026poster

Agent systems powered by large language models (LLMs) have demonstrated impressive performance on repository-level code-generation tasks. However, for tasks such as website codebase generation, which depend heavily on visual effects and user-interaction feedback, current code agents rely only on sim…

Cited by 0SourcecodeScholar
2026

World-Model Inspired Emotion-aware Token Refinement for Training-Free Multimodal Emotion Recognition

ICML 2026spotlight

Multimodal Large Language Models (MLLMs) show promise for Multimodal Emotion Recognition (MER) but often remain unreliable because sparse emotional cues could be easily overwhelmed and affected by redundant context. While fine-tuning is effective, it is usually costly when using large models. Traini…

Cited by 0SourceScholar
2025

A Variable Sensing Range Electrical Impedance Tomography Sensor for Robot Electric Skins

RA-L 2025

While stretchable and compressible sensors are commonly used to enhance the proprioception and exteroception of soft robots, their sensitivity and sensing range are constrained by their design and stiffness, often limiting their measurement capabilities. To overcome this limitation, this paper prese

Cited by 1SourceScholar
2025

Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding

CVPR 2025poster

Visual Document Understanding has become essential with the increase of text-rich visual content. This field poses significant challenges due to the need for effective integration of visual perception and textual comprehension, particularly across diverse document types with complex layouts. Moreove…

2025

Alignment with Fill-In-the-Middle for Enhancing Code Generation

EMNLP 2025

The code generation capabilities of Large Language Models (LLMs) have advanced applications like tool invocation and problem-solving. However, improving performance in code-related tasks remains challenging due to limited training data that is verifiable with accurate test cases. While Direct Prefer

2025

Beyond Spatial Domain: Cross-domain Promoted Fourier Convolution Helps Single Image Dehazing

AAAI 2025technical

Vanilla convolution and window-based self-attention have shown significant success in image dehazing. However, they are constrained by limited receptive fields and ignore frequency gaps between dehazed and clear images. The former hampers the modeling of global dependencies, while the latter impedes…

Cited by 0SourcePDFScholar
2025

COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

NeurIPS 2025poster

World models are critical for autonomous driving to simulate environmental dynamics and generate synthetic data. Existing methods struggle to disentangle ego-vehicle motion (perspective shifts) from scene evolvement (agent interactions), leading to suboptimal predictions. Instead, we propose to sepa…

Cited by 0SourcecodeScholar
2025

EFFOcc: Learning Efficient Occupancy Networks from Minimal Labels for Autonomous Driving

IROS 2025

3D occupancy prediction (3DOcc) is a rapidly rising and challenging perception task in the field of autonomous driving. Existing 3D occupancy networks (OccNets) are both computationally heavy and label-hungry. In terms of model complexity, OccNets are commonly composed of heavy Conv3D modules or tra

Cited by 7SourcecodeScholar
2025

EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning

ACL 2025long

Large Language Models (LLMs) have shown impressive reasoning capabilities in well-defined problems with clear solutions, such as mathematics and coding. However, they still struggle with complex real-world scenarios like business negotiations, which require strategic reasoning—an ability to navigate…

2025

Each Complexity Deserves a Pruning Policy

NeurIPS 2025poster

The established redundancy in visual tokens within large vision–language models (LVLMs) allows for pruning to effectively reduce their substantial computational demands. Empirical evidence from previous works indicates that visual tokens in later decoder stages receive less attention than shallow la…

Cited by 0SourcecodeScholar
2025

EquiBench: Benchmarking Large Language Models’ Reasoning about Program Semantics via Equivalence Checking

EMNLP 2025

As large language models (LLMs) become integral to code-related tasks, a central question emerges: Do LLMs truly understand program semantics? We introduce EquiBench, a new benchmark for evaluating LLMs through equivalence checking, i.e., determining whether two programs produce identical outputs fo

2025

Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces

ICML 2025poster

Modern scientific discovery increasingly relies on high-performance computing for complex modeling and simulation. A key challenge in improving parallel program performance is efficiently mapping tasks to processors and data to memory, a process dictated by intricate, low-level system code known as…

Cited by 0SourcePDFScholar
2025

LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding

EMNLP 2025

Recent progress in Large Language Models (LLMs) has opened new avenues for solving complex optimization problems, including Neural Architecture Search (NAS). However, existing LLM-driven NAS approaches rely heavily on prompt engineering and domain-specific tuning, limiting their practicality and sca

2025

LiNeS: Post-training Layer Scaling Prevents Forgetting and Enhances Model Merging

ICLR 2025poster

Fine-tuning pre-trained models has become the standard approach to endow them with specialized knowledge, but it poses fundamental challenges. In particular, (i) fine-tuning often leads to catastrophic forgetting, where improvements on a target domain degrade generalization on other tasks, and (ii)…

2025

MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs

NeurIPS 2025poster

Language models deployed in real-world systems often require post-hoc updates to incorporate new or corrected knowledge. However, editing such models efficiently and reliably—without retraining or forgetting previous information—remains a major challenge. Existing methods for lifelong model editing…

Cited by 0SourceScholar
2025

MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning

ACL 2025finding

Natural language image-caption datasets, widely used for training Large Multimodal Models, mainly focus on natural scenarios and overlook the intricate details of mathematical figures that are critical for problem-solving, hindering the advancement of current LMMs in multimodal mathematical reasonin…

2025

MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Code

ICLR 2025spotlight

Code has been shown to be effective in enhancing the mathematical reasoning abilities of large language models due to its precision and accuracy. Previous works involving continued mathematical pretraining often include code that utilizes math-related packages, which are primarily designed for fiel…

2025

NegRefine: Refining Negative Label-Based Zero-Shot OOD Detection

ICCV 2025poster

Recent advancements in Vision-Language Models like CLIP have enabled zero-shot OOD detection by leveraging both image and textual label information. Among these, negative label-based methods such as NegLabel and CSP have shown promising results by utilizing a lexicon of words to define negative labe…

2025

Nonlinear Anisotropic Diffusion-Based Channel Estimation in 5G Wireless Networks

ICASSP 2025accepted

In the context of the fifth-generation new radio downlink scenario, we introduce an innovative approach for channel estimation in this paper that circumvents the requirement for the prior dataset. We incorporate anisotropic diffusion and bit-plane decomposition to remove the noise in channel estimat…

Cited by 0SourceScholar
2025

Online Segment Any 3D Thing as Instance Tracking

NeurIPS 2025poster

Online, real-time, and fine-grained 3D segmentation constitutes a fundamental capability for embodied intelligent agents to perceive and comprehend their operational environments. Recent advancements employ predefined object queries to aggregate semantic information from Vision Foundation Models (VF…

Cited by 0SourcecodeScholar
2025

Probability-Consistent Preference Optimization for Enhanced LLM Reasoning

ACL 2025finding

Recent advances in preference optimization have demonstrated significant potential for improving mathematical reasoning capabilities in large language models (LLMs). While current approaches leverage high-quality pairwise preference data through outcome-based criteria like answer correctness or cons…

2025

RACQC: Advanced Retrieval-Augmented Generation for Chinese Query Correction

EMNLP 2025

In web search scenarios, erroneous queries frequently degrade users’ experience through irrelevant results, underscoring the pivotal role of Chinese Spelling Check (CSC) systems. Although large language models (LLMs) exhibit remarkable capabilities across many tasks, they face critical challenges in

2025

SAM Decoding: Speculative Decoding via Suffix Automaton

ACL 2025long

Speculative decoding (SD) has been demonstrated as an effective technique for lossless LLM inference acceleration.Retrieval-based SD methods, one kind of model-free method, have yielded promising speedup, but they often rely on single retrieval resources, inefficient retrieval methods, and are const…

2025

SATBench: Benchmarking LLMs’ Logical Reasoning via Automated Puzzle Generation from SAT Formulas

EMNLP 2025

We introduce SATBench, a benchmark for evaluating the logical reasoning capabilities of large language models (LLMs) through logical puzzles derived from Boolean satisfiability (SAT) problems.Unlike prior work that focuses on inference rule-based reasoning, which often involves deducing conclusions

2025

SDPO: Segment-Level Direct Preference Optimization for Social Agents

ACL 2025long

Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex social dialogues. Direct Preference Optimization (DPO) has proven effective in aligning LLM behavior with human preferences across various agent tasks. However, standard DPO f…

2025

Scaling and Taming Adversarial Training with Synthetic Data

ICCV 2025poster

Despite the success of adversarial training on small datasets, applying it to large-scale datasets like ImageNet remains challenging. Previous attempts using synthetic data show limited improvements. This work investigates the impact of synthetic data scaling, model scaling, and training strategies…

Cited by 0SourcePDFScholar
2025

SplatPose: Geometry-Aware 6-DoF Pose Estimation from Single RGB Image via 3D Gaussian Splatting

IROS 2025

6-DoF pose estimation is a fundamental task in computer vision with wide-ranging applications in augmented reality and robotics. Existing single RGB-based methods often compromise accuracy due to their reliance on initial pose estimates and susceptibility to rotational ambiguity, while approaches re

Cited by 6SourceScholar
2025

The Devil is in the Quality: Exploring Informative Samples for Semi-Supervised Monocular 3D Object Detection

ICRA 2025

This paper tackles the challenging problem of semi-supervised monocular 3D object detection with a general framework. In specific, having observed that the bottleneck of this task lies in lacking reliable and informative samples from unlabeled data for detector learning, we introduce a novel simple

Cited by 0SourceScholar
2025

VLIN-RL: A Unified Vision-Language Interpreter and Reinforcement Learning Motion Planner Framework for Robot Dynamic Tasks

IROS 2025

Recently, with the development of Large Language Models (LLMs), Embodied AI represented by Vision-Language-Action Models (VLAs) has played a significant role in realizing the natural language interaction between humans and robots. Current VLA models can process and understand visual information and

Cited by 0SourcecodeScholar
2025

Vision Transformers Beat WideResNets on Small Scale Datasets Adversarial Robustness

AAAI 2025technical

For an extensive period, Vision Transformers (ViTs) have been deemed unsuitable for attaining robust performance on small-scale datasets, with WideResNet models maintaining dominance in this domain. While WideResNet models have persistently set the state-of-the-art (SOTA) benchmarks for robust accur…

Cited by 0SourcePDFScholar
2025

WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch

NeurIPS 2025oral

LLM‑based agents have demonstrated great potential in generating and managing code within complex codebases. In this paper, we introduce WebGen-Bench, a novel benchmark designed to measure an LLM-based agent's ability to create multi-file website codebases from scratch. It contains diverse instructi…

Cited by 0SourcecodeScholar
2024

AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories

EMNLP 2024finding

Fine-tuning on agent-environment interaction trajectory data holds significant promise for surfacing generalized agent capabilities in open-source large language models (LLMs). In this work, we introduce AgentBank, by far the largest trajectory tuning data collection featuring more than 50k diverse…

2024

Aligning Logits Generatively for Principled Black-Box Knowledge Distillation

CVPR 2024poster

Black-Box Knowledge Distillation (B2KD) is a formulated problem for cloud-to-edge model compression with invisible data and models hosted on the server. B2KD faces challenges such as limited Internet exchange and edge-cloud disparity of data distributions. In this paper we formalize a two-step workf…

2024

An Electoral Approach to Diversify LLM-based Multi-Agent Collective Decision-Making

EMNLP 2024main

Modern large language models (LLMs) have exhibited cooperative synergy on complex task-solving, and collective decision-making (CDM) is a pivotal component in LLM-based multi-agent collaboration frameworks. Our survey on 52 recent such systems uncovers a severe lack of diversity, with a heavy relian…

2024

Enhancing the General Agent Capabilities of Low-Paramter LLMs through Tuning and Multi-Branch Reasoning

NAACL 2024findings

Open-source pre-trained Large Language Models (LLMs) exhibit strong language understanding and generation capabilities, making them highly successful in a variety of tasks. However, when used as agents for dealing with complex problems in the real world, their performance is far inferior to large co…

2024

FM-Fusion: Instance-Aware Semantic Mapping Boosted by Vision-Language Foundation Models

RA-L 2024

Semantic mapping based on the supervised object detectors is sensitive to image distribution. In real-world environments, the object detection and segmentation performance can lead to a major drop, preventing the use of semantic mapping in a wider domain. On the other hand, the development of vision

Cited by 11SourcecodeScholar
2024

FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents

EMNLP 2024finding

LLM-based agents have emerged as promising tools, which are crafted to fulfill complex tasks by iterative planning and action. However, these agents are susceptible to undesired planning hallucinations when lacking specific knowledge for expertise-intensive tasks. To address this, preliminary attemp…

2024

Localizing Task Information for Improved Model Merging and Compression

ICML 2024poster

Model merging and task arithmetic have emerged as promising scalable approaches to merge multiple single-task checkpoints to one multi-task model, but their applicability is reduced by significant performance loss. Previous works have linked these drops to interference in the weight space and erasur…

2024

MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

ICLR 2024poster

The recently released GPT-4 Code Interpreter has demonstrated remarkable proficiency in solving challenging math problems, primarily attributed to its ability to seamlessly reason with natural language, generate code, execute code, and continue reasoning based on the execution output. In this paper,…

2024

MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs

ACL 2024long

Large language models (LLMs) have exhibited great potential in mathematical reasoning. However, there remains a performance gap in this area between existing open-source models and closed-source models such as GPT-4. In this paper, we introduce MathGenie, a novel method for generating diverse and re…

2024

Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

NeurIPS 2024poster

Recent advancements in Large Multimodal Models (LMMs) have shown promising results in mathematical reasoning within visual contexts, with models exceeding human-level performance on existing benchmarks such as MathVista. However, we observe significant limitations in the diversity of questions and b…

Cited by 113SourcePDFScholar
2024

Multi-View Depth Estimation by Using Adaptive Point Graph to Fuse Single-View Depth Probabilities

RA-L 2024

Recently, some methods estimate depth maps by fusing several adjacent single-view depth probabilities. They have achieved promising performance in multi-view inconsistent areas, such as texture-less surfaces, reflective surfaces, and moving objects. However, these methods involve two new problems: t

Cited by 3SourceScholar
2024

OPUS: Occupancy Prediction Using a Sparse Set

NeurIPS 2024poster

Occupancy prediction, aiming at predicting the occupancy status within voxelized 3D environment, is quickly gaining momentum within the autonomous driving community. Mainstream occupancy prediction works first discretize the 3D environment into voxels, then perform classification on such dense grids…

2024

Pi-DUAL: Using privileged information to distinguish clean from noisy labels

ICML 2024poster

Label noise is a pervasive problem in deep learning that often compromises the generalization performance of trained models. Recently, leveraging privileged information (PI) -- information available only during training but not at test time -- has emerged as an effective approach to mitigate this is…

Cited by 2SourcePDFScholar
2024

Review-Enhanced Hierarchical Contrastive Learning for Recommendation

AAAI 2024technical

Designed to establish potential relations and distill high-order representations, graph-based recommendation systems continue to reveal promising results by jointly modeling ratings and reviews. However, existing studies capture simple review relations, failing to (1) completely explore hidden conne…

Cited by 9SourcePDFScholar
2024

SACNet: A Scattered Attention-Based Network With Feature Compensator for Visual Localization

RA-L 2024

Visual localization, an integral component of a vast array of computer applications, has been effectively resolved by scene coordinate regression (SCoRe) methods. However, due to the limited receptive field of convolutional neural networks (CNNs), current SCoRe methods have difficulty in distinguish

Cited by 4SourceScholar
2024

Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification

ICLR 2024poster

Recent progress in large language models (LLMs) like GPT-4 and PaLM-2 has brought significant advancements in addressing math reasoning problems. In particular, OpenAI's latest version of GPT-4, known as GPT-4 Code Interpreter, shows remarkable performance on challenging math datasets. In this paper…

Cited by 153SourcePDFScholar
2024

StreamingFlow: Streaming Occupancy Forecasting with Asynchronous Multi-modal Data Streams via Neural Ordinary Differential Equation

CVPR 2024highlight

Predicting the future occupancy states of the surrounding environment is a vital task for autonomous driving. However current best-performing single-modality methods or multi-modality fusion perception methods are only able to predict uniform snapshots of future occupancy states and require strictly…

2024

Towards Stable 3D Object Detection

ECCV 2024poster

"In autonomous driving, the temporal stability of 3D object detection greatly impacts the driving safety. However, the detection stability cannot be accessed by existing metrics such as mAP and MOTA, and consequently is less explored by the community. To bridge this gap, this work proposes (), a new…

2024

Watch Every Step! LLM Agent Learning via Iterative Step-level Process Refinement

EMNLP 2024main

Large language model agents have exhibited exceptional performance across a range of complex interactive tasks. Recent approaches have utilized tuning with expert trajectories to enhance agent performance, yet they primarily concentrate on outcome rewards, which may lead to errors or suboptimal acti…

2023

CO-Net: Learning Multiple Point Cloud Tasks at Once with A Cohesive Network

ICCV 2023poster

We present CO-Net, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains. CO-Net maintains the characteristics of high storage efficiency since models with the preponderance of shared parameters can be assembled into a single model. Specific…

Cited by 7PDFScholar
2023

Curricular Object Manipulation in LiDAR-Based Object Detection

CVPR 2023poster

This paper explores the potential of curriculum learning in LiDAR-based 3D object detection by proposing a curricular object manipulation (COM) framework. The framework embeds the curricular training strategy into both the loss design and the augmentation process. For the loss design, we propose the…

2023

Disambiguated Lexically Constrained Neural Machine Translation

ACL 2023findings

Lexically constrained neural machine translation (LCNMT), which controls the translation generation with pre-specified constraints, is important in many practical applications. Current approaches to LCNMT typically assume that the pre-specified lexicon constraints are contextually appropriate. This…

Cited by 4SourcePDFScholar
2023

Disentangled Representation for Causal Mediation Analysis

AAAI 2023technical

Estimating direct and indirect causal effects from observational data is crucial to understanding the causal mechanisms and predicting the behaviour under different interventions. Causal mediation analysis is a method that is often used to reveal direct and indirect effects. Deep learning shows prom…

2023

EAAINet: An Element-Wise Attention Network With Global Affinity Information for Accurate Indoor Visual Localization

RA-L 2023

Visual localization, a vital component of many visual applications, has been tackled by scene coordinates regression (SCoRe) methods that leverage neural networks to predict scene coordinates, followed by a PnP algorithm to recover camera pose. However, these methods do not consider the relationship

Cited by 19SourceScholar
2023

Easy Guided Decoding in Providing Suggestions for Interactive Machine Translation

ACL 2023long

Machine translation technology has made great progress in recent years, but it cannot guarantee error-free results. Human translators perform post-editing on machine translations to correct errors in the scene of computer aided translation. In favor of expediting the post-editing process, many works…

2023

Improving Neural Machine Translation by Multi-Knowledge Integration with Prompting

EMNLP 2023long findings

Improving neural machine translation (NMT) systems with prompting has achieved significant progress in recent years. In this work, we focus on how to integrate multi-knowledge, multiple types of knowledge, into NMT models to enhance the performance with prompting. We propose a unified framework, whi…

Cited by 0SourceScholar
2023

M$^3$Seg: A Maximum-Minimum Mutual Information Paradigm for Unsupervised Topic Segmentation in ASR Transcripts

EMNLP 2023short main

Topic segmentation aims to detect topic boundaries and split automatic speech recognition transcriptions (e.g., meeting transcripts) into segments that are bounded by thematic meanings. In this work, we propose M$^3$Seg, a novel Maximum-Minimum Mutual information paradigm for linear topic segmentat…

Cited by 0SourceScholar
2023

OFVL-MS: Once for Visual Localization across Multiple Indoor Scenes

ICCV 2023poster

In this work, we seek to predict camera poses across scenes with a multi-task learning manner, where we view the localization of each scene as a new task. We propose OFVL-MS, a unified framework that dispenses with the traditional practice of training a model for each individual scene and relieves…

Cited by 11PDFcodeScholar
2023

ORCHID: A Chinese Debate Corpus for Target-Independent Stance Detection and Argumentative Dialogue Summarization

EMNLP 2023long main

Dialogue agents have been receiving increasing attention for years, and this trend has been further boosted by the recent progress of large language models (LLMs). Stance detection and dialogue summarization are two core tasks of dialogue agents in application scenarios that involve argumentative di…

Cited by 0SourcecodeScholar
2023

On Human Grasping and Manipulation in Kitchens: Automated Annotation, Insights, and Metrics for Effective Data Collection

ICRA 2023poster

The advancement in robotic grasping and manipulation has elicited an increased research interest in the development of household robots capable of performing a plethora of complex tasks. These advancements require the shift of robotics research from a laboratory setting to dynamic and unstructured h…

Cited by 6SourceScholar
2023

PROSE: A Pronoun Omission Solution for Chinese-English Spoken Language Translation

EMNLP 2023long main

Neural Machine Translation (NMT) systems encounter a significant challenge when translating a pro-drop ('pronoun-dropping') language (e.g., Chinese) to a non-pro-drop one (e.g., English), since the pro-drop phenomenon demands NMT systems to recover omitted pronouns. This unique and crucial task, how…

Cited by 0SourceScholar
2023

Poly-MOT: A Polyhedral Framework For 3D Multi-Object Tracking

IROS 2023poster

3D Multi-object tracking (MOT) empowers mobile robots to accomplish well-informed motion planning and navigation tasks by providing motion trajectories of surrounding objects. However, existing 3D MOT methods typically employ a single similarity metric and physical model to perform data association…

Cited by 38SourcecodeScholar
2023

Poly-PC: A Polyhedral Network for Multiple Point Cloud Tasks at Once

CVPR 2023poster

In this work, we show that it is feasible to perform multiple tasks concurrently on point cloud with a straightforward yet effective multi-task network. Our framework, Poly-PC, tackles the inherent obstacles (e.g., different model architectures caused by task bias and conflicting gradients caused by…

Cited by 21SourcePDFScholar
2023

ResoNet: Noise-Trained Physics-Informed MRI Off-Resonance Correction

NeurIPS 2023poster

Magnetic Resonance Imaging (MRI) is a powerful medical imaging modality that offers diagnostic information without harmful ionizing radiation. Unlike optical imaging, MRI sequentially samples the spatial Fourier domain (k-space) of the image. Measurements are collected in multiple shots, or readout…

2023

Scalable. Intuitive Human to Robot Skill Transfer with Wearable Human Machine Interfaces: On Complex, Dexterous Tasks

IROS 2023poster

The advent of collaborative industrial and house-hold robotics has blurred the demarcation between the human and robot workspace. The capability of robots to function efficiently alongside humans requires new research to be conducted in dynamic environments as opposed to the traditional well-structu…

Cited by 6SourceScholar
2023

Semi-Supervised Parametric Real-World Image Harmonization

CVPR 2023poster

Learning-based image harmonization techniques are usually trained to undo synthetic global transformations, applied to a masked foreground in a single ground truth photo. This simulated data does not model many important appearance mismatches (illumination, object boundaries, etc.) between foregroun…

2022

A Deep Feature Aggregation Network for Accurate Indoor Camera Localization

RA-L 2022

As scene coordinate regression (SCoRe) methods become prevailing in the area of visual camera localization, the issue of repetitive or sparse texture scenes continues to be a concern. Specifically, they will suffer from performance degeneration due to ambiguous patterns caused by visual similarity.

Cited by 23SourceScholar
2022

Gated Mechanism Enhanced Multi-Task Learning for Dialog Routing

COLING 2022main

Currently, human-bot symbiosis dialog systems, e.g. pre- and after-sales in E-commerce, are ubiquitous, and the dialog routing component is essential to improve the overall efficiency, reduce human resource cost and increase user experience. To satisfy this requirement, existing methods are mostly h…

Cited by 0SourcePDFScholar
2022

Incorporating Item Frequency for Differentially Private Set Union

AAAI 2022technical

We study the problem of releasing the set union of users' items subject to differential privacy. Previous approaches consider only the set of items for each user as the input. We propose incorporating the item frequency, which is typically available in set union problems, to boost the utility of pri…

2022

Infergrad: Improving Diffusion Models for Vocoder by Considering Inference in Training

ICASSP 2022accepted

Denoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses the state-of-the-art generative models, which invariably results in slow inference speed. Previous approaches aim to opt…

Cited by 0SourceScholar
2022

On Wearable, Lightweight, Low-Cost Human Machine Interfaces for the Intuitive Collection of Robot Grasping and Manipulation Data

ICRA 2022poster

Robot grasping and manipulation allow robots to interact with their environments and execute a plethora of complex tasks that require increased dexterity (e.g., open a door, push buttons, collect and transpose objects, etc.). Collecting data of such activities is of paramount importance as it allows…

Cited by 3SourceScholar
2022

PANet: A Pixel-Level Attention Network for 6D Pose Estimation With Embedding Vector Features

RA-L 2022

In this work, we present PANet, a pixel-level attention network with embedding vector features, which addresses the challenge of 6D pose estimation from a single RGBD image under severe occlusion. PANet produces pixel-wise attention for strong representation learning and leverages a novel selection

Cited by 12SourceScholar
2022

Robust Learning against Relational Adversaries

NeurIPS 2022accept

Test-time adversarial attacks have posed serious challenges to the robustness of machine-learning models, and in many settings the adversarial perturbation need not be bounded by small $\ell_p$-norms. Motivated by attacks in program analysis and security tasks, we investigate $\textit{relational adv…

Cited by 8SourcePDFScholar
2022

XYLayoutLM: Towards Layout-Aware Multimodal Networks for Visually-Rich Document Understanding

CVPR 2022poster

Recently, various multimodal networks for Visually-Rich Document Understanding(VRDU) have been proposed, showing the promotion of transformers by integrating visual and layout information with the text embeddings. However, most existing approaches utilize the position embeddings to incorporate the s…

Cited by 105PDFScholar
2021

Benign Overfitting in Multiclass Classification: All Roads Lead to Interpolation

NeurIPS 2021poster

The growing literature on "benign overfitting" in overparameterized models has been mostly restricted to regression or binary classification settings; however, most success stories of modern machine learning have been recorded in multiclass settings. Motivated by this discrepancy, we study benign ov…

Cited by 64SourcePDFScholar
2021

Beyond Glass-Box Features: Uncertainty Quantification Enhanced Quality Estimation for Neural Machine Translation

EMNLP 2021finding

Quality Estimation (QE) plays an essential role in applications of Machine Translation (MT). Traditionally, a QE system accepts the original source text and translation from a black-box MT system as input. Recently, a few studies indicate that as a by-product of translation, QE benefits from the mod…

Cited by 5SourcePDFScholar
2021

Bridging the Domain Gap: Improve Informal Language Translation via Counterfactual Domain Adaptation

AAAI 2021technical

Despite the near-human performances already achieved on formal texts such as news articles, neural machine translation still has difficulty in dealing with "user-generated" texts that have diverse linguistic phenomena but lack large-scale high-quality parallel corpora. To address this problem, we pr…

Cited by 6SourcePDFScholar
2020

Design and Control of SLIDER: An Ultra-lightweight, Knee-less, Low-cost Bipedal Walking Robot

IROS 2020poster

Most state-of-the-art bipedal robots are designed to be anthropomorphic and therefore possess legs with knees. Whilst this facilitates more human-like locomotion, there are implementation issues that make walking with straight or near-straight legs difficult. Most bipedal robots have to move with a…

Cited by 28SourceScholar
2020

Differentially Private Top-k Selection via Stability on Unknown Domain

UAI 2020poster

We propose a new method that satisfies approximate differential privacy for top-$k$ selection with unordered output in the unknown data domain setting, not relying on the full knowledge of the domain universe. Our algorithm only requires looking at the top-$\bar{k}$ elements for any given $\bar{k} \…

Cited by 12SourcePDFScholar
2020

FlowNorm: A Learning-based Method for Increasing Convergence Range of Direct Alignment

ICRA 2020poster

Many approaches have been proposed to estimate camera poses by directly minimizing photometric error. However, due to the non-convex property of direct alignment, proper initialization is still required for these methods. Many robust norms (e.g. Huber norm) have been proposed to deal with the outlie…

Cited by 5SourceScholar
2020

HOPPITY: LEARNING GRAPH TRANSFORMATIONS TO DETECT AND FIX BUGS IN PROGRAMS

ICLR 2020spotlight

We present a learning-based approach to detect and fix a broad range of bugs in Javascript programs. We frame the problem in terms of learning a sequence of graph transformations: given a buggy program modeled by a graph structure, our model makes a sequence of predictions including the position of…

Cited by 272SourcecodeScholar
2020

SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference

ICLR 2020talk

We present a modern scalable reinforcement learning agent called SEED (Scalable, Efficient Deep-RL). By effectively utilizing modern accelerators, we show that it is not only possible to train on millions of frames per second but also to lower the cost. of experiments compared to current methods. We…

Cited by 166SourcecodeScholar
2019

Controllable Unsupervised Text Attribute Transfer via Editing Entangled Latent Representation

NeurIPS 2019poster

Unsupervised text attribute transfer automatically transforms a text to alter a specific attribute (e.g. sentiment) without using any parallel data, while simultaneously preserving its attribute-independent content. The dominant approaches are trying to model the content-independent attribute separa…

2019

Exact Gaussian Processes on a Million Data Points

NeurIPS 2019poster

Gaussian processes (GPs) are flexible non-parametric models, with a capacity that grows with the available data. However, computational constraints with standard inference procedures have limited exact GPs to problems with fewer than about ten thousand training points, necessitating approximations f…

2017

An underwater electrosensor for identifying objects of similar volume and aspect ratio using convolutional neural network

IROS 2017poster

Underwater electrosense is bio-inspired by weakly electric fishes that use an electric field to see the objects in the water. Current studies on engineering electrosense focus on designing sophisticated sensors and algorithms for emulating biological functions including localization and identificati…

Cited by 8SourceScholar