← Search

Bin Sun

26 accepted papers

2026

SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving

CVPR 2026

Recent end-to-end autonomous driving approaches have leveraged Vision-Language Models (VLMs) to enhance planning capabilities in complex driving scenarios. However, VLMs are inherently trained as generalist models, lacking specialized understanding of driving-specific reasoning in 3D space and time.

Cited by 0SourcecodeScholar
2026

Vec-QMDP: Vectorized POMDP Planning on CPUs for Real-Time Autonomous Driving

RSS 2026poster

Planning under uncertainty for real-world robotics tasks, such as autonomous driving, requires reasoning in enormous high-dimensional belief spaces, rendering the problem computationally intensive. While parallelization offers scalability, existing hybrid CPU-GPU solvers face critical bottlenecks du…

Cited by 0SourceScholar
2025

Bridging Layout and RTL: Knowledge Distillation based Timing Prediction

ICML 2025spotlight

Accurate and efficient timing prediction at the register-transfer level (RTL) remains a fundamental challenge in electronic design automation (EDA), particularly in striking a balance between accuracy and computational efficiency. While static timing analysis (STA) provides high-fidelity results thr…

2025

Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning

ICASSP 2025accepted

Direction reasoning is essential for intelligent systems to understand the real world. While existing work focuses primarily on spatial reasoning, compass direction reasoning remains underexplored. To address this, we propose the Compass Direction Reasoning (CDR) benchmark, designed to evaluate the…

Cited by 0SourceScholar
2025

DrVideo: Document Retrieval Based Long Video Understanding

CVPR 2025poster

Most of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The increased number of frames in long videos poses two main challenges: difficulty in locating key information and performing…

2025

Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic Guidance

IJCAI 2025

Open-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal assistants. The advent of large language models (LLMs) has greatly advanced this field by improving context understanding a

2024

Dynamic Stochastic Decoding Strategy for Open-Domain Dialogue Generation

ACL 2024findings

Stochastic sampling strategies such as top-k and top-p have been widely used in dialogue generation task. However, as an open-domain chatting system, there will be two different conversation scenarios, i.e. chit-chat and knowledge-based question answering. In the former situation, responses diversit…

2024

Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step Reasoning

ICLR 2024poster

Self-consistency (SC) has been a widely used decoding strategy for chain-of-thought reasoning. Despite bringing significant performance improvements across a variety of multi-step reasoning tasks, it is a high-cost method that requires multiple sampling with the preset size. In this paper, we propos…

2024

Exploring Conditional Variational Mechanism to Pinyin Input Method for Addressing One-to-Many Mappings in Low-Resource Scenarios

ACL 2024short

Pinyin input method engine (IME) refers to the transformation tool from pinyin sequence to Chinese characters, which is widely used on mobile phone applications. Due to the homophones, Pinyin IME suffers from the one-to-many mapping problem in the process of pinyin sequences to Chinese characters. T…

2024

Turning Dust into Gold: Distilling Complex Reasoning Capabilities from LLMs by Leveraging Negative Data

AAAI 2024technical

Large Language Models (LLMs) have performed well on various reasoning tasks, but their inaccessibility and numerous parameters hinder wide application in practice. One promising way is distilling the reasoning ability from LLMs to small models by the generated chain-of-thought reasoning paths. In so…

2023

Better Correlation and Robustness: A Distribution-Balanced Self-Supervised Learning Framework for Automatic Dialogue Evaluation

NeurIPS 2023poster

Turn-level dialogue evaluation models (TDEMs), using self-supervised learning (SSL) framework, have achieved state-of-the-art performance in open-domain dialogue evaluation. However, these models inevitably face two potential problems. First, they have low correlations with humans on medium coherenc…

Cited by 4SourcePDFScholar
2023

Hybrid Pixel-Unshuffled Network for Lightweight Image Super-resolution

AAAI 2023technical

Convolutional neural network (CNN) has achieved great success on image super-resolution (SR). However, most deep CNN-based SR models take massive computations to obtain high performance. Downsampling features for multi-resolution fusion is an efficient and effective way to improve the performance of…

2023

Large Language Models are Better Reasoners with Self-Verification

EMNLP 2023long findings

Recently, with the chain of thought (CoT) prompting, large language models (LLMs), e.g., GPT-3, have shown strong reasoning ability in several natural language processing tasks such as arithmetic, commonsense, and logical reasoning. However, LLMs with CoT require multi-step prompting and multi-token…

Cited by 0SourcecodeScholar
2023

Logo-Former: Local-Global Spatio-Temporal Transformer for Dynamic Facial Expression Recognition

ICASSP 2023accepted

Previous methods for dynamic facial expression recognition (DFER) in the wild are mainly based on Convolutional Neural Networks (CNNs), whose local operations ignore the long-range dependencies in videos. Transformer-based methods for DFER can achieve better performances but result in higher FLOPs a…

Cited by 0SourceScholar
2023

Towards Diverse, Relevant and Coherent Open-Domain Dialogue Generation via Hybrid Latent Variables

AAAI 2023technical

Conditional variational models, using either continuous or discrete latent variables, are powerful for open-domain dialogue response generation. However, previous works show that continuous latent variables tend to reduce the coherence of generated responses. In this paper, we also found that discre…

Cited by 6SourcePDFScholar
2023

Towards Fewer Hallucinations in Knowledge-Grounded Dialogue Generation via Augmentative and Contrastive Knowledge-Dialogue

ACL 2023short

Existing knowledge-grounded open-domain dialogue generation models often face the hallucination problem, i.e. the dialogue generative model will persist in an inappropriate knowledge and generate responses that inconsistent with the facts. We argue that this problem mainly stems from the polarized o…

Cited by 5SourcePDFScholar
2022

Diversifying Neural Dialogue Generation via Negative Distillation

NAACL 2022long

Generative dialogue models suffer badly from the generic response problem, limiting their applications to a few toy scenarios. Recently, an interesting approach, namely negative training, has been proposed to alleviate this problem by reminding the model not to generate high-frequency responses duri…

2022

Modeling Complex Dialogue Mappings via Sentence Semantic Segmentation Guided Conditional Variational Auto-Encoder

EMNLP 2022finding

Complex dialogue mappings (CDM), including one-to-many and many-to-one mappings, tend to make dialogue models generate incoherent or dull responses, and modeling these mappings remains a huge challenge for neural dialogue systems. To alleviate these problems, methods like introducing external inform…

Cited by 1SourcePDFScholar
2021

Generating Relevant and Coherent Dialogue Responses using Self-Separated Conditional Variational AutoEncoders

ACL 2021long

Conditional Variational AutoEncoder (CVAE) effectively increases the diversity and informativeness of responses in open-ended dialogue generation tasks through enriching the context vector with sampled latent variables. However, due to the inherent one-to-many and many-to-one phenomena in human dial…

Cited by 36SourcePDFScholar
2015

Blind bleed-through removal for scanned historical document images with conditional random fields

ICASSP 2015accepted

Due to the quality of paper and long-time preservation, the ink on one side of the historical documents often seeps through and appears on the other side. In this paper, a new blind ink bleed-through removal method is proposed to deal with the scanned historical document images. The scanned historic…

Cited by 0SourceScholar