← Search

Ting Zhang

46 accepted papers

2026

Autoregressive, Yet Revisable: In Decoding Revision for Secure Code Generation

ICML 2026poster

Large Language Model (LLM) based code generation is predominantly formulated as a strictly monotonic process, appending tokens linearly to an immutable prefix. This formulation contrasts to the cognitive process of programming, which is inherently interleaved with forward generation and on-the-fly r…

Cited by 0SourceScholar
2026

GeoLoom: High-quality Geometric Diagram Generation from Textual Input

ICML 2026poster

High-quality geometric diagram generation presents both a challenge and an opportunity: it demands strict spatial accuracy while offering well-defined constraints to guide generation. Inspired by recent advances in geometry problem solving that employ formal languages and symbolic solvers for enhanc…

Cited by 0SourceScholar
2026

On Multi-Step Theorem Prediction via Non-Parametric Structural Priors

ICML 2026poster

Multi-step theorem prediction is a central challenge in automated reasoning. Existing neural–symbolic approaches rely heavily on supervised parametric models, which exhibit limited generalization to evolving theorem libraries. In this work, we explore training-free theorem prediction through the len…

Cited by 0SourceScholar
2026

VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs

ICLR 2026poster

Large Multimodal Models have achieved remarkable progress in integrating vision and language, enabling strong performance across perception, reasoning, and domain-specific tasks. However, their capacity to reason over multiple, visually similar inputs remains insufficiently explored. Such fine-grain…

Cited by 0SourcecodeScholar
2025

CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China

EMNLP 2025

Minority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international standards. This discrepancy has led to a severe lack of relevant corpora, particularly for supervised tasks like headline

Cited by 0SourcePDFScholar
2025

Gaze-Language Alignment for Zero-Shot Prediction of Visual Search Targets from Human Gaze Scanpaths

ICCV 2025poster

Decoding human intent from eye gaze during a visual search task has become an increasingly important capability within augmented and virtual reality systems. However, gaze target prediction models used within such systems are constrained by the predefined target categories found within available gaz…

Cited by 0SourcePDFScholar
2025

Glance2Gaze: Efficient Vision-Language Models from Glance Fusion to Gaze Compression

NeurIPS 2025poster

Vision-language models heavily rely on visual representations, yet ensuring its efficiency remains a critical challenge. Most existing approaches focus on reducing visual tokens either at the visual encoder phase or during the LLM decoder stage. Inspired by human visual cognition, where an initial g…

Cited by 0SourceScholar
2025

High-Quality 3D Creation From a Single Image Using Subject-Specific Knowledge Prior

ICRA 2025

In this paper, we address the critical bottleneck in robotics caused by the scarcity of diverse 3D data by presenting a novel two-stage approach for generating high-quality 3D models from a single image. This method is motivated by the need to efficiently expand 3D asset creation, particularly for r

Cited by 6SourceScholar
2025

ILDiff: Generate Transparent Animated Stickers by Implicit Layout Distillation

ICASSP 2025accepted

High-quality animated stickers usually contain transparent channels, which are often ignored by current video generation models. To generate fine-grained animated transparency channels, existing methods can be roughly divided into video matting algorithms and diffusion-based algorithms. The methods…

Cited by 0SourceScholar
2025

LogRules: Enhancing Log Analysis Capability of Large Language Models through Rules

NAACL 2025findings

Currently, large language models (LLMs) have achieved impressive performance in natural language processing tasks. However, LLMs still exhibit many hallucinations when analyzing system logs, which is due to the implicit knowledge and rules in logs that LLMs cannot capture. Based on this, we propose…

Cited by 0SourcePDFScholar
2025

Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages

ACL 2025long

While multilingual language models like XLM-R have advanced multilingualism in NLP, they still perform poorly in extremely low-resource languages. This situation is exacerbated by the fact that modern LLMs such as LLaMA and Qwen support far fewer languages than XLM-R, making text generation models n…

2025

Pi-GPS: Enhancing Geometry Problem Solving by Unleashing the Power of Diagrammatic Information

ICCV 2025poster

Geometry problem solving has garnered increasing attention due to its potential applications in intelligent education field. Inspired by the observation that text often introduces ambiguities that diagrams can clarify, this paper presents Pi-GPS, a novel framework that unleashes the power of diagram…

2025

RingFormer: A Ring-Enhanced Graph Transformer for Organic Solar Cell Property Prediction

AAAI 2025technical

Organic Solar Cells (OSCs) are a promising technology for sustainable energy production. However, the identification of molecules with desired OSC properties typically involves laborious experimental research. To accelerate progress in the field, it is crucial to develop machine learning models capa…

2025

TACLR: A Scalable and Efficient Retrieval-based Method for Industrial Product Attribute Value Identification

ACL 2025long

Product Attribute Value Identification (PAVI) involves identifying attribute values from product profiles, a key task for improving product search, recommendation, and business analytics on e-commerce platforms.However, existing PAVI methods face critical challenges, such as inferring implicit value…

2025

Vehicle Drifting Planning and Control Framework for Flexible U-turns in Space-limited Environments

IROS 2025

Space-limited U-shape bend is a safety-critical scenario that requires the high maneuverability of vehicles. However, due to the non-holonomic nature of the vehicle, it is difficult to perform flexible U-turns without intricate adjustments, which is detrimental to the efficient execution of tasks. T

Cited by 0SourceScholar
2025

WalkVLM: Aid Visually Impaired People Walking by Vision Language Model

ICCV 2025poster

Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people.With the recent progress of vision-language models (VLMs), applying VLMs to offer walking guidance has bec…

Cited by 0SourcePDFScholar
2024

A Solid-Liquid Composite Flexible Bionic Three-Axis Tactile Sensor for Dexterous Hands

RA-L 2024

The sense of touch is fundamental to grasping and manipulating objects. Similarly, the interaction between the robot and the outside world also requires tactile sensors to provide tactile information to ensure that when grasping fragile objects, the object will not be broken because the grasping for

Cited by 5SourceScholar
2024

DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive Learning

NeurIPS 2024poster

Current techniques for detecting AI-generated text are largely confined to manual feature crafting and supervised binary classification paradigms. These methodologies typically lead to performance bottlenecks and unsatisfactory generalizability. Consequently, these methods are often inapplicable for…

2024

Design, Modeling and Analysis of a Spherical Parallel Continuum Manipulator for Nursing Robots

ICRA 2024poster

In the healthcare industry, nursing robots have made great contributions, assisting in the delivery of food and medicine as well as the movement and transfer of patients. However, the traditional continuum manipulator often has the problems of limited workspace and weak carrying capacity. Compared w…

Cited by 1SourceScholar
2024

InstructDiffusion: A Generalist Modeling Interface for Vision Tasks

CVPR 2024poster

We present InstructDiffusion a unified and generic framework for aligning computer vision tasks with human instructions. Unlike existing approaches that integrate prior knowledge and pre-define the output space (e.g. categories and coordinates) for each vision task we cast diverse vision tasks into…

Cited by 109SourcePDFScholar
2024

RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models

ECCV 2024poster

"We present RodinHD, which can generate high-fidelity 3D avatars from a portrait image. Existing methods fail to capture intricate details such as hairstyles which we tackle in this paper. We first identify an overlooked problem of catastrophic forgetting that arises when fitting triplanes sequentia…

2023

Make-It-3D: High-fidelity 3D Creation from A Single Image with Diffusion Prior

ICCV 2023poster

In this work, we investigate the problem of creating high-fidelity 3D content from only a single image. This is inherently challenging: it essentially involves estimating the underlying 3D geometry while hallucinating unseen textures. To address this challenge, we leverage prior knowledge in a well-…

Cited by 301PDFcodeScholar
2023

MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining

CVPR 2023poster

This paper presents a simple yet effective framework MaskCLIP, which incorporates a newly proposed masked self-distillation into contrastive language-image pretraining. The core idea of masked self-distillation is to distill representation from a full image to the representation predicted from a mas…

2023

Mimicking the Thinking Process for Emotion Recognition in Conversation with Prompts and Paraphrasing

IJCAI 2023poster

Emotion recognition in conversation, which aims to predict the emotion for all utterances, has attracted considerable research attention in recent years. It is a challenging task since the recognition of the emotion in one utterance involves many complex factors, such as the conversational cont…

2023

Model-enhanced Vector Index

NeurIPS 2023poster

Embedding-based retrieval methods construct vector indices to search for document representations that are most similar to the query representations. They are widely used in document retrieval due to low latency and decent recall performance. Recent research indicates that deep retrieval solutions o…

2023

Paint by Example: Exemplar-Based Image Editing With Diffusion Models

CVPR 2023poster

Language-guided image editing has achieved great success recently. In this paper, we investigate exemplar-guided image editing for more precise control. We achieve this goal by leveraging self-supervised training to disentangle and re-organize the source image and the exemplar. However, the naive ap…

2023

PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers

AAAI 2023technical

This paper explores a better prediction target for BERT pre-training of vision transformers. We observe that current prediction targets disagree with human perception judgment. This contradiction motivates us to learn a perceptual prediction target. We argue that perceptually similar images should…

Cited by 273SourcePDFScholar
2023

RODIN: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion

CVPR 2023highlight

This paper presents a 3D diffusion model that automatically generates 3D digital avatars represented as neural radiance fields (NeRFs). A significant challenge for 3D diffusion is that the memory and processing costs are prohibitive for producing high-quality results with rich details. To tackle thi…

Cited by 369SourcePDFScholar
2022

Bootstrapped Masked Autoencoders for Vision BERT Pretraining

ECCV 2022poster

"We propose bootstrapped masked autoencoders (BootMAE), a new approach for vision BERT pretraining. BootMAE improves the original masked autoencoders (MAE) with two core designs: 1) momentum encoder that provides online feature as extra BERT prediction targets; 2) target-aware decoder that tries to…

2022

General Facial Representation Learning in a Visual-Linguistic Manner

CVPR 2022oral

How to learn a universal facial representation that boosts all face analysis tasks This paper takes one step toward this goal. In this paper, we study the transfer performance of pre-trained models on face analysis tasks and introduce a framework, called FaRL, for general facial representation learn…

Cited by 199PDFcodeScholar
2022

Protecting Celebrities From DeepFake With Identity Consistency Transformer

CVPR 2022poster

In this work we propose Identity Consistency Transformer, a novel face forgery detection method that focuses on high-level semantics, specifically identity information, and detecting a suspect face by finding identity inconsistency in inner and outer face regions. The Identity Consistency Transforme…

Cited by 176PDFcodeScholar
2021

CoCosNet v2: Full-Resolution Correspondence Learning for Image Translation

CVPR 2021poster

We present the full-resolution correspondence learning for cross-domain images, which aids image translation. We adopt a hierarchical strategy that uses the correspondence from coarse level to guide the fine levels. At each hierarchy, the correspondence can be efficiently computed via PatchMatch tha…

Cited by 371PDFcodeScholar
2021

Prototypical Pseudo Label Denoising and Target Structure Learning for Domain Adaptive Semantic Segmentation

CVPR 2021poster

Self-training is a competitive approach in domain adaptive segmentation, which trains the network with the pseudo labels on the target domain. However inevitably, the pseudo labels are noisy and the target features are dispersed due to the discrepancy between source and target domains. In this paper…

Cited by 631PDFcodeScholar
2019

Interconnection and Damping Assignment Passivity-Based Impedance Control of a Compliant Assistive Robot for Physical Human-Robot Interactions

RA-L 2019

Interest in series elastic actuators (SEAs) for assistive robots has recently increased due to the unique properties of SEAs compared to those of rigid actuators, such as high force control accuracy, low output impedance, and tolerance to shocks. In this letter, an SEA with a clutch for an assistive

Cited by 23SourceScholar
2018

Interleaved Structured Sparse Convolutional Neural Networks

CVPR 2018poster

In this paper, we study the problem of designing efficient convolutional neural network architectures with the interest in eliminating the redundancy in convolution kernels. In addition to structured sparse kernels, low-rank kernels and the product of low-rank kernels,the product of structured spars…

Cited by 160SourcePDFScholar
2017

NREL-Exo: A 4-DoFs wearable hip exoskeleton for walking and balance assistance in locomotion

IROS 2017poster

In this paper, we presented a high-power, self-balancing, passively and software-controlled active compliant, and wearable hip exoskeleton to provide walking and balance assistance. The device features powered hip abduction/adduction (HAA) and hip flexion/extension (HFE) modules to provide assistanc…

Cited by 36SourceScholar