Skip to main content

Publications

Research papers in formal verification, trustworthy AI, and security.

Research Publications

2026

5 papers

Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization

Zihan Wang, Zhiyong Ma, Zhongkui Ma, Shuofeng Liu, Akide Liu, Derui Wang, Minhui Xue, Guangdong Bai

NeurIPS'26AcceptedCCF-A, CORE-A*

Recoding data in model-specific low-sensitivity directions to preserve designated-model utility, and studying the degradation that comes with misalignment against that subspace.

Abstract and citation

the Conference on Neural Information Processing Systems (NeurIPS)

Recent AI regulations increasingly emphasize the need for mechanisms that preserve the utility of data for AI innovation while preventing misuse, particularly by enforcing purpose limitation in downstream AI applications. In practice, enforcing this principle remains challenging, as released data can be trivially fed into arbitrary models beyond its declared intent. Existing approaches attempt to mitigate this risk by either perturbing data or retraining models to limit unintended use. These strategies, however, offer no protection against inference by unknown or externally trained models, or fundamentally rely on control over the training or deployment. In this work, we introduce non-transferable examples (NTEs), recoded data that act as a task-level "ciphertext" decodable only by a designated model. Whereas adversarial examples exploit directions of high model sensitivity, NTEs leverage the complementary insensitive subspace. We propose a training-free, data-agnostic method that recodes data within a model-specific low-sensitivity subspace, preserving outputs for the authorized model while degrading unauthorized ones through subspace misalignment. We establish formal bounds certifying authorized-model fidelity and showing that unauthorized degradation scales with measurable spectral misalignment between models. Empirically, NTEs preserve performance across diverse vision backbones and state-of-the-art vision-language models under common preprocessing, while unauthorized models collapse even under adaptive reconstruction attacks. These results establish NTEs as a practical means to preserve intended data utility while preventing unauthorized exploitation. The work was accepted as an oral presentation.

Non-Transferable Examples

Zihan Wang, Zhiyong Ma, Zhongkui Ma, Shuofeng Liu, Akide Liu, Derui Wang, Minhui Xue, Guangdong Bai

ECCV'26 LifeGenIP Workshop

The workshop version of the non-transferable example work, recoding data in a model-specific subspace so unauthorized models degrade without retraining or deployment control.

Abstract and citation

the ECCV Workshop on Life-Cycle Intellectual Property Governance of Visual Generative Models (LifeGenIP) · Malmo, Sweden

A companion workshop paper presenting non-transferable examples at the ECCV 2026 LifeGenIP Workshop on life-cycle intellectual property governance of visual generative models. The work recodes data inside a model-specific low-sensitivity subspace so that only the authorized model remains accurate, without retraining or control over deployment. It received a Best Paper Award Runner-Up at the workshop.

Re-Key-Free, Risky-Free: Adaptable Model Usage Control

Zihan Wang, Zhongkui Ma, Xinguo Feng, Chuan Yan, Dongge Liu, Ruoxi Sun, Derui Wang, Minhui Xue, Guangdong Bai

Euro S&P'26CCF-C, CORE-A

Proposes AdaLoc, which keeps key-based model usage control working as the model evolves instead of reissuing keys on every update.

Abstract and citation

the 11th IEEE European Symposium on Security and Privacy (EuroS&P) · Lisbon, Portugal

Deep neural networks (DNNs) have become valuable intellectual property of model owners, due to the substantial resources required for their development. To protect these assets in the deployed environment, recent research has proposed model usage control mechanisms to ensure models cannot be used without proper authorization. These methods typically lock the utility of the model by embedding an access key into its parameters. However, they often assume static deployment, and largely fail to withstand continual post-deployment model updates, such as fine-tuning or task-specific adaptation. In this paper, we propose AdaLoc, to endow key-based model usage control with adaptability during model evolution. It strategically selects a subset of weights as an intrinsic access key, which enables all model updates to be confined to this key throughout the evolution lifecycle. AdaLoc enables using the access key to restore the keyed model to the latest authorized states without redistributing the entire network (i.e., adaptation), and frees the model owner from full re-keying after each model update (i.e., lock preservation). We establish a formal foundation to underpin AdaLoc, providing crucial bounds such as the errors introduced by updates restricted to the access key. Experiments across six vision and language benchmarks and six modern architectures spanning CNNs and Transformers demonstrate that AdaLoc achieves high accuracy under significant updates while retaining robust protections. Specifically, authorized usages consistently achieve strong task-specific performance, while unauthorized usage accuracy drops to near-random guessing levels (e.g., 1.02% on CIFAR-100), compared to up to 87.01% under prior key-based defenses.

BibTeX

@inproceedings{wang2026rekeyfree,
  author = {Wang, Zihan and Ma, Zhongkui and Feng, Xinguo and Yan, Chuan and Liu, Dongge and Sun, Ruoxi and Wang, Derui and Xue, Minhui and Bai, Guangdong},
  title = {Re-Key-Free, Risky-Free: Adaptable Model Usage Control},
  booktitle = {2026 IEEE 11th European Symposium on Security and Privacy (EuroS\&P)},
  pages = {696--711},
  year = {2026},
  publisher = {IEEE},
  doi = {10.1109/EuroSP68448.2026.00051}
}

Mitigating Gradient Inversion Risks in Language Models via Token Obfuscation

Xinguo Feng, Zhongkui Ma, Zihan Wang, Alsharif Abuadbba, Guangdong Bai

Asia CCS'26CCF-C, CORE-A

Using token-level obfuscation to decouple gradient, embedding and token spaces, reducing gradient-based text reconstruction while preserving training utility in the evaluated settings.

Abstract and citation

the ACM Asia Conference on Computer and Communications Security (ASIA CCS) · Bangalore, India

Training and fine-tuning large-scale language models largely benefit from collaborative learning, but the approach has been proven vulnerable to gradient inversion attacks (GIAs), which allow adversaries to reconstruct private training data from shared gradients. Existing defenses mainly employ gradient perturbation techniques, e.g., noise injection or gradient pruning, to disrupt GIAs' direct mapping from gradient space to token space. However, these methods often fall short due to the retention of semantics similarity across gradient, embedding, and token spaces. In this work, we propose a novel defense mechanism named Ghost (gradient shield with obfuscated tokens), a token-level obfuscation mechanism that neutralizes GIAs by decoupling the inherent connections across gradient, embedding, and token spaces. Ghost is built upon an important insight: due to the large scale of the token space, there exist semantically distinct yet embedding-proximate tokens that can serve as the shadow substitutes of the original tokens, which enables a semantic disconnection in the token space while preserving the connection in the embedding and gradient spaces. Ghost comprises a searching step, which identifies semantically distinct candidate tokens using a multi-criteria searching process, and a selection step, which selects optimal shadow tokens to ensure minimal disruption to features critical for training by preserving alignment with the internal outputs produced by original tokens. Evaluation across diverse model architectures (from BERT to Llama) and datasets demonstrates the remarkable effectiveness of Ghost in protecting privacy (as low as 1% in recovery rate) and preserving utility (up to 0.92 in classification F1 and 5.45 in perplexity), in both classification and generation tasks against state-of-the-art GIAs and adaptive attack scenarios.

BibTeX

@inproceedings{feng2026mitigating,
  author = {Feng, Xinguo and Ma, Zhongkui and Wang, Zihan and Abuadbba, Alsharif and Bai, Guangdong},
  title = {Mitigating Gradient Inversion Risks in Language Models via Token Obfuscation},
  booktitle = {Proceedings of the ACM Asia Conference on Computer and Communications Security},
  pages = {1832--1848},
  year = {2026},
  publisher = {ACM},
  doi = {10.1145/3779208.3785389}
}

Shadow Queries for Private Retrieval in Vector Databases

Xinguo Feng, Zhongkui Ma, Zihan Wang, Chuan Yan, Guowei Yang, Alsharif Abuadbba, Guangdong Bai

arXiv:2609.04767Preprint

Studying generated shadow-query embeddings as an alternative retrieval representation for reducing document reconstruction risks.

Abstract and citation

arXiv preprint arXiv:2609.04767, submitted 4 September 2026

Large language models (LLMs) increasingly rely on information retrieval (IR) systems, such as Retrieval-Augmented Generation (RAG), to incorporate domain-specific knowledge without costly re-training. These systems often store pre-computed document embeddings in cloud-based vector databases. However, such embeddings are vulnerable to embedding inversion attacks (EIAs), which can reconstruct their underlying text. Existing defenses, such as adding noise or scaling embeddings, often provide limited privacy or significantly reduce retrieval utility. We propose SHAQ (shadow query generation), a semantic-decomposition and embedding-decoupling defense against EIAs. SHAQ is based on the insight that EIAs rely on the strong coupling between an embedding and its original text. Instead of storing document embeddings directly, SHAQ uses a generative language model to create diverse shadow queries that capture different semantic aspects of each document. These queries are then encoded and stored in place of the original document embeddings, thereby decomposing document semantics and decoupling stored embeddings from the source text. Experiments across diverse IR datasets show that SHAQ substantially improves privacy while preserving retrieval utility, achieving a recovery rate as low as 0.2104, defending up to 19.50% more tokens than baseline defenses, and reaching up to 0.7967 MAP@10 with up to 5.53% utility improvement. These results demonstrate that semantic decomposition and embedding decoupling provide an effective alternative to directly modifying embeddings for defending against EIAs.

BibTeX

@misc{feng2026shaq,
  author = {Feng, Xinguo and Ma, Zhongkui and Wang, Zihan and Yan, Chuan and Yang, Guowei and Abuadbba, Alsharif and Bai, Guangdong},
  title = {{Shadow Queries for Private Retrieval in Vector Databases}},
  year = {2026},
  eprint = {2609.04767},
  archivePrefix = {arXiv},
  primaryClass = {cs.AI},
  note = {Preprint},
  doi = {10.48550/arXiv.2609.04767},
  url = {https://arxiv.org/abs/2609.04767}
}

2025

2 papers

Convex Hull Approximation for Activation Functions

Zhongkui Ma, Zihan Wang, Guangdong Bai

OOPSLA'25CCF-A, CORE-A

Proposes WraAct, constructing tight over-approximations of activation function hulls efficiently, and evaluates it on Sigmoid, Tanh and MaxPool.

Abstract and citation

the ACM SIGPLAN Conference on Object-Oriented Programming, Systems, Languages, and Applications (OOPSLA'25) within SPLASH'25 · Singapore

The wide adoption of deep learning in safety-critical domains has driven the need for formally verifying the robustness of neural networks. A critical challenge in this endeavor lies in addressing the inherent non-linearity of activation functions. The convex hull of the activation function has emerged as a promising solution, as it effectively tightens variable ranges and provides multi-neuron constraints, which together enhance verification precision. Given that constructing exact convex hulls is computationally expensive and even infeasible in most cases, existing research has focused on over-approximating them. Several ad-hoc methods have been devised for specific functions such as ReLU and Sigmoid. Nonetheless, there remains a substantial gap in developing broadly applicable approaches for general activation functions. In this work, we propose WraAct, an approach to efficiently constructing tight over-approximations for activation function hulls. Its core idea is to introduce linear constraints to smooth out the fluctuations in the target function, by leveraging double-linear-piece (DLP) functions to simplify the local geometry. In this way, the problem is reduced to over-approximating DLP functions, which can be efficiently handled. We evaluate WraAct against SBLM+PDDM, the state-of-the-art (SOTA) multi-neuron over-approximation method. WraAct outperforms it on commonly-used functions like Sigmoid, Tanh, and MaxPool, offering superior efficiency (average 400X faster) and precision (average 150X) while constructing fewer constraints (average 50% reduction).

BibTeX

@article{10.1145/3763086,
  author = {Ma, Zhongkui and Wang, Zihan and Bai, Guangdong},
  title = {Convex Hull Approximation for Activation Functions},
  year = {2025},
  issue_date = {October 2025},
  publisher = {Association for Computing Machinery},
  volume = {9},
  number = {OOPSLA2},
  url = {https://doi.org/10.1145/3763086},
  doi = {10.1145/3763086},
  month = oct,
  articleno = {308},
  numpages = {27}
}

AI Model Modulation with Logits Redistribution

Zihan Wang, Zhongkui Ma, Xinguo Feng, Zhiyang Mei, Ethan Ma, Derui Wang, Minhui Xue, Guangdong Bai

WWW'25CCF-A, CORE-A*

Proposes AIM, redistributing logits so one trained model can serve several stakeholder-specific behaviours without retraining per user.

Abstract and citation

the ACM Web Conference 2025 (WWW'25) · Sydney, Australia

The substantial data and resource consumption of training deep neural networks has rendered the large-scale training accessible only to organizations with necessary infrastructure and massive datasets. We propose AIM, a novel model modulation paradigm that enables a single model to exhibit diverse behaviors meeting the specific needs of stakeholders. AIM introduces a logits redistribution strategy for modulating model behaviors in a training data-agnostic and retraining-free manner.

BibTeX

@inproceedings{wang2025ai,
  author = {Wang, Zihan and Ma, Zhongkui and Feng, Xinguo and Mei, Zhiyang and Ma, Ethan and Wang, Derui and Xue, Minhui and Bai, Guangdong},
  title = {{AI} Model Modulation with Logits Redistribution},
  booktitle = {Proceedings of the ACM Web Conference 2025},
  year = {2025},
  url = {https://openreview.net/forum?id=lOSomJvrc5}
}

2024

3 papers

Uncovering Gradient Inversion Risks in Practical Language Model Training

Xinguo Feng, Zhongkui Ma, Zihan Wang, Eu Joe Chegne, Mengyao Ma, Alsharif Abuadbba, Guangdong Bai

CCS'24CCF-A, CORE-A*

Presents GRAB, a gradient inversion attack for language models built on two alternating optimization processes, showing the threat survives realistic training setups.

Abstract and citation

the 31st ACM Conference on Computer and Communications Security (CCS'24) · Salt Lake City, USA

The gradient inversion attack has been demonstrated as a significant privacy threat to federated learning, particularly in continuous domains such as vision models. We propose GRAB (gradient inversion with hybrid optimization), a domain-specific gradient inversion attack featuring two alternating optimization processes. GRAB can recover a significant portion (up to 92.9% recovery rate) of the private training data.

BibTeX

@inproceedings{10.1145/3658644.3690292,
  author = {Feng, Xinguo and Ma, Zhongkui and Wang, Zihan and Chegne, Eu Joe and Ma, Mengyao and Abuadbba, Alsharif and Bai, Guangdong},
  title = {Uncovering Gradient Inversion Risks in Practical Language Model Training},
  year = {2024},
  publisher = {Association for Computing Machinery},
  doi = {10.1145/3658644.3690292},
  booktitle = {Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security},
  pages = {3525--3539},
  series = {CCS '24}
}

CoreLocker: Neuron-level Usage Control

Zihan Wang, Zhongkui Ma, Xinguo Feng, Ruoxi Sun, Hu Wang, Minhui Xue, Guangdong Bai

S&P'24CCF-A, CORE-A*

Proposes CoreLocker, which locks a model behind a small extracted subset of significant weights that acts as the access key.

Abstract and citation

the 45th IEEE Symposium on Security and Privacy (S&P'24) · San Francisco, USA

We propose CORELOCKER, employing the strategic extraction of a small subset of significant weights from the neural network. This subset serves as the access key to unlock the model's complete capability. Authorized users with the access key have full access to the model, while unauthorized users can have access to only part of its capability.

BibTeX

@inproceedings{wang2024corelocker,
  author = {Wang, Zihan and Ma, Zhongkui and Feng, Xinguo and Sun, Ruoxi and Wang, Hu and Xue, Minhui and Bai, Guangdong},
  title = {CoreLocker: Neuron-level Usage Control},
  booktitle = {2024 IEEE Symposium on Security and Privacy (SP)},
  pages = {222--222},
  year = {2024},
  publisher = {IEEE Computer Society}
}

ReLU Hull Approximation

Zhongkui Ma, Jiaying Li, Guangdong Bai

POPL'24CCF-A, CORE-A*

Proposes WraLU, which over-approximates the ReLU convex hull by reusing the function's own linear pieces, reporting large speedups and fewer constraints on the evaluated benchmarks.

Abstract and citation

the 51st ACM SIGPLAN Symposium on Principles of Programming Languages (POPL'24) · London, UK

Convex hulls are commonly used to tackle the non-linearity of activation functions in the verification of neural networks. We propose a fast and precise approach to over-approximating the convex hull of the ReLU function. Our key insight is to formulate a convex polytope that "wraps" the ReLU hull, by reusing the linear pieces of the ReLU function as the lower faces and constructing upper faces that are adjacent to the lower faces. We implement our approach as WraLU, and evaluate its performance. WraLU outperforms existing methods by generating fewer constraints to achieve tighter approximation in less time. It reduces the number of constraints by up to half, while delivering comparable or even superior results.

BibTeX

@article{ma2024relu,
  author = {Ma, Zhongkui and Li, Jiaying and Bai, Guangdong},
  title = {ReLU Hull Approximation},
  journal = {Proceedings of the ACM on Programming Languages},
  volume = {8},
  number = {POPL},
  pages = {2260--2287},
  year = {2024},
  publisher = {ACM New York, NY, USA}
}

2023

2 papers

Verifying Neural Networks by Approximating Convex Hulls

Zhongkui Ma

ICFEM'23CCF-C, CORE-C

Constructs a convex polytope over-approximating the ReLU hull in the multivariate case, extending the hull-approximation approach to more than one variable.

Abstract and citation

International Conference on Formal Engineering Methods (ICFEM'23), Doctoral Symposium · Brisbane, Australia

The increasing prevalence of neural networks necessitates their verification in order to ensure security. This work concentrates on approximating the convex hull of activation functions. An approach is proposed to construct a convex polytope to over-approximate the ReLU hull when considering multi-variables.

BibTeX

@inproceedings{ma2023verifying,
  author = {Ma, Zhongkui},
  title = {Verifying Neural Networks by Approximating Convex Hulls},
  booktitle = {International Conference on Formal Engineering Methods},
  pages = {261--266},
  year = {2023},
  publisher = {Springer}
}

Formalizing Robustness Against Character-Level Perturbations for Neural Network Language Models

Zhongkui Ma, Xinguo Feng, Zihan Wang, Shuofeng Liu, Mengyao Ma, Hao Guan, Mark Huasong Meng

ICFEM'23CCF-C, CORE-C

Defines a robustness specification for character-level perturbations of neural network language models, built on three metrics for generalizing text perturbations.

Abstract and citation

International Conference on Formal Engineering Methods (ICFEM'23) · Brisbane, Australia

This work focuses on formalizing robustness specification against character-level perturbations for neural network language models. We introduce a key principle of three metrics, namely probability distribution, density, and diversity, for generalizing neural network language model perturbations and formulate the robustness specification against character-level perturbed text inputs.

BibTeX

@inproceedings{ma2023formalizing,
  author = {Ma, Zhongkui and Feng, Xinguo and Wang, Zihan and Liu, Shuofeng and Ma, Mengyao and Guan, Hao and Meng, Mark Huasong},
  title = {Formalizing Robustness Against Character-Level Perturbations for Neural Network Language Models},
  booktitle = {International Conference on Formal Engineering Methods},
  pages = {100--117},
  year = {2023},
  publisher = {Springer}
}

Earlier Work

2018

2 papers

Does Truthfully-Stating Strategy Really Have its Reward? — Research on the Communication Strategies of Innovation Quality

Haixin Ding, Li Xie, Zhongkui Ma

Technology Intelligence Engineering

Model of Weibo Negative Public Opinion Communication in Colleges and Universities Based on Double-layer Network

Zhongkui Ma

Journal of Jiamusi Vocational Institute