SHAQ: Shadow Queries for Private Retrieval in Vector Databases
Xinguo Feng, Zhongkui Ma, Zihan Wang, Chuan Yan, Guowei Yang, Alsharif Abuadbba, Guangdong Bai
A shadow-query representation for reducing document-reconstruction risks from stored embeddings while retaining retrieval utility.
The problem
Access to a vector database’s embeddings can enable attempts to reconstruct source documents. SHAQ studies this threat while retaining the embeddings’ role in retrieval.
The method
A language model generates shadow queries about each document. Embeddings of these queries replace the direct document embeddings in the index. Retrieval maps a matching shadow-query vector back to its source document.
Evidence and scope
This changes the representation exposed to an attacker; it does not imply that shadow queries contain no information about the document. The paper evaluates reconstruction and retrieval under specified standard and adaptive attack settings. Reconstruction scores and retrieval metrics measure different outcomes and should be read with their dataset and baseline definitions.
Use the Paper and GitHub links above for the experiments. GHOST addresses a related privacy question during training, but its gradient-based threat model is different.
Citation
@misc{feng2026shaq,
author = {Feng, Xinguo and Ma, Zhongkui and Wang, Zihan and Yan, Chuan and Yang, Guowei and Abuadbba, Alsharif and Bai, Guangdong},
title = {{Shadow Queries for Private Retrieval in Vector Databases}},
year = {2026},
eprint = {2609.04767},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
note = {Preprint},
doi = {10.48550/arXiv.2609.04767},
url = {https://arxiv.org/abs/2609.04767}
}