Skip to main content
CCS'24Gradient InversionFederated LearningPrivacySecurity

GRAB: Gradient Inversion for Language Model Training

Xinguo Feng, Zhongkui Ma, Zihan Wang, Eu Joe Chegne, Mengyao Ma, Alsharif Abuadbba, Guangdong Bai

Gradient inversion in language-model training, combining dropout-mask estimation with discrete token-sequence optimisation.

The problem

Gradient inversion was already a demonstrated privacy threat in vision. Language models introduce different obstacles, including discrete tokens and practical training settings such as dropout. GRAB investigates whether these obstacles prevent reconstruction from shared gradients.

The method

GRAB alternates optimisation for token recovery, including estimation of dropout masks, with discrete optimisation of token ordering. The alternating procedure addresses both continuous optimisation and the discrete structure of a text sequence.

Evidence and scope

The paper reports reconstruction experiments under specified benchmark and practical training settings. Their success rates depend on the model, batch and training configuration, and the attacker’s access. They do not mean that every gradient reveals an entire training dataset.

The related GHOST project studies a token-level defence. See the paper and implementation linked above for the attack settings and reproduction procedure.

Citation

@inproceedings{10.1145/3658644.3690292,
  author = {Feng, Xinguo and Ma, Zhongkui and Wang, Zihan and Chegne, Eu Joe and Ma, Mengyao and Abuadbba, Alsharif and Bai, Guangdong},
  title = {Uncovering Gradient Inversion Risks in Practical Language Model Training},
  year = {2024},
  publisher = {Association for Computing Machinery},
  doi = {10.1145/3658644.3690292},
  booktitle = {Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security},
  pages = {3525--3539},
  series = {CCS '24}
}