Data That Only One Model Can Read
Adversarial examples exploit what a model notices; non-transferable examples use what it does not. How data can remain useful to a designated model without being equally useful to other models.
We usually treat a dataset as a general-purpose resource. Once it is published, any model may consume it, and the question we ask is who was allowed access. But there is a second question, asked only once the data has left: does it have to be equally useful to every model that reads it?
The intuition is that utility is a property of the data. It is not quite. Whether a piece of data carries information depends on the thing reading it — a block fits a hole, but what makes the fit is the pair, not either side alone. Ask whether data can be made to have the same property: still present, still perfectly valid input, still carrying everything it needs to carry for one particular reader, and no longer transferring that to anyone else.
This note is about the idea behind Catch-Only-One, which introduces non-transferable examples (NTEs): recoded data that preserves the outputs of a designated model while degrading other models, with no retraining of either and no control over deployment.
The same data, two readers
Start with a picture, because it carries most of the idea on its own.
A set of shapes has been recoded. One model was used to choose that recoding, and reading the recoded shapes it still returns the right answer for every one of them. A second model, which was never involved in choosing the recoding, is handed the same recoded shapes and asked the same question. It returns something that is not the shape.
Two details are what make this the right picture rather than an appealing one.
The data is not split. Both models receive one dataset. Nothing was given to the first model and withheld from the second; the relationship to draw is a fork in the readers, not a partition of the data.
Nothing is malformed. The shapes are well-formed, the input distribution still looks like the authorized one, and no model rejects the input for being the wrong shape or the wrong format. What has changed is the correspondence between the data and the reader: the shapes and the reading model no longer pair up, so the task utility is gone for the unentitled reader while everything else about the input is unchanged.
That is why this is different from noise injection or a corrupted file, and it is what makes it worth wanting. A model that is failing on damaged input is telling you its input was damaged. A model that is failing on well-formed input it simply cannot read is telling you something about the data’s relation to itself.
Where the interference has to go
The other intervention points are well trodden, and they are not weak. Make the weights refuse to run without a key, as CoreLocker does. Constrain what a model may learn, or what its updates may change, as AdaLoc does. Each of these is a real mechanism.
What they share is that each presupposes a party who is participating. A key is something you issue and somebody keeps. Constrained training means intervening before the model exists. And the party of interest is precisely the one who is not participating: someone who downloads the dataset and runs it through a model they control themselves. No mechanism placed downstream of that download can reach them, because there is no deployment to place it in.
So the intervention has to be in the data, before it leaves. That is the only place left, and the picture above is what such an intervention has to look like when it works.
Using the direction a model ignores
The obvious way to change a model’s behaviour on an input is an adversarial example: a perturbation inside the region where the model is most sensitive, arranged so the output flips. That literature works, and it needs no introduction here.
NTEs use the complementary geometry. If sensitive directions are the loud part of a model’s geometry, they are also the part where a small perturbation produces a large effect — which is the wrong place to hide anything, and the part a consumer of the data is most likely to notice and defend against.
The mirror image is the insensitive subspace: directions in the input space along which the model’s output barely moves. Move the data a long way along those directions and the designated model’s answers barely change. The perturbation is large in the input and small in the output, and that asymmetry is the mechanism.
Note what this does not require. The recoding is not something the designated model is taught to read, and there is no key and no training step. The direction is chosen to fit the structure the model already has, which is why the designated model is undisturbed and why nothing has to be taught.
Why small for one model is not small for another
This is the step where the idea stops being a picture.
The argument so far says: there is a subspace where model A barely responds, so move there and A is undisturbed. The obvious worry is that “insensitive” was only ever established about A, and B might be anything. If B is also insensitive in the directions chosen for A, the recoding is invisible to both and nothing has been achieved.
That is the case the construction cannot help, and getting the direction right matters because the intuition runs backwards. The chance comes from the opposite behaviour: B being sensitive along A’s quiet directions is the mechanism working, because it is what gives the same recoding a second reader whose answers it can move. The recoding is a large change in the input; the only reason A is undisturbed is that A does not notice, and a model that does notice is exactly a model whose answers change.
So the correct statement is the reverse of the naive one. Small response for the designated model does not imply small response for another model — and the construction succeeds precisely because that implication fails, in a way that can be measured rather than hoped for.
The quantity that makes it measurable is spectral misalignment: how far the consumer’s relevant subspace sits from the designated model’s. The paper’s bounds tie the two together, so authorized-model fidelity is certified and the degradation an unauthorized model suffers scales with how far misaligned that consumer is. The construction is training-free and data-agnostic on both sides — nothing is retrained and no model is modified — and the evaluation holds under common preprocessing, with unauthorized models collapsing even against adaptive reconstruction attacks rather than only against the easier ones.
That is also where the boundary is. Misalignment is a property of the pair, so “this data is non-transferable” is never a standalone fact — it is always “non-transferable to models in this class”, and the class is defined by measured misalignment. So what matters is not that two models differ in name; it is that they process the data differently. And the claim is not that every other model degrades: a third model with different parameters may share the designated model’s blindness and keep its accuracy. The data has to stay readable to somebody, so no construction in this family protects against a consumer whose geometry aligns with the designated model — that case is excluded by the requirement, not handled by it.
What makes this a bound rather than a demonstration is that a demonstration exhibits one model failing on recoded data. What is certified is that fidelity holds for the designated model under the assumptions just named, and that degradation for an unauthorized one follows from a quantity one can compute.
What this changes
The idea moves purpose limitation from a property of a deployment to a property of the artifact. That is a real shift in what has to be trusted: no key to issue, no cooperation from the model owner, and no retraining — which is what made the model-side mechanisms expensive in the first place.
Further Reading
- Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization — the paper, with the formal bounds and the empirical evaluation.
- Catch-Only-One — the project page, with the method and its stated scope.
- Where Control Enters: Weights, Updates, Outputs, and Data — the four intervention points this one occupies, and what each assumes about the other party.
- AdaLoc — the complementary question, controlling usage from the model side.
Where to go next
Related research
- Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization
The constraint placed upstream in the data.
Continue reading
- Where Control Enters: Weights, Updates, Outputs, and Data
The other model-control overview.
- What Gradients and Embeddings Can Reveal
The other leakage overview.