Skip to main content
Model Usage ControlAccess KeysPurpose LimitationSecurity

Where Control Enters: Weights, Updates, Outputs, and Data

Four places to intervene in the use of models and data, and the different assumptions behind CoreLocker, AdaLoc, AIM, and non-transferable examples.

A party can control a model’s parameters, the procedure that updates them, the outputs presented to users, or the data supplied to a model. These are different intervention points. They are useful for organising our collaborative work, but they do not define a ranking of security guarantees.

The distinction becomes concrete when we ask: who controls the intervention, what does the other party receive, and what behaviour is being sought?

Parameters: CoreLocker

CoreLocker extracts a small subset of significant weights and treats that subset as an access key. The distributed model is the same network with those weights missing, so restoring the key restores the original parameters rather than unlocking a separate object. See the paper and the project page.

This is not simply a claim that most weights can be pruned. A pruning objective retains performance after removing some weights; a locking objective seeks a small withheld component whose absence leaves an unauthorised holder with only part of the capability. The two questions can be motivated by model structure without being logically equivalent.

Security then depends on what an unauthorised holder can do with the distributed object and any auxiliary information. Do not replace the paper’s attack model with the assumption that the holder cannot modify the weights.

Updates: AdaLoc

AdaLoc considers adaptation over time. Its mechanism confines authorised updates to an intrinsic key, allowing an updated authorised state to be restored without redistributing the full network or performing full re-keying each time. Across six benchmarks and six architectures it reports unauthorised usage accuracy dropping to near-random guessing — 1.02% on CIFAR-100, against up to 87.01% under prior key-based defences. See the paper.

The important qualifier is under the specified update procedure. The construction is not a promise about arbitrary changes to every parameter. An error bound for a restricted update explains a particular cost or guarantee of that procedure; it should not be silently promoted into a theorem about every attack or every possible fine-tuning process.

CoreLocker and AdaLoc can therefore be read together around model access and evolution. That connection does not require asserting that every technical result of one is inherited by the other.

Outputs: AIM

AIM studies modulation of a trained model’s behaviour through logits redistribution, without retraining or access to its training data. Its two modes address utility and focus. See the paper.

This is a different goal from a binary access decision. Changing the level or character of service offered by a model is not, on its own, a mechanism that prevents someone from running an unmodified model. The placement of the modulation therefore matters: who applies it, and can the relevant party bypass it?

Neither “the output is changed” nor “the top-ranked label is changed” is a sufficient explanation of focus modulation: both are descriptions of an effect rather than a mechanism, and the mechanism is the objective being optimised and by whom.

Data: model-specific non-transferable examples

Catch-Only-One is the fourth point, and the only one that intervenes on the data rather than on a model or a deployment. Instead of constraining what a model may learn or how it is updated, it recodes the released data so that its task utility depends on which model reads it.

That is a different kind of constraint, because it presupposes no participating party at all — which is also what makes it the only one of the four that reaches someone who has already downloaded the data and runs it through a model they control. Data That Only One Model Can Read works through the mechanism, the measurable quantity it turns on, and the boundary of the guarantee. See the paper.

A comparison that does not erase the assumptions

WorkIntervention pointMain question
CoreLockerSelected model weightsCan capability depend on possession of a key?
AdaLocAuthorised update procedureCan usage control coexist with model adaptation?
AIMLogit-based behaviour modulationCan one trained model serve different behavioural requirements?
Catch-Only-OneReleased input representationCan task utility be tied to a designated model?

These are four intervention points, not four interchangeable locks. A model owner, a user adjusting behaviour and a data provider can have different goals. Their adversaries, observations and allowed operations need not coincide.

What a combined design would still need

Composing the four would be a new system design, and it would require checking compatible representations, owners of each secret or transformation, permitted updates, composition of utility effects and the combined attack surface.

The same caution applies to ranking them by “how few assumptions” they make. Assumptions about different objects are not generally ordered by a single scalar. The useful question is whether the assumptions fit the intended deployment.

Where this sits

The control research theme links the original work and full author lists. For reconstruction risks rather than usage goals, read What Gradients and Embeddings Can Reveal.

Related research

Continue reading