Technical appendix

  • Posterior sensitivity and attack optimization.
  • Moment matching and decision targets.
  • Exact defense models, practical approximations, and evaluation.

These are optional detours. Return to the main talk.

Why forward KL is convex in the weights

Let \ell(\theta) contain the observation log likelihoods. Up to a constant in w:

\begin{aligned} F(w)&=\log Z(w)-w^\top\mathbb E_{\pi_A}[\ell(\theta)],\\ \nabla_w^2F(w)&=\operatorname{Cov}_{\pi_w}\!\left(\ell(\theta),\ell(\theta)\right)\succeq0. \end{aligned}

  • The covariance matrix is positive semidefinite: the continuous objective is convex.
  • This does not remove integer constraints, sampler error or rounding effects.

Return to the attack algorithm

Influence is a covariance

For a posterior summary h(\theta):

\frac{\partial}{\partial w_i}\mathbb E_{\pi_w}[h(\theta)] =\operatorname{Cov}_{\pi_w}\!\left(h(\theta),\ell_i(\theta)\right).

  • Increasing a row’s weight raises the target expectation when its fit covaries positively with the target feature.
  • The same identity applies to moments and decision utilities.
  • It is a local derivative; large discrete changes require refitting.

Return to moment matching

MMD defines what “close” means

For independent draws within and across distributions:

\mathrm{MMD}^2(P,Q;k) =\mathbb E_{P,P}[k]-2\mathbb E_{P,Q}[k]+\mathbb E_{Q,Q}[k].

  • A characteristic kernel can distinguish full distributions.
  • A feature kernel k(\theta,\theta')=h(\theta)^\top h(\theta') compares only the selected feature expectations.
  • Equal feature expectations can give zero discrepancy even when the posteriors differ.

Return to moment matching

From a moment target to a decision

Let g_k(\theta)=u(a_k,\theta)-u(a_A,\theta) compare each alternative with the attacker’s desired action.

\left\|\mathbb E_{\pi_w}[g]+\gamma\mathbf1\right\|_2<\gamma \quad\Longrightarrow\quad \mathbb E_{\pi_w}[g_k]<0\ \text{for every }k.

  • With \gamma>0, this is a sufficient condition for the desired action to be uniquely optimal among finitely many actions.
  • It concerns exact expected utility gaps. Monte Carlo estimates need uncertainty bounds before claiming a certificate.

Return to the radon result

One spatial run and repeated performance

Spatial attack performance across deletion budgets and optimization methods

  • Five runs per method; error bars: \pm2 standard errors.
  • Main example: selected budget-20 Adam-R2 run.
  • Heuristics and rounding affect the returned attack.

Predictive entropy has two components

H(Y\mid x,D)= \underbrace{\mathbb E_{\theta\mid D}H(Y\mid x,\theta)}_{\text{average within-model uncertainty}} +\underbrace{I(Y;\theta\mid x,D)}_{\text{disagreement across models}}.

  • The second term measures dependence between predictions and parameters.
  • Both depend on the model and posterior approximation.
  • Neither term is a calibration test or a guaranteed attack detector.

The reactive update really is joint

p(x,\theta,\phi\mid x',D) \propto p(x'\mid x,\theta)\,p(x\mid\phi)\,p(\theta,\phi\mid D).

  • \phi parameterizes the distribution of clean inputs.
  • The received input updates plausible originals and global parameters.
  • The offline approximation suppresses the second update.

Two objectives that should not be confused

Generative channel likelihood

\log \mathbb E_{x'\mid x,\theta}p(y\mid x',\theta).

Practical surrogate contribution

\mathbb E_{x'\mid x,y,\theta}\log p(y\mid x',\theta).

The experiments use channel-averaged losses and a generalized Bayesian posterior.

What was the evaluation attacker allowed to do?

Component Manuscript evaluation
Access White-box, each defense’s own predictive distribution
Constraint L_2 perturbation budget \epsilon
PGD 50 likelihood-objective steps
PGD+ 25 likelihood steps + 25 entropy steps
Selective prediction Increase digit entropy; decrease FashionMNIST entropy

Where is the computational cost paid?

MNIST, reported configuration Standard Defended
Training, ms per batch iteration 2.48 4.69 (MIX)
Prediction, ms per input 1.60 2470 (onPure)
  • onPure: S=5 parameter draws, N=100 stored input samples.
  • Low-dimensional regression: roughly fivefold reactive overhead.
  • These are implementation measurements, not computational lower bounds.