Technical appendix
- Posterior sensitivity and attack optimization.
- Moment matching and decision targets.
- Exact defense models, practical approximations, and evaluation.
Why forward KL is convex in the weights
Let \ell(\theta) contain the observation log likelihoods. Up to a constant in w:
\begin{aligned}
F(w)&=\log Z(w)-w^\top\mathbb E_{\pi_A}[\ell(\theta)],\\
\nabla_w^2F(w)&=\operatorname{Cov}_{\pi_w}\!\left(\ell(\theta),\ell(\theta)\right)\succeq0.
\end{aligned}
- The covariance matrix is positive semidefinite: the continuous objective is convex.
- This does not remove integer constraints, sampler error or rounding effects.
Return to the attack algorithm
Influence is a covariance
For a posterior summary h(\theta):
\frac{\partial}{\partial w_i}\mathbb E_{\pi_w}[h(\theta)]
=\operatorname{Cov}_{\pi_w}\!\left(h(\theta),\ell_i(\theta)\right).
- Increasing a row’s weight raises the target expectation when its fit covaries positively with the target feature.
- The same identity applies to moments and decision utilities.
- It is a local derivative; large discrete changes require refitting.
Return to moment matching
MMD defines what “close” means
For independent draws within and across distributions:
\mathrm{MMD}^2(P,Q;k)
=\mathbb E_{P,P}[k]-2\mathbb E_{P,Q}[k]+\mathbb E_{Q,Q}[k].
- A characteristic kernel can distinguish full distributions.
- A feature kernel k(\theta,\theta')=h(\theta)^\top h(\theta') compares only the selected feature expectations.
- Equal feature expectations can give zero discrepancy even when the posteriors differ.
Return to moment matching
From a moment target to a decision
Let g_k(\theta)=u(a_k,\theta)-u(a_A,\theta) compare each alternative with the attacker’s desired action.
\left\|\mathbb E_{\pi_w}[g]+\gamma\mathbf1\right\|_2<\gamma
\quad\Longrightarrow\quad
\mathbb E_{\pi_w}[g_k]<0\ \text{for every }k.
- With \gamma>0, this is a sufficient condition for the desired action to be uniquely optimal among finitely many actions.
- It concerns exact expected utility gaps. Monte Carlo estimates need uncertainty bounds before claiming a certificate.
Return to the radon result
One spatial run and repeated performance
- Five runs per method; error bars: \pm2 standard errors.
- Main example: selected budget-20 Adam-R2 run.
- Heuristics and rounding affect the returned attack.
Predictive entropy has two components
H(Y\mid x,D)=
\underbrace{\mathbb E_{\theta\mid D}H(Y\mid x,\theta)}_{\text{average within-model uncertainty}}
+\underbrace{I(Y;\theta\mid x,D)}_{\text{disagreement across models}}.
- The second term measures dependence between predictions and parameters.
- Both depend on the model and posterior approximation.
- Neither term is a calibration test or a guaranteed attack detector.
The reactive update really is joint
p(x,\theta,\phi\mid x',D)
\propto
p(x'\mid x,\theta)\,p(x\mid\phi)\,p(\theta,\phi\mid D).
- \phi parameterizes the distribution of clean inputs.
- The received input updates plausible originals and global parameters.
- The offline approximation suppresses the second update.
Two objectives that should not be confused
Generative channel likelihood
\log \mathbb E_{x'\mid x,\theta}p(y\mid x',\theta).
Practical surrogate contribution
\mathbb E_{x'\mid x,y,\theta}\log p(y\mid x',\theta).
The experiments use channel-averaged losses and a generalized Bayesian posterior.
What was the evaluation attacker allowed to do?
| Access |
White-box, each defense’s own predictive distribution |
| Constraint |
L_2 perturbation budget \epsilon |
| PGD |
50 likelihood-objective steps |
| PGD+ |
25 likelihood steps + 25 entropy steps |
| Selective prediction |
Increase digit entropy; decrease FashionMNIST entropy |
Where is the computational cost paid?
| Training, ms per batch iteration |
2.48 |
4.69 (MIX) |
| Prediction, ms per input |
1.60 |
2470 (onPure) |
- onPure: S=5 parameter draws, N=100 stored input samples.
- Low-dimensional regression: roughly fivefold reactive overhead.
- These are implementation measurements, not computational lower bounds.