CUNEF University
École Centrale Nantes
Air Force Institute of Technology

Doctors compared two minimally invasive treatments for myomas to determine which:
Procedures Compared:
Doctors initially expected RFA to outperform the alternative treatment:
They entered the study already biased toward RFA.
Statistical analysis revealed no significant differences between RFA and the alternative treatment regarding length of hospital stay.
In other words, the data did not support the doctors’ initial preference for RFA.
One doctor asked:
“Can we perform a statistical adjustment to make RFA look better?”
And I replied
“What do you mean by ‘statistical adjustment’?”
Doctor:
“Look, if remove this patient with an unusually large myoma from the RFA group, the results become significant!”
How can one systematically select a minimal subset of data points to manipulate in order to alter a statistical conclusion?
In this project, we explore this question specifically within the context of Bayesian inference (aka Probabilistic Machine Learning).
Adversarial Machine Learning (AML), studies how data manipulations influence machine learning models.
However, there is a gap in the literature regarding adversarial attacks on Bayesian inference.
The types of attacks we consider involve manipulations of the training set and are known as poisoning attacks.
Posterior contains all relevant inferential information: \pi(\theta | \mathbf{X}) \propto \exp\Big(\textstyle\sum_{i=1}^n \log \pi(X_i | \theta)\Big) \pi(\theta)
Model used: y_i \sim \mathcal{N}(\beta_0 + \beta_1 x_i, \sigma^2)
Priors: \beta_0,\,\beta_1,\, \log(\sigma) \sim t(3,\,0,\,1000)

Attacker manipulates data by deleting or replicating points.
Represented by integer vector w \in \mathbb{Z}_{\geq 0}^n:
Resulting posterior (w-induced posterior): \pi_w(\theta | \mathbf{X}) = \frac{1}{Z(w)} \exp\left(\sum_{i=1}^n w_i \log \pi(X_i|\theta)\right) \pi(\theta)
Goal: alter statistical conclusions by manipulating just a few data points.
Just removing a strategically chosen 0.12 \% of the data points…
Goal: Find minimal data manipulations w \in \mathbb{Z}_{\geq 0}^n such that the resulting posterior \pi_w(\theta | \mathbf{X}) is as close as possible to a target distribution \pi_A(\theta).
Minimize forward KL divergence: \min_w \quad \text{KL}(\pi_A(\theta) \parallel \pi_w(\theta | X))
Subject to constraints: \|w - \mathbf{1}\|_1 \leq B, \quad \|w\|_\infty \leq L, \quad w \in \mathbb{Z}_{\geq 0}^n
Solve the continuous relaxation first using projected stochastic gradient-based optimization.
Then, project the solution to the integer space.
Equivalent simplified objective: -w^\top \mathbb{E}_{\pi_A(\theta)}[f_X(\theta)] + \log Z(w)
where f_X(\theta) = \log \pi(X|\theta) and \log Z(w) is the log of the normalization constant.
\nabla_w \log Z(w) - \mathbb{E}_{\pi_A}[f_X(\theta)]
\nabla_w \log Z(w) = \mathbb{E}_{\pi_w}[f_X(\theta)]
\mathbb{E}_{\pi_w}[f_X(\theta)] - \mathbb{E}_{\pi_A}[f_X(\theta)]
\nabla^2_w \log Z(w) = \text{Cov}_{\pi_w}(f_X(\theta), f_X(\theta)) \succeq 0
We use a two-stage heuristic (SGD-R2):
w_{\text{new}} \gets \Pi_{\mathcal{W}}\left(w_{\text{old}} - \gamma_t \hat{g}\right)
where \Pi_{\mathcal{W}} is the projection operator onto the feasible set
\mathcal{W} = \{w \in \mathbb{R}^n \mid w \succeq 0,\, \|w\|_{\infty}\le L,\; \|w - \mathbf{1}\|_1 \leq B\}
\hat{g} = \frac{1}{P}\sum_{i=1}^P f_X(\theta_i) - \frac{1}{Q}\sum_{j=1}^Q f_X(\theta_j)
With samples: \theta_i \sim \pi_w(\theta|\mathbf{X}) and \theta_j \sim \pi_A(\theta)
Interestingly, we do not need a closed-form expressions for neither the posterior nor the target distribution!
Solve constrained rounding problem to find integer feasible solution w^* close to relaxed solution:
w^* = \arg\min_{w' \in \mathcal{W}\cap\mathbb{Z}^n_{\geq 0}} \|w'-w\|^2_2
w^*_i = 1 + \text{sign}(w_i - 1)(\lfloor |w_i - 1| \rfloor + \alpha_i)
Adam-R2: Use Adam optimizer, scales gradients, faster practical convergence.
Second-order methods (2O-R2): Exploit Hessian information.
Start with initial feasible w = \mathbf{1}.
At each iteration, consider feasible neighbors w \pm e_i.
Select neighbor with best estimated improvement: j \gets \argmin_i \left\{-|\hat{g}_i|+\frac{1}{2}\hat{H}_{i,i}\right\}
Update: w \gets w - \text{sign}(\hat{g}_j)e_j
Infer parameters of a linear model for predicting housing prices from house characteristics (Boston Housing dataset):
y_i \sim \mathcal{N}(\alpha + x_i^\top \beta,\,\sigma^2), \quad i=1,\dots,n
Model chosen by the Honest Bayesian:
Linear model with sparsity-inducing Horseshoe prior on parameters (MCMC for inference).
Adversary’s goal: Manipulate data to steer inference about the parameter of the number of rooms (\beta_{RM}) toward 0.

KL Divergence vs Number of Manipulated Points

Posterior mean of \beta_{RM} vs Number of Manipulated Points


Attacks are precise: minimally affects other parameters
Our approach uses KL divergence to control precision, minimizing unwanted effects.
Develop methods for systematically designing adversarial targets:
Investigate scalable extensions for high-dimensional datasets and large Bayesian models.
Design manipulation stategies for situations in which the attacker has partial knowledge of the model.
Questions are welcome!
📧 roi.naveiro@cunef.edu
🌐 https://github.com/roinaveiro
