Secure Machine Learning

AIHUB CSIC
SUMMER SCHOOL · 2026
Secure
Machine Learning
A Bayesian perspective
Roi Naveiro · CUNEF University
September 2026

Amodei’s call to slow AI development

A typeset excerpt of a Washington Post headline reporting Amodei's call to allow time for AI safety.

What (supposedly) motivated this call?

Amodei’s stated concerns

Faster development cycle

AI helps develop its successors.

Agents crossed boundaries

A cybersecurity evaluation led to coordinated, unauthorized activity on real systems (Hugging Face Incident!)

What was the Hugging Face incident?

What was the Hugging Face incident?

The ExploitGym evaluation

OpenAI was measuring what its models could do in cybersecurity (to decide on safeguards before deployment).

  • Task. ExploitGym: exploit target program and recover flag proving success (graded by scorer).

  • Each agent works in a sandbox (no internet and no contact with other agents).

  • However, agents could install software (Artifactory), through internal package manager (with internet)

Agents crossed the sandbox boundary

198 of the 898 ExploitGym tasks had never been solved!

  • Agents never gave up.
  • Training rewarded finding answers (rather than deriving them).

Unintended uses of the one permitted service:

  • File written into Artifactory could be read by other agents (message board).

  • Manage to get Internet access throught Artifactory.

1,200 agents ended up communicating!!

Why Hugging Face?

Hugging Face: platform where researchers and companies share models, datasets, and applications.

Agents read ExploitGym paper and code, and concluded scorer also checks the vulnerability used (not just the flag!).

Tried to find benchmark’s reference solutions…

…Hugging Face is a plausible place to find it!

The official Hugging Face logo.

Seeking private benchmark data

Public repo → flag easy to reverse engineer. But necessary, not sufficient.

  • Agents needed proof of the intended vulnerability: reference exploits, patches, previous solved runs.

  • Not in the public release. Sat in gated datasets (credentials required) on Hugging Face.

A few Hugging Face write tokens, publicly exposed on the internet!!

Compromising the dataset pipeline

  • They could upload datasets to Hugging Face.

  • Hugging Face does not just store a dataset. It opens it: a server reads the description, finds the files, converts them.

  • Server runs inside production.

  • Idea: upload descriptions that lie about where the data is.

  • Agents managed to run code in production!!

The intrusion did not improve the score

  • They accessed five private datasets, the answer they came for…

  • …but OpenAI’s real scorer never checked which vulnerability was used (just the flag).

Fictional cartoon: two exhausted robot burglars open a huge vault, only to realize one was holding the flag all along.

Agents can pursue a task in unexpected ways, with harmful consequences.

AI-assisted model development

AI can contribute to the code, experiments and analysis used to develop later models.

→ Feedback Loop!

More capable models can help build the next generation.

Amodei’s concern: progress may shorten the time available for safety work.

Commercial incentives and alternative explanations

  • Anthropic and OpenAI are mid-IPO… Investors want a earnings, not a models races.

  • But staying at the frontier gets more (and more) expensive.

  • Both ask Washington for an antitrust waiver (cartel?).

  • Performance is plateauing anyway?

Risks of agent coordination

AI agents coordinations can be dangerous

But what is “AI agent coordination”?

AI agent coordination

  • Agent = language model + tools: read files, run code, query services. Choose an action, observe the result, repeat.

  • Coordination = agents exchange findings and divide the work.

  • Sometimes an orchestrator assigns tasks and decides who acts next.

  • Shared mistakes can spread through the system!

Coordination multiplies capability… and mistakes too!

Uncertainty about agent outcomes

Agent reports help us predict the consequences of possible actions. But their evidence may be unreliable.

  • Whose evidence is reliable? In which situations?

  • Independent findings, or the same source repeated?

  • An agent’s stated confidence need not be calibrated.

We need to quantify uncertainty about each action’s consequences, including task success and possible harm.

Choosing actions under uncertainty

Someone must choose (with uncertain and possibly unreliable evidence!)

  • Which agent acts next? With which permissions?

  • Act now, or run another check?

  • Escalate to a human?

  • Is that extra information worth its cost?

We need agents to make decisions according to a (coherent and auditable) framework for making decisions under uncertainty.

A Bayesian framework for decisions

Bayesian decision theory offers an answer to these questions.

  1. What to believe: update beliefs using evidence and its reliability.

  2. What to do: average each action’s consequences over those beliefs (and choose action maximizing expected utility.)

  3. Whether to ask: compare the value of new information with its cost.

Bayesian orchestration of AI agents

Title and first three author rows of Position: agentic AI orchestration should be Bayes-consistent, by Theodore Papamarkou and colleagues.

  • Position paper: the orchestration layer should be Bayes-consistent.

  • It keeps beliefs about the task, and updates them from agent messages and tool results.

  • It then picks the next action (in a Bayesian way), or pays for more information.

The coordinator can reason probabilistically even when its LLMs do not.

Bayesian decisions with manipulated evidence

If evidence is altered

This is what we will study in this talk!

  1. Foundations of Bayesian inference and decision theory.

  2. What breaks when someone alters the evidence.

  3. A Bayesian framework for protection.

Bayesian inference and decision theory

Bayesian inference
and decision theory

A decision about microcredit

Example:

A stakeholder must decide whether to expand a microcredit programme.

What does the decision depend on?

Suppose the stakeholder cares about two numbers only:

  • \beta_1: the average effect of offering credit on business profit (revenue minus expenses over two weeks)

  • c: the cost of expanding.

But \beta_1 is unknown. We need beliefs about it, and a rule for acting under them. Beliefs combine prior knowledge and observed data!

The prior beliefs

p(\theta): uncertainty before seeing the data.

  • A distribution over parameter values. Stakeholder’s belief.

  • Encodes which values are plausible, and how strongly.

A prior for the microcredit effect

Beliefs about \beta_1 before any data:

  • Gains and losses both plausible.

  • Centered at zero. Not certain of no effect.

  • Broad: room for large effects.

The actual Student-t prior on beta_1 used in the microcredit attack analysis: 3 degrees of freedom, location zero, scale 1000 in analysis profit units.

The likelihood

We collect dataset D and build a model (likelihood) that says how observations arise, given the parameters:

p(D\mid\theta)=\prod_{i=1}^{n}p(y_i\mid x_i,\theta).

Evidence in the microcredit study

Sonora, Mexico. Compartamos Banco offers small loans to women with limited access to banking.

  • 2009: credit access expanded in randomly selected communities.

  • 2011–2012: follow-up surveys. The attack analysis uses 16,560 records.

  • Recorded: credit-access assignment and reported business profit.

Panoramic photograph of Hermosillo, Sonora, at sunset, with the city and surrounding hills.

Hermosillo, Sonora

Evidence in the microcredit study

D=\{(x_i,y_i)\}_{i=1}^{n}, \qquad y_i\mid x_i,\theta\sim\mathcal N(\beta_0+\beta_1x_i,\sigma^2).

  • x_i=1: business offered credit access. x_i=0: control. y_i: reported profit.

  • Control average: \beta_0. Offered average: \beta_0+\beta_1.

  • So \beta_1 is the difference between the two group averages. The number we need.

  • \sigma: spread of individual profits around their group average.

Update prior knowledge with evidence

\underbrace{p(\theta\mid D)}_{\text{posterior}} \ \propto\ \underbrace{p(D\mid\theta)}_{\text{likelihood}} \ \underbrace{p(\theta)}_{\text{prior}}.

  • Start from what was plausible.

  • Give more weight to what explains the observations.

  • Normalize.

Colorful illustration of Thomas Bayes wearing a red carnation, with Madrid landmarks in the background.

Back to Mexico: beliefs about the effect

Original posterior treatment-effect plot for the Mexico case. The posterior is concentrated near a mean of minus 4.71 and its 95 percent credible interval spans zero.

  • Broad prior → much narrower beliefs.

  • Posterior mean effect: −4.71.

  • The 95% credible interval still includes zero.

Much less probability on a positive effect than a prior. But we are still uncertain…

What about predictions?

For a new outcome, use the posterior predictive distribution: p(\tilde y\mid x,D) =\int p(\tilde y\mid x,\theta)\,p(\theta\mid D)\,d\theta.

Gaussian-process regression on synthetic data: a curved posterior mean, observations, a 95 percent band for the latent function, and a wider 95 percent band for a new noisy observation. Both widen away from observed inputs.

Average over plausible functions and variation in a new outcome.

In practice, we work with samples

For most interesting models, the posterior is not available in closed form. We use MCMC to generate posterior draws \theta^{(1)},\ldots,\theta^{(S)}:

\mathbb E[h(\theta)\mid D] \ \approx\ \frac{1}{S}\sum_{s=1}^S h(\theta^{(s)}).

  • Mean effect: average the sampled \beta_1.

  • P(\beta_1>0\mid D): fraction of positive draws.

  • Expected utility: average an action’s consequences.

Beliefs are not a decision

To make a decision, we need:

  • Action a: what we choose.

  • State s: what matters but stays unknown (we have beliefs!).

  • Utility u(a,s): how desirable that consequence is.

Microcredit action Utility, with s=\beta_1
Expand the programme \beta_1-c
Do not expand 0

The Bayes decision rule

a^\star=\arg\max_{a\in\mathcal A} \int u(a,s)\,p(s\mid I)\,ds.

  • I: the data and other information available.

  • For each action: average utility over the uncertain state.

  • Choose the largest expected utility.

Optimal for the stated beliefs, actions, and preferences.

What does the rule choose in Mexico?

u(\mathrm{expand},\beta_1)=\beta_1-c, so

\mathbb E[u(\mathrm{expand},\beta_1)\mid D] =\mathbb E[\beta_1\mid D]-c.

With c=2 Expected utility
Expand -4.71-2=-6.71
Do not expand 0

Do not expand. The boundary sits at \mathbb E[\beta_1\mid D]=2.

Assumptions behind a Bayesian decision

  • A mean suffices for linear utility. Other losses need more of the distribution.

  • Every step assumed the model was appropriate.

  • And every step assumed the evidence was.

What if someone manipulated the evidence? Can a small perturbation affect the final decision?

How can we protect (decisions informed by) algorithms against deliberate manipulation of the evidence?

Adversarial machine learning

Adversarial machine learning is the study of the attacks on machine learning algorithms, and of the defenses against such attacks.

  1. Poisoning: alter training data to change inference and later decisions.
  2. Evasion: alter operational data to change predictions and decisions.
  3. Defenses: account for possible manipulation in inference and prediction.

We will examine each from a probabilistic perspective: beliefs, uncertainty, and decisions.

Poisoning probabilistic machine learning models

Poisoning probabilistic
machine learning models

Collaborating with gynecologists

  • A few years ago, I collaborated with doctors studying uterine fibroids, also called myomas.
  • They compared two minimally invasive treatments: radiofrequency ablation (RFA) and uterine artery embolization (UAE).
  • Which reduced fibroid size more? Which meant less blood loss and shorter hospital stays?

Doctor illustration accompanying the gynecologists anecdote

The doctors had a favorite

The doctors expected RFA to outperform UAE:

  • More effective at reducing fibroid size.
  • Less invasive for patients.
  • Shorter hospital stays and quicker recovery.

They already had a result they hoped the data would support.

Disappointing results…

  • The analysis found no significant difference in length of hospital stay.

  • For this outcome, the data did not support the doctors’ preference for RFA.

A “statistical adjustment”

One doctor asked:

“Can we perform a statistical adjustment to make RFA look better?”

“Look, if we remove this patient with an unusually large myoma from the RFA group, the results become significant!”

I said:

“I don’t think that’s how stats works…”

The issue: they were wanting to choose the evidence that lead to their preferred result….

Can data edits change a Bayesian analysis?

  • Selective data perturbations may change a statistical conclusion.

  • Their goal might be a favorable estimated effect, reassuring uncertainty, or inducing a different decision!

How little evidence must an attacker manipulate to change our conclusion?

Poisoning attacks

Poisoning probabilistic machine learning

We study how manipulating training data can alter a probabilistic model’s behavior after training.

  • Change the beliefs it learns about unknown quantities.
  • Change its predictions and reported uncertainty.
  • Ultimately, change the decisions made using those outputs.

The attacker’s control over the data

The defender fits a Bayesian model to the observations it receives.

The attacker knows

The data, model and prior, and the decision rule when targeting an action.

The attacker can

  • Delete or replicate existing observations.
  • Use at most B operations, with at most L copies per row.

Cartoon villain removes one data-record card and inserts identical copies of another.

How do deletions change a posterior?

  • Give each observation a weight: 0 deletes it, 1 keeps it, 2 includes it twice, and so on.

  • These weights change each row’s contribution to the likelihood:

\pi_w(\theta\mid D) =\frac{1}{Z(w)}\,p(\theta) \prod_{i=1}^{n}p(y_i\mid x_i,\theta)^{w_i}.

  • Z(w) normalizes the distribution. With w=\mathbf1, this is the original posterior.

The attacker

  • Let \pi_A be a desired posterior: for example, one favoring a positive treatment effect.
  • Find feasible weights that bring the defender’s posterior close to it:

\begin{aligned} \min_{w\in\mathbb Z_{\geq0}^{n}}\quad &\mathrm{KL}\!\left(\pi_A\;\Vert\;\pi_w\right)\\ \text{subject to}\quad &\|w-\mathbf1\|_1\leq B,\qquad \|w\|_\infty\leq L. \end{aligned}

  • B counts edits; L limits replication.

Optimizing a posterior discrepancy

  • Integer edits: choosing which rows to delete or replicate is combinatorial and NP-hard in general.
  • Intractable objective: the normalizer Z(w) usually has no analytic expression, so we cannot evaluate the KL exactly.
  • But the problem has useful mathematical structure! (convex objective, with gradients expressed as expectations).

We can build an effective heuristic using posterior samples!

A gradient from two expectations

Let \ell_i(\theta)=\log p(y_i\mid x_i,\theta) describe the fit of observation i.

\frac{\partial}{\partial w_i}\mathrm{KL}(\pi_A\Vert\pi_w) =\underbrace{\mathbb E_{\pi_w}[\ell_i]}_{\text{current beliefs}} -\underbrace{\mathbb E_{\pi_A}[\ell_i]}_{\text{desired beliefs}}.

  • Target fits the row better: increase its weight.
  • Current beliefs fit it better: reduce its weight.
  • Estimate the expectations with samples. No explicit Z(w) is needed.

Use stochastic gradient descent (SGD) to update data weights (rather than model parameters!) and project back into the edit budget.

From a gradient to actual edits

  1. Relax: temporarily allow fractional weights.

  2. Projected SGD: sample, estimate the gradient, update the weights, and project into the allowed budget.

  3. Round: obtain a feasible set of integer edits.

  4. Refit: evaluate the attacked posterior and the conclusion it produces.

Back to Mexico: the original conclusion

Original empirical posterior for the Mexico treatment effect

  • Recall the 16,560 observations from the Mexico microcredit study.
  • The treatment-effect posterior has mean -4.71.
  • With our illustrative expansion cost of 2, expected expansion utility is -6.71, versus zero for non-expansion.

The Bayes action is do not expand.

Data edits change the inferred effect

Original and attacked empirical treatment-effect posteriors for Mexico

  • 20 deletion operations chosen with our algorithm: about 0.12% of the sample size!

  • Posterior mean: -4.71\rightarrow+6.28. Attacked 95% interval: [0.02,12.43].

  • Expected expansion utility becomes +4.28.

New Bayes action: expand.

What should the desired posterior look like?

  • Attackers often care about a few posterior quantities, such as an effect or a decision.
  • Specifying an entire high-dimensional target distribution can be unnecessarily restrictive.
  • Could we target just some posterior moments?

Idea: MMD instead of KL

Replace KL with maximum mean discrepancy (MMD), keeping the same edit constraints:

\operatorname{MMD}(\pi_A,\pi_w;k) =\sup_{\substack{f\in\mathcal H_k\\\|f\|_{\mathcal H_k}\leq1}} \left\{\mathbb E_{\pi_A}[f(\theta)]-\mathbb E_{\pi_w}[f(\theta)]\right\}.

  • \mathcal H_k is the reproducing kernel Hilbert space (RKHS) defined by kernel k.
  • MMD finds the largest difference in expectations over these functions.

The kernel determines which differences between posteriors matter.

Targeting selected quantities with moment kernels

Choose features h(\theta)=(h_1(\theta),\ldots,h_q(\theta)) and desired means m:

\begin{aligned} k(\theta,\theta')&=h(\theta)^\top h(\theta'),\\ \operatorname{MMD}^2(\pi_A,\pi_w;k) &=\sum_{j=1}^{q}\left(\mathbb E_{\pi_w}[h_j(\theta)]-m_j\right)^2. \end{aligned}

  • Moment matching: h(\theta)=\beta_1 targets a mean; an indicator targets a probability.
  • Specify the desired expectations m, without a full target posterior.

Our algorithm uses posterior samples and stochastic gradients, then rounds and refits. No analytic posterior is needed.

A planning decision based on soil pollution

Imagine a regulator deciding whether to allow new homes near the Meuse river.

  • The regulator has soil measurements from the floodplain.
  • They want to understand how flooding and terrain relate to zinc concentration.

Is the evidence strong enough to restrict building in frequently flooded areas?

The evidence: 152 soil-sampling sites

Each observation is a floodplain topsoil sample, with its location and heavy-metal concentrations.

  • Response: zinc concentration (log scale).
  • Terrain: river distance, elevation and organic-matter content.
  • Flooding: indicator for flooding every 1–2 years, or less frequently.
  • Location: two map coordinates, in kilometers.

The question is how flooding relates to zinc after accounting for terrain and location.

A spatial model for log-zinc

\underbrace{y_i}_{\text{log-zinc}} =\underbrace{x_i^\top\beta}_{\text{measured features}} +\underbrace{v(s_i)}_{\text{spatial effect}} +\underbrace{\varepsilon_i}_{\text{remaining variation}}.

  • x_i contains flooding, river distance, elevation and organic matter.
  • A Gaussian process makes nearby spatial effects v(s_i) correlated.
  • \beta_{\mathrm{FFREQ}} describes the flooding association, conditional on the other model terms.

The clean evidence supports a restriction

For the flooding coefficient, under clean data:

\mathbb E[\beta_{\mathrm{FFREQ}}\mid D]\approx0.43, \qquad 95\%\text{ credible interval }[0.17,\,0.69].

Frequently flooded sites have higher zinc, conditional on the other model terms.

Suppose the regulator restricts building when this interval lies entirely above zero.

With the clean evidence: no permission to build in frequently flooded areas.

A developer wants permission to build

The developer’s land is in a frequently flooded area. The positive flooding association is an obstacle.

  • Goal: move the posterior mean of \beta_{\mathrm{FFREQ}} to zero.
  • Access: delete a few observations.
  • Method: our moment-matching algorithm, with h(\theta)=\beta_{\mathrm{FFREQ}} and m=0.

Can selected deletions remove the evidence behind the restriction?

Twenty deletions change the regulator’s conclusion

Original MMD manuscript map of 152 soil sites, with the 20 deleted observations circled in red.

Original flooding-coefficient panel: the gray clean posterior is positive; the blue attacked posterior is near zero and spans both signs.

Gray: clean. Blue: attacked.

Attacked mean 0.03, with 95% interval [-0.26,0.30].

Another attack changes the recorded locations

Original Bayesian Analysis figure: dashed lines connect original site circles to altered coordinate crosses, colored by flood class, with marker sizes indicating zinc concentration.

  • Here the attacker changes recorded coordinates, keeping the measurements and covariates fixed.

  • Dashed lines connect original sites to altered locations.

  • Changing distances changes spatial dependence and can weaken the fitted flooding association.

Target the decision directly

Suppose the attacker wants action a_A. Compare every alternative with it:

g_k(\theta)=u(a_k,\theta)-u(a_A,\theta).

  • If every expected gap is negative, a_A has the highest expected utility.
  • Use these gap functions as the moment-matching features introduced above.
  • Target negative expected gaps, so the attack aims across the decision boundary.

A new home in Lake County

A regulator must decide whether a new basement home needs measures to reduce radon exposure.

  • Radon is a radioactive gas. Exposure in homes increases lung-cancer risk.
  • The regulator has measurements from 919 homes in 85 Minnesota counties.
  • Each record gives the home’s county, floor level and radon measurement. County-level soil uranium supplies additional information.

Does the expected benefit justify the cost of remediation?

Counties share information in a hierarchical model

\begin{aligned} y_{ij}\mid\theta&\sim\mathcal N(\alpha_j+\beta_{\mathrm{floor}}x_{ij},\sigma_y^2),\\ \alpha_j&\sim\mathcal N(\mu_\alpha+\beta_{\mathrm{uranium}}U_j,\sigma_\alpha^2). \end{aligned}

  • y_{ij}: log-radon in home i of county j.
  • x_{ij}: basement (0) or first floor (1). U_j: county log-uranium.
  • Each county has its own level \alpha_j, informed by its observations and a shared uranium trend.

Lake County’s prediction depends partly on evidence from other counties.

The regulator’s utility

Let \tilde y be the new home’s log-radon, so its radon level is e^{\tilde y}.

Action Illustrative utility
Require remediation -2000
Take no action -700\,e^{\tilde y}
  • Remediation has a fixed cost.
  • Without it, the modeled cost increases with radon exposure.
  • Choose the action with the higher posterior expected utility.

The clean-data decision is to remediate

For the new basement home in Lake County:

Action Estimated expected cost
Require remediation 2000
Take no action 2179.9

Remediation is better by 179.9, but the decision is close to the boundary.

The developer cannot alter Lake County’s records

The developer wants no remediation for the Lake home (to avoid costs!)

  • They may delete observations only outside Lake County.
  • Lake’s own measurements remain intact.
  • They use our moment-matching algorithm.

Can changing evidence elsewhere reverse the decision here?

Six observations disappear

Original manuscript diagnostic showing six deleted homes in Carver, Cottonwood, Marshall, McLeod and St. Louis, against county cost and fitted log-radon residual.

6 of 919 homes (0.65%), from five counties. None is in Lake County.

Six deletions reverse the decision

Original Lake utility-gap posteriors with annotated clean mean 179.9 and attacked mean minus 11.4, crossing the zero decision threshold.

Mean utility gap: +179.9\ \longrightarrow\ -11.4.

The fitted Bayes action changes from remediate to take no action.

What did the attacker make us do?

Example What changed? Consequence
Mexico Treatment-effect posterior Expansion decision reverses
Meuse Flooding association and sign uncertainty Weaker evidence relevant to land use
Radon Predictive exposure cost Remediation decision reverses
  • These attacks also provide tools to audit selective-data sensitivity.
  • Protecting inference means asking which decisions depend on it.

Evasion attacks

Evasion attacks

Training data and operational data

Until now: poisoning. Change training data, which changes inference and the decisions based on it.

Now: evasion. Keep the trained model fixed. Change operational data, which changes predictions and the decisions based on them.

The attacker changes what the model sees when it is used.

A Bayesian (convolutional) neural network for image classification

Original handwritten seven

Predictive class probabilities concentrated on seven

Average over parameter uncertainty:

p(y\mid x,D)\approx\frac1S\sum_{s=1}^{S}p(y\mid x,\theta_s), \qquad\theta_s\sim p(\theta\mid D).

Attacker’s goal: slightly alter the pixels to change these predictions.

The attacker’s objective

Predictive distributions contain all full predictive information (not only point forecasts, also uncertainties).

More things can be the target of an attack!

  1. Selected quantities: a mean, a class probability, or uncertainty.
  2. The full distribution.

Choose a subtle perturbations that achieves the attacker’s goal.

\mathcal X=\{x':\|x'-x\|_2\leq\epsilon,\quad x'\text{ has valid pixel values}\}.

Targeting a predictive quantity

Choose a predictive quantity and its desired value G_A:

\min_{x'\in\mathcal X} \left\|\mathbb E_{Y\mid x',D}[g(x',Y)]-G_A\right\|_2^2.

  • g=Y: shift a predictive mean in regression.
  • g=\mathbf1\{Y=3\}, G_A=1: raise the probability of class 3.
  • g=-\log p(Y\mid x',D): target predictive entropy, the spread of the class probabilities.

We developed efficient solvers using posterior samples and projected SGD, without needing an analytic posterior.

Changing the prediction

Seven perturbed by a targeted image attack

Predictive probabilities after the attack, concentrated on class three

  • The goal was to make the network predict a 3.

Changing the uncertainty

Seven perturbed to increase predictive uncertainty

Predictive probability spread across several classes after attack

  • The image still resembles a 7, but probability is now spread across several classes!

Targeting the full predictive distribution

Choose a target distribution p_A, then bring the model’s prediction close to it:

\min_{x'\in\mathcal X} \operatorname{KL}\!\left(p_A\,\Vert\,p(\cdot\mid x',D)\right).

Our sampling-based algorithm again uses projected SGD to modify the input.

Distribution targets can also reshape uncertainty

MNIST digits: uniform target

Original UAI Figure 5a: predictive-entropy densities for MNIST digits shift upward under attacks targeting a uniform class distribution.

Predictive entropy increases.

notMNIST letters: concentrated target

Original UAI Figure 5b: predictive-entropy densities for unfamiliar notMNIST letters shift downward under attacks targeting a deterministic class distribution.

Predictive entropy decreases.

Uncertainty is very relevant for downstream tasks!

The intended rule

  • Low predictive entropy → accept.
  • High predictive entropy → review.

Would unfamiliar inputs be caught?

The attacker’s response

  • Increase entropy for familiar MNIST digits.
  • Decrease entropy for unfamiliar notMNIST letters.
  • Manipulate which cases pass the gate.

The attack changes which cases are accepted

Published selective accuracy on MNIST and notMNIST before and after attack

  • Rank cases by predictive entropy.
  • Keep the least uncertain fraction.
  • After attack, accuracy among accepted cases falls sharply.

The uncertainty score itself needs protection.

Bayesian defenses

Bayesian defenses

Probabilistic learning under attack

So far, attacks have changed inference or predictions, and therefore the decisions based on them.

Our goal is to build probabilistic algorithms that remain useful when the inputs they receive may be manipulated.

Model how attackers might behave, then integrate that knowledge into learning and prediction.

We develop defenses against evasion, assuming the training data are clean.

Learning how an attacker might behave

First investigate which manipulations change the model’s predictions or uncertainty.

What we need to model Where the knowledge comes from
Goals: induce wrong labels? increase uncertainty? Incentives and observed attacks
Access and constraints: inputs? strength of perturbation? The application and security tests
Knowledge: model details? just query access? Access information and attack experiments

Use this evidence to assign probabilities to plausible attacker behaviors.

Adversarial channels

All the collected info is integrated into a stochastic simulator (probabilistic model) of the attacker, called adversarial channel. \underbrace{x}_{\text{original input}} \quad\xrightarrow{\quad p(x'\mid x,\theta)\quad} \underbrace{x'}_{\text{received input}}

  • The channel describes which alterations are plausible, and our beliefs about how likely they are.

  • It can mix no attack, different objectives and different strengths.

  • Dependence on \theta allows an attacker to respond to the predictor.

We need to integrate this channel into the learning algorithm.

Test-time and training-time defenses

Test time: reactive defense

The defense is enacted during operations. Given a possibly attacked input, we infer the latent clean instance and average its possible predictions.

Training time: proactive defense

The defense is enacted during training. We change the learning procedure to anticipate the attacks the model may receive and learn to predict well under those alterations.

Test-time defense

Original Figure 1 from the defense manuscript: clean training data and a latent clean test input generating the response and corrupted observation

  • Shaded nodes are observed.
  • Train as usual on clean (x_i,y_i).
  • At test time, we receive x'_j; the original x_j is hidden.
  • The label y_j depends on that original input.

\phi: input-population parameters. \theta: predictor parameters.

Test-time defense

The reactive posterior predictive distribution:

p(y_j\mid \mathbf{x}'_j,\mathcal D) =\mathbb E_{(\theta,\phi)\mid \mathbf{x}'_j,\mathcal D} \left[ \mathbb E_{\mathbf{x}_j\mid \mathbf{x}'_j,\theta,\phi} \left[p(y_j\mid \mathbf{x}_j,\theta)\right] \right].

The inner expectation averages over plausible clean inputs, conditional on the parameters and the received input.

The outer expectation averages over parameter beliefs updated using that same received input.

Training-time defense

Original Figure 2 from the defense manuscript: latent corrupted training inputs connect clean inputs to observed responses, and the test response depends on the corrupted input

  • Observed training cases (x_i,y_i) are clean.
  • Hypothetical attacked inputs x'_i remain unobserved.
  • Train to predict y_i well from these possible alterations.
  • At operations, predict as usual from the received x'_j.

Predictions average over the parameter posterior learned during robust training.

Training-time defense

For each clean training case, average its prediction loss over possible attacks:

\begin{aligned} \bar\ell_i(\theta) &=\mathbb E_{x'_i\mid x_i,y_i,\theta} \left[-\log p(y_i\mid x'_i,\theta)\right],\\ p_{\mathrm{rob}}(\theta\mid D) &\propto p(\theta)\exp\!\left\{-\sum_i\bar\ell_i(\theta)\right\}. \end{aligned}

  • Fit this generalized Bayesian posterior by approximate inference.
  • At deployment, average predictions using the learned parameter distribution.

Standard defenses as limiting cases

The framework recovers familiar defenses as special cases or approximations:

  • Adversarial training
  • Randomized smoothing
  • Adversarial purification

Protection beyond the assumed attack

We train against one-step attacks, then test against stronger iterative attacks.

Without defense

PGD attack against the undefended classifier on an original digit two
Undefended predictive distribution under its own PGD attack: mass mainly on seven and eight

Probability moves to 7 and 8.

With our defense (MIX)

Separately optimized PGD attack against MIX on the same original digit two
MIX predictive distribution under its own PGD attack: most mass remains on class two

Most mass stays on 2.

Higher accuracy and better predictive probabilities

Original MNIST accuracy and NLL curves against a 50-step PGD attack

Our training-time defenses (MIX and NN50) give higher accuracy and better predictive probabilities than conventional adversarial training (AT).

Higher selective accuracy under attack

Selective accuracy against increasing attack strength on MNIST and FashionMNIST

Keep the 50% most confident predictions.

MIX and NN50 retain a more accurate set than conventional adversarial training under attack.

Decisions using the defended predictive distribution

Either defense gives a posterior predictive distribution that accounts for the modeled attacks.

Using this distribution, choose the action with the highest expected utility:

a^*=\arg\max_a\sum_y u(a,y)\, p_{\mathrm{def}}(y\mid x',D,I_A).

u(a,y) measures the consequence of taking action a when the outcome is y; I_A is our information about the attacker.

Starting point of a adversarially robust Bayesian Decision Theory.

From predictions to sequential play

What if both the world and the opponent keep changing?

From supervised learning to sequential decisions

Sequential decisions and feedback

  • Our evasion setting: the trained model is frozen; the attacker perturbs its input to change a prediction or decision; then we protect.

  • Reality is dynamic: upon observing defense, attacker changes strategy, and so on.

  • Thus, we need to update our believes about attacker, and redesign the defense…

ObserveUpdate beliefs
→
DecideChoose an action
→
ExperienceConsequences
↺

We must choose a sequence of decisions whose consequences unfold over time.

An opponent acts in parallel

Our agent

Observes the environment and chooses actions to improve its own utility.

The opponent

Acts in parallel, using its own information, preferences and learning rule.

Both actions affect the next state and what each player learns.

To choose our actions, we need to model

  • how the situation evolves
  • how the opponent behaves.

A partially observed stochastic game

Original two-player influence diagram from the sequential-play manuscript. Both actions affect the shared latent state, each player receives its own observation and utility, and dashed arrows connect successive stages.

X_t: the true, unobserved state of the environment.

D_t, A_t: our action and the opponent’s action.

Y_t, Z_t: our observation and the opponent’s private observation.

u_D, u_A: each player’s utility.

White: our agent. Gray: opponent. Dashed arrows connect successive stages.

How the world changes and what we observe

The state moves according to both players’ actions:

X_t=h(X_{t-1},D_t,A_t)+\varepsilon_{X,t}.

Each player observes a noisy version of that state:

Y_t=X_t+\varepsilon_{Y,t}, \qquad Z_t=X_t+\varepsilon_{Z,t}.

We specify the transition h and the noise distributions for the application.

After each stage, we observe A_t and Y_t. The opponent’s view Z_t remains hidden from us.

Which sequence of decisions should we choose?

A policy \pi chooses D_t using our history H_{t-1}. We want

\max_{\pi}\;\mathbb E^{\pi}\!\left[\sum_{t=1}^{T}u_D(D_t,A_t,Y_t)\right].

We control D_t. We must average over the opponent’s action and the uncertain consequences.

We need a probabilistic model for the opponent!

A behavioral model for a human opponent

Assume the opponent is human. Behavioral economics offers models of how people learn (and act) in repeated interactions.

We use Experience Weighted Attraction (EWA):

  • Actions that paid off become more attractive.
  • The player may also learn from payoffs it could have obtained with another action.
  • Past experience is gradually discounted.

These evolving attractions will determine probabilities over the next action.

What payoff does each action receive credit for?

Let j\in\mathcal A be a possible opponent action. Its modeled payoff is

u_A(D_t,j,Z_t\mid\beta),

where \beta describes the opponent’s preferences and Z_t is its private view of the state.

The payoff credited to action j is

r_{t,j}=\big[\delta+(1-\delta)\mathbf 1\{A_t=j\}\big]\, u_A(D_t,j,Z_t\mid\beta).

The chosen action receives full credit. An unchosen action receives a fraction \delta\in[0,1] of its foregone payoff.

Updating attraction and experience

\psi_{t,j} is the attraction of action j; \zeta_t is accumulated experience.

\zeta_t=\rho\zeta_{t-1}+1, \qquad \psi_{t,j}=\frac{\phi\zeta_{t-1}\psi_{t-1,j}+r_{t,j}}{\zeta_t}.

The numerator combines the previous attraction with the new payoff credit.

\phi\in[0,1] controls retention of past attractions; \rho\in[0,1] controls retention of experience.

From attractions to action probabilities

At the next stage, EWA assigns a probability to each action:

\Pr(A_{t+1}=j\mid\psi_t,\lambda) =\frac{\exp(\lambda\psi_{t,j})} {\sum_{k\in\mathcal A}\exp(\lambda\psi_{t,k})}.

More attractive actions are more likely!

EWA gives a probabilistic model of the opponent’s evolving behavior.

What is unknown about the opponent?

Collect the opponent’s parameters and evolving learning state into

\theta_t=(\phi_t,\delta_t,\rho_t,\lambda_t,\beta_t,\eta_t,\psi_t,\zeta_t).

  • \phi,\delta,\rho,\lambda: how it learns and chooses.
  • \beta: its preferences; \eta: uncertainty in its private observations.
  • \psi_t,\zeta_t: its current attractions and experience.

We start with priors over the initial physical state X_0 and opponent description \theta_0.

Updating beliefs about the opponent

Before stage t, our belief state is

b_t=p(X_{t-1},\theta_{t-1}\mid H_{t-1}).

After choosing D_t, we observe A_t,Y_t and add them to our history H_t.

Bayes’ rule updates the opponent beliefs:

p(\theta_{t-1}\mid H_t)\propto p(A_t,Y_t\mid D_t,H_{t-1},\theta_{t-1})\, p(\theta_{t-1}\mid H_{t-1}).

We also update the physical state, then advance EWA learning to obtain b_{t+1}.

The particle filter

Represent b_t by paired samples of the physical state and the opponent:

\mathcal P_t=\{(X_{t-1}^{(n)},\theta_{t-1}^{(n)})\}_{n=1}^{N}.

After choosing D_t and observing A_t,Y_t:

  1. Predict each X_t^{(n)} using the physical model and both actions.
  2. Weight by how well each particle explains the new evidence:

w_t^{(n)}\propto p(A_t\mid\theta_{t-1}^{(n)})\,p(Y_t\mid X_t^{(n)}).

  1. Resample according to the normalized weights.

  2. The resulting samples represent the next belief state:

\mathcal P_{t+1}\ \approx\ p(X_t,\theta_t\mid H_t).

From posterior samples to an action

At decision time t, \mathcal P_t contains plausible states and opponents given our history.

For each candidate action d, simulate an opponent action A_t^{(n)} and our next observation Y_t^{(n)} from each particle.

Then estimate its expected utility:

\widehat c_t(d)=\frac1N\sum_{n=1}^{N}u_D(d,A_t^{(n)},Y_t^{(n)}), \qquad D_t^{\mathrm{AMG}}\in\arg\max_d\widehat c_t(d).

This approximate myopic greedy policy chooses the best action for the current stage.

Looking ahead with simulated trajectories

For each short sequence d_t,\ldots,d_{t+H-1}, simulate possible futures from the current particles.

Each simulation evolves the physical state and the opponent’s EWA learning.

Choose the sequence with the largest estimated cumulative utility:

d^*\in\arg\max_{d_t,\ldots,d_{t+H-1}} \frac1N\sum_{n=1}^{N}\sum_{\tau=t}^{t+H-1} u_D(d_\tau,A_\tau^{(n,d)},Y_\tau^{(n,d)}).

Take only its first action, observe what happens, update the particles, and plan again. This is the H2S policy.

Learning the value of future consequences

Approximate dynamic programming (ADP) uses a learned value function instead of simulating every remaining action sequence online.

D_t^{\mathrm{ADP}}\in\arg\max_d \left\{\underbrace{\widehat c(S_t,d)}_{\text{immediate utility}} +\underbrace{\widehat V_t(\bar{\mathcal P}_t,d)}_{\text{future utility}}\right\}.

A neural network learns this value from the mean particle vector \bar{\mathcal P}_t and the proposed action.

A car must merge before its lane ends

Schematic highway merge: an automated vehicle occupies the lane that ends, while a human-driven car travels in the continuing lane. Both choose steering and acceleration.

Our car’s lane is closing. The human driver alongside it may change speed or direction while we try to merge.

Both choose steering and acceleration every 0.1 seconds. Our utility rewards staying within the road, keeping separation, and comfortable headings.

Can we merge successfully while learning how the other driver behaves?

Knowing the opponent does not complete the merge

Original myopic-clairvoyant simulation. The blue automated vehicle follows the closing road boundary and does not complete the merge. The orange opponent moves ahead. Their separation increases.

Myopic clairvoyant

Imagine and oracle tells you the opponent’s action probabilities…

… but we just optimize only the current stage. We do not merge in this episode!

Short look-ahead produces a late merge

Original horizon-simulation trajectory. The blue automated vehicle begins merging late and stays near the closing boundary, while its distance from the orange vehicle remains above the plotted threshold.

Horizon simulation

Uses the particle beliefs to anticipate a short sequence of consequences.

The car does merge, but stays close to the closing boundary and moves across late.

Planning further ahead gives a smoother merge

Original approximate-dynamic-programming simulation. The blue automated vehicle begins merging earlier and reaches the continuing lane while maintaining separation from the orange opponent.

ADP + particle filter

The car starts moving across earlier and completes a smoother merge.

Forecasting the opponent matters because it helps us plan useful actions over time.

Future work

What if the opponent is an AI agent?

Future work

Decisions made by AI agents

AI agents choose what to do: call a tool, gather evidence, delegate, or stop.

Bayesian decision theory connects these choices to beliefs about the task and the consequences of acting.

Another agent can manipulate the evidence or act in parallel, changing what happens next.

We need to account for the other agent’s behavior as well as uncertainty about the task.

Wya more complex when dealing with AI agents

In the driving example:

  • State was a compact vector of physical quantities.
  • Opponent was a human (we have behavioral models).

For AI agents

  • The state contains text, memory and tool results. Actions can themselves be messages or programs.

  • Behavioral models for AI agents?

Open questions for AI opponents

  1. Are behavioral-economics models useful for AI agents?

  2. What should a belief state retain from a conversation?
    Can we compress text without losing what changes the best action?

  3. When is a check worth its cost?

Tenure-track positions at CUNEF

Assistant Professor · Quantitative Methods

Madrid · 2026–2027 job market
Expected start: September 2027

Official CUNEF Universidad logo

  • Fields: Operations Research, Artificial Intelligence, Data Science, Computer Science and Engineering, and Robotics.
  • Profile: PhD completed or nearing completion; research in leading journals and teaching in English and Spanish.
  • Environment: international research, access to data and research funding, balanced teaching load, and highly competitive remuneration.

Send a CV and cover letter to jesusmaria.pinar@cunef.edu.

Papers and collaborators

Bayesian perspectives: Ríos Insua, Naveiro, Gallego & Poulos. JASA 2023.

Poisoning: Carreau, Naveiro & Caballero, AISTATS 2025; Caballero, Naveiro & Lunday, Bayesian Analysis. Naveiro, Caballero & Maroñas, Posterior Attraction via Moment Matching (work in progress).

Evasion and defenses: Arce, Naveiro & Ríos Insua. UAI 2025; A unifying Bayesian framework for adversarial robustness (work in progress).

Sequential play: Rafnson, Caballero, Naveiro & Marrero. Advancing Bayesian Sequential Play Against Boundedly Rational Opponents (local manuscript).

Agent coordination: Papamarkou et al., Bayesian orchestration, 2026 (position paper).

Thank you