Statistics in the Knowledge Economy
Roi Naveiro
CUNEF Universidad
Knowledge Economy
“Economic system in which the production of goods and services is based principally on knowledge-intensive activities that contribute to advancement in technical and scientific innovation.”
Knowledge Economy and Statistics
Statistics is (essentially) the discipline that turns data into knowledge.
If statistics is about extracting knowledge from data, then the knowledge economy is our natural home—but are we living up to that role?
No—otherwise, this session wouldn’t exist.
Why the disconnect?
The Shift in Drivers of Statistics
Tradition – Science as driver
- Domains: astronomy, genetics, medicine.
- Goal: diffuse and evolving → to generate knowledge, explain phenomena, refine theories.
Today – Industry as driver
Domains: tech, advertising, e-commerce, autonomous systems.
Goal: explicit and measurable → maximize profit.
- Proxy goals: engagement, retention, clicks, subscriptions, etc.
Key Differences in Goals
- Scientific goals are open-ended and hard to pin down.
- Industrial goals are singular and concrete (money).
The Problematic Disconnect
Causes of the Disconnect
Lack of deep collaboration with industry.
Lack of clear metrics to optimize.
Not enough emphasis on scalability.
Not enough emphasis on open-source software and open science.
A publishing system that (sometimes) values elegance over effectiveness.
A slow, peer-reviewed journal system that is out of sync with industry timelines.
…
Online Controlled Experiments
- Airbnb, Netflix… run thousands of A/B tests daily.
- Goal: maximize engagement and revenue.
- Needs: automation, reliable metrics, scalable experimentation platforms.
- Experiments are directly tied to business objectives.
Statistical Challenges
- Traditionally: small, clean trials in agriculture, medicine, psychology.
Today (industry-driven challenges):
- Many experiments running in parallel → interference.
- Millions of users → scalability.
- Real-time sequential testing and monitoring.
- Messy data: noncompliance, missingness …
- …
Large Language Models (LLMs)
- Foundation of today’s
knowledge intelligence economy.
- Used for search, chatbots, copilots, content generation.
- Progress is driven by scale and benchmarks.
- Key issues: reliability, hallucination, robustness, evaluation.
Statistical Challenges
- Quantify (and define) uncertainty in outputs.
- How can we callibrate such uncertainty?
- How can we certify LLMs outputs for high-stakes applications?
- Enhance robustness in LLMs (adversarial + distributional).
- …
Autonomous Driving
- Safety-critical AI system with massive societal impact.
- Combines sensors, perception, prediction, decision-making.
- Needs: real-time performance, reliability, scalability, robustness.
- Failures have catastrophic consequences.
Statistical Challenges
Deep learning perception + massive simulation environments.
Inference at scale (likelihood-free).
Robustness to distribution shifts and attacks.
Safety guarantees under uncertainty.
…
Common Threads Across Domains
Research is driven by clear metrics and scalability to massive datasets.
Real-time decisions require fast, robust methods—often without tractable likelihoods.
The world changes fast: distribution shifts and non-iid data are the norm.
- Happens in OCE (interventional shifts), LLMs (adversarial prompts), ADS (adversarial image manipulation).
Robustness and adaptability are essential for modern statistical practice.
To sum up…
Specially true in statistics, where the drivers of research have shifted from science to industry.
Many statisticians in academia working on industrial problems are disconnected from practice.
We need to adapt our methods and mindset to address the challenges of the knowledge intelligence economy.
Moving Forward: Learning from CS
Metrics-driven research: benchmarks and leaderboards.
Conference system: faster iteration, closer to industry timelines.
Software & reproducibility: code is a first-class research output.
Think in more dimensions: scalability, robustness, latency, interpretability.
This creates a stronger ecosystem between academia and industry.
Moving Forward
We don’t have to become computer scientists (please!), but rather take inspiration from their practices to make our research more relevant.
Computational Advertising
- Multi-billion-dollar industry: ad auctions, targeting, personalization.
- Decisions in milliseconds across billions of users.
- Goal: maximize clicks, conversions, and revenue.
- Research is metrics-driven: CTR, auction efficiency.
New Statistical Challenges
- Traditional discrete-choice / survey models don’t scale.
- Adversarial setting: agents manipulate bids, users game the system.
- Need causal inference at scale: bias correction for targeting.
- Inference under strategic behavior and non-iid data.
Statistical challenges
- Inference methods scalable to massive data.
- Fast inference for real-time decisions.
- Likelihood-free inference.
- Robustness to distribution shifts.
- Robustness to adversarial attacks.
- Causal inference under interference.