Notebooks

Estimating Functionals of Distributions

Last update: 25 Jun 2026 22:04
First version: 25 June 2026

Yet Another Inadequate Placeholder

Many interesting and important quantities in probability, information theory, causal inference, etc., are real- or vector- valued functions of probability distributions, say \( \phi(P) \) when the probability measure is \( P \). These are what we call "functionals". (Originally, as I understand it, a "functional" was a real-valued function of a function, such as a probability density function, and the name got fixed in our literature.) Often, though not always, these are integrals, \( \phi(P) = \int{f(x) dP(x)} \) for some \( f \). If we know (make up) the probability distribution \( P \), then finding \( \phi(P) \) is a calculation, though perhaps an annoying, tricky, or computationally-intractable one.

Suppose that we are doing statistics and not probability theory (i.e., we're in the real world and not the realm of mythology). Then we do not have the probability distribution \( P \), but at most a sample from it. Perhaps it's a sample of \( n \) independent draws from \( P \), say \( X_1, X_2, \ldots X_n \). How do we estimate a functional of the distribution from such data? An obvious approach is to form the empirical distribution \[ \hat{P}_n = \frac{1}{n}\sum_{i=1}^{n}{\delta_{X_i}} \] (Here, \( \delta_x \) is the distribution which puts probability 1 on the point \( x \).) We know, from the (uniform, infinite-dimensional) law of large numbers, that \( \hat{P}_n \rightarrow P \) as \( n \rightarrow \infty \), so if \( \phi \) is suitably continuous on the space of probability measures, \( \phi(\hat{P}_n) \rightarrow \phi(P) \). Alternatively, we could apply all sorts of density estimation tricks to get other estimates of the true distribution, and calculate \( \phi \) for them.

This obvious approach is called the "plug-in" method (because one plugs the estimated distribution into the functional). The drawback is that estimating distributions is hard, and errors made in estimating the distribution propagate into errors in approximating the functional. Since, in this problem, we don't really care about getting the whole distribution right, but only about getting the functional right, it'd be nice if there were ways to avoid having to estimate the distribution. (Then we can search over the low-dimensional space of possible functional values, rather than the infinite-dimensional space of possible probability distributions.) If this notebook were not an inadequate placeholder, I would now explain some of those ways.


Notebooks:   Powered by Blosxom