6  MMSE estimation

Updated

September 30, 2026

Conditional expectation is not only a way of updating a probability model; it is also the optimal estimate of an unknown random variable under mean-squared error. In this lecture, we develop that interpretation, work two non-Gaussian examples (BPSK and Poisson counts), and then specialize the theory to jointly Gaussian random vectors.

6.1 Conditional expectation as minimum mean square estimator

  1. Suppose we observe a random variable \(Y\) and wish to estimate a square-integrable random variable \(X\). Among all square-integrable functions \(h(Y)\), we seek one that minimizes the mean-squared error: \[ \EXP[(X - h(Y))^2]. \] The function \(h\) that minimizes this quantity is called the minimum mean-squared error (MMSE) estimator of \(X\) given \(Y\).

  2. The conditional expectation \(\EXP[X \mid Y]\) is the MMSE estimator of \(X\) given \(Y\).

    NoteProof

    Recall that in Example 2.23, we had shown that \[ \EXP[ (X-t)^2 ] = \EXP[ (X - μ)^2] + (t - μ)^2. \]

    By the same argument, we have \[ \EXP[ (X - h(Y))^2 \mid Y ] = \EXP[ (X - \EXP[X\mid Y])^2 \mid Y] + (h(Y) - \EXP[X\mid Y])^2 \] since \(h(Y)\) is \(Y\)-measurable.

    By smoothing property of conditional expectation, we have \[\begin{equation}\label{eq:MMSE-error} \EXP[(X - h(Y))^2] = \EXP[(X - \EXP[X \mid Y])^2] + \EXP[(h(Y) - \EXP[X \mid Y])^2]. \end{equation}\] This is minimized when \(h(Y) = \EXP[X\mid Y]\).

  3. The error \(X - \EXP[X \mid Y]\) is orthogonal to every square-integrable function of \(Y\): for any square-integrable \(g(Y)\), \[ \EXP[(X - \EXP[X \mid Y]) g(Y)] = 0. \]

    NoteProof

    By the smoothing property of conditional expectation, we have \[\begin{align*} \EXP[(X - \EXP[X \mid Y]) g(Y)] &= \EXP[\EXP[(X - \EXP[X \mid Y]) g(Y) \mid Y]] \\ &= \EXP[g(Y) \EXP[X - \EXP[X \mid Y] \mid Y]] \\ &= \EXP[g(Y) \cdot 0] = 0 \end{align*}\] where we used the fact that \(\EXP[X - \EXP[X \mid Y] \mid Y] = 0\).

  4. Conversely, if \(h(Y)\) is square-integrable and the error \(X-h(Y)\) is orthogonal to every square-integrable function of \(Y\), i.e., \(\EXP[(X - h(Y)) g(Y)] = 0\) for every square-integrable \(g(Y)\), then \(h(Y) = \EXP[X \mid Y]\) (almost surely), and hence \(h(Y)\) is the MMSE estimator.

    NoteProof

    Suppose \(\EXP[(X - h(Y)) g(Y)] = 0\) for every square-integrable function \(g(Y)\). In particular, taking \(g(Y) = h(Y) - \EXP[X \mid Y]\), we have: \[ \EXP[(X - h(Y)) (h(Y) - \EXP[X \mid Y])] = 0. \] Expanding this, we get \[\begin{align*} 0 &= \EXP[(X - h(Y)) (h(Y) - \EXP[X \mid Y])] \\ &= \EXP{ (X - \EXP[X \mid Y] - (h(Y) - \EXP[X \mid Y])) (h(Y) - \EXP[X \mid Y]) } \\ &= \EXP[(X - \EXP[X \mid Y]) (h(Y) - \EXP[X \mid Y])] - \EXP[(h(Y) - \EXP[X \mid Y])^2] \end{align*}\] The first term is zero because \(h(Y) - \EXP[X \mid Y]\) is a function of \(Y\) and, by the orthogonality property in point 3, \(X - \EXP[X \mid Y]\) is orthogonal to all functions of \(Y\). Therefore, \[ \EXP[(h(Y) - \EXP[X \mid Y])^2] = 0, \] which implies \(h(Y) = \EXP[X \mid Y]\) almost surely. By point 2, this means \(h(Y)\) is the MMSE estimator.

  5. The estimation error has zero mean: \[ \EXP[X - \EXP[X \mid Y]] = \EXP[X] - \EXP[\EXP[X \mid Y]] = 0, \] by the smoothing property.

  6. The total variance of \(X\) can be decomposed as: \[ \VAR(X) = \EXP[(X - \EXP[X])^2] = \EXP[(X - \EXP[X \mid Y])^2] + \EXP[(\EXP[X \mid Y] - \EXP[X])^2]. \] This follows by taking \(h(Y) = \EXP[X]\) (a constant function) in \(\eqref{eq:MMSE-error}\).

    Define the conditional variance by \[ \VAR(X \mid Y) := \EXP[(X-\EXP[X\mid Y])^2 \mid Y]. \] the smoothing property gives \[ \EXP[(X-\EXP[X\mid Y])^2] = \EXP[\VAR(X\mid Y)]. \] Therefore, the decomposition can also be written in its standard form: \[ \bbox[5pt,border: 1px solid] {\VAR(X) = \EXP[\VAR(X \mid Y)] + \VAR(\EXP[X \mid Y]).} \] The first term is the estimation error and the second term is the variance of the estimator.

  7. The same interpretation applies to random vectors. If \(X\in\reals^n\) is square-integrable, then \(\EXP[X\mid Y]\) minimizes \[ \EXP[\|X-h(Y)\|^2] \] over all square-integrable vector-valued functions \(h(Y)\). The error is orthogonal to every square-integrable vector-valued function \(g(Y)\) in the sense that \[ \EXP[(X-\EXP[X\mid Y])^\TRANS g(Y)]=0. \]

6.2 Soft estimation over a BPSK channel

Example 6.1 (MMSE estimation for BPSK over AWGN) Recall the BPSK over AWGN model: a random bit \(B\in\{+1,-1\}\) with \(\PR(B=+1)=\PR(B=-1)=\tfrac12\) is transmitted as the symbol \(X=aB\) with \(a>0\), and the receiver observes \[ Y=aB+W, \] where \(W\sim\mathcal N(0,σ^2)\) is independent of \(B\). Conditioned on the bit, the observation is Gaussian, \[ Y\mid(B=+1)\sim\mathcal N(a,σ^2), \qquad Y\mid(B=-1)\sim\mathcal N(-a,σ^2), \] so, with the standard normal density \(ϕ\), \[ f_{Y\mid B}(y\mid +1) =\frac{1}{σ}\,ϕ\Bigl(\frac{y-a}{σ}\Bigr), \qquad f_{Y\mid B}(y\mid -1) =\frac{1}{σ}\,ϕ\Bigl(\frac{y+a}{σ}\Bigr), \] and the marginal of \(Y\) is \[ f_Y(y) =\frac{1}{2σ}\Biggl[ ϕ\Bigl(\frac{y-a}{σ}\Bigr) +ϕ\Bigl(\frac{y+a}{σ}\Bigr) \Biggr]. \] What is the MMSE estimate of the bit given \(Y\)?

  1. By the general theory above, the MMSE estimator is the conditional expectation \(\EXP[B\mid Y]\). Since \(B\) is discrete, \[ \EXP[B\mid Y=y] =(+1)\,\PR(B=+1\mid Y=y) +(-1)\,\PR(B=-1\mid Y=y). \] Bayes’ rule and equal priors give \[ \PR(B=+1\mid Y=y) =\frac{ϕ((y-a)/σ)}{ϕ((y-a)/σ)+ϕ((y+a)/σ)}, \qquad \PR(B=-1\mid Y=y) =\frac{ϕ((y+a)/σ)}{ϕ((y-a)/σ)+ϕ((y+a)/σ)}. \] Therefore \[ \EXP[B\mid Y=y] =\frac{ϕ((y-a)/σ)-ϕ((y+a)/σ)}{ϕ((y-a)/σ)+ϕ((y+a)/σ)}. \]

  2. The Gaussian densities simplify the expression. Their ratio is \[ \frac{ϕ((y-a)/σ)}{ϕ((y+a)/σ)} =\exp\Bigl(\frac{2ay}{σ^2}\Bigr), \] so \[ \EXP[B\mid Y=y] =\frac{\exp(2ay/σ^2)-1}{\exp(2ay/σ^2)+1} =\tanh\Bigl(\frac{ay}{σ^2}\Bigr). \] Thus \[ \bbox[5pt,border: 1px solid]{ \EXP[B\mid Y]=\tanh\Bigl(\frac{aY}{σ^2}\Bigr). } \] The corresponding MMSE estimate of the transmitted amplitude is \[ \EXP[X\mid Y]=a\,\EXP[B\mid Y]=a\tanh\Bigl(\frac{aY}{σ^2}\Bigr). \] The estimate takes values in \((-1,1)\): for large \(|Y|\) it saturates near \(\pm 1\), while for \(Y\) near \(0\) it shrinks toward \(0\). Because \(B\) is discrete, the pair \((B,Y)\) is not jointly Gaussian, and \(\EXP[B\mid Y]\) is a nonlinear function of \(Y\).

Figure 6.1: MMSE estimate of the BPSK bit as a function of the received value \(y\), with \(σ=1\) and adjustable \(a/σ\).

6.3 Soft estimation of a Poisson count

Example 6.2 (MMSE estimation for two Poisson light sources) Poisson random variables are a standard model for photon counts; see Example 2.8. Suppose two independent light sources of intensities \(λ>0\) and \(μ>0\) illuminate the same device, producing independent counts \[ X\sim\text{Poisson}(λ), \qquad Y\sim\text{Poisson}(μ). \] The device observes only the total count \[ Z=X+Y. \] What is the MMSE estimate of \(X\) given \(Z\)?

  1. As shown in Example 3.15 (and again later via moment generating functions), \[ Z\sim\text{Poisson}(λ+μ). \] For integers \(0\le m\le n\), independence gives \[\begin{align*} \PR(X=m\mid Z=n) &=\frac{\PR(X=m)\PR(Y=n-m)}{\PR(Z=n)} \\ &=\frac{e^{-λ}λ^m/m!\cdot e^{-μ}μ^{n-m}/(n-m)!} {e^{-(λ+μ)}(λ+μ)^n/n!} \\ &=\binom{n}{m} \Bigl(\frac{λ}{λ+μ}\Bigr)^{m} \Bigl(\frac{μ}{λ+μ}\Bigr)^{n-m}. \end{align*}\] Thus, conditioned on \(Z=n\), the random variable \(X\) is \(\text{Binomial}\bigl(n,λ/(λ+μ)\bigr)\).

  2. By Table 2.1, the mean of a \(\text{Binomial}(n,p)\) random variable is \(np\). Taking \(p=λ/(λ+μ)\) gives \[ \EXP[X\mid Z=n] =n\,\frac{λ}{λ+μ}. \] Therefore the MMSE estimator is \[ \bbox[5pt,border: 1px solid]{ \EXP[X\mid Z] =\frac{λ}{λ+μ}\,Z. } \] Observe that the estimate is linear in the observation \(Z\).

6.4 MMSE estimation for jointly Gaussian random vectors

In the BPSK example, the MMSE estimator was a nonlinear function of the observation. We now specialize to the case where the unknown and the observation are jointly Gaussian. Then the MMSE estimator is affine, and there is an explicit formula for it in terms of means and covariances.

Suppose \(X\in\reals^n\) and \(Y\in\reals^m\) are jointly Gaussian, with \[ \MATRIX{X\\Y} \sim \mathcal N\left( \MATRIX{\mu_X\\\mu_Y}, \MATRIX{ \Sigma_{XX} & \Sigma_{XY}\\ \Sigma_{YX} & \Sigma_{YY} } \right), \] where \(\Sigma_{YX}=\Sigma_{XY}^\TRANS\) and \(\Sigma_{YY}\) is positive definite.

  1. Define the affine estimate \[ \widehat X =\mu_X+\Sigma_{XY}\Sigma_{YY}^{-1}(Y-\mu_Y) \] and the corresponding error \(E=X-\widehat X\). Its mean and its cross-covariance with \(Y\) are \[\begin{align*} \EXP[E]&=0,\\ \COV(E,Y) &=\Sigma_{XY} -\Sigma_{XY}\Sigma_{YY}^{-1}\Sigma_{YY} =0. \end{align*}\]

  2. Since \((E,Y)\) is an affine transformation of the jointly Gaussian vector \((X,Y)\), it is jointly Gaussian. For jointly Gaussian vectors, zero cross-covariance implies independence. Hence \(E\independent Y\), and therefore \[ \EXP[E\mid Y]=\EXP[E]=0. \] Using \(X=\widehat X+E\) gives \[ \bbox[5pt,border: 1px solid]{ \EXP[X\mid Y] =\mu_X+\Sigma_{XY}\Sigma_{YY}^{-1}(Y-\mu_Y). } \]

  3. The covariance of the estimation error is \[\begin{align*} \Sigma_E &=\COV\left(X-\Sigma_{XY}\Sigma_{YY}^{-1}Y\right)\\ &=\Sigma_{XX} -\Sigma_{XY}\Sigma_{YY}^{-1}\Sigma_{YX}. \end{align*}\] Because \(E\) is Gaussian and independent of \(Y\), conditioning on \(Y=y\) gives the complete conditional distribution: \[ \bbox[5pt,border: 1px solid]{ X\mid(Y=y) \sim \mathcal N\left( \mu_X+\Sigma_{XY}\Sigma_{YY}^{-1}(y-\mu_Y), \Sigma_{XX}-\Sigma_{XY}\Sigma_{YY}^{-1}\Sigma_{YX} \right). } \]

  4. For scalar jointly Gaussian random variables with \[ \COV(X,Y)=\rho\sigma_X\sigma_Y, \] the formulas reduce to \[\begin{align*} \EXP[X\mid Y] &=\mu_X+\rho\frac{\sigma_X}{\sigma_Y}(Y-\mu_Y),\\ \VAR(X\mid Y)&=\sigma_X^2(1-\rho^2). \end{align*}\]

Example 6.3 (Estimation using a sensor network) Let the unknown state be \[ X\sim\mathcal N(\mu_X,\Sigma_X). \] Sensor \(i\) observes \[ Y_i=C_iX+W_i, \qquad i=1,\ldots,N, \] where \(W_i\sim\mathcal N(0,R_i)\). Assume that the noises are mutually independent and independent of \(X\), and that \(\Sigma_X\) and every \(R_i\) are positive definite. Find the MMSE estimate of \(X\) using all the sensor measurements.

Stack the measurements and noises as \[ Y=\MATRIX{Y_1\\\vdots\\Y_N}, \qquad W=\MATRIX{W_1\\\vdots\\W_N}, \qquad C=\MATRIX{C_1\\\vdots\\C_N}. \] Then \[ Y=CX+W, \qquad R=\COV(W)=\operatorname{diag}(R_1,\ldots,R_N). \] Consequently, \[ \mu_Y=C\mu_X, \qquad \Sigma_{XY}=\Sigma_XC^\TRANS, \qquad \Sigma_{YY}=C\Sigma_XC^\TRANS+R. \] The MMSE estimate and its error covariance are therefore \[ \widehat X =\EXP[X\mid Y] =\mu_X+\Sigma_XC^\TRANS (C\Sigma_XC^\TRANS+R)^{-1}(Y-C\mu_X) \] and \[ P_N =\COV(X\mid Y) =\Sigma_X-\Sigma_XC^\TRANS (C\Sigma_XC^\TRANS+R)^{-1}C\Sigma_X. \]

Equivalently, these expressions can be written in information form as \[\begin{align*} P_N^{-1} &=\Sigma_X^{-1}+\sum_{i=1}^N C_i^\TRANS R_i^{-1}C_i,\\ \widehat X &=P_N\left( \Sigma_X^{-1}\mu_X +\sum_{i=1}^N C_i^\TRANS R_i^{-1}Y_i \right). \end{align*}\] Each additional sensor contributes the positive-semidefinite term \(C_i^\TRANS R_i^{-1}C_i\) to the information matrix. Thus, adding a sensor cannot increase the conditional covariance: \(P_N\preceq P_{N-1}\).