A random variable (r.v.) \(X\) is a real-valued function defined on the sample space \(S\) of a random experiment, i.e. \(X : S \to \mathbb{R}\). Each outcome \(\omega \in S\) is assigned a real number \(X(\omega)\).
Toss a coin twice; \(S = \{HH, HT, TH, TT\}\). Define \(X\) = number of heads.
Then \(X(HH)=2,\;X(HT)=X(TH)=1,\;X(TT)=0\). \(X\) takes values 0, 1, 2.
Time (in minutes) a customer waits in a queue is a random variable \(T\). It can be any non-negative real number.
A random variable is a function \(X(w)\) with domain \(S\) and range \((-\infty,\infty)\) such that for every real number \(a\), the event \(\{w : X(w)\le a\}\in B\), where \(B\) is the \(\sigma\)-field of events on \(S\).
That condition is what makes \(P(X\le a)\) meaningful at all: it says every set the distribution function will ever be asked about is a set the probability function can measure.
Note. The triplet \((S,B,P)\) is called the probability space, where
An experiment consists of rolling a die, so \(S=\{1,2,3,4,5,6\}\). Let the random variable \(X\) be the number of points turned on the face of the die. Then \(X(1)=1,\ X(2)=2,\ X(3)=3,\ X(4)=4,\ X(5)=5,\ X(6)=6\), and \(X\) takes the values \(1,2,3,4,5,6\).
The heights of a group of persons is a random variable taking values between 4 feet and 6 feet (say).
There are two types of random variable: discrete and continuous.
If a random variable takes at most a countable number of values it is called a discrete random variable. "At most countable" is the whole of it: finitely many values, or countably infinitely many, are both allowed.
Examples. The number of heads in tossing three coins; the number of points turned on the face of a die; the number of births in a city; the number of telephone calls received in a particular day in a government office; the number of accidents occurring at a junction of a city.
If \(X\) is a discrete r.v. taking values \(x_1, x_2, \ldots\), the function
\[ p(x_i) \;=\; P(X = x_i) \]is called its probability mass function. It must satisfy:
Probability distribution of a discrete random variable. The set of pairs \(\big(x_i,\,P(x_i)\big)\) is called the probability distribution of the discrete random variable — the values paired with their probabilities, nothing more.
For a fair die, \(X\) = number on top face has PMF \(p(x) = 1/6\) for \(x = 1,2,\ldots,6\). Sum = 1 ✓.
Let \(p(x) = kx\) for \(x = 1,2,3,4\). Then \(\sum p(x) = k(1+2+3+4) = 10k = 1\) ⇒ \(k = 1/10\).
Hence \(p(1)=0.1,p(2)=0.2,p(3)=0.3,p(4)=0.4\).
If a random variable takes all the possible values between certain limits, it is called a continuous random variable.
Examples. The age of a group of persons; the weights of competitors in a game; the heights of soldiers in the country; the temperature recorded in a city; the rainfall observed in a day in a hill area.
The probability density function of a continuous random variable is denoted by \(f_X(x)\) or \(f(x)\). The function \(f(x)\) is called a probability density function if it satisfies:
Note: \(P(X = c) = 0\) for any single point \(c\) when \(X\) is continuous.
\(f(x) = kx,\; 0 \le x \le 2\). Find \(k\).
\(\int_0^2 kx\,dx = k \cdot 2 = 1 \Rightarrow k = 1/2\).
For \(f(x) = 3x^2\) on \([0,1]\): check \(\int_0^1 3x^2 dx = 1\) ✓.
\(P(X \le 0.5) = \int_0^{0.5} 3x^2 dx = x^3|_0^{0.5} = 0.125\).
The distribution function of a random variable \(X\) is denoted by \(F_X(x)\) or \(F(x)\) and is defined as
\[ F_X(x) \;=\; P(X \le x), \quad -\infty < x < \infty. \]For a discrete random variable,
\[ F_X(x) = P(X\le x) = \sum_{-\infty}^{x} P(X = x), \]and for a continuous random variable,
\[ F_X(x) = P(X\le x) = \int_{-\infty}^{x} f(x)\,dx . \]The list above is the working summary. Below is the same ground covered properly: each property stated and then proved, in the order a textbook derives them. Everything rests on one move — writing an event as a union of disjoint events and adding the probabilities.
1. If \(F(x)\) is the distribution function of a random variable \(X\) and \(a < b\), then \(P(a < X \le b) = F(b) - F(a)\).
Proof. Express the event \(X\le b\) as the union of the two disjoint events \(X\le a\) and \(a < X\le b\):
\[ X\le b = (X\le a)\cup(a < X\le b). \]Applying probabilities to both sides,
\[ P(X\le b) = P\big[(X\le a)\cup(a < X\le b)\big] = P(X\le a) + P(a < X\le b) \]since the events are disjoint. Therefore
\[ P(a < X\le b) = P(X\le b) - P(X\le a) = F(b) - F(a). \]2. \(P(a < X < b) = F(b) - F(a) - P(X=b)\).
Proof. \(P(a < X < b) = P(a < X\le b) - P(X=b) = F(b) - F(a) - P(X=b)\).
3. \(P(a \le X \le b) = P(X=a) + F(b) - F(a)\).
Proof. \(P(a \le X\le b) = P(X=a) + P(a < X\le b) = P(X=a) + F(b) - F(a)\).
4. \(P(a \le X < b) = F(b) - F(a) - P(X=b) + P(X=a)\).
Proof. \(P(a \le X < b) = P(X=a) - P(X=b) + P(a < X\le b) = P(X=a) - P(X=b) + F(b) - F(a)\).
Properties 1 to 4 are one fact seen four ways: whether each endpoint is included costs you exactly \(P(X=a)\) or \(P(X=b)\). For a continuous random variable those are zero and all four collapse into one.
5. If \(F(x)\) is the distribution function of a random variable \(X\), then \(0 \le F(x) \le 1\).
Proof. By definition \(F(x) = P(X\le x)\). Since a probability lies between 0 and 1, the distribution function lies between 0 and 1:
\[ 0 \le P(X\le x) \le 1 \;\Longrightarrow\; 0 \le F(x) \le 1 . \]6. If \(F(x)\) is the distribution function of \(X\), then \(F(x)\le F(y)\) for \(x < y\).
Proof. For \(x < y\), property 1 gives \(P(x < X\le y) = F(y)-F(x)\). Since a probability is always \(\ge 0\),
\[ P(x < X\le y)\ge 0 \;\Longrightarrow\; F(y)-F(x)\ge 0 \;\Longrightarrow\; F(y)\ge F(x), \]that is \(F(x)\le F(y)\) for \(x < y\). The distribution function never decreases.
7. If \(F(x)\) is the distribution function of \(X\), then \(F(-\infty)=0\) and \(F(\infty)=1\).
Proof. Express the whole sample space \(S\) as a countable union of disjoint events:
\[ S = \left[\bigcup_{n=1}^{\infty}(-n < X \le -n+1)\right] \cup \left[\bigcup_{n=0}^{\infty}(n < X \le n+1)\right]. \]Taking probabilities on both sides, and using that the events are disjoint together with countable additivity,
\[ 1 = P\left[\bigcup_{n=1}^{\infty}(-n < X\le -n+1)\right] + P\left[\bigcup_{n=0}^{\infty}(n < X\le n+1)\right] \] \[ = \sum_{n=1}^{\infty} P(-n < X\le -n+1) + \sum_{n=0}^{\infty} P(n < X\le n+1) \] \[ = \sum_{n=1}^{\infty}\big[F(-n+1)-F(-n)\big] + \sum_{n=0}^{\infty}\big[F(n+1)-F(n)\big] \]Both sums telescope:
\[ = \big[F(0)-F(-1)+F(-1)-F(-2)+F(-2)-F(-3)+\cdots\big] + \big[F(1)-F(0)+F(2)-F(1)+\cdots\big] \] \[ = F(\infty)-F(-\infty) \qquad\Longrightarrow\qquad F(\infty)-F(-\infty) = 1 . \tag{1} \]From property 6, since \(-\infty < \infty\),
\[ F(-\infty) \le F(\infty). \tag{2} \]From property 5,
\[ 0 \le F(-\infty) \le 1, \qquad 0 \le F(\infty) \le 1, \tag{3} \]so in particular
\[ F(-\infty)\ge 0 \quad\text{and}\quad F(\infty)\le 1 . \tag{4} \]From (2) and (3),
\[ 0 \le F(-\infty) \le F(\infty) \le 1 . \tag{5} \]From (1), (4) and (5) together, the only possibility is
\[ F(-\infty) = 0, \qquad F(\infty) = 1 . \quad\blacksquare \]Worth noticing what did the work: (1) says the two values differ by exactly 1, and (5) says both sit inside \([0,1]\). Only the endpoints satisfy both.
Toss two fair coins; \(X\) = number of heads. PMF: \(p(0)=1/4, p(1)=1/2, p(2)=1/4\).
CDF: \(F(0) = 1/4,\; F(1) = 3/4,\; F(2) = 1\). \(F\) is a step function.
For \(f(x) = 3x^2\) on [0,1]: \(F(x) = \int_0^x 3t^2 dt = x^3\) for \(0 \le x \le 1\).
\(F(x) = 0\) for \(x < 0\) and \(F(x) = 1\) for \(x > 1\).
Check: \(P(0.2 < X \le 0.5) = F(0.5) - F(0.2) = 0.125 - 0.008 = 0.117\).
If \(Y = g(X)\) is a function of an r.v. \(X\), then \(Y\) is also an r.v.
If \(X\) takes values \(x_i\) with PMF \(p(x_i)\) and \(Y = g(X)\), then \(P(Y=y) = \sum_{x_i:\,g(x_i)=y} p(x_i)\).
If \(g\) is a one-to-one differentiable function and \(Y = g(X)\):
Work through the distribution function. If \(g\) is increasing, \(Y \le y \iff X \le g^{-1}(y)\), so
\[ F_Y(y) = P\big(g(X)\le y\big) = P\big(X \le g^{-1}(y)\big) = F_X\!\big(g^{-1}(y)\big). \]
Differentiating with the chain rule gives \(f_Y(y) = f_X(g^{-1}(y))\,\dfrac{d}{dy}g^{-1}(y)\). If \(g\) is decreasing the inequality flips and the derivative is negative, so the two cases combine into the single \(|\,dx/dy\,|\) formula — the absolute value simply keeps the density positive. (Example 2 below is exactly this with \(g(x) = -\ln x\).)
\(X\) takes values \(-1,0,1,2\) each with probability \(1/4\). Let \(Y = X^2\). Then \(Y\) takes 0,1,4 with probabilities \(P(Y=0) = 1/4,\; P(Y=1) = 1/4 + 1/4 = 1/2,\; P(Y=4) = 1/4\).
\(X \sim\) Uniform(0,1) so \(f_X(x) = 1\) on (0,1). Let \(Y = -\ln X\), so \(X = e^{-Y}\), \(dX/dY = -e^{-y}\).
\(f_Y(y) = 1 \cdot e^{-y} = e^{-y}\) for \(y > 0\). Hence \(Y \sim\) Exponential(1).
For \(p(x) = x/10\), \(x = 1,2,3,4\):
\(\mu'_1 = \sum x p(x) = 1(0.1)+2(0.2)+3(0.3)+4(0.4) = 3.0\).
\(\mu'_2 = \sum x^2 p(x) = 1(0.1)+4(0.2)+9(0.3)+16(0.4) = 10.0\).
\(\sigma^2 = 10 - 9 = 1\); SD = 1.
\(f(x) = 3x^2\) on [0,1]:
\(\mu'_1 = \int_0^1 x \cdot 3x^2 dx = 3/4\).
\(\mu'_2 = \int_0^1 x^2 \cdot 3x^2 dx = 3/5\).
\(\sigma^2 = 3/5 - (3/4)^2 = 0.6 - 0.5625 = 0.0375\). SD = 0.1936.
Moments about origin (non-central moments)
\[ \mu_1' = \sum x\,P(x), \quad \mu_2' = \sum x^{2}P(x), \quad \mu_3' = \sum x^{3}P(x), \quad \mu_4' = \sum x^{4}P(x) \]The harmonic mean \(H\) is given by
\[ \frac1H = \int \frac1x\,f(x)\,dx . \]The geometric mean \(G\) is given by
\[ \log G = \int \log x\; f(x)\,dx . \]The median \(M\) is given by (in the range \(a,b\))
\[ \int_{a}^{M} f(x)\,dx = \int_{M}^{b} f(x)\,dx = \frac12, \qquad\text{i.e.}\qquad \int_{a}^{M} f(x)\,dx = \frac12 \quad\text{or}\quad \int_{M}^{b} f(x)\,dx = \frac12 . \]Mean deviation about the mean
\[ \text{M.D.} = \int |x - \text{mean}|\,f(x)\,dx . \]The \(r\)th moment about origin
\[ \mu_r' = \int x^{r} f(x)\,dx . \]The absolute value in the mean deviation is why that one integral always splits in two, at the mean: below it the bracket is negative and the sign flips. Problem 3 below does exactly that.
Symmetric PMF: \(p(-1) = p(0) = p(1) = 1/3\).
\(\mu = 0,\; \mu_2 = (1+0+1)/3 = 2/3,\; \mu_3 = 0,\; \mu_4 = (1+0+1)/3 = 2/3\).
\(\beta_1 = 0\) (symmetric); \(\beta_2 = (2/3)/(2/3)^2 = 1.5\) → platykurtic.
For \(f(x) = e^{-x},\; x \ge 0\) (Exponential(1)).
\(\mu'_r = \int_0^\infty x^r e^{-x} dx = r!\). So \(\mu'_1 = 1,\; \mu'_2 = 2,\; \mu'_3 = 6,\; \mu'_4 = 24\).
\(\mu_2 = 2 - 1 = 1;\; \mu_3 = 6 - 3(2)(1) + 2(1)^3 = 2;\; \mu_4 = 24 - 4(6)(1) + 6(2)(1) - 3 = 9\).
\(\beta_1 = 4,\; \gamma_1 = 2\) (strongly right-skewed).
\(\beta_2 = 9,\; \gamma_2 = 6\) (strongly leptokurtic, heavy tails).
Fifteen problems in the order the textbook sets them. Between them they use every formula in §6.4 at least once, which is the point of reading them as a run rather than picking one.
Source note. These are numbered 1 to 15 here. The textbook numbers them 14 to 28, because its Problems 1 to 13 work the discrete cases. The chapter's closing Exercise is unsolved in the source and is therefore not reproduced here.
For the density function \(f(x)=Cx^{2}(1-x),\ 0<x<1\), find (i) the constant \(C\), (ii) the mean.
(i) By the definition of a p.d.f., \(\int f(x)\,dx = 1\):
\[ \int_{0}^{1} Cx^{2}(1-x)\,dx = 1 \quad\Longrightarrow\quad C\int_{0}^{1}(x^{2}-x^{3})\,dx = 1 \] \[ C\left[\frac{x^{3}}{3}-\frac{x^{4}}{4}\right]_{0}^{1} = 1 \quad\Longrightarrow\quad C\left[\frac13-\frac14\right] = 1 \quad\Longrightarrow\quad C\left(\frac{1}{12}\right) = 1 \quad\therefore\quad C = 12 . \](ii)
\[ \text{Mean} = \int x f(x)\,dx = \int_{0}^{1} x\cdot 12\,x^{2}(1-x)\,dx = 12\int_{0}^{1}(x^{3}-x^{4})\,dx = 12\left[\frac{x^{4}}{4}-\frac{x^{5}}{5}\right]_{0}^{1} = 12\cdot\frac{1}{20} = \frac35 . \]A continuous random variable \(X\) has the p.d.f. \(f(x)=A+Bx,\ 0\le x\le 1\). If the mean of the distribution is \(\tfrac12\), find \(A\) and \(B\).
Condition 1 — it is a density.
\[ \int_{0}^{1}(A+Bx)\,dx = 1 \quad\Longrightarrow\quad \left[Ax+\frac{Bx^{2}}{2}\right]_{0}^{1} = 1 \quad\Longrightarrow\quad A+\frac{B}{2}=1 \quad\Longrightarrow\quad 2A+B = 2 . \tag{1} \]Condition 2 — the mean is \(\tfrac12\).
\[ \int_{0}^{1} x(A+Bx)\,dx = \frac12 \quad\Longrightarrow\quad \left[A\frac{x^{2}}{2}+\frac{Bx^{3}}{3}\right]_{0}^{1} = \frac12 \quad\Longrightarrow\quad \frac{A}{2}+\frac{B}{3}=\frac12 \quad\Longrightarrow\quad 3A+2B = 3 . \tag{2} \]Solving, \(2\times(1)\) gives \(4A+2B=4\); subtracting (2) gives \(A=1\), and then \(2+B=2\), so \(B=0\).
So the density is the uniform \(f(x)=1\) on \([0,1]\) — the only linear density on that interval with mean \(\tfrac12\) is the flat one.
The diameter of an electric cable, say \(X\), is a continuous random variable with p.d.f. \(f(x)=6x(1-x),\ 0\le x\le 1\). (i) Check that it is a p.d.f. (ii) Determine \(b\) such that \(P(X<b)=P(X>b)\). (iii) Find the mean deviation about the mean.
(i)
\[ \int_{0}^{1}6x(1-x)\,dx = \int_{0}^{1}(6x-6x^{2})\,dx = \left[\frac{6x^{2}}{2}-6\cdot\frac{x^{3}}{3}\right]_{0}^{1} = 3-2 = 1 . \]So the given \(f(x)\) is a p.d.f.
(ii) \(P(X<b)=P(X>b)\) means \(b\) is the median, so each side is \(\tfrac12\):
\[ \int_{0}^{b}6x(1-x)\,dx = \frac12 \quad\Longrightarrow\quad 6\left[\frac{b^{2}}{2}-\frac{b^{3}}{3}\right] = \frac12 \quad\Longrightarrow\quad 2(3b^{2}-2b^{3}) = 1 \] \[ \Longrightarrow\quad 4b^{3}-6b^{2}+1 = 0 \quad\Longrightarrow\quad b = -0.37,\ 1.37,\ \tfrac12 . \]Since a probability lies between 0 and 1, \(\ b=\tfrac12\).
(iii) First the mean:
\[ \text{mean} = \int_{0}^{1}x\cdot 6x(1-x)\,dx = 6\int_{0}^{1}(x^{2}-x^{3})\,dx = 6\left[\frac13-\frac14\right] = 6\cdot\frac{1}{12} = \frac12 . \]Then, splitting at the mean because of the absolute value,
\[ \text{M.D.} = \int_{0}^{1}\left|x-\tfrac12\right|6x(1-x)\,dx = 6\left\{\int_{0}^{1/2}\left(\frac{1-2x}{2}\right)(x-x^{2})\,dx + \int_{1/2}^{1}\left(\frac{2x-1}{2}\right)(x-x^{2})\,dx\right\} \] \[ = 3\left\{\int_{0}^{1/2}(x-3x^{2}+2x^{3})\,dx + \int_{1/2}^{1}(3x^{2}-2x^{3}-x)\,dx\right\} \] \[ = 3\left\{\left(\frac{x^{2}}{2}-\frac{3x^{3}}{3}+\frac{2x^{4}}{4}\right)_{0}^{1/2} + \left(\frac{3x^{3}}{3}-\frac{2x^{4}}{4}-\frac{x^{2}}{2}\right)_{1/2}^{1}\right\} \] \[ = 3\left(\frac{1}{32}+\frac78-\frac{15}{32}-\frac38\right) = 3\left(\frac{-14}{32}+\frac12\right) = \frac{3}{16} . \]A continuous random variable \(X\) has the p.d.f. \(f(x)=3x^{2},\ 0\le x\le 1\). Find \(a\) and \(b\) such that (i) \(P(X\le a)=P(X>a)\), (ii) \(P(X>b)=0.05\).
(i) By the property of the median, \(P(X\le a)=P(X>a)=\tfrac12\):
\[ \int_{0}^{a}3x^{2}\,dx = \frac12 \quad\Longrightarrow\quad 3\left(\frac{x^{3}}{3}\right)_{0}^{a} = \frac12 \quad\Longrightarrow\quad a^{3} = \frac12 \quad\Longrightarrow\quad a = \sqrt[3]{\tfrac12} = 0.7937 . \](ii)
\[ \int_{b}^{1}3x^{2}\,dx = 0.05 \quad\Longrightarrow\quad 1-b^{3} = 0.05 \quad\Longrightarrow\quad b^{3} = 0.95 \quad\Longrightarrow\quad b = \sqrt[3]{0.95} = 0.9830 . \]A continuous random variable \(X\) has the p.d.f.
\[ f(x) = \begin{cases} \tfrac{1}{16}(3+x)^{2}, & -3\le x\le -1\\[2pt] \tfrac{1}{16}(6-2x^{2}), & -1\le x\le 1\\[2pt] \tfrac{1}{16}(3-x)^{2}, & 1\le x\le 3 \end{cases} \](i) Verify that the area under the curve is unity. (ii) Find the mean and variance.
(i) Integrating piece by piece,
\[ \int_{-3}^{3} f(x)\,dx = \frac{1}{16}\left\{\left(18+\frac{26}{3}-24\right)+\left(12-\frac43\right) +\left(18+\frac{26}{3}-24\right)\right\} = \frac{1}{16}\left[\frac{52}{3}-\frac43\right] = \frac{1}{16}\cdot\frac{48}{3} = 1 . \]So the given \(f(x)\) is a p.d.f.
(ii) The density is symmetric about 0, and the arithmetic bears that out:
\[ \text{Mean} = \frac{1}{16}\{(-36-20+52)+(0-0)+(36+20-52)\} = 0 . \]With the mean at 0 the variance is just \(\int x^{2}f(x)\,dx\):
\[ \text{Variance} = \frac{1}{16}\left\{\left(78+\frac{242}{5}-120\right)+\left(4-\frac45\right) +\left(78+\frac{242}{5}-120\right)\right\} = \frac{1}{16}\left\{\frac{480}{5}-80\right\} = \frac{1}{16}[96-80] = \frac{16}{16} = 1 . \]The printed page writes the middle line of this variance calculation with a factor \(\tfrac16\); it is \(\tfrac{1}{16}\), as the line before and the line after both have.
Let \(X\) be a continuous random variable with p.d.f.
\[ f(x) = \begin{cases} ax, & 0\le x\le 1\\[2pt] a, & 1\le x\le 2\\[2pt] -ax+3a, & 2\le x\le 3 \end{cases} \](i) Determine \(a\). (ii) Compute \(P(X\le 1.5)\).
(i) By the definition of a p.d.f.,
\[ \int_{0}^{1}ax\,dx + \int_{1}^{2}a\,dx + \int_{2}^{3}(-ax+3a)\,dx = 1 \] \[ \left[a\frac{x^{2}}{2}\right]_{0}^{1} + a\big[x\big]_{1}^{2} - a\left[\frac{x^{2}}{2}\right]_{2}^{3} + 3a\big[x\big]_{2}^{3} = 1 \] \[ \frac{a}{2}+a-\frac{5a}{2}+3a = 1 \quad\Longrightarrow\quad 4a-\frac{4a}{2} = 1 \quad\Longrightarrow\quad 2a = 1 \quad\therefore\quad a = \frac12 . \](ii) The point 1.5 falls in the flat middle piece, so the integral splits at 1:
\[ P(X\le 1.5) = \int_{0}^{1}\frac12 x\,dx + \int_{1}^{1.5}\frac12\,dx = \frac12\cdot\frac12 + \frac12(1.5-1) = \frac14+\frac14 = \frac12 . \]If a random variable \(X\) has the density function \(f(x)=\tfrac14\) for \(-2<x<2\) and 0 otherwise, find (i) \(P(X<1)\), (ii) \(P(|X|>1)\), (iii) \(P(2X+3>5)\).
\[ \text{(i)}\quad P(X<1) = \int_{-2}^{1}\frac14\,dx = \frac14\big[x\big]_{-2}^{1} = \frac34 \] \[ \text{(ii)}\quad P(|X|>1) = 1-P(|X|\le 1) = 1-P(-1\le X\le 1) = 1-\int_{-1}^{1}\frac14\,dx = 1-\frac12 = \frac12 \] \[ \text{(iii)}\quad P(2X+3>5) = P(2X>2) = P(X>1) = \int_{1}^{2}\frac14\,dx = \frac14 \]Part (iii) is the useful habit: rearrange the inequality into a statement about \(X\) first, then integrate. Nothing about the density changes.
A continuous random variable \(X\) has the p.d.f. \(f(x)=y_0(x-x^{2}),\ 0\le x\le 1\), where \(y_0\) is a constant. Find (i) the arithmetic mean, (ii) the harmonic mean, (iii) the median, (iv) the mode, and (v) show that the distribution is symmetrical.
Finding \(y_0\).
\[ \int_{0}^{1} y_0(x-x^{2})\,dx = 1 \quad\Longrightarrow\quad y_0\left[\frac{x^{2}}{2}-\frac{x^{3}}{3}\right]_{0}^{1} = 1 \quad\Longrightarrow\quad \frac{y_0}{6} = 1 \quad\Longrightarrow\quad y_0 = 6, \]so \(f(x)=6(x-x^{2}),\ 0\le x\le 1\).
(i) Arithmetic mean.
\[ \text{Mean} = \int_{0}^{1}x\cdot 6(x-x^{2})\,dx = 6\int_{0}^{1}(x^{2}-x^{3})\,dx = 6\left(\frac13-\frac14\right) = \frac{6}{12} = \frac12 . \](ii) Harmonic mean.
\[ \frac1H = \int_{0}^{1}\frac1x\cdot 6(x-x^{2})\,dx = 6\int_{0}^{1}(1-x)\,dx = 6\left[x-\frac{x^{2}}{2}\right]_{0}^{1} = 6\cdot\frac12 = 3 \quad\Longrightarrow\quad H = \frac13 . \](iii) Median.
\[ \int_{0}^{M}6(x-x^{2})\,dx = \frac12 \quad\Longrightarrow\quad 6\left[\frac{M^{2}}{2}-\frac{M^{3}}{3}\right] = \frac12 \quad\Longrightarrow\quad 4M^{3}-6M^{2}+1 = 0 \] \[ \Longrightarrow\quad M = -0.37,\ 1.37,\ \tfrac12 . \]Since the median lies between the limits 0 and 1 of \(X\), the median is \(\tfrac12\).
(iv) Mode. Setting \(f'(x)=0\):
\[ \frac{d}{dx}\big[6(x-x^{2})\big] = 0 \quad\Longrightarrow\quad 6(1-2x) = 0 \quad\Longrightarrow\quad x = \frac12, \]and \(x\) is the mode provided \(f''(x)<0\); here \(f''(x) = 6(-2) = -12 < 0\), so the mode is \(x=\tfrac12\).
(v) Symmetry about \(\tfrac12\) means \(f\big(\tfrac12 + t\big) = f\big(\tfrac12 - t\big)\), that is \(f(1-x) = f(x)\). Here \(f(1-x) = 6(1-x)\big(1-(1-x)\big) = 6x(1-x) = f(x)\), so the distribution is symmetrical about \(\tfrac12\). Parts (i), (iii) and (iv) agree with this — mean = median = mode = \(\tfrac12\) — but that equality on its own would not prove it: an asymmetric distribution can have all three equal.
The p.d.f. of a random variable \(X\) is \(f(x)=e^{-x},\ x\ge 0\). Find (i) the mean and variance, (ii) the \(r\)th moment about the origin, (iii) the distribution function.
(i) Mean.
\[ \text{Mean} = \int_{0}^{\infty}xe^{-x}\,dx = \left[x\left(\frac{e^{-x}}{-1}\right)-1\cdot(e^{-x})\right]_{0}^{\infty} = -(0-0)-(0-1) = 1 . \]Variance.
\[ \text{Variance} = \int_{0}^{\infty}x^{2}e^{-x}\,dx - (1)^{2} = \left[x^{2}\left(\frac{e^{-x}}{-1}\right)-2x(e^{-x}) +2\left(\frac{e^{-x}}{-1}\right)\right]_{0}^{\infty} - 1 = 2-1 = 1 . \](ii) \(r\)th moment about origin. The gamma integral does it in one line:
\[ \mu_r' = \int_{0}^{\infty}x^{r}e^{-x}\,dx = \int_{0}^{\infty}e^{-x}x^{(r+1)-1}\,dx = \Gamma(r+1) = r! \]which gives the mean and variance again for nothing: \(\mu_1'=1!=1\), \(\mu_2'=2!=2\), so variance \(=\mu_2'-\mu_1'^{2}=2-1=1\).
(iii) Distribution function.
\[ F(x) = P(X\le x) = \int_{0}^{x}e^{-t}\,dt = \left[\frac{e^{-t}}{-1}\right]_{0}^{x} = -(e^{-x}-1) = 1-e^{-x} . \]For a continuous random variable \(X\), \(f(x)=Kx^{2}e^{-x},\ x\ge 0\). Find (i) \(K\), (ii) the mean, (iii) the variance, (iv) the standard deviation.
(i)
\[ K\int_{0}^{\infty}x^{2}e^{-x}\,dx = 1 \quad\Longrightarrow\quad K\{-2(0-1)\} = 1 \quad\Longrightarrow\quad 2K = 1 \quad\Longrightarrow\quad K = \frac12 . \](ii)
\[ \text{Mean} = \int_{0}^{\infty}x\cdot\frac12 x^{2}e^{-x}\,dx = \frac12\int_{0}^{\infty}x^{3}e^{-x}\,dx = \frac12[-6(0-1)] = 3 . \](iii)
\[ \text{Variance} = \frac12\int_{0}^{\infty}x^{4}e^{-x}\,dx - (3)^{2} = \frac12[-24(0-1)] - 9 = 12-9 = 3 . \](iv) S.D. \(=\sqrt{\text{Variance}} = \sqrt3\).
Verify under what conditions \(f(x)=Ke^{ax},\ x>0\) is a frequency function, and find \(K\).
Solution. Since the total probability is unity,
\[ \int_{0}^{\infty}Ke^{ax}\,dx = 1 \quad\Longrightarrow\quad K\left[\frac{e^{ax}}{a}\right]_{0}^{\infty} = 1 . \]The integral can be defined only for negative values of \(a\). Let \(a=-b\) with \(b>0\):
\[ K\left[\frac{e^{-bx}}{-b}\right]_{0}^{\infty} = 1 \quad\Longrightarrow\quad \frac{-K}{b}[0-1] = 1 \quad\Longrightarrow\quad \frac{K}{b} = 1 \quad\Longrightarrow\quad K = b = -a , \]which means \(K\) is the negative of \(a\). The condition and the constant arrive together: the density exists only when \(a<0\), and then \(K=-a\).
The p.d.f. of \(X\) is \(f(x)=e^{-x},\ x\ge 0\). Find (i) the coefficient of skewness \(\beta_1\), (ii) the coefficient of kurtosis \(\beta_2\).
Solution. From Problem 9, \(\mu_r' = \Gamma(r+1) = r!\), so substituting \(r=1,2,3,4\):
\[ \mu_1' = 1! = 1, \quad \mu_2' = 2! = 2, \quad \mu_3' = 3! = 6, \quad \mu_4' = 4! = 24 . \]Now by the interrelations, the central moments:
\[ \mu_1 = 0, \qquad \mu_2 = \mu_2'-\mu_1'^{2} = 2-(1)^{2} = 1, \] \[ \mu_3 = \mu_3'-3\mu_2'\mu_1'+2\mu_1'^{3} = 6-3\times2\times1+2\times(1)^{3} = 2, \] \[ \mu_4 = \mu_4'-4\mu_3'\mu_1'+6\mu_2'\mu_1'^{2}-3\mu_1'^{4} = 24-4\times6\times1+6\times2\times1-3\times(1)^{4} = 9 . \] \[ \text{(i)}\quad \beta_1 = \frac{\mu_3^{2}}{\mu_2^{3}} = \frac{2^{2}}{1^{3}} = 4 \]Since \(\beta_1>0\), the distribution is positively skewed.
\[ \text{(ii)}\quad \beta_2 = \frac{\mu_4}{\mu_2^{2}} = \frac{9}{1^{2}} = 9 \]Since \(\beta_2>3\), the distribution is leptokurtic.
If a random variable \(X\) has the density function \(f(x)=\tfrac14\) for \(-2<x<2\) and 0 otherwise, find the first four moments about the mean.
Solution. The \(r\)th moment about the origin is
\[ \mu_r' = \int_{-2}^{2}x^{r}\cdot\frac14\,dx = \frac14\left[\frac{x^{r+1}}{r+1}\right]_{-2}^{2} = \frac{2^{\,r+1}-(-2)^{\,r+1}}{4(r+1)} . \]First four moments about origin.
\[ \mu_1' = \frac{2^{2}-(-2)^{2}}{4(2)} = 0, \qquad \mu_2' = \frac{2^{3}-(-2)^{3}}{4(3)} = \frac43, \] \[ \mu_3' = \frac{2^{4}-(-2)^{4}}{4(4)} = 0, \qquad \mu_4' = \frac{2^{5}-(-2)^{5}}{4(5)} = \frac{16}{5} . \]The odd ones vanish, which is symmetry showing up in the arithmetic.
Central moments, using the relations between central and non-central:
\[ \mu_1 = 0, \qquad \mu_2 = \mu_2'-\mu_1'^{2} = \frac43-0 = \frac43, \] \[ \mu_3 = \mu_3'-3\mu_2'\mu_1'+2\mu_1'^{3} = 0-3\left(\frac43\right)(0)+2(0) = 0, \] \[ \mu_4 = \mu_4'-4\mu_3'\mu_1'+6\mu_2'\mu_1'^{2}-3\mu_1'^{4} = \frac{16}{5}-4(0)(0)+6\left(\frac43\right)(0)-3(0) = \frac{16}{5} . \]A continuous random variable \(X\) has the p.d.f. \(f(x)=3x^{2},\ 0\le x\le 1\). Find the first four moments about the mean.
\(r\)th moment about origin.
\[ \mu_r' = \int_{0}^{1}x^{r}\cdot 3x^{2}\,dx = 3\int_{0}^{1}x^{\,r+2}\,dx = 3\left[\frac{x^{\,r+3}}{r+3}\right]_{0}^{1} = \frac{3}{r+3} . \] \[ \mu_1' = \frac34 = 0.75, \quad \mu_2' = \frac35 = 0.6, \quad \mu_3' = \frac36 = 0.5, \quad \mu_4' = \frac37 = 0.43 . \]Moments about the mean, using the relations:
\[ \mu_1 = 0, \qquad \mu_2 = 0.6-(0.75)^{2} = 0.0375, \] \[ \mu_3 = 0.5-3(0.6)(0.75)+2(0.75)^{3} = -0.00625, \] \[ \mu_4 = 0.43-4(0.5)(0.75)+6(0.6)(0.75)^{2}-3(0.75)^{4} = 0.00578 . \]A negative \(\mu_3\) says the distribution leans left, which matches a density that rises towards \(x=1\).
One caution on that last figure. Carrying \(\mu_4'=3/7\) exactly instead of the rounded \(0.43\) gives \(\mu_4 = 0.004353\), not \(0.00578\). The printed answer is not a misprint but a rounding artefact: \(\mu_4\) is the small difference of four numbers near 1 or 2, so two decimal places in \(\mu_4'\) are nowhere near enough. Both values are shown because the exam answer is the book's; the lesson is to keep fractions until the last line.
For a continuous random variable \(X\), \(f(x)=\tfrac12 x^{2}e^{-x},\ x\ge 0\). Find (i) the coefficient of skewness, (ii) the coefficient of kurtosis.
\(r\)th moment about origin.
\[ \mu_r' = \int_{0}^{\infty}x^{r}\cdot\frac12 x^{2}e^{-x}\,dx = \frac12\int_{0}^{\infty}e^{-x}x^{\,r+2}\,dx = \frac12\int_{0}^{\infty}e^{-x}x^{\,(r+3)-1}\,dx = \frac12\,\Gamma(r+3) \]by the gamma integral.
First four moments about origin (using \(\Gamma(r+1)=r!\)):
\[ \mu_1' = \frac{\Gamma(4)}{2} = \frac{3!}{2} = 3, \qquad \mu_2' = \frac{\Gamma(5)}{2} = \frac{4!}{2} = 12, \] \[ \mu_3' = \frac{\Gamma(6)}{2} = \frac{5!}{2} = 60, \qquad \mu_4' = \frac{\Gamma(7)}{2} = \frac{6!}{2} = 360 . \]Central moments (by using the relations):
\[ \mu_1 = 0, \qquad \mu_2 = \mu_2'-\mu_1'^{2} = 12-3^{2} = 3, \] \[ \mu_3 = \mu_3'-3\mu_2'\mu_1'+2\mu_1'^{3} = 60-3(12)(3)+2(3)^{3} = 6, \] \[ \mu_4 = \mu_4'-4\mu_3'\mu_1'+6\mu_2'\mu_1'^{2}-3\mu_1'^{4} = 360-4(60)(3)+6(12)(3)^{2}-3(3)^{4} = 45 . \] \[ \text{(i)}\quad \beta_1 = \frac{\mu_3^{2}}{\mu_2^{3}} = \frac{6^{2}}{3^{3}} = 1.33 \]The distribution is positively skewed.
\[ \text{(ii)}\quad \beta_2 = \frac{\mu_4}{\mu_2^{2}} = \frac{45}{3^{2}} = 5 > 3 \]Since \(\beta_2>3\), the kurtosis of the distribution is leptokurtic.
The printed page gives \(\mu_2 = 12-3^{2} = 34\). It is 3, as above — and the same page then uses \(\mu_2^{3}=27\) and \(\mu_2^{2}=9\) to reach \(\beta_1=1.33\) and \(\beta_2=5\), so only the one printed digit is wrong.