PYQ Vault

NDA Mathematics · Formula sheet

Statistics formulas

29 formulas and 38 common traps for NDA Mathematics Statistics, grouped by subtopic.

Full notes with worked examples

Foundations + Measures of Central Tendency

Learn this subtopic in the notes

Frequency and tabulation

Total frequency

N=∑i=1kfiN = \sum_{i=1}^{k} f_i
  • kknumber of distinct values or class intervals
  • fif_ifrequency of the ii-th value/class
  • NNtotal number of observations

Class marks and class width (grouped data)

Class mark and class width

xmark=L+U2h=U−Lx_{\text{mark}} = \dfrac{L + U}{2} \qquad h = U - L

Summation notation Σ

Definition + two identities

∑i=1nxi=x1+⋯+xn,∑i=1nc=nc,∑(axi+b)=a∑xi+nb\sum_{i=1}^{n} x_i = x_1 + \cdots + x_n,\quad \sum_{i=1}^{n} c = nc,\quad \sum (a x_i + b) = a\sum x_i + nb

Arithmetic Mean (raw data)

Arithmetic Mean

xˉ=1n∑i=1nxi=x1+x2+⋯+xnn\bar{x} = \dfrac{1}{n}\sum_{i=1}^{n} x_i = \dfrac{x_1 + x_2 + \cdots + x_n}{n}
  • xˉ\bar{x}the arithmetic mean
  • xix_ithe ii-th observation
  • nnthe total number of observations

Arithmetic Mean (frequency / grouped data)

Frequency-weighted Mean

xˉ=∑fixi∑fi\bar{x} = \dfrac{\sum f_i x_i}{\sum f_i}
  • xix_ivalue (or class mark for grouped data)
  • fif_ifrequency of xix_i
  • ∑fi\sum f_itotal frequency = total observations

Linear Transformation of the Mean

Linear transformation rule

yi=a xi+b⟹yˉ=a xˉ+by_i = a\,x_i + b \quad\Longrightarrow\quad \bar{y} = a\,\bar{x} + b
  • aascale factor (multiplied)
  • bbshift (added)

Replacement and Wrong-Value Correction of the Mean

Replacement rule (single observation, n unchanged)

Mnew=M+y−xnM_{\text{new}} = M + \dfrac{y - x}{n}
  • MMoriginal mean
  • nnnumber of observations (unchanged in pure replacement)
  • xxthe value being removed (or wrongly recorded)
  • yythe value taking its place (or the correct one)

Special-Case Means — Consecutive Integers, Squares, AP, Binomial

Closed-form means for common sequences

xˉa..b=a+b2k2‾∣1n=(n+1)(2n+1)6xˉAP=a1+an2\bar{x}_{a..b} = \dfrac{a+b}{2} \qquad \overline{k^2}\big|_{1}^{n} = \dfrac{(n+1)(2n+1)}{6} \qquad \bar{x}_{\text{AP}} = \dfrac{a_1 + a_n}{2}
  • a,ba, bfirst and last integer of an arithmetic run
  • nnnumber of terms (for the squares formula, the upper index)
  • a1,ana_1, a_nfirst and last term of an AP

Combined mean of two groups

M12=n1M1+n2M2n1+n2M_{12} = \dfrac{n_1 M_1 + n_2 M_2}{n_1 + n_2}
  • n1,n2n_1, n_2sizes of the two groups
  • M1,M2M_1, M_2means of the two groups
  • M12M_{12}combined mean of the pooled dataset

Median — Middle Value

Median (raw and grouped)

Raw: M={x(n+1)/2n oddxn/2+xn/2+12n evenGrouped: M=L+n2−Ff h\text{Raw: } M = \begin{cases} x_{(n+1)/2} & n \text{ odd} \\[4pt] \dfrac{x_{n/2} + x_{n/2+1}}{2} & n \text{ even} \end{cases} \qquad \text{Grouped: } M = L + \dfrac{\tfrac{n}{2} - F}{f}\,h
  • LLlower bound of the median class
  • FFcumulative frequency before the median class
  • fffrequency of the median class
  • hhclass width

Mode — Most Frequent Value

Mode (grouped data)

M0=L+f1−f02f1−f0−f2 hM_0 = L + \dfrac{f_1 - f_0}{2f_1 - f_0 - f_2}\,h
  • LLlower bound of the modal class
  • f1f_1frequency of the modal class
  • f0f_0frequency of the class before
  • f2f_2frequency of the class after
  • hhclass width

Geometric Mean (GM)

Geometric Mean

GM=x1 x2 ⋯ xnn=(∏i=1nxi)1/n\text{GM} = \sqrt[n]{x_1 \, x_2 \, \cdots \, x_n} = \left(\prod_{i=1}^{n} x_i\right)^{1/n}
  • nnnumber of observations (all positive)

Harmonic Mean (HM)

Harmonic Mean

HM=n∑i=1n1xi=n1x1+1x2+⋯+1xn\text{HM} = \dfrac{n}{\displaystyle\sum_{i=1}^{n} \dfrac{1}{x_i}} = \dfrac{n}{\dfrac{1}{x_1} + \dfrac{1}{x_2} + \cdots + \dfrac{1}{x_n}}
  • nnnumber of observations (all positive)

Sum of Deviations & Empirical Relations

Identities to memorise

∑i=1n(xi−xˉ)=0,Mode≈3 Median−2 Mean,MD≈45 SD\sum_{i=1}^{n}(x_i - \bar{x}) = 0, \quad \text{Mode} \approx 3\,\text{Median} - 2\,\text{Mean}, \quad \text{MD} \approx \tfrac{4}{5}\,\text{SD}

Common traps

Outliers move the mean — sometimes a lot

A single very large or very small value shifts the mean noticeably. If you suspect skew, ask the question whether mean or median is the right choice.

Divide by ∑fi\sum f_i, not by the number of classes

If marks are 20, 30, 30, 20 students across four classes, the divisor is 100 — not 4. This is the single most common arithmetic error on grouped-mean PYQs.

Shift moves the mean, but not the SD

Adding a constant bb shifts xˉ\bar{x} by bb but leaves the standard deviation unchanged. Multiplying by aa scales both. Don't apply the mean rule to dispersion questions.

Divide by nn, not by 1

Students often subtract x−yx - y directly from MM. The mistake: only ONE of the nn terms changed, so the shift in the average is the change in that one term divided by nn — not the full change.

Discards: work with totals nMnM, not the rule directly

When kk observations are discarded, nn itself changes. Don't try to force the single-replacement formula. Instead: original total =nM= nM, new total =(n−k)Mnew= (n-k)M_{\text{new}}, the difference is the sum of the discarded values.

AP shortcut fails for GPs and other non-uniform spacings

The mean (a1+an)/2(a_1 + a_n)/2 works only because in an AP every term sits at equal distance around the centre. For 1,2,4,8,…1, 2, 4, 8, \ldots (GP) the shortcut gives the wrong answer — you must sum properly or use the GP sum formula.

Binomial-weighted means use ∑(nk)=2n\sum \binom{n}{k} = 2^n

When asked the mean of 1,2,…,n+11, 2, \ldots, n+1 with frequencies (n0),(n1),…,(nn)\binom{n}{0}, \binom{n}{1}, \ldots, \binom{n}{n}, the denominator is 2n2^n (sum of one row of Pascal's triangle) — not the number of distinct values. Use ∑k(nk)=n⋅2n−1\sum k \binom{n}{k} = n \cdot 2^{n-1} for the numerator.

Plain average of the two means is wrong unless n1=n2n_1 = n_2

Students average M1M_1 and M2M_2 directly. That gives the correct combined mean ONLY when both groups are the same size. For unequal sizes the larger group pulls the combined mean toward its own mean — which is exactly what the weighted formula encodes.

Reverse-solve: combined + group means give the size ratio

If M12, M1, M2M_{12},\ M_1,\ M_2 are given and you need n1:n2n_1 : n_2, rearrange the formula to n1n2=M2−M12M12−M1\dfrac{n_1}{n_2} = \dfrac{M_2 - M_{12}}{M_{12} - M_1}. PYQs use this shape with concrete totals (150 students, combined 60 kg, boys 70, girls 55) to test whether you recognise it as one equation in one unknown.

Always sort before reading off the middle

The median of an unsorted list is not the middle of the original order. PYQs sometimes hand you data in random order to catch this.

Mode can be undefined or multimodal — don't force one answer

If every value occurs exactly once, there is no mode. If two values tie for highest frequency, the data is bimodal and the answer is both values. PYQs use this to test understanding.

GM is only defined for positive numbers

Zero or negative observations break the geometric mean — the product vanishes or the root becomes imaginary. If a PYQ throws a zero or negative into the set, GM is not the right measure.

Order is always AM≥GM≥HM\text{AM} \geq \text{GM} \geq \text{HM}

For any set of positive numbers, this inequality is strict unless every observation is equal. If your computed HM exceeds GM or AM, you made an arithmetic error.

GM2=AM×HM\text{GM}^2 = \text{AM} \times \text{HM} for two numbers

For exactly two positive numbers, the geometric mean is the geometric mean of the arithmetic and harmonic means: GM2=AM⋅HM\text{GM}^2 = \text{AM} \cdot \text{HM}. When a PYQ gives you two of {AM,GM,HM}\{\text{AM}, \text{GM}, \text{HM}\} for a pair (e.g. 5 HM=4 GM5\,\text{HM} = 4\,\text{GM}), use this identity to recover the third without solving for the original numbers — much faster than setting up two equations in m,nm, n.

Sum of deviations is zero only about the mean

About any other reference point cc, the sum equals ∑xi−nc=n(xˉ−c)\sum x_i - nc = n(\bar{x} - c) — non-zero unless c=xˉc = \bar{x}. PYQs often plant a non-mean reference point to test exactly this.

Empirical relation is approximate, not exact

It works for moderately skewed unimodal data. For symmetric data (mean = median = mode) it is trivially true. For multimodal or heavily skewed data it can be misleading.

Dispersion — Standard Deviation, Variance, Mean Deviation

Learn this subtopic in the notes

Mean Deviation

Mean Deviation about A

MD(A)=1n∑i=1n∣xi−A∣\text{MD}(A) = \dfrac{1}{n}\sum_{i=1}^{n}|x_i - A|
  • AAreference point (typically mean or median)
  • ∣xi−A∣|x_i - A|absolute deviation of xix_i from AA

Variance

Variance — two equivalent forms

σ2=1n∑i=1n(xi−xˉ)2=∑xi2n−xˉ2\sigma^2 = \dfrac{1}{n}\sum_{i=1}^{n}(x_i - \bar{x})^2 = \dfrac{\sum x_i^2}{n} - \bar{x}^2
  • σ2\sigma^2variance
  • xˉ\bar{x}arithmetic mean
  • ∑xi2\sum x_i^2sum of squares of observations

Standard Deviation

σ=σ2=1n∑i=1n(xi−xˉ)2\sigma = \sqrt{\sigma^2} = \sqrt{\dfrac{1}{n}\sum_{i=1}^{n}(x_i - \bar{x})^2}
  • σ\sigmastandard deviation (always ≥0\geq 0)

Linear Transformation of SD and Variance

Variance and SD under Y = aX + b

Var(Y)=a2 Var(X)σY=∣a∣ σX\text{Var}(Y) = a^2\,\text{Var}(X) \qquad \sigma_Y = |a|\,\sigma_X
  • aascale factor
  • bbshift — irrelevant for dispersion

Special-Case Variance & SD — Natural Numbers and AP

Variance of the first n natural numbers and of an AP

σ1..n2=n2−112σAP2=d2 n2−112\sigma^2_{1..n} = \dfrac{n^2 - 1}{12} \qquad \sigma^2_{\text{AP}} = d^2\,\dfrac{n^2 - 1}{12}
  • nnnumber of terms
  • ddcommon difference of the AP (=1=1 for natural numbers)

Coefficient of Variation (CV)

Coefficient of Variation

CV=σxˉ×100%\text{CV} = \dfrac{\sigma}{\bar{x}} \times 100 \%
  • σ\sigmastandard deviation
  • xˉ\bar{x}arithmetic mean

Computational Identity & Minimum-SSE Property

Two load-bearing identities

∑xi2n=xˉ2+σ2andarg⁡min⁡a∑i(xi−a)2=xˉ\dfrac{\sum x_i^2}{n} = \bar{x}^2 + \sigma^2 \qquad \text{and} \qquad \arg\min_{a}\sum_{i}(x_i - a)^2 = \bar{x}

Common traps

Mean deviation about median is always ≤\leq about the mean

Among all reference points AA, the median minimises ∑∣xi−A∣\sum|x_i - A|. If a PYQ asks for the minimum mean deviation, the answer uses the median, not the mean.

Computational form saves time on Σxi2\Sigma x_i^2-style PYQs

When you are given ∑xi\sum x_i and ∑xi2\sum x_i^2 directly, use σ2=x2‾−xˉ2\sigma^2 = \overline{x^2} - \bar{x}^2 — not the original definition. NDA papers favour this shape because it tests whether you remember the identity.

SD and mean deviation share units; variance does not

If the data is measured in cm, SD and MD are also in cm but variance is in cm². A PYQ asks "which has the same unit as the mean?" — answer is SD or MD, not variance.

Squaring aa for variance, taking absolute value for SD

Students often write σY2=a σX2\sigma_Y^2 = a\,\sigma_X^2 (forgetting the square) or σY=a σX\sigma_Y = a\,\sigma_X (forgetting the modulus). If aa is negative, ∣a∣|a| is the correct scale for SD — SD is non-negative by definition.

It is (n2−1)/12(n^2 - 1)/12, not n2/12n^2/12 or (n2+1)/12(n^2 + 1)/12

The exact constant is n2−112\dfrac{n^2 - 1}{12} — memorise the −1-1. A frequent slip is writing n2/12n^2/12 (forgetting the −1-1) or confusing it with the mean-of-squares formula.

For an AP, only the common difference matters — not the starting value

Adding a constant shifts every term but not the spread, so {3,6,…,60}\{3, 6, \ldots, 60\} and the natural numbers {1,…,20}\{1, \ldots, 20\} scaled by 3 have the SAME variance (32×202−112=299.25)\left(3^2 \times \dfrac{20^2-1}{12} = 299.25\right). Only dd and nn enter the formula.

CV is unitless — that's the entire point

Some answers give CV with a unit attached. Wrong. CV is a percentage. If the question compares two datasets with different units (e.g. height in cm vs weight in kg), only CV makes a fair comparison — not SD.

x2‾≠(xˉ)2\overline{x^2} \neq (\bar{x})^2 — they differ by exactly σ2\sigma^2

Mean of the squares is NOT the square of the mean. Their difference is the variance: x2‾−xˉ2=σ2≥0\overline{x^2} - \bar{x}^2 = \sigma^2 \geq 0. PYQs plant this trap by asking for "mean of squares" or "M2+σ2M^2 + \sigma^2" and expecting you to recognise it as ∑xi2/n\sum x_i^2/n.

Scaling inside the deviation moves the minimiser too

For S(a)=∑(c xi−a)2S(a) = \sum (c\,x_i - a)^2, expand and minimise: the minimum is at a=c xˉa = c\,\bar{x}, NOT a=xˉa = \bar{x}. PYQs commonly use c=2c = 2 (e.g. S=∑(2xi−a)2S = \sum (2x_i - a)^2) and expect you to identify the minimiser as 2xˉ2\bar{x} — twice the mean, not the mean.

Regression and Correlation

Learn this subtopic in the notes

Correlation Coefficient and Its Properties

Correlation Coefficient and Invariance Rule

r=Cov(X,Y)σX σYr(aX+b, cY+d)=sign(ac) rXYr = \dfrac{\text{Cov}(X,Y)}{\sigma_X\,\sigma_Y} \qquad r_{(aX+b,\,cY+d)} = \text{sign}(ac)\,r_{XY}
  • Cov(X,Y)\text{Cov}(X,Y)covariance of X and Y
  • σX,σY\sigma_X,\sigma_Ystandard deviations of X and Y
  • sign(ac)\text{sign}(ac)+1+1 if a,ca, c have same sign, −1-1 otherwise

Lines of Regression

Lines of Regression (point-slope form)

y−yˉ=byx(x−xˉ)x−xˉ=bxy(y−yˉ)y - \bar{y} = b_{yx}(x - \bar{x}) \qquad x - \bar{x} = b_{xy}(y - \bar{y})
  • byxb_{yx}slope of yy on xx line =r σy/σx= r\,\sigma_y/\sigma_x
  • bxyb_{xy}slope of xx on yy line =r σx/σy= r\,\sigma_x/\sigma_y
  • (xˉ,yˉ)(\bar{x},\bar{y})the only point on BOTH regression lines

Regression Coefficients and Their Link to r

Product Identity

byx⋅bxy=r2,r=±byx bxyb_{yx} \cdot b_{xy} = r^2, \qquad r = \pm\sqrt{b_{yx}\,b_{xy}}
  • Sign of rrsame as the common sign of byxb_{yx} and bxyb_{xy}

Identifying Which Regression Line is Which

Sieve Inequality

Correct pairing satisfies byx bxy≤1; wrong pairing gives >1.\text{Correct pairing satisfies } b_{yx}\,b_{xy} \leq 1; \text{ wrong pairing gives } > 1.

Angle Between the Two Regression Lines

Angle between two lines (applied to regression)

tan⁡θ=∣m1−m21+m1 m2∣\tan\theta = \left|\dfrac{m_1 - m_2}{1 + m_1\,m_2}\right|
  • m1,m2m_1, m_2slopes of the two regression lines in the (x,y)(x,y) plane
  • θ\thetaacute angle between the lines

Common traps

rr is bounded by −1-1 and +1+1 — always

If a calculation gives ∣r∣>1|r| > 1, the arithmetic is wrong. Use this as a sanity check at the end of any correlation computation.

Shift does not change rr; only scale-with-negative-sign flips it

Adding constants to either variable is invisible to rr. Multiplying by a positive constant is also invisible. Only a negative multiplier flips the sign — and even then, the magnitude is preserved.

Both regression lines always pass through (xˉ,yˉ)(\bar{x}, \bar{y})

If a PYQ gives you two regression lines and asks for the means, solve the two equations simultaneously — their intersection is exactly (xˉ,yˉ)(\bar{x}, \bar{y}). No need to compute anything from raw data.

From raw bivariate data, compute byxb_{yx} via the Pearson form

When given nn raw paired points (e.g. four (xi,yi)(x_i, y_i) values), use the computational formula byx=n∑xiyi−∑xi∑yin∑xi2−(∑xi)2b_{yx} = \dfrac{n\sum x_i y_i - \sum x_i \sum y_i}{n\sum x_i^2 - (\sum x_i)^2} with xˉ,yˉ\bar{x}, \bar{y} read straight from the column sums. The regression line is then y−yˉ=byx(x−xˉ)y - \bar{y} = b_{yx}(x - \bar{x}). Faster than the deviation-from-mean form because it works directly off the column totals.

byx⋅bxy≤1b_{yx} \cdot b_{xy} \leq 1 is non-negotiable

If your computed product exceeds 1, you have assigned the wrong line to yy on xx. Swap the assignment and recompute — the inequality byxbxy=r2≤1b_{yx} b_{xy} = r^2 \leq 1 picks the correct pairing every time.

Both slopes share the sign of rr

You cannot have byx>0b_{yx} > 0 and bxy<0b_{xy} < 0 — if a problem seems to suggest this, the lines have been labelled wrong.

Try the inequality before doing anything else

Always compute both candidate pairings of byx⋅bxyb_{yx} \cdot b_{xy}. The pairing that satisfies the ≤1\leq 1 condition is the correct one. Don't try to reason geometrically from the slopes — the inequality is mechanical and unambiguous.

Slope of the xx-on-yy line is NOT bxyb_{xy} in the (x,y)(x,y) plane

The slope of the yy-on-xx line in the (x,y)(x,y) plane is byxb_{yx}. But the slope of the xx-on-yy line (written x=a+bxyyx = a + b_{xy} y) in the (x,y)(x,y) plane is 1/bxy1/b_{xy}, NOT bxyb_{xy}. When reading slopes off the line equation directly (solve for yy, take the coefficient of xx), you sidestep this trap.

Acute angle only — take absolute value

A negative tangent would correspond to the obtuse supplement. The formula's absolute value guarantees θ≤90∘\theta \leq 90^\circ. PYQs always ask the acute angle.

Frequency Distributions and Graphical Representation

Learn this subtopic in the notes

Histograms, Frequency Polygons & Ogives

Frequency Density (for unequal class widths)

Density=FrequencyClass width\text{Density} = \dfrac{\text{Frequency}}{\text{Class width}}
  • Class widthupper bound − lower bound of the class

Pie Charts

Sector Angle in a Pie Chart

θi=fiN×360∘∑iθi=360∘\theta_i = \dfrac{f_i}{N} \times 360^\circ \qquad \sum_i \theta_i = 360^\circ
  • fif_ifrequency / count of category ii
  • NNtotal frequency ∑fi\sum f_i

Reading Frequency Tables — Mode, Cumulative, Median

Median from a Grouped Frequency Distribution

M=L+n2−Ff hM = L + \dfrac{\tfrac{n}{2} - F}{f}\,h
  • LLlower bound of the median class
  • FFcumulative frequency BEFORE the median class
  • fffrequency of the median class
  • hhclass width

Common traps

Bar height ≠\neq frequency when class widths differ

If two classes have the same frequency but different widths, the wider class has the SHORTER bar — because density (height) divides frequency by width. Students draw bars of equal height for equal frequencies; correct histograms make AREAS equal, not heights.

All angles MUST sum to 360∘360^\circ

If your computed angles don't add to 360, you have an arithmetic error. PYQs that give angle relations ("9p=3q=2r=6s9p = 3q = 2r = 6s") use p+q+r+s=360p + q + r + s = 360 as the closing equation — without that, the system is underdetermined.

Cumulative frequency is RUNNING total, not class total

The cumulative frequency at class kk is the sum of frequencies from class 1 through class kk — not the frequency of class kk alone. Tripping on this turns every median-from-grouped-data question into nonsense.

Identify the median class FIRST, then plug into the formula

The median class is the class where the cumulative frequency first reaches or exceeds n/2n/2. Don't pick the class with the highest frequency (that's the modal class) or the middle row of the table.

More NDA Mathematics formula sheets