L
LLLOS.ai
LLOS.ai
L
Class 10 Mathematics Chapter 14 of 15

Chapter 14 — Statistics

Overview

This chapter introduces Statistics as the mathematics of collecting, organising, presenting and interpreting numerical data. You will learn to distinguish between raw (ungrouped) data and grouped (frequency) data, and to summarise data using frequency distributions and class intervals. The chapter develops graphical tools — bar graphs, histograms, frequency polygons and ogives — that visually display patterns and help compare data sets. Central to the chapter are measures of central tendency: mean, median and mode. You will learn how to compute these measures for ungrouped data and for grouped (continuous) data, using direct and assumed-mean methods for the mean and the standard techniques (including formulae) to estimate the median and mode from a frequency distribution. Importance: Statistics provides simple, powerful ways to summarise large amounts of information so you can make comparisons, spot trends, and draw informed conclusions in science, economics, social studies and everyday life. The chapter emphasises interpretation — not just computation — so you can choose suitable graphical forms and the most representative measure (mean, median or mode) depending on the context.…

Learning Objectives

  • Define frequency distribution, class intervals and class boundaries for grouped data.
  • Explain the concepts of class mark (mid-point) and class width and their use in calculations.
  • Construct ungrouped and grouped frequency distributions from raw data and organize data into classes.
  • Represent data graphically using histograms and frequency polygons and interpret these graphs.
  • Plot ogives (less-than and more-than cumulative frequency curves) for grouped data.
  • Determine the median for ungrouped data by arranging data and for grouped data using ogive or formula.
  • Compute the mode for ungrouped data and for grouped data using the modal-class formula.
  • Apply the direct (raw-data) method to calculate the arithmetic mean for ungrouped and grouped data.

Topics in this chapter

8 topics · tap a topic title to jump straight to it.

📊1

Introduction to Statistics

What is Statistics? Statistics is the science of collecting, organising, presenting, analysing and interpreting numerical data to make decisions or draw conclusions. It converts raw data into useful information.

Why study Statistics? To summarise large amounts of data compactly, to compare groups, to identify patterns or trends, and to make informed decisions in fields such as education, business, health, government and science.

Basic steps in a statistical study

  • Collection of data (primary or secondary)
  • Organisation and classification (tables, frequency distributions)
  • Presentation (graphs and charts)
  • Analysis (measures of central tendency and dispersion)
  • Interpretation and conclusion

Types of data

  • Primary data: collected first-hand (surveys, experiments)
  • Secondary data: obtained from existing sources (books, reports)
  • Qualitative (categorical) vs Quantitative (numerical)
  • Discrete (countable values) vs Continuous (measured values)

Raw data and frequency distribution: When many observations exist, we group them into classes (class intervals) and record frequencies (fi). For grouped data use class boundaries and class mark (xi) = (lower limit + upper limit)/2. Cumulative frequency (cf) is the running total of frequencies up to a class.

Measures introduced in Class 10: Measures of central tendency — mean, median and mode — are used to describe the centre of a data set. For grouped data we use class marks and standard formulas (including the assumed mean method) to compute these measures.

Important points

  • Choose appropriate class width (equal width classes are usual).
  • Label axes and give a title when drawing graphs.
  • Use cumulative frequency (ogive) to read median graphically.

📌 Examples
  • Example 1 (Ungrouped data): Marks scored by 7 students: 12, 15, 12, 18, 20, 15, 12. Mean = (12+15+12+18+20+15+12)/7 = 104/7 = 14.857 ≈ 14.86. Arrange data: 12,12,12,15,15,18,20 => Median (middle value) = 15. Mode (most frequent) = 12.
  • Example 2 (Forming grouped frequency table): Raw ages: 12, 14, 15, 11, 13, 18, 16, 12, 14, 15. Group into classes 10–13 (2), 14–17 (5), 18–21 (1) etc., then record frequencies and cumulative frequencies to prepare tables and graphs.
  • Example 3 (Grouped data: mean, median, mode): Classes and frequencies: 0–9 (f=5), 10–19 (f=12), 20–29 (f=7), 30–39 (f=6). Class marks xi: 4.5, 14.5, 24.5, 34.5. Σfi xi = 575, N = 30. Mean = 575/30 = 19.17. Median class is 10–19 because cumulative frequencies 5,17,...; using l=9.5, cf(before)=5, f=12, h=10 gives Median = 9.5 + ((15-5)/12)*10 = 17.83. Modal class is 10–19 (fm=12, f1=5, f2=7): Mode = 9.5 + ((12-5)/(2*12 -5 -7))*10 = 15.33.
🧮 Formulas
  1. Mean (ungrouped): x̄ = (Σxi)/n
  2. Median (position for ungrouped): arrange data in order; if n odd, median is ((n+1)/2)th value; if n even, average of (n/2)th and (n/2 +1)th values.
  3. Mode (ungrouped): value with maximum frequency (may be more than one).
  4. Class mark (grouped): xi = (lower limit + upper limit)/2
  5. Mean (grouped): Mean = (Σfi xi)/N where xi is class mark and N = Σfi
  6. Assumed mean method: Mean = A + (Σfi di)/N where di = xi - A (A = chosen assumed mean)
📊 Visual ideas
Bar graph: For discrete or categorical data. X-axis: categories, Y-axis: frequency. Bars separated and labelled with frequencies.
Histogram: For grouped continuous data. X-axis: class intervals (use class boundaries), Y-axis: frequency. Draw contiguous bars with heights equal to frequencies. Useful to see distribution shape.
Frequency polygon: Plot points at class marks (xi, fi) and join them with straight lines. Useful to compare two distributions on same axes.
Ogive (cumulative frequency curve): Plot cumulative frequency against upper class boundaries (or class limits). Use to read median, quartiles graphically.
🔢2

Frequency Distribution

What is Frequency Distribution?
A frequency distribution is a way to organize raw data into classes or categories showing how often (frequency) each value or class occurs. It simplifies large data sets to reveal patterns (like peaks, spread) and is the first step in making statistical graphs.

Types

  • Ungrouped frequency distribution — used for discrete or small data sets; lists each distinct value with its frequency.
  • Grouped frequency distribution — used for large or continuous data; data are grouped into class intervals (e.g., 10–14, 15–19) with frequency for each class.

Important terms

  • Class limits — lower and upper values of a class (e.g., 10 and 14)
  • Class boundaries — true boundary values used for continuous data (often lower limit − 0.5 and upper limit + 0.5 when data are whole numbers)
  • Class width (size) — difference between upper and lower limits of a class
  • Class mark (mid-point) — (lower limit + upper limit) / 2
  • Frequency (f) — number of observations in that class
  • Cumulative frequency (CF) — running total of frequencies up to a class
  • Relative/percentage frequency — frequency divided by total observations (or ×100 for percent)

How to construct a grouped frequency distribution (steps)

  1. Find the range = maximum − minimum.
  2. Choose number of classes (k). Typical choice: 5–15 classes; Sturges' rule gives k ≈ 1 + 3.3 log10 N as a guideline.
  3. Compute class width ≈ range / k and round to a convenient value.
  4. Set non-overlapping class intervals of equal width, then tally and count frequencies.

When preparing tables — always label columns (class limits/intervals, class mark, frequency, cumulative frequency, relative frequency).

Uses
Frequency distributions are used to summarize exam marks, heights, incomes, ages, temperatures, etc., and are the base for histograms, frequency polygons and ogives.

📌 Examples
  • Marks of 50 students grouped into class intervals 0–9, 10–19, 20–29, ... and frequencies showing how many students fall in each range.
  • Daily high temperatures for a month grouped into intervals (e.g., 20–22°C, 23–25°C, 26–28°C) to analyze the weather pattern.
  • Income distribution in a survey: incomes grouped as 0–9999, 10000–19999, 20000–29999, ... with frequencies to study economic classes.
  • Heights of 100 students grouped into intervals like 140–144 cm, 145–149 cm, 150–154 cm to find the most common height range.
🧮 Formulas
  1. Range = Maximum value − Minimum value
  2. Class width (continuous) = upper class limit − lower class limit
  3. Class width (for inclusive integer classes) ≈ (upper limit − lower limit + 1)
  4. Class mark (mid-point) m_i = (lower limit + upper limit) / 2
  5. \[Cumulative frequency CF_k = Σ (frequencies up to class k) = Σ_{i=1..k} f_i\]
  6. Relative frequency r_i = f_i / N (N = total number of observations)
📊 Visual ideas
Histogram: Bar-like contiguous rectangles whose base spans class intervals and height equals frequency. Use for continuous/grouped data. Ensure equal-width classes for direct comparison (if widths differ, use frequency density on the vertical axis). Label axes: class intervals (x-axis) and frequency (y-axis).
Frequency polygon: Plot class marks on x-axis and corresponding frequencies on y-axis, then join points with straight lines. Useful to compare two distributions on the same graph.
Ogive (cumulative frequency graph): Plot cumulative frequency against upper class boundaries (or class marks) and join points. Use ogive to read medians, quartiles and percentiles graphically.
Bar graph: Use for ungrouped discrete data; bars are separated (gap between bars) and height equals frequency.
📈3

Graphical Representation of Data

Graphical Representation of Data is the process of converting numerical data (raw or tabulated) into visual forms so patterns, trends and comparisons become easy to see and interpret. In Class 10 Statistics the common graphs are: bar graph, histogram, frequency polygon, and ogive (cumulative frequency graph). These represent frequency distributions (ungrouped or grouped).

Key steps before graphing: collect raw data, decide if you need an ungrouped or grouped frequency distribution, create class intervals (for grouped data), compute class-marks (mid-points) and frequencies, and compute cumulative frequencies if you will draw an ogive.

Choice of graph and main features:

  • Bar graph: for categorical or discrete numerical data. Bars are separated by gaps; height ∝ frequency.
  • Histogram: for continuous grouped data. Rectangular bars touch each other; width = class width and area ∝ frequency.
  • Frequency polygon: join mid-points (class-marks) of top of histogram bars by straight lines; good for comparing distributions.
  • Ogive (less-than or more-than): plot cumulative frequency against class boundary (or upper/lower class limits) to find median, quartiles and percentiles graphically.

Interpretation: read frequencies, proportions, central tendency (mean/median/mode) and spread from graphs; use ogive to estimate median and percentiles; compare distributions using shape (symmetry, skewness) and peaks (modes).

📌 Examples
  • Marks of 30 students (ungrouped): 45, 62, 78, 34, 56, ... — construct an ungrouped frequency table, draw a bar graph of marks ranges, and identify the modal class.
  • Heights (grouped): 140-144,145-149,150-154,... with given frequencies — draw a histogram (bars touching), then construct the corresponding frequency polygon by joining class-mark points.
  • Given a grouped frequency distribution with classes and frequencies, compute cumulative frequencies and draw both 'less than' and 'more than' ogives. Use the ogives to approximate the median and quartiles.
  • Monthly sales of four products — represent data as a pie chart to show percentage contribution of each product to total sales, and as a bar graph to compare absolute sales values.
🧮 Formulas
  1. Class mark (mid-point) of class [l, u): x_m = (l + u) / 2
  2. Cumulative frequency (CF): CF_k = f_1 + f_2 + ... + f_k
  3. Mean (ungrouped): x̄ = (Σx_i) / n
  4. \[Mean (grouped\]
    \[using mid-points x_m): x̄ = (Σ f_i x_{m,i}) / (Σ f_i)\]
  5. \[Mean (grouped\]
    \[assumed mean A and class width h): x̄ = A + (Σ f_i d_i / Σ f_i) × h\]
    \[where d_i = (x_{m,i} - A)/h\]
  6. Median (ungrouped): middle value when data are ordered (or average of two middle values if n is even)
📊 Visual ideas
Histogram: x-axis = class intervals (continuous, boundaries), y-axis = frequency. Draw contiguous bars whose widths equal class widths and heights equal frequency. Ensure bars touch and use equal scales. Suitable for grouped continuous data.
Frequency polygon: plot points at (class-mark, frequency). Start and end with a zero-frequency class-mark (optional) and join points with straight lines. Useful for comparing two distributions on same axes.
Ogive (Less-than): x-axis = upper class boundaries, y-axis = cumulative frequency (less than). Plot (upper boundary, cumulative freq up to that class) and join points by smooth/straight line. To find median, draw horizontal line at N/2 and read corresponding x.
Ogive (More-than): similar but use lower class boundaries and cumulative frequencies counted from top (more than). Intersection of less-than and more-than ogives gives median.
🔢4

Arithmetic Mean (Average)

The arithmetic mean (commonly called the average) of a data set is a measure of central tendency that represents the centre or typical value of the data. For a set of numerical observations, the mean is obtained by adding all observations and dividing by the number of observations. The mean gives a single value that summarises the data but is sensitive to extreme values (outliers).

Types and methods:

  • Ungrouped data (individual observations): Mean = (Sum of observations) / (Number of observations).
  • Grouped data (frequency distribution): Use class mid-points. Mean = (Sum of frequency × mid-point) / (Total frequency).
  • Assumed mean / Step-deviation methods: Useful to simplify calculations for grouped data when class mid-points are large. These methods reduce arithmetic by using an assumed mean A and, in the step-deviation method, dividing deviations by class width h.

Important properties to remember:

  • Sum of deviations from the mean is zero: Σ(x - x̄) = 0.
  • Mean is unique for a given data set.
  • Mean is affected by extreme values (outliers).
  • Linear transformation: If y = a x + b, then mean(y) = a·mean(x) + b.

When to use mean: for numerical data where you need a measure of central value and when the data have no large outliers. For skewed data or data with outliers, median may be preferred.

📌 Examples
  • Example 1 (Ungrouped data): Find the mean of marks: 72, 65, 80, 90, 83. Sum = 72+65+80+90+83 = 390. Number of observations n = 5. Mean x̄ = 390 / 5 = 78. So average mark = 78.
  • Example 2 (Grouped data using mid-points): Classes (10–20, 20–30, 30–40, 40–50) with frequencies f = (5, 8, 12, 5). Mid-points x = (15, 25, 35, 45). Compute Σf = 30. Compute Σ(f·x) = 5·15 + 8·25 + 12·35 + 5·45 = 75 + 200 + 420 + 225 = 920. Mean x̄ = Σ(f·x) / Σf = 920 / 30 ≈ 30.67.
  • Example 3 (Assumed mean and step-deviation): Classes 100–109, 110–119, 120–129, 130–139 with frequencies 4, 6, 10, 5. Mid-points x = 104.5, 114.5, 124.5, 134.5. Take assumed mean A = 124.5 and class width h = 10. Compute u = (x − A)/h: u = (−2, −1, 0, 1). Compute f·u = 4·(−2) + 6·(−1) + 10·0 + 5·1 = −8 −6 + 0 +5 = −9. Σf = 25. Then x̄ = A + h*(Σf·u)/Σf = 124.5 + 10*(−9)/25 = 124.5 − 3.6 = 120.9 (approx).
🧮 Formulas
  1. Ungrouped data (n observations): x̄ = (Σx) / n
  2. Grouped data (with mid-points x_i and frequency f_i): x̄ = (Σ f_i x_i) / (Σ f_i)
  3. Assumed mean method: x̄ = A + (Σ f_i d_i) / (Σ f_i), where d_i = x_i − A
  4. Step-deviation method: x̄ = A + h * (Σ f_i u_i) / (Σ f_i), where u_i = (x_i − A) / h and h is class width
  5. Linear transformation: If y = a x + b then mean(y) = a·mean(x) + b
  6. Property: Σ (x_i − x̄) = 0
📊 Visual ideas
Histogram for grouped continuous data: plot class intervals on x-axis and frequency on y-axis. Mark the mean on the histogram as a vertical line to show the balancing point of the distribution.
Frequency polygon (line joining mid-points at heights equal to frequency): useful to compare distributions and visually estimate central tendency; mark the mean on the horizontal axis.
Bar diagram for discrete/ungrouped numeric data: heights represent frequencies or values; place a line at the mean value to show where the centre lies relative to bars.
Box plot (box-and-whisker): although it shows median and quartiles rather than mean, plotting the mean as a point on the box plot helps visualise the effect of skewness and outliers (difference between mean and median).
🔢5

Median

Definition: The median of a data set is the value that divides the ordered data into two equal parts — half the observations are less than or equal to it and half are greater than or equal to it. It is a measure of central tendency and is robust to outliers.

For ungrouped (individual) data:

  • Arrange the data in ascending order.
  • If the number of observations n is odd, the median is the value at position (n+1)/2.
  • If n is even, the median is the average of the two middle values (positions n/2 and n/2 + 1).

Example (ungrouped): Data: 7, 3, 5, 9, 10. Ordered: 3, 5, 7, 9, 10. n = 5 (odd), median = value at (5+1)/2 = 3rd position = 7.

For grouped frequency distributions (continuous classes): we find the median class — the class where cumulative frequency first becomes ≥ N/2 (N = total frequency) — then use linear interpolation inside that class assuming uniform distribution of observations within the class.

Grouped data formula:

Median = l + ((N/2 − c.f.) / f) × h

where

  • l = lower boundary of the median class (use true class boundary, e.g., 9.5 for class 10–19)
  • N = total frequency
  • c.f. = cumulative frequency of the class preceding the median class
  • f = frequency of the median class
  • h = width (class size) of the median class

Worked grouped example: Classes: 0–10: 5, 10–20: 8, 20–30: 12, 30–40: 5. N = 5+8+12+5 = 30. N/2 = 15. Cumulative freqs: 5, 13, 25, 30. Median class is 20–30 (cumulative becomes ≥ 15 here). l = 20, c.f. (before median class) = 13, f = 12, h = 10. Median = 20 + ((15 − 13)/12) × 10 = 20 + (2/12) × 10 ≈ 21.67.

Notes:

  • Use true class boundaries when classes are continuous (e.g., 10–20 has boundaries 9.5 and 20.5 if data are continuous).
  • Open-ended classes (like "50 and above") prevent exact median calculation unless additional information is given.
  • Median is preferred over mean when data are skewed or contain outliers because it is less affected by extreme values.
📌 Examples
  • Ungrouped (odd n): Data = [7, 3, 5, 9, 10]. Ordered = [3, 5, 7, 9, 10]. n=5 ⇒ median position=(5+1)/2=3 ⇒ median=7.
  • Ungrouped (even n): Data = [4, 1, 7, 2]. Ordered = [1, 2, 4, 7]. n=4 ⇒ middle values at positions 2 and 3 are 2 and 4 ⇒ median = (2+4)/2 = 3.
  • Grouped: Classes 0–10:5, 10–20:8, 20–30:12, 30–40:5. N=30 ⇒ N/2=15. Cumulative freqs: 5,13,25,... Median class = 20–30. l=20, c.f.=13, f=12, h=10. Median = 20 + ((15−13)/12)×10 = 21.67 (approx).
🧮 Formulas
  1. Ungrouped (position): median position = (n + 1) / 2. If n is odd take that position; if n is even take average of values at positions n/2 and n/2 + 1.
  2. Grouped (continuous classes): Median = l + ((N/2 − c.f.) / f) × h, where l = lower boundary of median class, N = total frequency, c.f. = cumulative frequency before median class, f = frequency of median class, h = class width.
  3. For a probability distribution: median m satisfies P(X ≤ m) ≥ 1/2 and P(X ≥ m) ≥ 1/2 (smallest m with cumulative probability ≥ 1/2).
📊 Visual ideas
Ogive (cumulative frequency curve): Plot cumulative frequency vs upper class boundaries. To find median, draw a horizontal line at N/2, meet the ogive, then drop vertically to the x-axis — the x-coordinate is the median. Include labels for N/2 and intersection point.
Histogram with median interpolation: Draw histogram, identify median class (where cumulative reaches N/2). Mark the median class on the histogram and annotate the interpolation step inside that class to show how the median is estimated.
Box-and-whisker plot: Useful for visualizing median as the line inside the box. It also shows quartiles and spread; good for comparing medians across groups.
Stem-and-leaf or ordered dot plot: Show raw data in order and highlight the middle value(s) directly — best for small ungrouped data sets.
🔢6

Mode

Definition: The mode of a data set is the observation(s) that occur(s) most frequently. It is a measure of central tendency representing the most common value(s).

Types and existence: A data set may be unimodal (one mode), bimodal (two modes), multimodal (more than two), or have no mode if all values occur with the same frequency.

Mode for ungrouped (raw) data: For a list of individual observations, count frequencies and pick the value(s) with the highest frequency. Example: in {4, 7, 2, 4, 9, 4, 7} the mode is 4 (occurs 3 times).

Mode for discrete frequency distribution (tables): Identify the class or value with maximum frequency; that value (or class midpoint for grouped presentation of discrete values) is the mode.

Mode for continuous grouped data (class intervals): If data are given in class intervals with frequencies, the modal class is the class interval with the highest frequency. The mode can be estimated using the formula below (step-by-step):

  1. Find the modal class (class with largest frequency).
  2. Let l = lower boundary of the modal class (use class boundary, not the discrete limit), fm = frequency of modal class, f1 = frequency of the class before modal class, f2 = frequency of the class after modal class, and h = class width (upper boundary − lower boundary).
  3. Apply the formula: Mode = l + ((fm - f1) / (2fm - f1 - f2)) × h.

Notes: For continuous grouped data you must use class boundaries (e.g., for class 10–20 boundaries are 9.5 and 20.5 if class limits are inclusive). The formula gives an estimate of the most likely location of the mode inside the modal class. If two or more classes share the maximum frequency, the distribution may be bimodal or multimodal and the formula is not directly applicable.

Interpretation and use: Mode is useful for categorical data (most common category), for identifying the most typical value in a dataset (e.g., most common shoe size, most sold product), and when the most frequent occurrence is more informative than average values. It is not affected by extreme values but may be unstable for small samples.

📌 Examples
  • Ungrouped data: Data = {4, 7, 2, 4, 9, 4, 7}. Frequencies: 2→1, 4→3, 7→2, 9→1. Mode = 4 (most frequent).
  • Grouped (continuous) data: Classes and frequencies: 0-10:5, 10-20:8, 20-30:12, 30-40:7, 40-50:3. Modal class = 20-30 (fm = 12). Here f1 = 8 (previous), f2 = 7 (next), h = 10, and l = 20 (lower boundary). Mode = l + ((fm - f1) / (2fm - f1 - f2)) × h = 20 + ((12 - 8) / (24 - 8 - 7)) × 10 = 20 + (4 / 9) × 10 ≈ 24.44.
  • Real-life: In a survey of daily transport choices (bus, car, bike, walk) if 'bus' is chosen by the largest number of respondents, then 'bus' is the modal category — useful for transport planning.
🧮 Formulas
  1. Mode for ungrouped data: value(s) with maximum frequency.
  2. Mode for grouped continuous data (estimated): Mode = l + ((fm - f1) / (2fm - f1 - f2)) × h, where l = lower boundary of modal class, fm = frequency of modal class, f1 = frequency of previous class, f2 = frequency of next class, h = class width.
  3. If two or more distinct values/classes share the highest frequency → data is bimodal or multimodal; if all frequencies equal → no mode.
📊 Visual ideas
Histogram with classes on x-axis and frequency on y-axis; highlight or shade the modal class (the tallest bar). For continuous data use class boundaries on the x-axis.
Frequency polygon (line joining midpoints of class tops); the peak indicates the modal region.
Bar chart (for categorical or discrete data) showing frequencies; the tallest bar corresponds to the mode.
Dot plot (for small ungrouped data) to visually see which value occurs most often.
🔢7

Relationship between Mean, Median and Mode

Overview: Mean, median and mode are measures of central tendency. For symmetric distributions they are equal. For skewed distributions they follow a typical order: in a positively skewed (right-skewed) distribution, Mode < Median < Mean; in a negatively skewed (left-skewed) distribution, Mean < Median < Mode.

Empirical (Pearson) relation: For many moderately skewed frequency distributions an approximate relation holds:

Mode ≈ 3 × Median − 2 × Mean

This relation comes from Pearson's coefficients of skewness and is an empirical (approximate) formula — it is useful for quick checks but is not an exact identity for every dataset.

Alternate forms:

  • Mean − Mode ≈ 3 × (Mean − Median)
  • Skewness (Pearson's second) ≈ 3(Mean − Median)/SD (where SD is standard deviation)

When to use and interpretation: Use the empirical relation when the distribution is unimodal and not highly irregular. If Mean > Median the data is right-skewed (tail to right); if Mean < Median the data is left-skewed (tail to left). The mode is the peak of the distribution and tends to be pulled opposite to the tail.

Grouped data formulas (to compute each quantity):

  • Mean (using assumed mean a): Mean = a + (Σf d)/Σf, where d = x − a for class marks x and f = frequency.
  • Median (grouped): Median = L + ((N/2 − C_f)/f_m) × h, where L = lower boundary of median class, N = total frequency, C_f = cumulative frequency before median class, f_m = frequency of median class, h = class width.
  • Mode (grouped): Mode = L + ((f_m − f_1)/(2f_m − f_1 − f_2)) × h, where f_m = frequency of modal class, f_1 and f_2 are frequencies of the preceding and succeeding classes respectively, L = lower boundary of modal class, h = class width.

Important note: The Pearson relation is a rule of thumb. For multimodal distributions, or when the distribution is highly skewed or irregular, the relation may not hold well.

📌 Examples
  • Simple ungrouped example: Data = {2, 3, 3, 4, 4, 4, 5, 6}. Mean = (2+3+3+4+4+4+5+6)/8 = 31/8 = 3.875. Median = average of 4th and 5th values = (4+4)/2 = 4. Mode = 4 (most frequent). Check relation: Mode ≈ 3×Median − 2×Mean = 3×4 − 2×3.875 = 12 − 7.75 = 4.25 (close to 4; good approximation).
  • Grouped-data example: Classes (10-20,20-30,30-40,40-50) with frequencies (5, 8, 12, 5). N=30. Median class: cumulative frequencies 5,13,25,30 so N/2=15 → median class is 30-40. Using L=30, C_f=13, f_m=12, h=10 → Median = 30 + ((15−13)/12)×10 = 30 + (2/12)×10 = 31.666.... Modal class = 30-40 (f_m=12, f_1=8, f_2=5) → Mode = 30 + ((12−8)/(2×12−8−5))×10 = 30 + (4/(24−13))×10 = 30 + (4/11)×10 ≈ 33.64. If mean computed (by class-mark method) ≈ 32.2, then Mode ≈ 3×Median − 2×Mean = 3×31.666 − 2×32.2 ≈ 95 − 64.4 = 30.6 (approximate; shows empirical nature).
  • Real-life example: Income distribution in a population is usually right-skewed (a few very high incomes pull the mean to the right). So typically Mode (most common income) &lt; Median (middle income) &lt; Mean (average income). This explains why median income is often reported when describing typical earnings.
🧮 Formulas
  1. Ungrouped mean: Mean = (Σx)/n
  2. Grouped mean (assumed mean method): Mean = a + (Σf d)/Σf, where d = x − a
  3. Ungrouped median (odd n): Median = middle value; (even n): median = average of two middle values
  4. Grouped median: Median = L + ((N/2 − C_f)/f_m) × h
  5. Ungrouped mode: most frequent value(s)
  6. Grouped mode: Mode = L + ((f_m − f_1)/(2f_m − f_1 − f_2)) × h
📊 Visual ideas
Histogram showing a right-skewed distribution: mark and label Mode (peak), Median (position dividing area in half), and Mean (pulled toward tail) on the horizontal axis to visualize Mode &lt; Median &lt; Mean.
Histogram showing a left-skewed distribution with labeled Mean, Median, Mode to show Mean &lt; Median &lt; Mode.
Overlay a frequency polygon on a histogram and draw vertical lines for Mean, Median and Mode to compare their positions.
Box plot (box-and-whisker): median is shown as the central line; a long whisker to the right indicates positive skew (mean will lie to the right of median).
🔢8

Practical Applications and Interpretation

What it means: Practical applications and interpretation in statistics means using collected data (scores, incomes, measurements, counts) to calculate summary measures and to draw meaningful conclusions for decision making. It involves choosing appropriate measures (mean, median, mode), visualizing data (graphs), checking spread (range, quartiles, standard deviation) and being aware of limitations (outliers, biased samples, misleading graphs).

How to approach a real dataset:

  • Step 1: Identify type of data — quantitative (continuous or discrete) or qualitative (categorical).
  • Step 2: Choose summary measures — mean (typical numeric average), median (middle value, robust to outliers), mode (most frequent, useful for category and discrete data).
  • Step 3: Measure spread — range, interquartile range (IQR), variance and standard deviation to understand variability.
  • Step 4: Visualize — histogram/frequency polygon for distributions, ogive for medians/percentiles, box‑plot for spread and outliers, bar/pie charts for categorical proportions.
  • Step 5: Interpret in context — compare groups, note skewness, consider effects of extreme values, and check if data collection was unbiased.

Key interpretation guidelines:

  • Use median instead of mean when data are skewed or contain outliers (e.g., incomes).
  • Use mode for categorical or most‑likely outcomes (e.g., most common shoe size).
  • If two groups have similar means, compare spread (standard deviation or IQR) to decide which is more consistent.
  • Ogives let you read medians, quartiles and percentiles directly — useful for ranking and cutoff decisions.
  • Watch graph axes and scales — truncated axes or unequal class widths can mislead.

Common pitfalls: Small or biased samples give unreliable conclusions. Outliers can distort the mean. Grouped data require class‑wise methods (midpoints, class boundaries) for accurate estimates.

📌 Examples
  • Class test scores of 40 students: compute mean to get average performance, median to see the middle performer and compare to mean to spot skewness; use histogram to check if most students clustered around a score or spread out.
  • Household incomes in a locality: use median income instead of mean because a few very high incomes (outliers) inflate the mean and do not represent the typical household.
  • A shoe store records sizes sold: mode identifies the most frequently sold size so the store can stock more of it; a bar chart shows proportions across sizes.
  • Using an ogive (less‑than cumulative frequency graph) for marks grouped into classes to find the median and quartiles quickly — useful for setting pass marks or selecting top X% students.
🧮 Formulas
  1. Mean (ungrouped): x̄ = (Σx_i) / n
  2. Mean (grouped, using midpoints): x̄ = (Σf_i m_i) / Σf_i, where m_i = class midpoint
  3. Mean (grouped, assumed mean method): x̄ = a + (h * Σf_i u_i) / Σf_i, where u_i = (m_i - a)/h and a = assumed mean, h = class width
  4. Median (grouped): Median = l + [(N/2 − C) / f] × h, where l = lower boundary of median class, C = cumulative frequency before median class, f = frequency of median class, h = class width, N = total frequency
  5. Mode (grouped): Mode = l + [(f1 − f0) / (2f1 − f0 − f2)] × h, where f1 = frequency of modal class, f0 = frequency of previous class, f2 = frequency of next class, l = lower boundary of modal class, h = class width
  6. Variance (ungrouped): σ^2 = (Σ(x_i − x̄)^2) / n (population) or s^2 = (Σ(x_i − x̄)^2) / (n−1) (sample)
📊 Visual ideas
Histogram: plot class intervals on x-axis and frequency on y-axis. Use for continuous grouped data to show shape (symmetry, skewness) and relative concentrations.
Ogive (cumulative frequency curve): two curves (less‑than and/or greater‑than). Use to find median, quartiles and percentiles by reading corresponding cumulative frequencies.
Frequency polygon: join midpoints of class tops with straight lines. Good for comparing two distributions on the same axes.
Bar chart: categories on x-axis, heights show frequencies. Use for qualitative/categorical data (e.g., brand preference).

Key Concepts

Statistics
Branch of mathematics dealing with collection, organization, presentation, analysis and interpretation of numerical data.
Data
Facts or figures collected for analysis; can be numerical or categorical.
Primary data
Data collected first-hand by the investigator for a specific purpose.
Secondary data
Data obtained from already published or collected sources.
Variable
A characteristic or attribute that can take different values among individuals.
Frequency
Number of times a particular value or class of values occurs in the data.
Class interval
A range of values used to group continuous data in a frequency distribution.
Class mark (midpoint)
Midpoint of a class interval, calculated as (lower limit + upper limit)/2.
Discrete data
Data that can take only specific separated values (often integers).
Continuous data
Data that can take any value in an interval (measurable quantities).
Ungrouped data
Raw data presented as individual observations without forming classes.
Grouped data
Data organized into classes (intervals) with corresponding frequencies.
Frequency distribution
Tabular summary showing classes and their frequencies.
Cumulative frequency
Running total of frequencies up to a given class or value.
Relative frequency
Frequency of a class divided by total number of observations (often expressed as fraction or percentage).
Mean (Arithmetic mean)
Sum of observations divided by the number of observations; measure of central tendency.
Median
Middle value that divides ordered data into two equal parts; for even n it's average of two middle values.
Mode
Value(s) that occur most frequently in the data.
Histogram
Bar graph representing frequency distribution for continuous data where adjacent bars touch; area of bar ∝ frequency.
Ogive (Cumulative frequency curve)
Graph of cumulative frequency against class boundaries used to determine medians and percentiles.

End-of-Chapter Trial Paper & Test Questions

Topic-wise questions to test your understanding of every concept in this chapter.

  1. Define class mark (mid-point) of a class interval and state its formula. / वर्ग चिह्न (मध्य-बिंदु) को परिभाषित कीजिए तथा इसका सूत्र लिखिए।
    Show answer

    The class mark is the midpoint of a class interval, given by (lower limit + upper limit)/2; it represents the whole class in mean calculations. / वर्ग चिह्न किसी वर्ग अंतराल का मध्य-बिंदु होता है, जो (निम्न सीमा + उच्च सीमा)/2 से प्राप्त होता है; यह माध्य की गणना में पूरे वर्ग का प्रतिनिधित्व करता है।

  2. Find the mean of the ungrouped data: 72, 65, 80, 90, 83. / असमूहीकृत आँकड़ों 72, 65, 80, 90, 83 का माध्य ज्ञात कीजिए।
    Show answer

    Sum = 390 and n = 5, so mean = 390/5 = 78. / योग = 390 तथा n = 5, अतः माध्य = 390/5 = 78।

  3. Why is the median preferred over the mean when data contain extreme values (outliers)? / जब आँकड़ों में चरम मान (बाह्य मान) हों तो माध्यक को माध्य की अपेक्षा क्यों वरीयता दी जाती है?
    Show answer

    Because the median depends only on the middle position of ordered data and is not affected by extreme values, while the mean is pulled toward outliers. / क्योंकि माध्यक केवल क्रमित आँकड़ों की मध्य स्थिति पर निर्भर करता है और चरम मानों से प्रभावित नहीं होता, जबकि माध्य बाह्य मानों की ओर खिंच जाता है।

  4. For a grouped distribution N = 30, median class 20-30 with l = 20, c.f. = 13, f = 12, h = 10, find the median. / एक समूहीकृत बंटन में N = 30, माध्यक वर्ग 20-30 जिसमें l = 20, c.f. = 13, f = 12, h = 10 है, माध्यक ज्ञात कीजिए।
    Show answer

    Median = l + ((N/2 − c.f.)/f) × h = 20 + ((15 − 13)/12) × 10 = 20 + 1.67 = 21.67 (approx). / माध्यक = l + ((N/2 − c.f.)/f) × h = 20 + ((15 − 13)/12) × 10 = 20 + 1.67 = 21.67 (लगभग)।

  5. State the modal-class formula for grouped data and name each symbol used. / समूहीकृत आँकड़ों के लिए बहुलक वर्ग का सूत्र लिखिए तथा प्रयुक्त प्रत्येक प्रतीक का नाम बताइए।
    Show answer

    Mode = l + ((fm − f1)/(2fm − f1 − f2)) × h, where l = lower boundary of modal class, fm = modal class frequency, f1 = previous class frequency, f2 = next class frequency, h = class width. / बहुलक = l + ((fm − f1)/(2fm − f1 − f2)) × h, जहाँ l = बहुलक वर्ग की निम्न सीमा, fm = बहुलक वर्ग की बारंबारता, f1 = पूर्ववर्ती वर्ग की बारंबारता, f2 = अगले वर्ग की बारंबारता, h = वर्ग चौड़ाई।

  6. Using the empirical relation, if median = 24 and mean = 21, estimate the mode. / सांख्यिकीय संबंध का उपयोग करते हुए, यदि माध्यक = 24 और माध्य = 21 हो तो बहुलक का आकलन कीजिए।
    Show answer

    Mode ≈ 3 × Median − 2 × Mean = 3 × 24 − 2 × 21 = 72 − 42 = 30. / बहुलक ≈ 3 × माध्यक − 2 × माध्य = 3 × 24 − 2 × 21 = 72 − 42 = 30।

  7. How is the median estimated graphically from an ogive? / तोरण (ओजाइव) से माध्यक का आलेखीय आकलन कैसे किया जाता है?
    Show answer

    Plot cumulative frequency against upper class boundaries, draw a horizontal line at N/2 to meet the ogive, then drop a vertical line to the x-axis; the x-coordinate is the median. / संचयी बारंबारता को उच्च वर्ग सीमाओं के सापेक्ष आलेखित करें, N/2 पर क्षैतिज रेखा खींचकर तोरण से मिलाएँ, फिर लंबवत रेखा x-अक्ष तक गिराएँ; वह x-निर्देशांक माध्यक है।

  8. Distinguish between a histogram and a bar graph. / आयतचित्र (हिस्टोग्राम) तथा दंड आलेख (बार ग्राफ) में अंतर बताइए।
    Show answer

    A histogram is used for continuous grouped data with bars touching each other, whereas a bar graph is used for discrete or categorical data with gaps between the bars. / आयतचित्र सतत समूहीकृत आँकड़ों के लिए प्रयुक्त होता है जिसमें दंड एक-दूसरे को स्पर्श करते हैं, जबकि दंड आलेख विविक्त या श्रेणीगत आँकड़ों के लिए प्रयुक्त होता है जिसमें दंडों के बीच अंतराल होता है।

Related Laws & Principles

Explore all

Foundational laws & principles behind this chapter. Each one opens a full page — what it says, why it matters, five practice questions and the mistakes to avoid.

Loading related laws…
Sourced from 117 content files · LLOS Learn · browse all chapters