L
LLLOS.ai
LLOS.ai
L
Class 11 Mathematics Chapter 15 of 16

Chapter 15 — Statistics

Overview

Chapter "Statistics" (Class 11 NCERT Mathematics) introduces basic tools for describing and summarising univariate data. It explains how to present data in tabular and graphical form (histograms, frequency polygons, ogives) and develops the main measures used to locate and spread a distribution — measures of central tendency (mean, median, mode) and measures of dispersion (range, quartile deviation, mean deviation, variance, standard deviation). The chapter also treats grouped and ungrouped data separately, gives practical computation techniques (class marks, assumed-mean/step-deviation methods) and formulas for combined data sets, and shows how simple linear transformations affect mean and variance. This content builds students' ability to analyse data, compare distributions and lay a foundation for probability and inferential statistics.

Learning Objectives

  • Define raw data, frequency distribution, class interval, class mark and cumulative frequency.
  • Explain the difference between grouped and ungrouped data and when to use each representation.
  • Construct a frequency distribution table (discrete and continuous) and compute class marks and cumulative frequencies.
  • Compute the arithmetic mean for ungrouped data using direct summation (Σ notation).
  • Apply the assumed‑mean and step‑deviation methods to compute the mean of grouped data efficiently.
  • Determine the median for ungrouped data and for a continuous grouped frequency distribution using linear interpolation.
  • Calculate the mode for ungrouped data and estimate the modal value for grouped data using the modal class formula.
  • Compute mean deviation about the mean and about the median for grouped and ungrouped data.

Topics in this chapter

13 topics · tap a topic title to jump straight to it.

📊1

Introduction to Statistics

What is Statistics? Statistics is the branch of mathematics that deals with collection, organization, presentation, analysis and interpretation of numerical data to make decisions. It converts raw data into useful information through suitable methods.

Objectives / Steps in a Statistical Study

  • Collection of data (primary or secondary)
  • Organization of data (classification and tabulation)
  • Presentation of data (graphs, charts, tables)
  • Analysis and interpretation (measures of central tendency, dispersion, etc.)
  • Drawing conclusions and making decisions

Types of Data

  • Qualitative (Categorical): Non-numeric attributes (e.g., blood group, city names).
  • Quantitative (Numerical): Numeric values. Further divided into:
    • Discrete (countable values: number of students)
    • Continuous (measurable values: height, weight, temperature)

Primary vs Secondary Data: Primary data are obtained by direct observation or survey. Secondary data are collected from published sources or other studies.

Classification and Frequency Distribution: For quantitative data (especially large data), we group into class intervals and make a frequency distribution giving class limits, class boundaries, class width (h), class mark (mid-point) and frequency (f).

Key Terms

  • Class width (h): upper limit − lower limit of a class
  • Class mark (mid-point) x_i: (lower limit + upper limit) / 2
  • Frequency (f_i): number of observations in a class
  • Cumulative frequency (cf): running total of frequencies up to a class
  • Relative frequency: f_i / n (often expressed as a fraction or percentage)

Presentation of Data: Data are presented using tables and graphs for clarity. Common graphs: bar charts (for categorical/discrete data), histograms and frequency polygons (for continuous/grouped data), pie charts (for categorical proportions), and ogives (cumulative frequency curves).

Uses and Limitations: Statistics helps summarize large data sets and make informed decisions (policy, business, science). Limitations include sensitivity to biased sampling, poor data collection, and misleading graphics if used improperly.

📌 Examples
  • Recording daily maximum temperatures of a city for a month (continuous data) and creating a frequency distribution to study temperature variation.
  • Survey of marks obtained by 100 students in an exam, then computing mean, median and mode to describe class performance.
  • Counting the number of cars of different colours in a parking lot (discrete/categorical) and showing results by a bar chart or pie chart.
  • Population census data (secondary data) classified by age groups to analyze age distribution using histograms and ogives.
  • A health researcher measuring weights of newborns (continuous data), grouping them into classes and estimating typical weight using the class mid-point.
🧮 Formulas
  1. Sample / Raw (unGrouped) arithmetic mean: x̄ = (Σ x_i) / n
  2. Grouped data arithmetic mean: x̄ = (Σ f_i x_i) / Σ f_i where x_i is class mark (mid-point) and f_i is class frequency
  3. Class mark (mid-point): x_i = (lower limit + upper limit) / 2
  4. Class width (h): h = upper limit − lower limit
  5. \[Cumulative frequency (CF) up to k-th class: CF_k = Σ_{i=1}^k f_i\]
  6. Relative frequency: r_i = f_i / n
📊 Visual ideas
Histogram (for continuous/grouped data): x-axis = class intervals (use class boundaries), y-axis = frequency. Adjacent bars touch to indicate continuity. Useful to see shape and spread.
Frequency polygon: Plot class marks on x-axis and frequencies on y-axis, join consecutive points by straight lines. Useful to compare two distributions on the same graph.
Ogive (cumulative frequency curve): x-axis = upper (or lower) class boundaries, y-axis = cumulative frequency. Two types: less-than and more-than ogives. Use to estimate median and percentiles graphically.
Bar chart (for discrete/categorical data): x-axis = categories, y-axis = frequencies. Bars are separated. Good for comparing counts across categories.
📊2

Collection of Data

What is data? Data are raw facts and figures collected for analysis. In Statistics, collecting data is the first step: you record observations about a variable(s) from a population or a sample.

Types of data

  • Qualitative (Attribute): non-numeric categories (e.g., blood group, colour).
  • Quantitative: numeric measurements.
    • Discrete: countable values (e.g., number of students).
    • Continuous: measurable on a continuum (e.g., height, weight).

Scales of measurement: nominal, ordinal, interval, ratio. These determine what analyses and graphs are appropriate.

Sources of data

  • Primary data: collected first-hand for the study (surveys, experiments, observations).
  • Secondary data: obtained from existing sources (books, articles, government reports).

Methods of collecting primary data

  • Observation: record what is seen (useful when respondents can’t report reliably).
  • Questionnaire / Interview: ask participants via written forms or oral interviews.
  • Schedule: structured form filled by investigator after interviewing the respondent.
  • Experiment / Measurement: controlled conditions, often for scientific studies.

Sampling methods (brief): simple random sampling, stratified sampling, systematic sampling, cluster sampling, convenience sampling. Choose based on study goals, resources and need for representativeness.

Classification and tabulation: Raw data are grouped into classes (for continuous data) or categories (for discrete/qualitative data) to make frequency distributions (tables of class/cateogry vs frequency). Steps: 1) Decide variable type and scale, 2) Choose number of classes (k), 3) Determine class width/intervals, 4) Tally frequencies, 5) Compute relative frequencies/cumulative frequencies if needed.

Coding data: Often raw values are transformed by a linear change to simplify calculations: y = (x − a)/b (shift by a and scale by b). Recover original by x = b*y + a.

Important practical points

  • Design questionnaires to avoid ambiguity and bias; pilot-test them.
  • Ensure sample represents the population if generalisation is required.
  • Watch for non-response and measurement errors; document data source and method.

From collection to visualization: After classifying and tabulating, data are summarized with graphs (bar charts, histograms, pie charts, ogives, scatter plots). Choice of graph depends on data type: qualitative → bar/pie; quantitative (continuous) → histogram, frequency polygon, ogive; bivariate quantitative → scatter plot.

📌 Examples
  • Classroom example: Teacher records marks of 40 students (quantitative, discrete). She may list raw marks, create class intervals (e.g., 0–9, 10–19, …), and compute a frequency distribution to study performance.
  • Health survey: A nurse measures the heights of 200 children (quantitative, continuous). Heights are measured (primary data), grouped into class intervals (e.g., 110–114.9 cm) and plotted in a histogram to observe distribution.
  • Market research: A company collects customer preferences for four product colours (qualitative). Results are tabulated and shown as a bar chart or pie chart to find the most popular colour.
  • Census data: Government uses secondary data from previous censuses and combines with a new household survey (primary) to estimate population changes and plan resources.
🧮 Formulas
  1. Number of classes (Sturges' rule, guideline): k ≈ 1 + 3.3 log10(n) (n = sample size)
  2. Class width / size (h): h ≈ (Max − Min) / k (choose convenient value; make intervals equal-width when possible)
  3. Class limits and class boundaries: if class is 10–19, lower limit = 10, upper limit = 19; for continuous data class boundaries = lower − 0.5 unit and upper + 0.5 unit (if measurement precision is 1 unit).
  4. Class mark / midpoint (x_c): x_c = (lower limit + upper limit) / 2
  5. Frequency (f): count of observations in a class/category
  6. Relative frequency: f_r = f / n
📊 Visual ideas
Histogram: For continuous quantitative data. Plot class intervals on x-axis and frequencies (or frequency density) on y-axis. Use equal-width bins when possible.
Bar chart: For qualitative or discrete quantitative data. Bars separated (nominal) or adjacent (ordinal). Height = frequency or percentage.
Pie chart: For showing percentage composition of categories (qualitative). Keep categories limited (≤6) for clarity.
Ogive (cumulative frequency curve): Plot cumulative frequency vs upper class boundaries. Useful to read medians and percentiles.
📊3

Classification and Tabulation of Data

Definition: Classification is the process of arranging raw data into groups or classes according to some rule. Tabulation is the systematic presentation of classified data in the form of tables showing frequencies and other summary measures.

Types of data (brief):

  • By source: Primary (collected first-hand), Secondary (from existing records).
  • By nature: Qualitative (categorical) and Quantitative (numerical).
  • Quantitative can be Discrete (take specific values, e.g., number of children) or Continuous (measured, e.g., height, weight).
  • Measurement scales (useful to know): Nominal, Ordinal, Interval, Ratio.

Why classify and tabulate? To simplify large data sets, reveal patterns, allow computation of statistical measures (mean, median, mode, dispersion), and prepare data for graphical presentation.

Key terms:

  • Class interval: a range of values (e.g., 30–39).
  • Class limits: lower and upper limits of the interval.
  • Class boundaries: exact end points when intervals are made continuous (e.g., 29.5–39.5 for integer data).
  • Class width (h): difference between upper and lower limits of a class (constant for equal-width classes).
  • Class mark (midpoint) xi: (lower limit + upper limit) / 2.
  • Frequency fi: number of observations in a class.
  • Relative frequency: fi / N (N = total observations).
  • Cumulative frequency: running total of frequencies up to (or from) a class.

Steps to create a grouped frequency distribution:

  1. Arrange raw data in ascending order (optional but helpful).
  2. Decide number of classes k. Common rules: 1 + 3.3 log10 n (Sturges' rule) or k roughly between 5 and 15; sometimes k ≈ √n.
  3. Compute class width h ≈ (max − min) / k and round to a convenient value.
  4. Choose the first lower limit; build k mutually exclusive, exhaustive class intervals of equal width.
  5. Tally observations into classes and obtain frequencies fi.
  6. Compute class marks, relative frequencies, percentage frequencies and cumulative frequencies as needed.

Guidelines for classes:

  • Classes should not overlap and must cover the entire range.
  • Prefer equal class widths for easier interpretation and graphical work (histogram, frequency polygon).
  • For discrete integer data, use boundaries like 9.5–19.5 to avoid gaps when drawing histograms.
  • Use open-ended classes only when necessary (e.g., "60 and above").

Inclusive vs exclusive class limits: In inclusive form (used for integer data) an interval 10–14 means values 10,11,12,13,14. For continuous data or for drawing histograms, use class boundaries (e.g., 9.5–14.5).

Common mistakes to avoid:

  • Overlapping classes or leaving gaps.
  • Unequal or odd class widths without reason when comparison is required.
  • Incorrect cumulative counting direction for ogives (be clear whether you use "less than" or "more than" ogive).

Summary: Classification groups observations into meaningful categories; tabulation arranges these groups neatly with frequencies and related measures so that further statistical analysis and graphical representation become easy and reliable.

📌 Examples
  • Real-life examples: (a) Heights of students grouped into intervals for a school health report. (b) Daily temperature readings grouped by ranges to study climate patterns. (c) Monthly household incomes grouped into income brackets for socio-economic analysis. (d) Marks of students grouped into class intervals to study grade distribution.
  • Worked example (marks of 20 students): Raw data: 35, 42, 47, 51, 56, 58, 60, 62, 64, 67, 69, 70, 72, 75, 77, 79, 81, 83, 85, 88 Group into classes: 30–39, 40–49, 50–59, 60–69, 70–79, 80–89 Frequency table: Class | Frequency fi | Class mark xi | Cumulative F 30–39 | 1 | 34.5 | 1 40–49 | 2 | 44.5 | 3 50–59 | 3 | 54.5 | 6 60–69 | 5 | 64.5 | 11 70–79 | 5 | 74.5 | 16 80–89 | 4 | 84.5 | 20 Total N = 20.
  • Example of qualitative tabulation: Survey of favourite fruit among 100 people -> Table with categories (Apple, Banana, Mango, Others) and their frequencies, relative frequencies and percentages; display as a bar chart or pie chart.
🧮 Formulas
  1. Class mark (midpoint): xi = (lower limit + upper limit) / 2
  2. Class width: h = (maximum value − minimum value) / number of classes (k). Round to a convenient value.
  3. Relative frequency: ri = fi / N, where N is total number of observations.
  4. Percentage frequency: (fi / N) × 100
  5. Cumulative frequency (up to class j): CFj = Σ (fi) for classes 1 to j
  6. Sturges' rule (suggested number of classes): k ≈ 1 + 3.3 log10(N) (useful guideline, not strict)
📊 Visual ideas
Histogram: For quantitative continuous/grouped data. X-axis = class intervals (use class boundaries, no gaps between bars), Y-axis = frequency. Height of each bar = frequency. Useful to see shape of distribution.
Frequency polygon: Plot class marks on x-axis and corresponding frequencies on y-axis, join successive points by straight lines. Helpful to compare distributions (overlay polygons).
Ogive (cumulative frequency curve): 'Less-than' ogive plots upper class boundaries against cumulative frequencies; 'more-than' ogive plots lower boundaries against cumulative frequencies. Used to read percentiles and median graphically.
Bar chart / Pie chart: For qualitative (categorical) data. Bars separated for categories; pie shows proportion of each category.
📊4

Presentation of Data

What is presentation of data? Presentation of data means arranging and displaying collected numerical or categorical observations in tables and graphs so that important features (patterns, trends, comparisons) become clear and easy to interpret.

Why present data? To summarise large data sets, reveal patterns (central tendency, spread, skewness), enable comparisons, and support decision making.

Types of data: qualitative (categorical) and quantitative (numerical). Quantitative can be discrete (countable) or continuous (measured).

Tabular presentation — frequency distribution. Two main kinds:

  • Ungrouped frequency distribution: list distinct values and their frequencies (used when number of distinct values is small).
  • Grouped frequency distribution: data grouped into class intervals (used for large/continuous data).

Steps to make a grouped frequency distribution (typical):

  1. Arrange raw data; find n (total observations), minimum and maximum.
  2. Compute range R = max − min.
  3. Choose number of classes k (common rules: 5–15; sometimes k ≈ √n).
  4. Class width (approx) h = R / k (round up to convenient number).
  5. Form k consecutive non-overlapping class intervals of width h covering all data.
  6. Tally frequencies fi for each class; compute cumulative frequencies (CF), class marks xi, relative frequencies (fi/n) and percentages.

Important terms:

  • Class limits: lower and upper limits of a class.
  • Class boundaries: adjust limits by half the least measurement unit to remove gaps (e.g., if data are integers, class 10–19 and 20–29 have boundaries 9.5–19.5 and 19.5–29.5).
  • Class mark (midpoint): xi = (lower limit + upper limit) / 2.
  • Cumulative frequency (CF): running total of frequencies up to a class.
  • Relative frequency: fi / n; percentage = (fi / n) × 100%.

Graphical presentation — choose depending on data and purpose:

  • Bar diagram: categorical or discrete data comparisons (gaps between bars).
  • Histogram: continuous data grouped into classes (contiguous bars; area of bar ∝ frequency; height = frequency / class width for unequal widths).
  • Frequency polygon: plot class marks against frequencies and join by straight lines (useful to compare distributions).
  • Ogive (cumulative frequency curve): plot upper class limits (or boundaries) vs cumulative frequency; useful to read medians/percentiles.
  • Pie chart: show parts of a whole for categorical data (use when number of categories is small).
  • Pictogram/stem-and-leaf: alternative visual summaries (pictogram for simple categories; stem-and-leaf retains raw-data information).

Rules & tips:

  • Choose class width and k so classes are neither too few (loss of detail) nor too many (noisy).
  • Use equal class widths if possible; if unequal widths are necessary, draw histogram with bar area proportional to frequency (height = frequency / width).
  • Label axes, give title, and include units and legends where appropriate.
📌 Examples
  • Example 1 (Grouped frequency distribution): Marks of 30 students (out of 100). Suppose min = 12, max = 88 so range R = 76. Choose k = 8 ⇒ h ≈ R/k = 9.5 → take class width 10. Classes: 10–19, 20–29, …, 80–89. Tally frequencies fi for each class, compute class marks xi = (lower+upper)/2, relative frequency fi/30, cumulative frequency CF, then draw a histogram or frequency polygon.
  • Example 2 (Histogram vs. Bar chart): Heights of people (continuous) are best shown with a histogram (contiguous bars). Types of fruit sold (apples, bananas, mangoes) are categorical — use a bar chart or pie chart.
  • Example 3 (Ogive to get median): From a grouped distribution, plot upper class boundaries on x-axis and cumulative frequencies on y-axis, draw the ogive. To find median (50th percentile), draw a horizontal line at n/2 and read x where it meets the ogive.
🧮 Formulas
  1. n = total number of observations
  2. Range: R = maximum − minimum
  3. Approx. number of classes: k ≈ √n (or choose 5–15 depending on data)
  4. Class width (approx): h = R / k (round to convenient value)
  5. Class mark (midpoint): xi = (lower limit + upper limit) / 2
  6. Frequency: fi (count in class i)
📊 Visual ideas
Bar diagram — use for qualitative/categorical data or discrete counts. Draw bars with equal widths and gaps between them. Label categories on x-axis and frequency on y-axis.
Histogram — use for continuous quantitative data grouped into equal class intervals. Draw contiguous rectangles; height = frequency (if widths equal) or height = frequency / width (if widths unequal).
Frequency polygon — plot class marks (xi) on x-axis and frequencies fi on y-axis, join the points with straight lines; close polygon at ends with x before first and after last class with zero frequency.
Ogive (cumulative frequency curve) — plot upper class boundaries vs cumulative frequency and join smoothly or with straight segments. Use it to read medians and percentiles.
📈5

Graphical Representation of Data

Graphical Representation of Data is the use of diagrams and graphs to present statistical data visually so patterns, trends and comparisons become clear at a glance. It is a key part of Statistics in Class 11 because it helps interpret raw and grouped data without lengthy calculations.

Major types of graphs and when to use them:

  • Bar graph — for qualitative (categorical) data or discrete numerical data (gaps between bars).
  • Pie chart — to show percentage/share of a whole for categories (useful for market-share, budget splits).
  • Histogram — for continuous grouped data; contiguous bars show frequency per class interval. If class widths differ, use frequency density (frequency/width) for bar height.
  • Frequency polygon — join class-midpoint frequencies with straight lines; useful to compare distributions and show smoothing of a histogram.
  • Frequency curve — smooth curve drawn through points of a frequency polygon (approximate continuous distribution).
  • Ogive (cumulative frequency curve) — plots cumulative frequency vs class boundary; two types: ‘less than’ and ‘greater than’. Useful for reading medians, quartiles and percentiles graphically.
  • Scatter plot (XY-plot) — for bivariate data to examine relationship (correlation/causation) between two variables.
  • Time-series (line graph) — for data indexed by time to show trends, seasonal patterns.

General steps to construct common graphs:

  • Histogram: determine class intervals and class boundaries, compute frequencies (or frequency density if widths unequal), draw contiguous rectangles with base equal to class width and height equal to frequency or frequency density.
  • Frequency polygon: compute class midpoints, plot midpoint vs frequency, join successive points with straight lines; close polygon by joining first and last points to the horizontal axis at the class-endpoints.
  • Ogive: compute cumulative frequencies (less-than: cumulative up to upper class boundary); plot boundary vs cumulative frequency and join by a smooth/straight curve.
  • Pie chart: convert each frequency to angle = (frequency/total)×360°, draw sectors accordingly.

Interpreting graphs: identify central tendency (peak of histogram or polygon), spread (width of distribution), skewness (tail left/right), modality (uni-/bi-/multi-modal), outliers, and trends (in time series or scatter plots).

Advantages: quick visual comparison, intuitive pattern recognition. Limitations: can hide exact values, depend on choice of class width and origin (histogram) or angle rounding (pie chart).

📌 Examples
  • Marks distribution of 200 students in an exam: use a histogram to view how marks are spread and an ogive to estimate median and quartiles.
  • Monthly sales of a store for a year: use a time-series line graph to show upward/downward trends and seasonality.
  • Market share of 5 companies: use a pie chart to present percentage shares; angles = (share/total)*360°.
  • Survey of preferred transport mode (bus, train, car, bike): use a bar graph to compare categorical frequencies.
  • Relationship between hours studied and marks obtained: use a scatter plot to check correlation and fit a trend line.
🧮 Formulas
  1. Class midpoint (xi) = (lower class boundary + upper class boundary) / 2
  2. Frequency density (for unequal width classes) = frequency / class width
  3. Angle for pie chart sector = (frequency / total frequency) × 360°
  4. Mean for grouped data = (Σ f_i x_i) / Σ f_i, where x_i are class midpoints and f_i frequencies
  5. Median for grouped data = l + ((N/2 − C) / f) × h, where l = lower boundary of median class, N = total frequency, C = cumulative frequency before median class, f = frequency of median class, h = class width
  6. Mode for grouped data (approximate) = l + [(f_m − f_1) / (2f_m − f_1 − f_2)] × h, where f_m is frequency of modal class, f_1 previous class freq, f_2 next class freq, l lower boundary of modal class, h class width
📊 Visual ideas
Histogram — Visual suggestion: horizontal axis = class intervals (continuous, contiguous bars), vertical axis = frequency (or frequency density). For equal widths draw bar heights proportional to frequency; for unequal widths use height = frequency / width. Use this for marks, income ranges, ages.
Frequency Polygon — Visual suggestion: plot class midpoints on the x-axis and corresponding frequencies on y-axis; join points with straight lines. Close polygon to x-axis at ends. Useful to compare multiple distributions on same axes.
Ogive (Cumulative Frequency Curve) — Visual suggestion: plot upper class boundaries (less-than ogive) vs cumulative frequency; join points smoothly. Use to read median (where ogive meets N/2), quartiles and percentiles graphically.
Bar Graph — Visual suggestion: separate bars for each category on x-axis with gaps; height = frequency. Label bars and axes. Use for qualitative data like transport modes or categories.
🔢6

Cumulative Frequency and Ogives

Cumulative Frequency

Cumulative frequency (CF) for a class or value is the running total of frequencies up to that class/value. If f1, f2, ..., fk are frequencies of successive classes, then the cumulative frequencies are CF1 = f1, CF2 = f1 + f2, ..., CFk = f1 + f2 + ... + fk. Cumulative frequency organizes data to show how many observations lie below (or above) a given value.

Types of Cumulative Frequency

  • Less-than cumulative frequency: CF at an upper class boundary = total frequency of all classes with values ≤ that upper boundary. Used to make the "less-than" ogive.
  • Greater-than cumulative frequency: CF at a lower class boundary = total frequency of all classes with values ≥ that lower boundary. Used to make the "greater-than" ogive. It can be computed as (Total N) − (less-than CF before that boundary).

Preparing data (grouped data)

  1. Start with a frequency distribution table of class intervals and frequencies.
  2. Decide correct class boundaries (for continuous data use the class limits; for integer data adjust boundaries by 0.5 if classes are written as whole numbers: e.g. 10–19 implies boundaries 9.5–19.5).
  3. Compute cumulative (less-than) frequencies using upper class boundaries. Optionally compute greater-than CF using lower boundaries.

Ogives (Cumulative Frequency Curves)

An ogive is a graph of cumulative frequency versus class boundary/value.

  • Less-than ogive: Plot points (upper class boundary, cumulative frequency up to that boundary). Start at the lower-most boundary with CF = 0 (optional) and end at the highest boundary with CF = N.
  • Greater-than ogive: Plot points (lower class boundary, cumulative frequency of values greater than or equal to that boundary). It typically starts at N and ends at 0.

Reading information from ogives

  • Median: draw a horizontal line at N/2 on the cumulative-frequency axis and read the corresponding x-value on the ogive. If you draw both less-than and greater-than ogives, their intersection gives the median.
  • Quartiles/percentiles: draw horizontal lines at N/4, N/2, 3N/4 (or other percentile positions) and read the corresponding x-values on the ogive.
  • Proportion below/above a value: read CF at that value and divide by N to get proportion or percentage.

Why ogives are useful

Ogives provide a clear picture of cumulative distribution: they show how quickly observations accumulate with increasing value, enable quick estimation of medians and percentiles, and help compare distributions.

📌 Examples
  • Real-life: Exam scores — use cumulative frequency to answer ‘how many students scored less than 60?’ or find the median score and quartiles for the class.
  • Real-life: Income brackets — cumulative frequencies show how many people earn up to a certain income level (useful for medians and poverty thresholds).
  • Worked numerical example: Suppose grouped data (class interval : frequency) are 0–10: 5, 10–20: 9, 20–30: 12, 30–40: 4. Total N = 30. Less-than cumulative frequencies at upper boundaries are: 10→5, 20→5+9=14, 30→14+12=26, 40→26+4=30. To find median: N/2 = 15. Median class is 20–30 (CF before class = 14, f = 12, class width h = 10). Using grouped median formula median = l + ((N/2 − CF_prev)/f) × h = 20 + ((15−14)/12)×10 ≈ 20.83.
  • Using ogives: plot points (10,5), (20,14), (30,26), (40,30) for the less-than ogive. Draw a horizontal line at CF = 15; where it intersects the ogive read x ≈ 20.83 (median).
🧮 Formulas
  1. \[Cumulative frequency (less-than) at k-th class: CF_k = sum_{i=1}^{k} f_i\]
  2. \[Cumulative frequency (greater-than) at k-th class: GCF_k = sum_{i=k}^{m} f_i = N - CF_{k-1} (where CF_{k-1} is less-than CF before class k)\]
  3. Median (grouped data) formula: Median = l + ((N/2 − CF_prev)/f) × h, where l = lower class boundary of median class, CF_prev = cumulative frequency before median class, f = frequency of median class, h = class width
  4. Quartile/percentile by ogive: draw horizontal line at p×N (e.g. p=1/4,1/2,3/4) and read x; equivalently use interpolation in the appropriate class similar to median formula
  5. If class limits are integers and classes are contiguous, class boundaries = class limits ± 0.5 (e.g. interval 10–19 ⇒ boundary 9.5–19.5).
📊 Visual ideas
Less-than ogive: x-axis = class boundaries (use upper boundaries for each class), y-axis = cumulative frequency (CF). Plot points (upper boundary, CF) and join points with a smooth or polygonal curve. Optionally start at the smallest lower boundary with CF=0.
Greater-than ogive: x-axis = class boundaries (use lower boundaries), y-axis = cumulative frequency of values ≥ boundary. Plot and join similarly. The intersection of less-than and greater-than ogives gives the median.
Mark horizontal lines at N/4, N/2, 3N/4 (or other percentiles). From each horizontal line drop a vertical to read the corresponding x-value (quartiles/percentiles).
Overlay suggestion: draw the ogive next to a histogram or frequency polygon for comparison. In software (Excel/Google Sheets/Desmos), plot points and use 'line' or 'smooth line' to join points; annotate median and quartiles for clarity.
📏7

Measures of Central Tendency — Overview

What are Measures of Central Tendency?

Measures of central tendency are single values that describe the centre or typical value of a data set. They summarise a distribution by identifying a central point around which data cluster. The three commonly used measures are Arithmetic Mean, Median and Mode.

1. Arithmetic Mean (Average)

  • Definition: Sum of all observations divided by the number of observations.
  • Use: Best for quantitative data without extreme outliers; widely used in science, finance and education.
  • Properties: Unique, uses all values, sensitive to extreme values (not robust).

2. Median

  • Definition: Middle value when observations are arranged in ascending order. For even n, median is the average of the two middle values.
  • Use: Preferred for skewed data (e.g. incomes, house prices) because it is not affected by extreme values.
  • Properties: Robust (resistant to outliers), may not use all data values.

3. Mode

  • Definition: Value that occurs most frequently in the data set. A distribution may be unimodal, bimodal, or multimodal.
  • Use: Useful for categorical data (e.g. most common shoe size, favourite colour) and for identifying peaks in a distribution.
  • Properties: Can be non-unique; simple to identify for qualitative data.

Grouped Data vs Raw Data

For raw (individual) data the mean, median and mode are computed directly from the list of numbers. For grouped data (data in class intervals) we use formulas and approximations:

  • Grouped mean: use class mid-points and frequencies; or use the assumed mean method to simplify calculation.
  • Grouped median: approximate position n/2 and interpolate inside the median class using class width and cumulative frequency.
  • Grouped mode: identify the modal class (highest frequency) and use formula to approximate the mode by interpolation inside that class.

Relation and Skewness

In practice, the relative position of mean, median and mode indicates skewness:

  • Symmetric distribution: Mean ≈ Median ≈ Mode.
  • Right (positive) skew: Mode < Median < Mean.
  • Left (negative) skew: Mean < Median < Mode.

Choosing the Appropriate Measure

  • Use mean for balanced quantitative data without outliers and when mathematical manipulation is needed.
  • Use median for skewed distributions or when a resistant measure is required (e.g. income data).
  • Use mode for categorical data or to identify the most frequent category/value.

Summary

Mean, median and mode give complementary information about the centre of a distribution. Understanding their definitions, calculation methods (raw and grouped), properties and behavior under skewness helps in choosing the most meaningful measure for a given real-life situation.

📌 Examples
  • Class test scores: use arithmetic mean to report the average score of the class.
  • Household income in a city: use median to represent a typical household income because incomes are usually right-skewed.
  • Most sold shoe size in a store: use mode to find the most common shoe size to stock.
  • Age distribution grouped in intervals: compute grouped mean and median using class mid-points and cumulative frequencies.
  • Customer ratings (categorical): use mode to find the most frequent rating category.
🧮 Formulas
  1. Arithmetic mean (raw data): x̄ = (Σx_i) / n
  2. Weighted mean: x̄ = (Σ w_i x_i) / (Σ w_i)
  3. Grouped mean (using mid-points): x̄ = (Σ f_i m_i) / Σ f_i where m_i is class midpoint
  4. Assumed mean method (grouped): x̄ = A + [Σ f_i d_i / Σ f_i] where d_i = (m_i - A)/h and h is class width
  5. Median (raw): position = (n + 1) / 2; median = value at that position (or average of two middle values if n even)
  6. Median (grouped): Median = L + [(n/2 - c.f) / f] × h where L = lower boundary of median class, c.f = cumulative freq before median class, f = freq of median class, h = class width
📊 Visual ideas
Histogram with a vertical line for mean, a vertical line for median, and a marker for mode — useful to visualise differences and skewness.
Ogive (cumulative frequency curve) — useful to locate the median graphically (intersection with n/2).
Box-and-whisker plot — shows median and spread, highlights outliers and skewness (does not show mean).
Frequency polygon or smoothed density curve — overlays mean/median/mode to display their relative positions for symmetric or skewed distributions.
🔢8

Arithmetic Mean

Definition: The arithmetic mean (often called simply the mean) of a data set is the sum of all observations divided by the number of observations. It is a measure of central tendency that represents the typical value of the data.

When to use: Use the arithmetic mean for quantitative (numerical) data — both ungrouped (raw) data and grouped (binned) data. It is best when data have no extreme outliers and when every value contributes equally.

Formulas & key ideas:

  • Ungrouped data: mean x̄ = (Σx)/n, where Σx is the sum of observations and n is the number of observations.
  • Grouped data (with class marks): x̄ = (Σ f x_i)/Σ f, where f is class frequency and x_i is the class mark (midpoint) x_i = (lower limit + upper limit)/2.
  • Assumed-mean method (useful for grouped data): choose an assumed mean A, let d_i = x_i − A, then x̄ = A + (Σ f d_i)/Σ f.
  • Step-deviation method (reduces big numbers): let h be class width and d'_i = (x_i − A)/h, then x̄ = A + h·(Σ f d'_i)/Σ f.
  • Combined mean of two groups: x̄ = (n1 x̄1 + n2 x̄2)/(n1 + n2).
  • Important properties: Σ(x_i − x̄) = 0 and Σ x_i = n x̄. For linear transformation y = a x + b, mean(y) = a·mean(x) + b.

Limitations: The mean is sensitive to extreme values (outliers) and may not represent a skewed distribution as well as the median.

How to compute (steps): For raw data, add all values and divide by count. For grouped data, replace each class by its class mark, multiply each class mark by the class frequency, sum these products, and divide by total frequency. For large grouped data, use assumed-mean or step-deviation to simplify arithmetic.

📌 Examples
  • Ungrouped example: Data = {12, 15, 20, 22, 18}. Mean = (12 + 15 + 20 + 22 + 18)/5 = 87/5 = 17.4.
  • Grouped example: Classes 0–9 (f=2), 10–19 (f=5), 20–29 (f=3). Class-marks: 4.5, 14.5, 24.5. Σf x_i = 2·4.5 + 5·14.5 + 3·24.5 = 155. Total f = 10. Mean = 155/10 = 15.5.
  • Assumed-mean / step-deviation (same grouped data): Choose A = 14.5. d_i = x_i − A = {−10, 0, 10}. Σf d_i = 2(−10) + 5(0) + 3(10) = 10. Mean = A + (Σf d_i)/Σf = 14.5 + 10/10 = 15.5. Using h=10, d'_i={−1,0,1}, Σf d'_i = 1, so x̄ = 14.5 + 10·(1/10) = 15.5.
🧮 Formulas
  1. Ungrouped data: x̄ = (Σ x_i) / n
  2. Grouped data (class marks): x̄ = (Σ f_i x_i) / (Σ f_i), where x_i = (l_i + u_i)/2
  3. Assumed-mean method: x̄ = A + (Σ f_i d_i) / (Σ f_i), where d_i = x_i − A
  4. Step-deviation method: x̄ = A + h·(Σ f_i d'_i) / (Σ f_i), where d'_i = (x_i − A)/h and h is class width
  5. Combined mean (two groups): x̄ = (n1 x̄1 + n2 x̄2) / (n1 + n2)
  6. Properties: Σ(x_i − x̄) = 0 and Σ x_i = n x̄
📊 Visual ideas
Histogram of grouped data with a vertical line marking the arithmetic mean (helps visualize where the mean falls relative to the bulk of data).
Frequency polygon with the mean marked — useful to see how the mean compares to the mode and shape of the distribution.
Bar chart for discrete/ungrouped data showing heights for each value and a horizontal/vertical marker for the mean.
Ogive (cumulative frequency curve) annotated with the mean (though median is read from ogive, you can project mean on the horizontal axis to compare central tendency).
🔢9

Median

Definition: The median of a dataset is the value that separates the ordered data into two equal parts — 50% of observations are less than or equal to it and 50% are greater than or equal to it. It is the 50th percentile and a measure of central tendency that is robust to extreme values.

General steps to find the median:

  1. Arrange the data in ascending order.
  2. Find the middle position (depends on the number of observations).
  3. If needed (for grouped data), use interpolation inside the median class.

1. Ungrouped (raw) data:

  • If n (number of observations) is odd: median is the value at position (n + 1)/2 in the ordered list.
  • If n is even: median is the average of values at positions n/2 and (n/2) + 1.

2. Discrete frequency distribution (individual values with frequencies):

  • Compute cumulative frequencies. Let N be total frequency. Locate the value (or class) where cumulative frequency ≥ N/2. That value (or category) is the median (or the median class if values are grouped by categories).

3. Continuous grouped frequency distribution (class intervals with frequencies):

  • Find total frequency N and cumulative frequencies up to each class.
  • Identify the median class: the first class whose cumulative frequency ≥ N/2.
  • Use linear interpolation inside the median class (assumes uniform distribution within the class) with the formula: Median = l + ((N/2 - cfbefore) / f_med) * h
  • Here, l = lower boundary of the median class (use class boundaries, not limits), cfbefore = cumulative frequency before the median class, f_med = frequency of the median class, h = class width.

Important properties:

  • Median is unaffected by extreme outliers (robust).
  • For symmetric distributions mean = median = mode (approximately).
  • Median minimizes the sum of absolute deviations: it is the value m that minimizes sum |x_i - m|.
  • Median is the 50th percentile (P50).

Worked numeric examples (briefly explained):

  1. Ungrouped odd n: Data = {12, 15, 11, 20, 18}. Sorted = {11,12,15,18,20}, n=5, position (5+1)/2 = 3 → Median = 15.
  2. Ungrouped even n: Data = {10,20,30,40}. Sorted same, n=4, positions 2 and 3: (20+30)/2 = 25 → Median = 25.
  3. Grouped continuous example: Classes: 0–10 (5), 10–20 (7), 20–30 (12), 30–40 (6). N = 30, N/2 = 15. Cumulative freqs: 5,12,24,30 → median class is 20–30 (cfbefore = 12, f_med = 12, l = 20, h = 10). Median = 20 + ((15 - 12)/12)*10 = 20 + (3/12)*10 = 20 + 2.5 = 22.5.
📌 Examples
  • Median of test scores: To report a typical student score when outliers exist (e.g., a few very low or high scores), use the median — it shows the middle performance.
  • Median household income: Less affected by extremely high incomes than the mean, so it better represents a typical family's income.
  • Median house price in a locality: Used by real-estate markets because a few very expensive houses would otherwise inflate the mean.
  • Working numerical example (odd n): {12, 15, 11, 20, 18} → Median = 15.
  • Working numerical example (continuous grouped): Classes 0–10:5, 10–20:7, 20–30:12, 30–40:6 → Median = 22.5 (calculation shown in explanation).
🧮 Formulas
  1. Ungrouped (odd n): position = (n + 1) / 2 → median = value at this position.
  2. Ungrouped (even n): median = (value at n/2 + value at n/2 + 1) / 2.
  3. Discrete frequency (using cumulative frequency): locate smallest value/category where cumulative frequency ≥ N/2 → median value/category.
  4. Grouped continuous: Median = l + ((N/2 - cfbefore) / f_med) * h, where l = lower class boundary of median class, cfbefore = cumulative frequency before median class, f_med = frequency of median class, h = class width.
  5. Relation to percentiles: Median = 50th percentile (P50).
📊 Visual ideas
Less-than and greater-than ogives: draw cumulative frequency curves (one for '≤' and one for '≥'); their intersection on the x-axis gives the median. This is a standard graphical method.
Box plot (box-and-whisker): the line inside the box marks the median; useful to compare medians across groups and visualize skewness.
Histogram or frequency polygon with a vertical line at the computed median: helps visualize where the median falls relative to distribution shape.
Cumulative frequency curve (single ogive): draw cumulative frequency vs. class boundary and read x at y = N/2 to get median (interpolate on the graph).
🔢10

Mode

Definition: The mode of a data set is the value (or values) that occur most frequently. It is the observation with the maximum frequency. Mode applies to numerical and categorical data.

Types: unimodal (one mode), bimodal (two modes), multimodal (more than two modes), or no mode (all values occur equally).

Ungrouped (raw) data: For a list of observations, identify the value(s) with the highest count. Example: in {2, 3, 5, 3, 8, 3, 5}, the mode is 3 because it appears most often.

Grouped (continuous) data: For a frequency distribution with class intervals, find the modal class (the class with largest frequency). The mode is estimated by interpolation using the formula:

Mode ≈ l + [(f_m - f_{1}) / (2f_m - f_{1} - f_{2})] × h

where l = lower limit of modal class, f_m = frequency of modal class, f_{1} = frequency of the class before modal class, f_{2} = frequency of the class after modal class, and h = class width (length of modal class interval).

Notes / Cautions:

  • If class widths are unequal, compare frequency densities (frequency ÷ class width) to choose the modal class; the standard interpolation formula assumes equal widths.
  • Mode is the only measure of central tendency applicable to categorical (nominal) data (e.g., most common colour).
  • Mode can be less stable than mean/median for small samples but is intuitive and simple to find.

Procedure to compute modal value for grouped data: identify modal class → note l, f_m, f_{1}, f_{2}, h → substitute into the formula → calculate the estimated mode.

Relation with other measures: Mode, median and mean are different summary measures; for symmetric unimodal distributions they coincide, but in skewed distributions the mode is the peak (most frequent value) and may be quite different from mean/median.

📌 Examples
  • Ungrouped single mode: Data = {2, 3, 5, 3, 8, 3, 5} → mode = 3 (occurs 3 times).
  • Bimodal data: Data = {1, 2, 2, 3, 3, 4} → modes = 2 and 3 (each occurs twice).
  • Grouped data (estimate by formula): Classes and frequencies: 0–10: 2, 10–20: 5, 20–30: 9, 30–40: 6, 40–50: 3. Modal class = 20–30 (f_m = 9), f1 = 5, f2 = 6, l = 20, h = 10. Mode ≈ 20 + [(9 − 5)/(2×9 − 5 − 6)]×10 = 20 + (4/7)×10 ≈ 25.714.
  • Categorical example: Survey of favourite fruit {apple: 40, banana: 25, mango: 60, orange: 30} → mode = mango (most chosen).
🧮 Formulas
  1. Ungrouped data: Mode = value(s) with maximum frequency.
  2. Grouped data (equal class widths): Mode ≈ l + [(f_m - f_1) / (2f_m - f_1 - f_2)] × h, where l = lower limit of modal class, f_m = frequency of modal class, f_1 = frequency of previous class, f_2 = frequency of next class, h = class width.
  3. If class widths unequal: compute frequency density = frequency / class width; choose class with highest density as modal class (interpolation formula requires modification or equal-width conversion).
📊 Visual ideas
Histogram for grouped numerical data: draw bars for class intervals; highlight (colour) the modal class bar and mark the estimated mode on the x-axis (peak of the top of modal bar).
Frequency polygon: join class mid-points by straight lines; the highest peak corresponds to the modal region — mark the peak and label the modal value.
Bar chart for categorical/ungrouped discrete data: tallest bar shows the mode (use labels with frequencies).
Stem-and-leaf plot for small numerical data: show which leaf value appears most often to read the mode directly.
📏11

Measures of Position

Definition: Measures of position locate a particular value in a data set relative to other observations. Common measures are quartiles, deciles and percentiles (including the median as the 50th percentile). They answer questions such as “What value separates the lowest 25% from the rest?”

Key ideas:

  • Order the data from smallest to largest for ungrouped (raw) data.
  • A percentile P (0<P<100) is the value below which P% of the observations lie. Quartiles and deciles are special percentiles: Q1 = 25th percentile, Q2 = 50th (median), Q3 = 75th; decile Dj = j*10th percentile.
  • If the computed position/index is not an integer, interpolate between neighbouring observations (for raw data) or use linear interpolation inside the class (for grouped data).

Ungrouped data (discrete list) — procedure and interpolation:

  • Order values x1 ≤ x2 ≤ ... ≤ xn.
  • Position (index) for the Pth percentile: i = (P/100)·(n + 1). For quartiles: Q1 at (n+1)/4, Q2 at (n+1)/2, Q3 at 3(n+1)/4. For decile Dj: position j(n+1)/10.
  • If i is integer, the value at that index is the percentile. If i = k + d (k integer, 0<d<1), percentile = x_k + d·(x_{k+1} − x_k).

Grouped data (continuous class intervals):

Find the class that contains the required cumulative frequency position (kN/100). Then use linear interpolation inside that class:

Value = L + [ (kN/100 − c.f.) / f ] · h

where L = lower boundary of the class containing the percentile, c.f. = cumulative frequency before that class, f = frequency of that class, h = class width, N = total frequency, k = percentile number (e.g., k = 25 for Q1, k = 90 for 90th percentile).

Interpretation: Saying a student is at the 85th percentile means they scored better than 85% of the group. Quartiles summarize location and spread and are used in boxplots to show central tendency and variability.

📌 Examples
  • Ungrouped data example: Data = [12, 15, 18, 20, 22, 24, 30, 35, 40]. n = 9. Q1 position = (n+1)/4 = 10/4 = 2.5 → between 2nd (15) and 3rd (18) values. Interpolate: Q1 = 15 + 0.5*(18−15) = 16.5. Q2 (median) position = (n+1)/2 = 5 → 5th value = 22. Q3 position = 3(n+1)/4 = 7.5 → between 7th (30) and 8th (35): Q3 = 30 + 0.5*(35−30) = 32.5.
  • Grouped data example: Classes and frequencies: 0–10:5, 10–20:10, 20–30:12, 30–40:8, 40–50:5. Total N = 40. Find the 90th percentile (P90). kN/100 = 90·40/100 = 36. Cumulative frequencies: 5, 15, 27, 35, 40. The 36th observation lies in class 40–50 (L = 40), c.f. before = 35, f = 5, h = 10. P90 = 40 + ((36 − 35)/5)·10 = 40 + (1/5)·10 = 42.
  • Interpretation example: If a child is at the 75th percentile for height, about 75% of children of the same age are shorter; the child is taller than three-quarters of peers.
  • Decile example: For n = 19 observations, position of the 3rd decile (D3) = 3(n+1)/10 = 3·20/10 = 6 → the 6th ordered value is D3.
🧮 Formulas
  1. \[Ungrouped (percentile position): i = (P/100)·(n + 1)\]
    \[If i = k + d (k integer, 0 ≤ d &lt\]
    \[1)\]
    \[value = x_k + d·(x_{k+1} − x_k).\]
  2. Quartile positions (ungrouped): Q1 at (n+1)/4, Q2 at (n+1)/2, Q3 at 3(n+1)/4.
  3. Decile position (ungrouped): Dj at j(n+1)/10, where j = 1,2,...,9.
  4. Grouped (percentile or kth position): value = L + [ (kN/100 − c.f.) / f ] · h, where L = lower class boundary of the containing class, c.f. = cumulative frequency before that class, f = class frequency, h = class width, N = total frequency, k = percentile number.
  5. Median for grouped data: Median = L + [ (N/2 − c.f.) / f ] · h (special case with k = 50).
📊 Visual ideas
Box-and-whisker plot (boxplot): displays Q1 (left side of box), median (line in box), Q3 (right side), whiskers to min/max; useful for visualising quartiles, spread and outliers.
Ogive (cumulative frequency curve): plot cumulative frequency against class boundaries; draw a horizontal line at k% of N and drop to the x-axis to read the kth percentile.
Histogram with vertical lines at Q1, median and Q3: helps see where these position measures fall relative to the distribution shape.
Empirical cumulative distribution function (ECDF): a step function for raw data — percentiles read directly from the y value (P/100) and the corresponding x.
⚙️12

Formulas, Techniques and Worked Procedures

Overview: This topic collects the key formulas and standard procedures used to summarise and analyse univariate statistical data (both ungrouped and grouped). It explains how to compute measures of central tendency (mean, median, mode), dispersion (variance, standard deviation), and practical techniques (assumed mean, step-deviation, interpolation for median, modal-class formula). The emphasis is on reliable step-by-step procedures you can apply to real datasets.

Key concepts and notation:

  • x or x_i : observed values (raw data)
  • f_i : frequency of the i-th class or value
  • N = Σ f_i : total frequency (sample size)
  • c_i : class boundaries or class intervals; h : class width (for equal-width classes)
  • m_i : class-mark (midpoint) = (lower bound + upper bound)/2
  • A : assumed mean (a convenient class mark used to simplify arithmetic)
  • u_i : step-deviation = (m_i − A)/h (usually an integer)

Worked procedures (step-by-step):

  • Arithmetic mean — Ungrouped data: sort data (optional), compute Σx and N, mean x̄ = Σx/N.
  • Arithmetic mean — Grouped data (class intervals): compute class-marks m_i, then x̄ = Σ(f_i m_i)/N. If numbers are large, choose an assumed mean A and use step-deviation: x̄ = A + h*(Σ f_i u_i)/N, where u_i = (m_i − A)/h.
  • Median — Ungrouped data: sort data and pick middle value (if N odd) or average of two middle values (if N even).
  • Median — Grouped data (interpolation): find cumulative frequencies and locate the median class (where cumulative frequency ≥ N/2). Use interpolation formula: Median = L + ((N/2 − c_f)/f_m) * h, where L = lower boundary of median class, c_f = cumulative frequency of preceding class, f_m = frequency of median class, h = class width.
  • Mode — Ungrouped data: the value with the highest frequency (if multimodal, there may be more than one).
  • Mode — Grouped data (modal class interpolation): identify modal class (class with largest f). Use formula: Mode = L + ((f_1 − f_0)/(2f_1 − f_0 − f_2)) * h, where f_1 is frequency of modal class, f_0 previous class frequency, f_2 next class frequency, L is lower boundary of modal class, h is class width.
  • Variance & Standard Deviation — Grouped/ungrouped: variance σ^2 = (1/N) Σ f_i (x_i − x̄)^2. Standard deviation σ = √σ^2. Use step-deviation to simplify: σ^2 = h^2 * [Σ f_i u_i^2 / N − (Σ f_i u_i / N)^2].
  • Combined (pooled) mean and variance: For two groups of sizes N1, N2 with means x̄1, x̄2: combined mean x̄ = (N1 x̄1 + N2 x̄2)/(N1 + N2). For variance, use: combined σ^2 = [N1(σ1^2 + x̄1^2) + N2(σ2^2 + x̄2^2)]/(N1+N2) − x̄^2.

Practical tips & techniques:

  • Always check if classes are continuous (use class boundaries) and of equal width; if not equal, use actual widths in formulas and class-marks carefully.
  • Choose an assumed mean A near the center to make u_i small integers — reduces arithmetic errors.
  • For median interpolation, use cumulative frequencies (either ‘less than’ ogive or cumulative Σf) to locate N/2 quickly.
  • When using modal formula, ensure modal class is clearly the highest frequency; if adjacent classes tie, mode may be ambiguous.
  • Report final answers with appropriate rounding and units (marks, rupees, kg, etc.).

When to use which method:

  • Use raw formulas for small ungrouped datasets.
  • Use grouped formulas and interpolation when data are presented in class intervals (large datasets).
  • Use step-deviation for grouped mean/variance when class-marks are large or spread widely.
📌 Examples
  • Mean (ungrouped): Given marks 72, 85, 90, 68, 95. Compute Σx = 410, N = 5, mean = 410/5 = 82.
  • Mean (grouped with assumed mean): Classes 10–20(5), 20–30(8), 30–40(12), 40–50(5). Class-marks m_i = 15,25,35,45. Choose A = 35, h = 10. Compute u_i = (m_i − 35)/10 = −2, −1, 0, 1. Then Σ f_i u_i = 5(−2)+8(−1)+12(0)+5(1) = −10−8+0+5 = −13; N = 30. Mean = 35 + 10*(−13)/30 = 35 − 4.333 = 30.667.
  • Median (grouped interpolation): Data in classes with frequencies → cumulative freq reaches N/2 inside class 30–40. Let L = 30, c_f (before) = 18, f_m = 12, h = 10, N = 50. Median = 30 + ((25 − 18)/12)*10 = 30 + (7/12)*10 = 35.83 (approx).
  • Mode (grouped): If modal class is 30–40 with f1 = 20, previous f0 = 12, next f2 = 8, L = 30, h = 10: Mode = 30 + ((20−12)/(2*20 − 12 − 8))*10 = 30 + (8/(40 − 20))*10 = 30 + (8/20)*10 = 34.
  • Variance using step-deviation: With same grouped data and u_i as above, compute Σ f_i u_i^2 and Σ f_i u_i; use σ^2 = h^2[Σ f_i u_i^2 / N − (Σ f_i u_i / N)^2] to get variance, then σ = √σ^2.
🧮 Formulas
  1. Mean (ungrouped): x̄ = (Σ x_i)/N
  2. Mean (grouped, direct): x̄ = (Σ f_i m_i)/N, where m_i = class-mark
  3. Mean (assumed mean / step-deviation): x̄ = A + h*(Σ f_i u_i)/N, u_i = (m_i − A)/h
  4. Median (ungrouped): if N odd → middle value; if N even → average of two middle values
  5. Median (grouped interpolation): Median = L + ((N/2 − c_f)/f_m) * h
  6. Mode (ungrouped): value with maximum frequency
📊 Visual ideas
Histogram: plot class intervals on x-axis and frequency on y-axis. Use continuous class boundaries, equal-width bars; useful to visualise distribution shape (symmetry, skewness, modes).
Ogive (less-than cumulative frequency curve): plot upper class boundaries on x-axis and cumulative frequencies on y-axis. The point where ogive crosses N/2 gives the median (interpolate visually).
Frequency polygon: plot class-marks vs frequency and join with straight lines (helps compare distributions and see peaks).
Box-and-whisker plot (boxplot): from five-number summary (min, Q1, median, Q3, max) — useful to visualise median, spread and outliers.
📊13

Applications and Interpretation of Statistical Results

What this topic covers: How numerical summaries and graphs (mean, median, mode, dispersion, correlation, regression, and plots) are used to describe real data, draw conclusions, make comparisons and inform decisions. It also covers common pitfalls in interpreting results (bias, outliers, correlation vs causation) and how to present results so they are meaningful and honest.

Why it matters: Statistics turn raw data into actionable information. For example, policy makers, businesses and researchers use averages, variability and relationships to decide resource allocation, set prices, evaluate programs and make predictions.

Key ideas and interpretation:

  • Central tendency (mean, median, mode): indicate a typical value. Choose median when data are skewed or contain outliers; mean is useful for further algebraic work (e.g., total/averages).
  • Dispersion (range, interquartile range, variance, standard deviation, coefficient of variation): show how spread out values are. High dispersion means more variability and less predictability.
  • Shape and outliers: skewness (right/left) affects relationship between mean and median. Outliers can distort the mean but have less effect on the median and IQR.
  • Association (correlation and regression): correlation measures strength and direction of linear association; regression provides a predictive linear model. Remember: correlation ≠ causation.
  • Sampling and bias: ensure samples are representative. Non-random sampling, response bias, or measurement error can lead to misleading summaries.
  • Comparisons: use coefficient of variation (CV) to compare relative variability of datasets with different units or scales.
  • Communicating results: choose clear graphs and report context (sample size, units, measures used). Avoid cherry-picking or mis-scaling axes that mislead.

How to interpret typical outputs:

  • Mean ± standard deviation (e.g., 50 ± 8): roughly describes the typical spread if distribution is near normal.
  • Median and IQR (e.g., median = 48, IQR = 10): gives the middle and spread of the central 50% and is robust to outliers.
  • Correlation coefficient r (e.g., r = 0.85): strong positive linear relationship; check scatter plot for nonlinearity or influential points.
  • Regression equation (e.g., y = a + bx): use for prediction within the observed range; report goodness-of-fit (r^2) and caution about extrapolation.

Common cautions:

  • Do not infer causation from correlation without controlled experiments or additional evidence.
  • Watch for hidden averages (Simpson's paradox) where aggregated trends reverse when data are disaggregated.
  • Report sample size and measures of uncertainty where possible; small samples can give unstable estimates.
📌 Examples
  • Median vs Mean: A town's house prices (one luxury villa much larger than others) — the mean is pulled up by the villa, while the median better represents a typical buyer's price.
  • Outlier detection: In students' test scores, a single 0 (absent student) lowers the class mean substantially but has less effect on the median and IQR; a boxplot helps visualize this outlier.
  • Correlation but not causation: A strong positive correlation between ice cream sales and drowning incidents during summer — both rise due to a lurking variable (temperature), not because ice cream causes drowning.
  • Using regression for prediction: Using linear regression of weight (y) on height (x) to estimate expected weight for a given height, reporting the regression equation and r^2, and warning against predicting for heights far outside the sample.
  • Comparing variability with CV: Two machines produce bolts with mean diameters 10 mm (SD = 0.5) and 50 mm (SD = 1.0). CVs (5% vs 2%) show the first machine has larger relative variability despite smaller absolute SD.
🧮 Formulas
  1. Mean (ungrouped): x̄ = (Σx_i)/n
  2. Mean (grouped): x̄ = (Σ f_i x_i)/N where f_i = frequency, x_i = class mark, N = Σ f_i
  3. Median (grouped): Median = l + ((N/2 − C_f)/f) × h where l = lower boundary of median class, C_f = cumulative freq before median class, f = freq of median class, h = class width
  4. \[Mode (grouped): Mode = l + ((f_m − f_{prev})/(2f_m − f_{prev} − f_{next})) × h where f_m is freq of modal class\]
  5. Variance (population, grouped): σ^2 = (Σ f_i x_i^2)/N − (x̄)^2
  6. Sample variance: s^2 = Σ(x_i − x̄)^2/(n − 1)
📊 Visual ideas
Histogram with class boundaries and a superimposed frequency polygon — to show distribution shape, skewness and modality.
Box-and-whisker plot (boxplot) — to show median, IQR, and outliers clearly; useful for comparing groups side-by-side.
Ogive (cumulative frequency curve) — to find medians, quartiles and percentiles graphically.
Scatter plot with best-fit regression line and reporting r and r^2 — to visualize linear association and make predictions; mark influential points if any.

Key Concepts

Statistics
The science of collecting, organizing, presenting, analysing and interpreting numerical data to make decisions.
Population
The complete set of items or individuals under study.
Sample
A subset of the population selected for analysis, ideally representative of the population.
Variable
A characteristic or quantity that can take different values (qualitative or quantitative).
Attribute
A qualitative variable describing a quality or category rather than a number.
Frequency
The number of times a particular value or class occurs in the data.
Frequency distribution
A tabular representation showing classes or values and their corresponding frequencies.
Class interval
A continuous range of values used to group data in a frequency distribution.
Class mark (Midpoint)
The value halfway between the lower and upper limits of a class: (lower + upper)/2.
Cumulative frequency
The running total of frequencies up to and including a given class or value.
Relative frequency
The fraction or proportion of the total corresponding to a particular frequency (frequency/total).
Frequency density
Used with unequal class widths; frequency per unit class width = frequency / class width.
Histogram
A bar-type graph showing class intervals on x-axis and frequency (or density) on y-axis; bars touch each other.
Frequency polygon
A line graph obtained by joining midpoints of class tops (class marks) plotted against frequencies.
Ogive (Cumulative frequency curve)
A graph of cumulative frequency against class boundaries (or marks), used to find medians/quartiles graphically.
Mean (Arithmetic mean)
Sum of observations divided by number of observations; for grouped data use sum(f × class mark)/sum(f).
Median
The middle value that divides ordered data into two equal parts; for grouped data use interpolation in the median class.
Mode
The value (or class) with the highest frequency; for grouped data use modal class and interpolation formula.
Range
The difference between the maximum and minimum observations; a simple measure of spread.
Standard deviation
Square root of the mean of squared deviations from the mean; measures spread around the mean (population form).

End-of-Chapter Trial Paper & Test Questions

Topic-wise questions to test your understanding of every concept in this chapter.

  1. Define the class mark of a class interval and find the class mark of the interval 40-49. / वर्ग अंतराल के वर्ग चिह्न को परिभाषित कीजिए और अंतराल 40-49 का वर्ग चिह्न ज्ञात कीजिए।
    Show answer

    The class mark is the mid-point of a class, given by (lower limit + upper limit)/2; for 40-49 it is (40+49)/2 = 44.5. / वर्ग चिह्न किसी वर्ग का मध्य-बिंदु होता है, जो (निम्न सीमा + उच्च सीमा)/2 से दिया जाता है; 40-49 के लिए यह (40+49)/2 = 44.5 है।

  2. Why is the median preferred over the mean for representing household income data? / पारिवारिक आय के आँकड़ों को दर्शाने के लिए माध्य की तुलना में माध्यिका को क्यों प्राथमिकता दी जाती है?
    Show answer

    Income data is usually right-skewed with a few very large values, and the mean is sensitive to such extreme outliers, whereas the median is robust and not affected by them. / आय के आँकड़े प्रायः दाईं ओर विषम होते हैं जिनमें कुछ बहुत बड़े मान होते हैं, और माध्य ऐसे चरम बाह्यमानों के प्रति संवेदनशील होता है, जबकि माध्यिका दृढ़ होती है और उनसे प्रभावित नहीं होती।

  3. Compute the mean of the ungrouped data 12, 15, 20, 22, 18. / अवर्गीकृत आँकड़ों 12, 15, 20, 22, 18 का माध्य ज्ञात कीजिए।
    Show answer

    Mean = (12+15+20+22+18)/5 = 87/5 = 17.4. / माध्य = (12+15+20+22+18)/5 = 87/5 = 17.4।

  4. For grouped data with classes 0-10:5, 10-20:7, 20-30:12, 30-40:6, find the median. / वर्गीकृत आँकड़ों जिनके वर्ग 0-10:5, 10-20:7, 20-30:12, 30-40:6 हैं, की माध्यिका ज्ञात कीजिए।
    Show answer

    N=30, N/2=15; cumulative frequencies are 5,12,24,30 so median class is 20-30 (l=20, cf=12, f=12, h=10); Median = 20 + ((15-12)/12)×10 = 22.5. / N=30, N/2=15; संचयी बारंबारताएँ 5,12,24,30 हैं अतः माध्यिका वर्ग 20-30 है (l=20, cf=12, f=12, h=10); माध्यिका = 20 + ((15-12)/12)×10 = 22.5।

  5. Explain how the relative positions of mean, median and mode indicate the skewness of a distribution. / माध्य, माध्यिका और बहुलक की सापेक्ष स्थितियाँ किसी बंटन की विषमता को कैसे दर्शाती हैं, समझाइए।
    Show answer

    For a symmetric distribution Mean ≈ Median ≈ Mode; for right (positive) skew Mode < Median < Mean; for left (negative) skew Mean < Median < Mode. / सममित बंटन के लिए माध्य ≈ माध्यिका ≈ बहुलक; दाईं (धनात्मक) विषमता के लिए बहुलक < माध्यिका < माध्य; बाईं (ऋणात्मक) विषमता के लिए माध्य < माध्यिका < बहुलक।

  6. For the grouped data 0-10:2, 10-20:5, 20-30:9, 30-40:6, 40-50:3, estimate the mode. / वर्गीकृत आँकड़ों 0-10:2, 10-20:5, 20-30:9, 30-40:6, 40-50:3 के लिए बहुलक का आकलन कीजिए।
    Show answer

    Modal class is 20-30 with f_m=9, f_1=5, f_2=6, l=20, h=10; Mode = 20 + [(9-5)/(2×9-5-6)]×10 = 20 + (4/7)×10 ≈ 25.71. / बहुलक वर्ग 20-30 है जिसमें f_m=9, f_1=5, f_2=6, l=20, h=10; बहुलक = 20 + [(9-5)/(2×9-5-6)]×10 = 20 + (4/7)×10 ≈ 25.71।

  7. How is the median of a grouped distribution found graphically using ogives? / तोरण (ogives) का उपयोग करके वर्गीकृत बंटन की माध्यिका आलेखीय रूप से कैसे ज्ञात की जाती है?
    Show answer

    Draw both the less-than and greater-than ogives on the same axes; the x-coordinate of their point of intersection gives the median (equivalently, draw a horizontal line at N/2 to the single ogive and read the x-value). / एक ही अक्षों पर 'से कम' और 'से अधिक' दोनों तोरण खींचिए; उनके प्रतिच्छेदन बिंदु का x-निर्देशांक माध्यिका देता है (समतुल्य रूप से, एकल तोरण पर N/2 पर एक क्षैतिज रेखा खींचकर x-मान पढ़िए)।

  8. State the step-deviation method formula for the mean and explain why it is useful. / माध्य के लिए पद-विचलन विधि का सूत्र लिखिए और बताइए कि यह उपयोगी क्यों है।
    Show answer

    x̄ = A + h·(Σf_i d'_i)/(Σf_i), where A is the assumed mean, h is the class width and d'_i = (x_i − A)/h; it is useful because dividing by h reduces large numbers, simplifying the arithmetic for grouped data. / x̄ = A + h·(Σf_i d'_i)/(Σf_i), जहाँ A कल्पित माध्य है, h वर्ग चौड़ाई है और d'_i = (x_i − A)/h; यह उपयोगी है क्योंकि h से भाग देने पर बड़ी संख्याएँ छोटी हो जाती हैं, जिससे वर्गीकृत आँकड़ों की गणना सरल हो जाती है।

Related Laws & Principles

Explore all

Foundational laws & principles behind this chapter. Each one opens a full page — what it says, why it matters, five practice questions and the mistakes to avoid.

Loading related laws…
Sourced from 177 content files · LLOS Learn · browse all chapters