Overview
Introduction: This chapter introduces the role and application of statistical tools in economics. It explains what statistics is, why economists use it, and how statistical tools convert raw facts into meaningful information for description, comparison, interpretation and policy making. Importance: The chapter emphasises statistics as an essential language of economic analysis — for summarising large volumes of data, detecting patterns and trends, making comparisons across time and space, and supporting decisions in business, government and research. Key themes: types and sources of data (primary and secondary), methods of data collection (census and sampling, surveys and observation), organisation and presentation of data (frequency distributions, tables, diagrams and graphs such as bar charts, histograms, ogives and pie charts), and summary measures (measures of central tendency and dispersion). The chapter also discusses practical uses of these tools in real-life economic contexts (price and output analysis, income distribution, unemployment statistics, inflation measurement) and cautions about limitations and misuse of statistics. What the student will learn: how to classify…
Learning Objectives
- Define statistics, population, sample, variable and distinguish between them
- Explain primary and secondary data and describe common methods of data collection (census, sample survey, experiments)
- Construct frequency distribution tables for discrete and continuous data from raw data
- Draw and interpret bar diagrams, pie charts, histogram, frequency polygon and ogive for given data
- Calculate arithmetic mean, median and mode for ungrouped and grouped data using direct and short‑cut methods
- Compute weighted mean for grouped data and apply it to solve economic problems
- Compute range, quartile deviation, mean deviation, variance and standard deviation for grouped and ungrouped data
- Calculate coefficient of variation and use it to compare variability of different data sets
Topics in this chapter
10 topics · tap a topic title to jump straight to it.
Introduction to Statistics
Introduction to Statistics
Key Point: Mean (ungrouped): x̄ = (Σx_i) / n, where x_i are observations and n is number of observations.
What is Statistics? Statistics is a branch of mathematics that deals with collection, presentation, analysis and interpretation of numerical data to make decisions in the presence of uncertainty. In economics it helps summarize economic phenomena, draw conclusions and support policy and business decisions.
Main Functions of Statistics
- Collection: Gathering relevant data (surveys, experiments, records).
- Presentation: Organising data into tables, charts and diagrams for clarity.
- Analysis: Summarising data with measures (mean, median, mode, dispersion).
- Interpretation: Drawing conclusions, making forecasts and supporting decisions.
Types of Data
- Qualitative (categorical): e.g., gender, occupation (nominal/ordinal).
- Quantitative (numerical): e.g., income, age (discrete or continuous).
Sources of Data
- Primary data: Collected firsthand (surveys, experiments).
- Secondary data: Collected earlier by others (official records, publications).
Population and Sample — Population is the full set of items/people under study; a sample is a subset chosen for study. Proper sampling methods ensure representativeness.
Steps in a Statistical Investigation
- Define objective and variables.
- Decide source and method of data collection.
- Collect data.
- Classify and tabulate data (frequency distribution).
- Summarise and analyse using measures and graphs.
- Interpret results and report findings.
Presentation of Data — Raw data are made meaningful by arranging into frequency distributions, tables, and by drawing diagrams and graphs: bar charts, histograms, pie charts, frequency polygons, ogives.
Measures Used — Two broad kinds: measures of central tendency (mean, median, mode) and measures of dispersion (range, variance, standard deviation, coefficient of variation). These describe the typical value and spread of data.
Uses of Statistics (especially in Economics)
- Summarise large amounts of economic data (national income, prices).
- Compare different groups or periods (growth rates, inflation).
- Forecast and inform policy (demand estimation, resource allocation).
- Test hypotheses (evaluate claims about populations).
Limitations and Precautions
- Quality depends on the quality of data; biased or incomplete data produce misleading results.
- Care needed in sample selection, question design, and interpretation.
- Correlation does not imply causation.
- Class test marks: A teacher collects marks of 40 students to compute class average (mean), median and identify most frequent score (mode) to assess performance.
- Household income survey: A government samples households to estimate average monthly income and income distribution to design welfare programmes.
- Market research: A company surveys customers on preferred product features (categorical data) and purchase frequency (numerical) to plan production.
- Weather analysis: Meteorological department uses daily rainfall data to compute monthly averages and variability for crop planning and flood warnings.
- \[Mean (ungrouped): x̄ = (Σx_i) / n\]\[where x_i are observations and n is number of observations.\]
- \[Mean (grouped): x̄ = Σ(f_i * m_i) / Σf_i\]\[where f_i = class frequency\]\[m_i = class midpoint.\]
- \[Median (ungrouped): For ordered data\]\[if n is odd median = value at (n+1)/2\]\[if even\]\[median = average of values at n/2 and (n/2 + 1).\]
- \[Median (grouped): Median = L + [(N/2 - c.f) / f] * h\]\[where L = lower boundary of median class\]\[N = total frequency\]\[c.f = cumulative frequency before median class\]\[f = frequency of median class\]\[h = class width.\]
- \[Mode (ungrouped): Value with highest frequency\]\[If bimodal/multimodal there are multiple modes.\]
- \[Mode (grouped): Mode = L + [(f1 - f0) / (2f1 - f0 - f2)] * h\]\[where f1 = frequency of modal class\]\[f0 = freq of previous class\]\[f2 = freq of next class\]\[L = lower boundary of modal class\]\[h = class width.\]
Collection of Data
Collection of Data
Key Point: Total frequency: N = Σ f_i (sum of all class/sample frequencies)
What is Collection of Data? Collection of data is the process of gathering information needed to answer a research question or to compile statistics. In economics and statistics it is the first and most important step in any empirical study because the quality of results depends on the quality of data collected.
Types of data by source:
- Primary data: Data collected first-hand for a specific purpose (e.g., through surveys, interviews, observation, experiments).
- Secondary data: Data already collected by others for some purpose (e.g., government publications, books, journals, administrative records, online databases).
Methods of collecting primary data
- Questionnaire / Schedule: Structured set of written questions. A questionnaire is usually filled by the respondent; a schedule is filled by the enumerator.
- Interview: Face-to-face, telephone, or online interviews—useful for detailed or qualitative responses.
- Observation: Recording behaviour/events directly (participant or non‑participant observation).
- Experimentation: Controlled collection of data by manipulating variables (used less often in basic economic surveys).
Sampling vs Census
- Census: Collecting data from every unit of the population — accurate but time-consuming and costly.
- Sample survey: Collecting data from a representative subset — less costly, quicker, with sampling error.
Important considerations and steps
- Define objectives clearly (what to measure and why).
- Define the population and sampling frame.
- Choose a method (census or sample; decide on sampling technique: simple random, stratified, systematic, cluster, etc.).
- Design instrument (questionnaire/schedule): simple language, clear order, pre-test (pilot survey).
- Train enumerators and carry out data collection ethically (consent, confidentiality).
- Validate, code, and record data carefully to avoid errors (checking for missing/illogical values).
Quality aspects: Reliability (consistency of measurement), validity (measuring what is intended), accuracy (closeness to true value), timeliness, and cost-effectiveness. Non-sampling errors (response bias, recording errors, non-response) are often more serious than sampling errors.
- Household income survey: A local NGO uses a stratified random sample of households in a town and a structured questionnaire to estimate average household income and sources of livelihood.
- Product-market research: A company conducts online questionnaires and in-store observations to collect consumer preferences before launching a new product.
- Government census: The national population census collects demographic and socio-economic information from every household (census method).
- School performance data: A school collects students' test scores (secondary data from school records) and runs a small survey among parents (primary data) to study factors affecting performance.
- Health surveillance: A public health department uses routine administrative records (secondary) and special field surveys (primary) to monitor the spread of a disease.
- \[Total frequency: N = Σ f_i (sum of all class/sample frequencies)\]
- \[Relative frequency of class i: r_i = f_i / N\]
- \[Percentage of class i: p_i = (f_i / N) × 100\]
- \[Cumulative frequency (up to class k): CF_k = Σ_{i=1 to k} f_i\]
- \[Response rate (for surveys): Response rate (%) = (Number of completed responses / Number of eligible units contacted) × 100\]
- \[Non-response rate (%) = 100 - Response rate (%)\]
Sampling Methods
Sampling Methods
Key Point: Sample mean: x̄ = (1/n) Σ xi
What is sampling? Sampling is the process of selecting a part (sample) of a population to draw inferences about the whole. Sampling is used when a census is impractical, costly, or time-consuming.
Key concepts: population, sampling frame (list from which sample is drawn), sample size, sampling unit, sampling error (difference between sample estimate and true population value), and bias (systematic error).
Major types of sampling
- Probability sampling — every unit has a known (non-zero) chance of selection. Common methods:
- Simple random sampling (SRS): every sample of size n from N has equal chance. Use random numbers or lotteries.
- Systematic sampling: choose every k-th unit from an ordered list after a random start (k = N/n).
- Stratified sampling: divide population into homogeneous strata (groups) and draw samples from each stratum. Improves precision when strata differ.
- Cluster sampling: divide population into clusters (usually naturally occurring groups like villages), randomly select clusters, and study all or a sample of units within chosen clusters.
- Multistage sampling: a generalization of cluster sampling where selection occurs in stages (for example, districts → schools → students).
- Probability proportional to size (PPS): clusters or units are selected with probability proportional to their size; useful when cluster sizes vary.
- Non-probability sampling — selection is not based on known probabilities. Common methods:
- Convenience sampling: choose easily available units (fast but high bias).
- Judgment/purposive sampling: researcher selects units judged to be typical or knowledgeable.
- Quota sampling: sample reflects certain characteristics of the population (age, sex) but selection within quotas is non-random.
- Snowball sampling: existing respondents recruit further respondents (useful for hard-to-reach populations).
How to choose a method? Use probability methods when you need unbiased estimates and the sampling frame exists. Use stratification to increase precision when subgroups differ. Use cluster or multistage sampling for wide geographic areas to reduce cost. Non-probability methods are suitable for exploratory or quick studies where inference to a larger population is not required.
Precision and sample size: precision improves with larger n. For many estimators the sample mean has an approximately normal sampling distribution (Central Limit Theorem) for moderately large n, enabling confidence intervals and hypothesis tests.
Advantages and disadvantages (brief)
- Probability sampling: less bias, allows estimation of sampling error; but requires frame and can be costlier.
- Non-probability sampling: cheaper and quicker; results are subject to selection bias and cannot support reliable measures of sampling error.
Common pitfalls: incomplete or biased sampling frame, nonresponse bias, improper use of systematic or quota selection without random start, treating non-probability samples as if they were probability samples.
- Simple random sampling: From a school roll of 1,000 students, select 100 students using random numbers to study average study hours.
- Systematic sampling: From a production line, inspect every 20th item after a random start to estimate defect rate.
- Stratified sampling: To study household income in a city, divide population by wards (strata) and sample proportionally from each ward to ensure representation.
- Cluster sampling: For a national education survey, randomly select 30 schools (clusters) and survey all students in selected schools.
- Multistage sampling: Select districts at random, then randomly select schools within chosen districts, then randomly select students within schools.
- PPS sampling: When villages vary in population, select villages with probability proportional to village population so larger villages are more likely to be selected.
- \[Sample mean: x̄ = (1/n) Σ xi\]
- \[Sample variance (unbiased): s^2 = (1/(n-1)) Σ (xi - x̄)^2\]
- \[Sample proportion: p̂ = x/n where x is number of successes in sample\]
- \[Standard error of the mean (SRS\]\[large population approx): SE(x̄) = σ / √n (σ is population SD\]\[estimate with s when σ unknown)\]
- \[Standard error of proportion: SE(p̂) = √(p̂(1 - p̂) / n)\]
- \[Finite population correction (FPC) factor: FPC = √((N - n) / (N - 1))\]
Classification and Tabulation
Classification and Tabulation
Key Point: Range = Maximum value − Minimum value
What is Classification?
Classification is the process of arranging raw data into groups or classes based on common characteristics so that it becomes easier to study, compare and draw conclusions. It reduces complexity and reveals patterns.
Types of Classification
- By nature of data: Qualitative (categorical) vs Quantitative (numerical).
- By measurement: Discrete (countable) vs Continuous (measurable).
- By number of attributes: One-way (univariate), Two-way (bivariate) and Multi-way (multivariate) classification.
Important principles of good classification
- Mutually exclusive: an observation must belong to one class only.
- Exhaustive: all observations must be covered.
- Homogeneous within classes and heterogeneous between classes.
- Simple and convenient for analysis.
- Prefer equal class-widths for continuous data (unless there is reason not to).
What is Tabulation?
Tabulation is the systematic arrangement of classified data in rows and columns (tables) with clear headings. A good table has a title, labeled rows/columns, units, and totals where appropriate. Tabulation summarizes data and prepares it for graphical presentation and calculation of statistical measures.
Steps to classify and tabulate numerical data
- Decide the objective and the variable(s) to be classified.
- For quantitative continuous data determine range = Max – Min.
- Choose number of classes (k) and class width (h).
- Construct class intervals (mutually exclusive and exhaustive).
- Assign each observation to a class and tally frequencies.
- Compute relative frequency, percentage frequency and cumulative frequency.
- Present results in a clear table with totals.
Simple example table (student marks)
| Class Interval | Frequency (f) | Relative Frequency (f/N) | Cumulative Frequency |
|---|---|---|---|
| 0 – 19 | 1 | 0.0667 | 1 |
| 20 – 39 | 2 | 0.1333 | 3 |
| 40 – 59 | 5 | 0.3333 | 8 |
| 60 – 79 | 4 | 0.2667 | 12 |
| 80 – 99 | 3 | 0.2000 | 15 |
| Total | 15 | 1.0000 | — |
Uses of classification & tabulation
- Simplify large datasets for interpretation.
- Prepare data for graphs (histogram, pie chart, ogive, etc.).
- Calculate central tendency, dispersion and other statistical measures.
- Compare subgroups (e.g., income classes, age groups, regions).
- Classifying households by monthly income ranges (e.g., <5000, 5000–14999, 15000–29999, 30000+), then tabulating the number of households in each range to study poverty and consumption patterns.
- Grouping students by marks (0–19, 20–39, 40–59, 60–79, 80–100) and creating a frequency table to determine pass/fail rates and the distribution of performance.
- Classifying firms by number of employees (small, medium, large) and tabulating counts by industry to analyze employment concentration.
- Tabulating survey responses (qualitative) such as choice of transport (bus, train, car, bicycle) into a frequency table and converting to percentages for presentation.
- \[Range = Maximum value − Minimum value\]
- \[Sturges' rule (suggested number of classes): k ≈ 1 + 3.322 log10 N (N = sample size)\]
- \[Class width (approx): h ≈ Range / k (round up to convenient value)\]
- \[Class mark (midpoint) = (Lower limit + Upper limit) / 2\]
- \[Relative frequency = f / N (where f = class frequency\]\[N = total observations)\]
- \[Percentage frequency = (f / N) × 100\]
Diagrammatic and Graphic Presentation
Diagrammatic and Graphic Presentation
Key Point: Angle for pie chart (degrees) = (Category frequency / Total frequency) × 360
Diagrammatic and Graphic Presentation refers to methods of representing statistical data visually so patterns, trends and comparisons become easy to understand. Good diagrams/graphs make complex numerical information clear at a glance and are essential for communication in economics.
Key principles and conventions
- Every diagram must have a clear title, labeled axes (with units), an appropriate scale, and source/footnote if necessary.
- Choose the type of diagram that suits the data: qualitative categories → bar/column/pie; grouped frequency data → histogram, frequency polygon, ogive; paired numerical data → scatter diagram.
- Maintain correct proportions, avoid distorted scales, and include a legend when more than one series is shown.
Common types and how to construct them
- Bar (Column) Diagram: Represents categorical data by rectangular bars. Use gaps between bars for discrete categories. Bars can be simple, multiple (side-by-side for comparison), component (stacked parts of a total) or sub-divided.
- Pie Chart: Shows parts of a whole. Convert each category's share to an angle: angle = (frequency / total) × 360°. Use when categories sum to a meaningful total (e.g., budget composition).
- Histogram: For continuous grouped data. Draw contiguous rectangles whose heights equal class frequency (or frequency density if class widths are unequal). No gaps between bars.
- Frequency Polygon: Plot class mid-points on x-axis and join successive points by straight lines. Useful for comparing distributions.
- Ogive (Cumulative Frequency Curve): Plot cumulative frequency against upper (or lower) class boundaries; connects points to show cumulative distribution and to read medians/percentiles.
- Frequency Curve: A smoothed version of the frequency polygon; used to show general shape of distribution.
- Scatter Diagram: Plot paired observations (x,y) to study relationship or correlation. Add a line of best fit to indicate trend.
Advantages
- Quick visual comparison and pattern recognition (trends, peaks, gaps).
- Accessible to non-technical audiences.
- Highlights relationships between variables (with scatter diagrams) and composition (with pie/component bars).
Limitations
- Can be misleading if scales or proportions are manipulated.
- Lose precise numerical detail—diagrams summarize rather than show raw data.
- Some graphs (like pie charts) are unsuitable for many categories or very similar shares.
When to use which type (quick guide)
- Comparing categories: Bar/column diagram.
- Showing parts of a whole: Pie chart or stacked bar.
- Distribution of continuous data: Histogram, frequency polygon, frequency curve.
- Cumulative analysis (median, quartiles): Ogive.
- Studying relationship between two numerical variables: Scatter diagram.
Using these methods correctly helps students and analysts present economic data (like income distribution, expenditure composition, production trends, price movements) clearly and honestly.
- Monthly household expenditure composition: use a pie chart to show shares of food, rent, education, transport. Calculate each share as (category expenditure / total expenditure) × 100 and convert to angles for the pie.
- Student marks distribution in an exam: group marks into class intervals and draw a histogram to show how many students fall in each range; use an ogive to read the median mark.
- Comparing unemployment rates across three states for two years: use a multiple (side-by-side) bar chart with states on the x-axis and unemployment rate on the y-axis, with different colored bars for each year.
- Relationship between advertising spend and sales: plot advertising expenditure on the x-axis and sales on the y-axis in a scatter diagram; add a trend line to see correlation.
- Population by age groups: use a component/stacked bar diagram to show male and female populations within each age group (useful for age-sex pyramids or comparative population structure).
- \[Angle for pie chart (degrees) = (Category frequency / Total frequency) × 360\]
- \[Percentage share (%) = (Category frequency / Total frequency) × 100\]
- \[Class mid-point = (Lower class limit + Upper class limit) / 2\]
- \[Class width = Upper class limit − Lower class limit (for equal-width classes)\]
- \[Frequency density (when class widths differ) = Frequency / Class width (use height = frequency density for histogram)\]
- \[Cumulative frequency (CF) at a boundary = Sum of frequencies up to that boundary\]
Measures of Central Tendency
Measures of Central Tendency
Key Point: Arithmetic mean (ungrouped): x̄ = Σx / n
Definition: Measures of central tendency are statistical values that describe the center or typical value of a dataset. The main measures are Arithmetic Mean, Median and Mode. They summarize data with a single representative number and are used to compare datasets, study distribution and make decisions.
Why they matter (Uses):
- Simplify large datasets to a single typical value (e.g., average income, average marks).
- Compare different groups (e.g., regions, years).
- Serve as inputs for other statistical measures (variance, standard deviation).
Desirable properties: They should be: representative, simple to compute, based on all observations (in case of mean), and stable (not too sensitive to small changes).
1. Arithmetic Mean (Average)
Definition: Sum of all observations divided by the number of observations. Useful when all values are equally important and data are not highly skewed or contain extreme outliers.
For ungrouped data (individual observations): x̄ = Σx / n
For grouped data (class intervals with mid-points x_i and frequencies f_i):
x̄ = (Σ f_i x_i) / N, where N = Σ f_i
Step-deviation method (useful for large class widths): choose a convenient origin a and class size h, set u_i = (x_i − a)/h, then
x̄ = a + h*(Σ f_i u_i / N)
2. Median
Definition: The middle value that divides the distribution into two equal parts (50% below and 50% above). Preferred for skewed distributions and when outliers are present.
For ungrouped (sorted) data:
- If n is odd → median is the ((n+1)/2)-th value.
- If n is even → median is the average of (n/2)-th and (n/2+1)-th values.
For grouped data (continuous classes): use linear interpolation within the median class:
Median = l + [((N/2) − c.f.) / f_m] * h
where l = lower boundary of median class, N = total frequency, c.f. = cumulative frequency before median class, f_m = frequency of median class, h = class width.
3. Mode
Definition: The value(s) that occur most frequently. Useful to identify the most common category or value (e.g., most common family size, most sold product).
For ungrouped data: the observation with highest frequency (can be unimodal, bimodal, multimodal).
For grouped data (continuous classes): estimate mode by interpolation using the modal class (class with maximum frequency):
Mode = l + [ (f1 − f0) / (2f1 − f0 − f2) ] * h
where l = lower boundary of modal class, f1 = frequency of modal class, f0 = frequency of class before, f2 = frequency of class after, h = class width.
Comparison & selection:
- Use mean when distribution is symmetric and data have no extreme outliers.
- Use median when distribution is skewed or contains outliers (it resists extremes).
- Use mode for categorical data or to find the most frequent value.
Limitations: Mean is sensitive to outliers; median ignores detailed values and only depends on rank; mode may be non-unique and unstable for small samples.
Practical tips for computation:
- Always check the data type: quantitative (mean, median, mode) vs categorical (mode only).
- For grouped data ensure class intervals are continuous and of equal width (or note the width in formulas).
- When using formulas for grouped data, use class mid-points for mean and linear interpolation for median and mode.
- Example 1 (Ungrouped numeric data — marks): Data: 45, 67, 78, 45, 56, 67, 89, 67, 45, 72. n = 10. Sum = 631 → Mean = 631/10 = 63.1. Sorted values: 45,45,45,56,67,67,67,72,78,89. Median = average of 5th & 6th = (67+67)/2 = 67. Mode = 45 and 67 (both occur 3 times) → bimodal.
- Example 2 (Grouped data — income classes): Classes (in 000s): 0–10(10), 10–20(25), 20–30(40), 30–40(20), 40–50(5). Mid-points x_i = 5,15,25,35,45; compute Σ f_i x_i and N = 100. Mean = (Σ f_i x_i)/100 gives average income. For skewed incomes use median from cumulative frequencies (find class where cumulative ≥ 50).
- Example 3 (Median in real life): When reporting house prices in a city with a few very expensive homes, median price is preferred to indicate a typical house price because it is not affected by the extreme high values.
- Example 4 (Mode in real life): A retail store analyzes shoe sizes sold; the mode (most frequently sold size) helps stock inventory. For categorical data such as favorite color, mode is the only central measure.
- \[Arithmetic mean (ungrouped): x̄ = Σx / n\]
- \[Arithmetic mean (grouped): x̄ = (Σ f_i x_i) / N\]\[where N = Σ f_i and x_i are class mid-points\]
- \[Step-deviation mean: x̄ = a + h*(Σ f_i u_i / N)\]\[with u_i = (x_i − a)/h\]\[a = assumed mean\]\[h = class size\]
- \[Median (ungrouped): If n odd → median is value at position (n+1)/2\]\[if n even → median = average of values at positions n/2 and n/2 + 1\]
- \[Median (grouped): Median = l + [((N/2) − c.f.) / f_m] * h\]\[where l = lower boundary of median class\]\[c.f. = cumulative frequency before median class\]\[f_m = frequency of median class\]\[h = class width\]
- \[Mode (ungrouped): value(s) with highest frequency\]
Measures of Dispersion
Measures of Dispersion
Key Point: Range = Maximum value − Minimum value
What are Measures of Dispersion?
Measures of dispersion describe how spread out or scattered the values of a data series are around a central value (mean, median or mode). While measures of central tendency (mean, median, mode) tell us the centre of data, measures of dispersion tell us the variability or consistency of the data. Common measures: Range, Quartile Deviation (QD), Mean Deviation (MD), Variance and Standard Deviation (SD).
Why they matter
Two distributions can have the same mean but very different spreads. Dispersion helps compare variability (e.g., marks of two classes, incomes of two cities, temperature changes), decide reliability of averages and assess risk (e.g., stock volatility).
Short descriptions
- Range: Difference between maximum and minimum values. Easiest but most affected by extremes.
- Quartile Deviation (QD) / Semi-interquartile range: Half the difference between third and first quartiles; resistant to outliers and shows spread of middle 50% of data.
- Mean Deviation (MD): Average of absolute deviations from a chosen central value (mean or median). MD about median gives minimum average absolute deviation.
- Variance: Average of squared deviations from the mean; gives squared units.
- Standard Deviation (SD): Square root of variance; same units as data and widely used because of good mathematical properties.
- Coefficient of Variation (CV): (SD/mean) × 100% — a relative measure useful to compare variability of series with different units or means.
Properties & Practical points
- All measures ≥ 0. Zero dispersion means all observations equal.
- Range and QD are simple and robust (QD more robust than range).
- Variance and SD use all observations and are sensitive to outliers (extreme values).
- SD has preferable algebraic properties (useful in statistical inference and probability).
- CV is dimensionless; use it to compare variability across datasets with different units or widely different means.
- For grouped data, use class mid-points (xi) and frequencies (fi).
How to compute (conceptual steps)
- Range: sort data, subtract min from max.
- QD: find Q1 (25th percentile) and Q3 (75th percentile) from sorted data (or from ogive for grouped data); QD = (Q3 - Q1)/2.
- MD: compute deviations from chosen center, take absolute values, average them (use frequencies for grouped data).
- SD: find mean, compute squared deviations, average (population) or average adjusted by (n-1) for sample, then square-root.
When to use which
- Use range or QD for quick, robust summary (QD when outliers present).
- Use MD when interest is in average absolute deviation (simpler interpretation).
- Use SD and variance for most statistical work, modelling and comparisons; use CV to compare across different units or scales.
Class 11 (CBSE) tips
- For grouped data, always use class mid-points to represent class values.
- When classes are open or unequal, take care in choosing representative values and class width h in computational shortcuts.
- Remember that median minimizes the sum of absolute deviations; mean minimizes the sum of squared deviations.
- Comparing marks: Two sections have the same average score (70). Section A scores are tightly clustered around 70 (low dispersion). Section B has many very high and very low scores (high dispersion). SD will be much larger for Section B.
- Income distribution: City X and City Y have equal mean incomes, but City Y has a small number of extremely high incomes and many low incomes. Range and SD for City Y will be larger; QD will show spread of middle incomes.
- Daily temperatures: Meteorologists use SD to quantify how variable daily temperatures are in a month. A small SD implies steady weather; a large SD implies frequent extremes.
- Stock returns: Investors use coefficient of variation (CV) to compare risk (volatility) relative to expected return. A stock with higher CV is riskier relative to its mean return.
- \[Range = Maximum value − Minimum value\]
- \[Quartile Deviation (QD) = (Q3 − Q1) / 2\]
- \[Mean (ungrouped) x̄ = Σxi / n\]
- \[Mean Deviation (about mean\]\[ungrouped) MD = (1/n) Σ |xi − x̄|\]
- \[Mean Deviation (grouped) MD = (1/N) Σ fi |xi − A| (use A = mean or median\]\[xi = class mid-point)\]
- \[Variance (population\]\[ungrouped) σ² = (1/n) Σ (xi − x̄)²\]
Correlation
Correlation
Key Point: Pearson's coefficient (deviation form): r = [Σ(x - x̄)(y - ȳ)] / sqrt{Σ(x - x̄)^2 · Σ(y - ȳ)^2}
Definition: Correlation is a statistical measure that describes the degree and direction of association between two quantitative variables. In economics it shows how one economic variable changes when another changes (for example, income and consumption).
Types of correlation:
- Positive correlation: Both variables move in the same direction (e.g., study hours and marks).
- Negative correlation: Variables move in opposite directions (e.g., price and quantity demanded for a normal good).
- No correlation: No discernible relationship; points are scattered randomly.
Methods of studying correlation:
- Scatter diagram (graphical): Plot paired observations (x,y). The pattern gives a visual idea of direction and strength.
- Karl Pearson’s coefficient of correlation (r): Measures the degree of linear relationship between two variables measured on an interval/ratio scale.
- Spearman’s rank correlation (ρ): Used when data are ordinal or when we use ranks instead of raw values; useful for monotonic but not necessarily linear relationships.
Interpretation: The correlation coefficient (r or ρ) ranges from −1 to +1. Sign indicates direction (+ for positive, − for negative). Magnitude shows strength (values close to ±1 indicate strong association; close to 0 indicate weak or no linear association). Important caution: correlation does not imply causation. A high correlation can arise from coincidence, confounding variables, or indirect relationships.
Key properties (brief): Pearson's r is unitless, lies between −1 and +1, is unchanged by change of origin and scale (linear transformations), and measures only linear association; it is sensitive to outliers. Spearman’s ρ is less sensitive to outliers and measures monotonic association.
Use in economics: Correlation helps analyse relationships such as consumption vs income, advertising vs sales, education vs earnings, price vs demand, and to check multicollinearity among explanatory variables before regression analysis.
- Marks in Mathematics and Physics for a class: typically show positive correlation — students who score high in one tend to score high in the other.
- Advertising expenditure and sales revenue: usually a positive correlation (more advertising → higher sales), other factors controlled.
- Price of a normal good and quantity demanded: negative correlation — higher price tends to reduce demand.
- Years of education and individual income: generally positive correlation — more education often associates with higher earnings.
- Number of working days absent and exam performance: negative correlation — more absenteeism tends to be associated with lower marks.
- \[Pearson's coefficient (deviation form): r = [Σ(x - x̄)(y - ȳ)] / sqrt{Σ(x - x̄)^2 · Σ(y - ȳ)^2}\]
- \[Pearson's coefficient (computational/shortcut form): r = [nΣxy - (Σx)(Σy)] / sqrt{[nΣx^2 - (Σx)^2] · [nΣy^2 - (Σy)^2]}\]
- \[Spearman's rank correlation (for no tied ranks): ρ = 1 - [6 Σd^2] / [n(n^2 - 1)]\]\[where d = difference between ranks of each pair\]
- \[Range and interpretation: −1 ≤ r (or ρ) ≤ +1\]\[Sign = direction\]\[magnitude ≈ strength (e.g., |r|>0.7 strong, 0.3–0.7 moderate, <0.3 weak).\]
- \[Note: Correlation ≠ Causation — a high r does not mean one variable causes the other.\]
Index Numbers
Index Numbers
Key Point: Simple aggregate index (prices): Index = (Σp1 / Σp0) × 100, where p0 = base prices, p1 = current prices.
What is an Index Number?
An index number is a statistical measure that shows the relative change in a variable (or group of related variables) over time or between places, expressed with respect to a base value (usually 100). It simplifies complex data into a single figure that facilitates comparison.
Purpose & Uses
- Measure inflation (e.g., Consumer Price Index).
- Compare price or quantity changes over time or across regions.
- Adjust monetary figures for price changes (real vs nominal values).
- Summarize movements in baskets of goods (e.g., cost of living, wholesale prices).
Types of Index Numbers
- By variable: Price index (price changes), Quantity index (output or consumption changes), Value index (value change).
- By weighting: Simple (unweighted) index, Weighted index.
- By formula: Aggregate (ratio of sums), Price relatives (averages of item-wise ratios).
Common Methods / Formulas (conceptual)
- Simple Aggregate Price Index: (Sum of current prices / Sum of base year prices) × 100
- Weighted Aggregate Price Index (Laspeyres-type): (Sum of base year quantities × current prices) / (Sum of base year quantities × base prices) × 100
- Price Relatives: For each item, (Current price / Base price) × 100. Then average these relatives (simple or weighted).
- Laspeyres Index (L): uses base period weights (quantities).
- Paasche Index (P): uses current period weights.
- Fisher Ideal Index: geometric mean of Laspeyres and Paasche: sqrt(L × P).
Desirable Tests (Properties)
- Time reversal test: If you swap base and current periods, the reciprocal relation should hold.
- Factor reversal test: For price index P and quantity index Q, P × Q should equal the value index (if properly constructed).
- Circular test: Index from A→B→C should be consistent with A→C.
Step-by-step construction (typical Laspeyres price index)
- Choose base year and current year.
- Select a representative basket of goods and their base-year quantities.
- Collect base-year prices and current-year prices for each item.
- Compute weighted sum: Σ(q0 × p1) and Σ(q0 × p0).
- Index = [Σ(q0 × p1) / Σ(q0 × p0)] × 100.
Short numerical example
Consider three goods A, B, C with base-year (0) quantities and prices and current-year (1) prices:
Item q0 p0 p1 A 10 5 6 B 20 2 3 C 30 1 1.2
Compute Laspeyres price index:
Σ(q0×p1) = 10×6 + 20×3 + 30×1.2 = 60 + 60 + 36 = 156 Σ(q0×p0) = 10×5 + 20×2 + 30×1 = 50 + 40 + 30 = 120 Index = (156 / 120) × 100 = 130 → Prices up by 30% since base year
Interpretation
An index of 130 (base =100) means an average increase of 30% in the price level of the chosen basket since the base year. Choice of weights and items affects results—hence different indices (CPI, WPI, GDP deflator) can give different percentages.
Limitations
- Choice of base year and basket may bias results (substitution bias, quality changes).
- New goods and changes in quality require adjustments.
- Different formulae give different answers; no single perfect index.
Practical examples and uses
Governments and central banks use index numbers to measure inflation (CPI, WPI), to adjust wages, pensions, tax brackets, and to deflate nominal GDP to obtain real GDP.
- Consumer Price Index (CPI): Measures changes in the cost of living by tracking prices of a fixed basket of consumer goods and services. Used to adjust salaries and pensions for inflation.
- Wholesale Price Index (WPI): Tracks price changes at the wholesale level (producers/wholesalers) and gives early signals of inflationary pressures.
- Stock market indices (e.g., Sensex, Nifty): Special kinds of indices that measure price movement of a selected group of stocks to indicate market performance.
- Construction cost index: Measures changes in prices of materials and labor used in building — used by builders and governments to adjust project budgets and contracts.
- \[Simple aggregate index (prices): Index = (Σp1 / Σp0) × 100\]\[where p0 = base prices\]\[p1 = current prices.\]
- \[Simple average of price relatives: Index = (1/n) Σ (p1 / p0 × 100).\]
- \[Weighted aggregate index (Laspeyres): L = [Σ (q0 × p1) / Σ (q0 × p0)] × 100\]\[where q0 are base-year quantities (weights).\]
- \[Paasche index: P = [Σ (q1 × p1) / Σ (q1 × p0)] × 100\]\[where q1 are current-year quantities (weights).\]
- \[Fisher Ideal index: F = sqrt(L × P) (geometric mean of Laspeyres and Paasche).\]
- \[Relating price\]\[quantity and value indexes (factor reversal idea): Price Index × Quantity Index ≈ Value Index.\]
Interpretation, Limitations and Common Statistical Errors
Interpretation, Limitations and Common Statistical Errors
Key Point: Arithmetic mean: x̄ = (Σ xi) / n
What interpretation means
Interpretation of statistics is the process of drawing meaningful conclusions from numerical data. Good interpretation requires attention to context (what is measured, units, time period), the type of data (sample or population), measures used (mean, median, percentages), variability and reliability (sample size, standard error, significance) and the possibility of alternative explanations.
Principles for correct interpretation
- Context: Always ask who collected the data, when and why. Units and base years matter for indices and percentages.
- Understand the measure: Mean, median and mode can tell different stories; choose the one appropriate to the distribution.
- Look at variability: Two series with the same mean can have very different spreads—use variance/SD and range.
- Sample versus population: If data come from a sample, check sample size, sampling method and margin of error before generalising.
- Beware of outliers: Single extreme values can distort means and correlations.
- Time patterns: Distinguish trend from seasonal or cyclical fluctuations in time-series data.
- Correlation ≠ causation: A statistical association does not prove one variable causes the other—look for confounders, time order and plausible causal mechanism.
- Aggregation effects: Aggregating data (e.g., averages over regions) can hide important within-group differences (ecological fallacy).
Limitations of statistical methods
- Incomplete picture: Statistics describe measurable aspects; qualitative or unmeasured factors (preferences, culture, policy details) may be important.
- Sampling and measurement errors: Poor questionnaire design, non-response, misreporting and recording errors lead to wrong conclusions.
- Bias: Selection bias, non-random samples and survivorship bias produce misleading estimates.
- Time-lag and dynamics: Statistics are often backward looking and may not reflect rapid changes.
- Choice-dependence: Results can change with definitions, base year, or choice of index formula and averaging method.
- Over-simplification: Single summary statistics (like a mean) hide distributional issues such as inequality.
Common statistical errors (what to watch for)
- Confusing correlation with causation — assuming A causes B because they move together.
- Using an inappropriate average — using mean for highly skewed income data instead of median.
- Biased sampling — using a sample that is not representative (e.g., internet poll for national opinion without accounting for access differences).
- Misleading graphs — truncated axes, distorted scales, 3D effects, unclear labels or omitted baselines.
- Ignoring confounding variables — apparent effects disappear when a third variable is controlled for.
- Simpson’s paradox and ecological fallacy — aggregated data can reverse or hide relationships present in subgroups.
- Cherry-picking and data dredging — selective reporting of favorable results or searching until a significant result appears.
- Misuse of percentages and indices — failing to note base sizes or base-year effects (e.g., a 100% increase from 1 to 2 is small in absolute terms).
How to avoid mistakes
- Always check source, sample method and sample size.
- Compare mean, median and mode and examine dispersion (variance/SD, box-plot).
- Plot data—histograms, boxplots and scatter plots reveal structure that summary numbers hide.
- Test robustness—see if results hold under different definitions, subgroups and time periods.
- Look for alternative explanations and control for likely confounders where possible.
Summary
Good interpretation combines numerical computation with judgement about data quality, context and plausible causal stories. Awareness of limitations and common errors helps avoid misleading conclusions.
- Correlation vs causation: Ice-cream sales and drowning incidents both rise in summer. The common factor (temperature) causes both, not ice-cream causing drownings.
- Misleading average: The mean monthly income in a village is Rs. 50,000 because one industrial owner earns Rs. 10 lakh; the median income (Rs. 6,000) better represents most households.
- Biased sample (Literary Digest 1936): A poll that used automobile registration lists predicted a wrong US election outcome because the sample over-represented wealthy readers.
- Truncated-axis graph: A line chart showing GDP growth that starts the vertical axis at 2% (not 0) exaggerates perceived variation between years.
- Simpson's paradox: Two hospitals A and B: A has higher survival rates for both simple and complex surgeries separately, but when combined B shows a higher overall survival rate because B treated a much larger share of easy cases.
- \[Arithmetic mean: x̄ = (Σ xi) / n\]
- \[Weighted mean: x̄w = (Σ wi xi) / (Σ wi)\]
- \[Median: middle value when data are ordered (or average of two middle values if n is even)\]
- \[Variance (population): σ^2 = (Σ (xi - μ)^2) / N\]
- \[Standard deviation: σ = sqrt(σ^2)\]
- \[Sample variance: s^2 = (Σ (xi - x̄)^2) / (n - 1)\]
Key Concepts
- Statistics
- A branch of mathematics dealing with collection, presentation, analysis and interpretation of numerical data.
- Statistical tools
- Methods and techniques like tabulation, diagrams and measures used to summarize and analyze data.
- Primary data
- Data collected first-hand by the researcher for a specific purpose.
- Secondary data
- Data already collected by others and used for a new analysis.
- Population
- The complete set of units or observations under study.
- Sample
- A subset of the population selected for analysis.
- Variable
- A characteristic or attribute that can take different values.
- Qualitative variable
- A variable that describes qualities or categories, not numerical magnitude.
- Quantitative variable
- A variable that is numerical and measurable.
- Frequency distribution
- A tabular arrangement showing classes or values and their corresponding frequencies.
- Class interval
- A range of values grouped together in a frequency distribution for continuous data.
- Cumulative frequency
- The running total of frequencies up to a given class or value.
- Tabulation
- Organizing raw data into tables to simplify presentation and analysis.
- Arithmetic mean
- The sum of all observations divided by the number of observations; a measure of central tendency.
- Median
- The middle value in an ordered data set that divides it into two equal parts.
- Mode
- The value that occurs most frequently in a data set.
- Range
- The difference between the maximum and minimum values in a data set; a simple measure of dispersion.
- Variance
- The average of squared deviations of observations from the mean; measures spread of data.
- Standard deviation
- The square root of variance; indicates average deviation of values from the mean in original units.
- Correlation
- A statistical measure that shows the degree and direction of association between two variables.
Practice Questions
-
Define statistics and distinguish between a population and a sample. / सांख्यिकी की परिभाषा दीजिए तथा समष्टि और प्रतिदर्श में अंतर कीजिए।
Show answer
Statistics is the branch of mathematics dealing with collection, presentation, analysis and interpretation of numerical data; a population is the complete set of units under study, while a sample is a representative subset selected from it. / सांख्यिकी गणित की वह शाखा है जो संख्यात्मक आँकड़ों के संग्रह, प्रस्तुतीकरण, विश्लेषण व निर्वचन से संबंधित है; समष्टि अध्ययनाधीन इकाइयों का पूर्ण समूह है, जबकि प्रतिदर्श उसमें से चुना गया प्रतिनिधि उपसमूह है।
-
Differentiate between primary and secondary data with one example of each. / प्राथमिक और द्वितीयक आँकड़ों में एक-एक उदाहरण सहित अंतर कीजिए।
Show answer
Primary data are collected first-hand for a specific purpose (e.g., a household income survey conducted by a researcher), whereas secondary data are already collected by others (e.g., census figures from government publications). / प्राथमिक आँकड़े किसी विशेष उद्देश्य हेतु प्रत्यक्ष रूप से एकत्र किए जाते हैं (जैसे शोधकर्ता द्वारा किया गया घरेलू आय सर्वेक्षण), जबकि द्वितीयक आँकड़े पहले से दूसरों द्वारा एकत्रित होते हैं (जैसे सरकारी प्रकाशनों के जनगणना आँकड़े)।
-
When should the median be preferred over the arithmetic mean as a measure of central tendency? / केंद्रीय प्रवृत्ति के माप के रूप में माध्यिका को समांतर माध्य पर कब प्राथमिकता दी जानी चाहिए?
Show answer
The median is preferred when the distribution is skewed or contains extreme outliers, because, unlike the mean, it is not affected by extreme values and better represents the typical value. / माध्यिका को तब प्राथमिकता दी जाती है जब वितरण विषम हो या उसमें चरम बहिरमान हों, क्योंकि माध्य के विपरीत यह चरम मानों से प्रभावित नहीं होती और प्रतिनिधि मान को बेहतर दर्शाती है।
-
Find the arithmetic mean and median of the data: 12, 15, 10, 18, 20. / आँकड़ों 12, 15, 10, 18, 20 का समांतर माध्य और माध्यिका ज्ञात कीजिए।
Show answer
Mean = (12+15+10+18+20)/5 = 75/5 = 15. Sorted data: 10,12,15,18,20; with n=5 (odd), median is the (5+1)/2 = 3rd value = 15. / माध्य = (12+15+10+18+20)/5 = 75/5 = 15। क्रमबद्ध आँकड़े: 10,12,15,18,20; n=5 (विषम) होने पर माध्यिका (5+1)/2 = तीसरा मान = 15 है।
-
Why is the coefficient of variation (CV) used to compare the variability of two data sets? / दो आँकड़ा समूहों की परिवर्तनशीलता की तुलना के लिए विचरण गुणांक (CV) का प्रयोग क्यों किया जाता है?
Show answer
The coefficient of variation, CV = (SD/Mean) × 100, is a relative and unitless measure, so it allows fair comparison of variability between data sets that have different units or widely different means. / विचरण गुणांक, CV = (मानक विचलन/माध्य) × 100, एक सापेक्ष और इकाई-रहित माप है, इसलिए यह भिन्न इकाइयों या बहुत भिन्न माध्यों वाले आँकड़ा समूहों की परिवर्तनशीलता की उचित तुलना की अनुमति देता है।
-
Calculate the Laspeyres price index for goods with base prices p0 = 5, 2 and current prices p1 = 6, 3, with base quantities q0 = 10, 20. / आधार कीमतों p0 = 5, 2 तथा वर्तमान कीमतों p1 = 6, 3 और आधार मात्राओं q0 = 10, 20 वाली वस्तुओं के लिए लैस्पीयर्स कीमत सूचकांक की गणना कीजिए।
Show answer
Σ(q0×p1) = 10×6 + 20×3 = 60 + 60 = 120; Σ(q0×p0) = 10×5 + 20×2 = 50 + 40 = 90; Index = (120/90)×100 ≈ 133.3, meaning prices rose about 33.3% over the base year. / Σ(q0×p1) = 10×6 + 20×3 = 60 + 60 = 120; Σ(q0×p0) = 10×5 + 20×2 = 50 + 40 = 90; सूचकांक = (120/90)×100 ≈ 133.3, अर्थात आधार वर्ष की तुलना में कीमतें लगभग 33.3% बढ़ीं।
-
Explain why 'correlation does not imply causation' with a suitable example. / उपयुक्त उदाहरण सहित समझाइए कि 'सहसंबंध कारणता का संकेत नहीं देता'।
Show answer
Two variables can be statistically associated without one causing the other; for example, ice-cream sales and drowning incidents both rise in summer due to a common factor (temperature), not because ice-cream causes drowning. / दो चर सांख्यिकीय रूप से संबद्ध हो सकते हैं बिना एक के दूसरे का कारण बने; उदाहरणतः, आइसक्रीम की बिक्री और डूबने की घटनाएँ दोनों गर्मियों में बढ़ती हैं, जो साझा कारक (तापमान) के कारण है, न कि इसलिए कि आइसक्रीम डूबने का कारण है।
-
How can a graph with a truncated vertical axis mislead the reader? / प्रच्छिन्न (कटे हुए) ऊर्ध्व अक्ष वाला आरेख पाठक को कैसे भ्रमित कर सकता है?
Show answer
When the y-axis does not start at zero, small differences between values appear exaggerated, making variation or growth look much larger than it actually is, thereby misleading interpretation. / जब y-अक्ष शून्य से शुरू नहीं होता, तो मानों के बीच छोटे अंतर बढ़े-चढ़े दिखते हैं, जिससे परिवर्तन या वृद्धि वास्तविकता से कहीं अधिक प्रतीत होती है और निर्वचन भ्रामक हो जाता है।
Related Laws & Principles
Explore allFoundational laws & principles connected to this chapter — tap to open in the Laws Explorer.