Overview
This unit introduces statistics as a tool for economic analysis. It covers collection, classification and presentation of data, construction of frequency distributions and diagrams, and the main numerical measures used to summarise data. Students will learn central tendency (mean, median, mode) and measures of dispersion (range, quartile deviation, mean deviation, variance and standard deviation), and how to interpret coefficient of variation. The unit also introduces correlation and regression as methods to examine relationships between two variables, including graphical methods, Karl Pearson's coefficient, Spearman's rank correlation, and the equations of regression lines. The unit explains sampling and sources of data, limitations of statistics, and uses of statistical methods in economics such as forecasting, policy evaluation and comparison. Learning this unit helps students read economic reports, summarize large data sets, make comparisons, and interpret relationships between economic variables. A clear grasp of statistical techniques is essential for higher study in economics, business and social sciences, and for understanding empirical evidence in newspapers and research papers.
Learning Objectives
- Explain the meaning, scope and importance of statistics in economics.
- Differentiate between types of data and describe methods of data collection and sources.
- Classify and tabulate raw data and present it using suitable diagrams and graphs.
- Construct frequency distributions for grouped and ungrouped data and draw histograms, frequency polygons and ogives.
- Compute and interpret measures of central tendency: mean, median and mode for different types of data.
- Calculate measures of dispersion including range, quartile deviation, mean deviation, variance and standard deviation.
- Determine and interpret coefficient of variation and explain its use in comparing variability.
- Examine the relationship between two variables using correlation and regression techniques, and compute correlation coefficients.
- Recognise limitations of statistical methods and apply statistics critically to economic problems.
Topics in this chapter
18 topics · tap a topic title to jump straight to it.
Introduction: Meaning and Scope of Statistics
What is Statistics?
Statistics is a systematic collection of methods used to organise, summarise, present and interpret numerical facts. It is not merely numbers but a way of thinking about data arising from economic activities — production, consumption, prices, employment and income. In practical terms statistics converts mass data into manageable summaries that reveal patterns, differences, trends and relationships which are essential for policy-making and business decisions.
Why Statistics Matters in Economics
Economics deals with broad aggregations and trends — national income, inflation, unemployment, trade balances — where raw observations are too numerous to examine individually. Statistical methods help condense such information to reveal overall direction and intensity of economic changes. For instance, producing an index of prices simplifies thousands of price quotes into a single measure that can be compared over time.
Descriptive and Inferential Branches
The discipline of statistics is divided into descriptive and inferential branches. Descriptive statistics summarise data through tables, graphs and summary measures like averages and dispersion. Inferential statistics allow conclusions about larger populations using sample data; this involves probability, estimation and hypothesis testing. In Class 11 economics you begin with descriptive methods and elementary correlation and regression that form the foundation for later inferential techniques.
Scope of Applications
Within economics, statistics is used to estimate national income, construct price and quantity indices, compare regional or sectoral performance, measure inequality, test hypotheses about relationships (for example, between income and consumption), forecast future values and evaluate policy impacts. Firms use statistical forecasting to plan production and inventories; governments use statistical indicators to design fiscal and monetary policies. Thus statistical literacy allows students to interpret economic reports and contributes to informed citizenship.
Key Principles
Effective statistical work depends on clear definitions, careful data collection and awareness of limitations. Terms must be defined consistently (e.g., what counts as employment), sampling must aim for representativeness when a census is impossible, and choice of summarising measures must match data types. Throughout this unit students learn not only procedures but also how to judge the reliability and relevance of statistical results.
- A table showing monthly inflation rates turned into a line diagram to show trend.
- Using sample survey of households to estimate average consumption in a town.
Functions and Limitations of Statistics
Key Functions
Statistics serves several important functions in the study and practice of economics. First, it organises data: raw observations are classified and put into tables that make information accessible. Second, it describes data: through averages, dispersion measures and graphs, statistics summarises the central tendency and spread of economic variables. Third, it compares: indices, ratios and relative figures help compare regions, sectors or time periods. Fourth, it forecasts: historical patterns in data allow short-term predictions useful for policy and business planning. Finally, statistics aids control and evaluation: trends and indicators show whether policies meet targets, and sampling helps monitor implementation across locations.
Support to Decision Making
Statistical evidence provides empirical backing for economic decisions. Budget allocations use statistical estimates of needs; central banks use inflation statistics to set interest rates; businesses rely on sales forecasts and consumer surveys to plan production. Without quantitative evidence, decisions would be guesswork; statistics gives structure and plausibility to choices by quantifying expected benefits, costs and risks.
Limitations and Cautions
Despite its strengths, statistics has limitations. The quality of conclusions depends critically on data quality: measurement errors, non-response in surveys, misreporting and biased data collection can all produce misleading summaries. Even with correct data, inappropriate method selection — such as using an arithmetic mean for highly skewed income data — can misrepresent reality. Another limitation is that statistical association does not prove causation. Correlation between two series could arise from a third common factor or mere coincidence. Additionally, aggregated statistics may hide sub-group differences. For example, national averages can conceal regional inequality.
Design and Interpretation Issues
Errors may also arise from poor sampling design. Non-random samples or small sample sizes reduce representativeness and increase sampling error. Visual presentation can be misleading when graphs use distorted scales or omit context. Ethically, analysts must report methods, limitations, and uncertainty so readers can judge validity. Critical interpretation requires looking beyond headline numbers to definitions, sources, time periods and procedures used.
Balance Between Use and Skepticism
Students should therefore learn two attitudes: confidence in the powerful insights statistics offers, and a healthy skepticism towards claims unsupported by transparent methodology. Understanding both uses and limits prepares learners to apply statistical tools responsibly in economic analysis and to read others' statistical claims critically.
- Comparing per capita income of two states using averages and percentages.
- Predicting next quarter's sales using trend from past quarters and noting the uncertainty.
Types of Data: Qualitative and Quantitative, Discrete and Continuous
Basic Distinction: Qualitative vs Quantitative
Understanding data types is the first step in choosing appropriate statistical methods. Qualitative (or categorical) data describe qualities or categories that cannot be measured numerically in a meaningful arithmetic sense: examples are industry type, employment status, or ownership form. Such data may be nominal (no order, e.g., industry names) or ordinal (ordered categories, e.g., low, medium, high). Quantitative data are numerical and measure quantities such as income, production, price or age; arithmetic operations like addition and averaging are meaningful for these.
Discrete and Continuous Quantitative Data
Quantitative data split into discrete and continuous types. Discrete data take specific separate values, often counts: number of employees, number of firms, number of defects. Continuous data can take any value within a range and are often results of measurement: weight, income measured precisely, time. Continuous variables are typically grouped into class intervals for summarised presentation, while discrete variables are often presented as frequency distributions of individual values.
Scales of Measurement
Four measurement scales are relevant: nominal (labels), ordinal (ranked categories), interval (numeric differences meaningful but no true zero), and ratio (numeric with meaningful zero). In economics most variables are ratio scale (e.g., income, price), allowing full range of statistical operations. Ordinal data permit median and percentiles but not arithmetic mean in a strict sense. Therefore, choice of averages depends on scale: median and mode suit ordinal data, mean requires interval or ratio.
Implications for Presentation and Analysis
Data type dictates presentation: qualitative data are best summarised by frequency tables, bar charts or pie charts. Quantitative discrete data can be shown as stem-and-leaf displays or frequency tables; continuous data are grouped and shown via histograms, frequency polygons and ogives. For analysis, correlation and regression need quantitative numeric data; for ordinal data Spearman's rank is appropriate. Understanding data type prevents misuse: using arithmetic mean for strongly skewed ratio data without noting outliers can mislead, while ignoring order in ordinal data wastes information.
Converting and Coding Data
Sometimes qualitative data must be coded as numbers for analysis (e.g., education levels coded 1–4). This should be done carefully to preserve meaning — arbitrary coding can mislead if treated as numeric. When converting continuous measurements into groups, be aware of information loss and choose class widths that balance clarity and detail.
- Qualitative: Classifying firms by industry sector and counting each sector's firms.
- Quantitative discrete: Number of employees in a firm; Quantitative continuous: Monthly household income.
Collection of Data: Primary and Secondary Sources, Sampling
Primary Data Sources
Primary data are collected directly for the purpose of the study. Methods include surveys using questionnaires, structured or unstructured interviews, direct observation and experiments. In economic studies primary data might be household expenditure surveys, firm-level cost accounts, price observations in markets or experimental studies on consumer choice. Primary data allow the researcher to define variables, design instruments and control measurement procedures. However, they require time, trained enumerators and funding.
Secondary Data Sources
Secondary data are existing data compiled by others: government statistical releases, central bank reports, industry publications, census data, research databases and company reports. Secondary sources are cost-effective and often cover long periods, but they may not perfectly match the current research needs; definitions, units and methods used by the source should be examined. When using secondary data, always cite the source and check for revisions or known quality issues.
Census versus Sample
A census attempts to collect information from the entire population—useful when population size is small or completeness is needed. Most large-scale economic studies use sampling because a census is often impractical. Sampling draws a subset to make inferences about the whole. Proper sampling reduces cost and time while maintaining acceptable accuracy if designed correctly.
Sampling Methods
Important sampling methods include random sampling (every unit has equal chance), stratified sampling (population divided into homogeneous strata and sampled within each), cluster sampling (selection of groups or clusters), systematic sampling (every k-th unit), and purposive or convenience sampling (non-random, used for exploratory work). Random and stratified methods are preferred for statistical validity; stratification improves precision when strata are internally similar but different from each other.
Sample Size and Representativeness
Sample size should be large enough to capture variability: very small samples give unreliable estimates while unnecessarily large samples waste resources. Representativeness means the sample should mirror the population structure in key characteristics — if not, results will be biased. Weighting adjustments are sometimes applied to correct for non-representative samples.
Questionnaire Design and Data Quality
Clear definitions, pilot testing, simple and unambiguous questions, consistent units, and careful training of enumerators improve primary data quality. For secondary data, investigate collection methods, coverage, periodicity and revisions. Document limitations and potential biases when reporting results so users can interpret findings responsibly.
- Conducting a household survey on monthly expenditure using a structured questionnaire.
- Using government labour force reports as secondary data to analyse unemployment trends.
Classification and Tabulation of Data
Purpose of Classification
Classification organises raw observations into categories so that patterns and comparisons become visible. Without classification, a large list of figures is difficult to interpret. Classification reduces complexity by grouping similar items together, which helps in summarising data for analysis and presentation. Effective classification follows the principles of exhaustiveness (all observations fit into categories) and mutual exclusiveness (no observation belongs to more than one category).
Types of Classification
Classification may be chronological (year, month), geographical (state, district), qualitative (industry, occupation) or quantitative (income ranges, age groups). Quantitative classification often uses class intervals for continuous variables. The choice of classification depends on the research question: time series classification emphasises trends, while cross-sectional classification supports comparisons across groups at a point in time.
Constructing Tables
Tabulation summarises classified data into tables. A simple frequency table lists categories (or class intervals) and their frequencies. Important columns include class limits, class marks (mid-points), frequency, relative frequency (percentage), and cumulative frequency. Headings should be clear and units specified. A well-constructed table allows the reader to quickly understand distributions and compute summary measures.
Contingency Tables
Contingency (cross) tables present joint distributions of two variables, showing how frequencies are distributed across combinations of categories. They include row totals, column totals and overall totals. Contingency tables are useful for exploring relationships—for example, employment by education level and gender—and for computing conditional distributions and percentages.
Design Choices and Practical Tips
When creating classes for quantitative data, choose the number and width of classes to balance clarity and information loss; too many classes fragment the data, too few hide patterns. Use equal widths where possible and round limits to convenient numbers. Always include total sample size, and where relevant provide percentages or rates. For comparative tables, keep formats consistent so readers can compare across tables easily.
From Tables to Further Analysis
Tabulated data form the basis for graphical displays and numerical measures such as mean and variance. Accurate tabulation, with attention to boundaries and cumulative counts, ensures that subsequent calculations (median, quartiles, regression) produce valid results. Proper documentation of classification rules and any excluded observations helps others reproduce and trust the analysis.
- A frequency table showing number of households by monthly income ranges.
- A contingency table of employment by sector and gender with marginal totals.
Frequency Distributions: Ungrouped and Grouped
Ungrouped Frequency Distribution
Ungrouped frequency distributions list each distinct observation and the number of times it occurs. They are suited when data have few distinct values or when preserving each value is important, such as counts of firm sizes or ratings on a small scale. An ungrouped table may include columns for the value, its frequency, relative frequency (percentage) and cumulative frequency. This presentation is straightforward and retains full information on each observation.
Grouped Frequency Distribution
Grouped distributions are necessary when dealing with continuous data or large datasets with many distinct values. Data are arranged into class intervals, and the frequency of observations within each interval is recorded. Classes should be mutually exclusive and exhaustive, with clear lower and upper boundaries. Grouping simplifies large datasets and enables drawing of histograms and ogives; however, it approximates individual values by class marks, which introduces small errors in numerical summaries.
Choosing Number and Width of Classes
Choose the number of classes depending on sample size; typical guidance suggests between 5 and 20 classes. Class width is roughly (maximum − minimum)/number of classes, but choose a convenient, rounded interval for clarity. Equal class widths are preferable because they simplify interpretation and plotting. Ensure that class boundaries are selected to avoid ambiguity (e.g., use 0–9, 10–19 rather than overlapping boundaries).
Class Marks and Cumulative Frequencies
Class mark (mid-point) is the average of lower and upper class limits and is used as representative value for calculations like grouped mean and variance. Cumulative frequencies accumulate frequencies up to a class and are useful for determining medians, percentiles and quartiles graphically or by interpolation. Two types of cumulative frequency exist: ‘less than’ and ‘more than’ cumulative frequencies, both useful in different contexts.
Advantages and Limitations
Grouped distributions reduce complexity and make graphical representation and computation feasible for large datasets. They hide individual observations so some precision is lost; results like mean and median are approximations. The impact of grouping decreases with narrower classes. Good practice requires stating class intervals, total frequency and any assumptions made when approximating values with class marks.
- Ungrouped: Number of factories producing different distinct units listed with frequencies.
- Grouped: Household incomes grouped into ranges Rs.0-4999, 5000-9999, etc., with frequency counts.
- Class mark (x) = (Lower limit + Upper limit) / 2
- Cumulative frequency: running sum of frequencies up to a class
Diagrammatic Presentation: Bar Charts, Pie Charts and Pictograms
Purpose of Diagrammatic Presentation
Graphs and diagrams transform numerical tables into visual displays that communicate patterns quickly and effectively. They are useful for non-technical audiences and for highlighting comparisons, trends and compositions. Choosing the right diagram depends on the type of data and the message intended: comparison across categories, composition of a total, or change over time.
Bar Charts
Bar charts display categorical data or discrete numerical data with bars whose lengths are proportional to frequencies or percentages. Bars may be vertical or horizontal and should be of equal width with equal spacing between them unless a grouped bar chart is used for multiple series. Bar charts are excellent for comparing quantities across sectors, regions or time periods; stacking or side-by-side bars can compare components. Label axes, include units and provide a clear title to avoid misinterpretation.
Pie Charts
Pie charts represent parts of a whole by dividing a circle into sectors whose angles correspond to category shares. To construct a pie chart compute each category’s percentage of the total and convert it into degrees by multiplying by 360°. Pie charts are effective when the number of categories is small and when the intention is to show composition, such as budget shares. They become hard to read with many small categories; in such cases group small items into ‘others’.
Pictograms
Pictograms use repeated pictures or icons to represent quantities; for example, one symbol may represent 100 units. They are visually appealing and good for presentations but can be misleading if scaling is inconsistent or if viewers misinterpret partial symbols. Ensure that the symbol scaling is linear and include a legend showing the unit represented by each symbol.
Design Principles and Pitfalls
Good design avoids distortion: do not start the vertical axis at a value other than zero unless clearly noted, avoid 3D effects that change perceived sizes, and use consistent colours and scales. Provide legends for multi-series charts and annotate important points. Misleading visuals can arise from truncated axes, irregular intervals, or omission of sample sizes. Always pair diagrams with the underlying numeric table so precise values are accessible.
Choosing the Right Diagram
Use bar charts for comparing categories, line charts for time series trends, pie charts for showing composition, and pictograms for simple counts in public presentations. For continuous distributions prefer histograms and frequency polygons. Selecting the appropriate diagram and designing it carefully ensures that statistical evidence is communicated accurately and persuasively.
- A bar chart comparing monthly retail sales across five months.
- Constructing a pie chart showing government expenditure shares across sectors.
- Angle for pie sector = (Category frequency / Total frequency) × 360°
Histogram, Frequency Polygon and Ogive
Histogram: Visualising Continuous Data
A histogram displays the frequency distribution of continuous data using adjacent rectangles. The horizontal axis shows class intervals and the vertical axis shows class frequencies. For equal class widths the height of each rectangle is proportional to frequency. When class widths differ, the area of rectangles must be proportional to frequency so that visual area correctly represents frequency. Histograms are useful to study the shape of distribution — whether it is symmetric, skewed, peaked or flat — and to identify modes and gaps.
Frequency Polygon: Connecting Class Marks
A frequency polygon plots class marks (mid-points) on the horizontal axis against their frequencies and joins successive points by straight lines. To close the polygon, add a point at each end on the horizontal axis with zero frequency. Frequency polygons are particularly helpful when comparing two or more distributions on the same axes because overlapping lines are easier to compare than multiple histograms.
Ogive: Cumulative Frequency Curve
An ogive shows cumulative frequency and is drawn using either 'less than' or 'more than' cumulative frequencies. For a 'less than' ogive plot the upper class boundaries against cumulative frequencies and join the points by a smooth curve or straight lines. Ogives are powerful for reading off medians, quartiles and percentiles graphically: the value corresponding to a chosen cumulative frequency gives the desired percentile.
Constructing the Graphs
To draw a histogram use correct class limits and mark equal widths on the X-axis; draw contiguous rectangles up to the frequency value on Y-axis. For a frequency polygon compute class marks and plot frequency against each mid-point; join points and close polygon at ends. For an ogive calculate cumulative frequencies and plot them against class boundaries, then join the points. Label axes, indicate class boundaries and include a title and source.
Interpreting Shapes and Economic Meaning
Examine histograms and polygons for skewness: a right-skewed income distribution often shows a long right tail reflecting a small number of high incomes; a left-skewed distribution indicates concentration at higher values. Multimodal shapes suggest distinct sub-groups in data (e.g., different consumer segments). In economic analysis such shapes inform policy: a highly skewed income distribution may call for redistribution policies, while multimodal sales distributions could indicate different market segments requiring tailored strategies.
- Drawing a histogram for household incomes grouped into class intervals and noting skewness.
- Using an ogive to estimate median monthly expenditure from grouped data.
- Class mark = (Lower limit + Upper limit) / 2
- Cumulative frequency for a class = sum of frequencies up to that class
Measures of Central Tendency: Mean (Individual and Discrete Series)
Why Measure Central Tendency?
Measures of central tendency summarise a distribution by a single representative value. In economics these measures help understand the average level of variables such as income, consumption, price or output. The three common measures are arithmetic mean, median and mode; each has advantages and contexts where it is most appropriate. The arithmetic mean is widely used because it employs all observations and integrates well with further algebraic calculations.
Arithmetic Mean for Individual Observations
For n individual observations x1, x2, ... , xn the arithmetic mean is x̄ = Σxi / n. The mean is simple to interpret as the per-unit average and is useful when individual data are available and measured on an interval or ratio scale. It is also the balancing point of the distribution: the sum of deviations from the mean equals zero. However, the mean is sensitive to extreme values (outliers) which can skew the mean away from the central position experienced by most observations.
Arithmetic Mean for Discrete Frequency Distribution
When data occur with frequencies, replace the sum of individual values with sum of value times frequency. If values xi have frequencies fi, mean is x̄ = Σfi xi / Σfi. This weighted mean recognises repeated observations without listing them individually. For large values or mental calculation an assumed mean (A) helps: compute di = xi − A and use x̄ = A + (Σfi di) / Σfi. The step-deviation method divides di by the class width when xi are class marks of grouped data to further simplify arithmetic.
Grouped Data and Approximation
For grouped continuous data exact values are unknown within classes, so class marks mi (mid-points) serve as representative values: x̄ ≈ Σfi mi / Σfi. This yields an approximate mean whose accuracy improves with narrower class widths. Always state that the mean for grouped data is an approximation and show class marks and calculations for transparency.
Properties and Practical Considerations
The mean is unique and uses all data, making it useful for variance and regression computations. But for skewed income distributions, the median may better represent a ‘typical’ individual. When reporting means, also provide measures of dispersion (standard deviation or quartile deviation) so readers know how spread-out values are around the mean. In policy contexts, choice between mean and median affects interpretation: an increase in mean income might be driven by a few high-income earners and not indicate improvement for the majority.
- Individual mean: Average monthly income of five households with incomes Rs. 8000, 12000, 5000, 15000, 10000 is (8000+12000+5000+15000+10000)/5 = Rs.10000.
- Discrete series mean: Values 1,2,3 with frequencies 5, 7, 8 give mean = (1×5 + 2×7 + 3×8)/(5+7+8) = (5+14+24)/20 = 43/20 = 2.15.
- Mean (individual) x̄ = Σxi / n
- Mean (discrete) x̄ = Σfi xi / Σfi
- Assumed mean method: x̄ = A + (Σfi di) / Σfi where di = xi − A
Median and Mode (Ungrouped and Grouped Data)
Median: Concept and Calculation for Ungrouped Data
Median is the middle value of an ordered data set, dividing observations into two equal parts. For an odd number of observations the median is the central value after ordering; for an even number it is the average of the two middle values. The median is a positional average and is especially useful with skewed data because it is not affected by extreme values. For example, median income better represents the experience of a typical household when incomes are highly unequal.
Median for Grouped Data
For grouped frequency distributions the exact middle value is interpolated within the median class using cumulative frequencies. The formula Median = L + [(N/2 − cf) / f] × h estimates the position within the median class: L is the lower boundary of the median class, N total frequency, cf cumulative frequency before the median class, f frequency of the median class and h class width. This linear interpolation assumes uniform distribution within the class and gives a reasonable estimate when class widths are not too large.
Mode: Concept and Use
Mode is the most frequently occurring value or class in a distribution. For categorical data the mode is often the only meaningful average (e.g., most common occupation). In numerical distributions the mode marks the value with highest concentration. A distribution may be unimodal, bimodal or multimodal depending on the number of peaks present. Mode is easy to identify for ungrouped data: choose the value with highest frequency.
Mode for Grouped Data
When data are grouped, mode is estimated using the modal class (class with highest frequency). The formula Mode = L + [(fm − f1) / (2fm − f1 − f2)] × h interpolates within the modal class: L is lower class boundary of modal class, fm its frequency, f1 the frequency of preceding class, f2 the frequency of succeeding class and h the class width. This method assumes a roughly triangular shape around the modal class and yields a better estimate than simply taking class mid-point when frequencies around the peak are uneven.
Choosing Between Median and Mode
Which measure to use depends on data type and purpose. Use median for ordinal data and skewed numerical distributions; mode for categorical data and to identify the most common category. For symmetric distributions mean, median and mode coincide; differences among them signal skewness. Reporting more than one measure often gives a fuller picture: mean indicates arithmetic average, median indicates middle point, mode indicates most common value.
- Ungrouped median: For ordered incomes Rs. 2000, 3000, 5000, 7000, 9000, median = 5000.
- Grouped mode: For class frequencies where modal class 50–60 has fm=30, previous f1=20, next f2=15 and h=10, mode = 50 + [(30−20)/(2×30−20−15)]×10 = 50 + (10/25)×10 = 54.
- Median (grouped) = L + [(N/2 − cf) / f] × h
- Mode (grouped) = L + [(fm − f1) / (2fm − f1 − f2)] × h
Measures of Dispersion: Range, Mean Deviation and Quartile Deviation
Why Dispersion Matters
Measures of central tendency do not describe how spread out values are around the centre. Dispersion measures quantify variability and risk, which are crucial in economics: two regions may have the same average income but very different income distributions, leading to different policy implications. Dispersion helps compare stability of series such as prices, returns or production levels.
Range: Simple but Crude
Range is the difference between maximum and minimum values. It is quick to compute and gives a first glimpse of spread, but because it uses only two observations it is highly sensitive to outliers and provides little information about the rest of the distribution. Range is most useful for preliminary descriptions or quality checks.
Quartile Deviation (Semi-Interquartile Range)
Quartile deviation measures the spread of the middle 50% of data and is defined as Q.D. = (Q3 − Q1)/2 where Q1 and Q3 are the first and third quartiles. Being based on quartiles, Q.D. is robust to extreme values and provides a measure of central spread. For grouped data quartiles are obtained by interpolation using cumulative frequencies at 25% and 75% levels, similar to the median method.
Mean Deviation
Mean deviation (MD), or average absolute deviation, is computed as MD = Σ|xi − a| / n where a is a central value (mean or median). MD uses information from all observations but avoids squaring deviations, making its units the same as the original variable. MD about median is often smaller and more robust to outliers than MD about mean. For grouped data substitute class marks for xi and use Σfi|mi − a| / Σfi as the grouped form.
Comparison and Choice
Range is easiest but least informative. Quartile deviation balances resistance to outliers with interpretability and is useful for skewed distributions. Mean deviation uses all observations and gives a measure responsive to overall variation, though it involves absolute values which complicate algebraic manipulation. These measures prepare the ground for variance and standard deviation which are algebraically convenient and widely used in further statistical modelling and analysis.
- Range: For incomes 2000, 5000, 12000, 8000, range = 12000 − 2000 = Rs.10000.
- Quartile deviation: If Q1=3000 and Q3=8000, then Q.D.=(8000−3000)/2 = Rs.2500.
- Range = Maximum − Minimum
- Quartile Deviation (Q.D.) = (Q3 − Q1) / 2
- Mean Deviation MD = Σ|xi − a| / n (or for frequency data Σfi |xi − a| / Σfi)
Variance and Standard Deviation
Understanding Variance
Variance measures the average squared deviation of observations from their mean and is a central concept in statistics. Because deviations can be positive or negative, squaring ensures that positive and negative deviations do not cancel out. Variance thus captures overall dispersion in squared units of the variable. In economics variance is used to measure volatility in prices, returns or growth rates and underlies many inferential techniques.
Population and Sample Distinctions
For a full population of N observations xi with population mean μ, population variance is σ2 = Σ(xi − μ)2 / N and the corresponding standard deviation is σ = sqrt(σ2). In practice researchers often work with samples. For a sample of size n, sample variance uses denominator (n − 1): s2 = Σ(xi − x̄)2 / (n − 1). The (n − 1) adjustment (Bessel’s correction) produces an unbiased estimator of population variance when sample mean x̄ is used in place of unknown μ.
Computation Formulas and Shortcuts
Direct computation of Σ(xi − mean)2 requires many subtractions and squares. A practical shortcut is the computational formula Σ(xi − x̄)2 = Σxi2 − n x̄2 which reduces computations to sums of xi and xi2. For grouped data replace individual observations by class marks mi and use Σfi mi and Σfi mi2 to compute mean and variance approximately. The assumed mean and step-deviation methods further simplify calculations when values are large or classes many.
Standard Deviation and Interpretation
Standard deviation is the square root of variance and has the same units as the original data, making it easier to interpret. A larger standard deviation implies greater average deviation from the mean and thus more variability. For symmetric bell-shaped distributions, the standard deviation defines concentration: about 68% of observations lie within ±1σ of the mean and about 95% within ±2σ — this rule aids intuition though formal probability foundations extend beyond this class.
Economic Relevance and Cautions
Variance and standard deviation permit quantifying risk and comparing stability across series, especially via the coefficient of variation. They are essential inputs to regression analysis and hypothesis testing. However, they are sensitive to outliers, and squared units can make interpretation less direct; always accompany them with robust measures (median, quartile deviation) in reports. When comparing across datasets, ensure scale and units are compatible or use relative measures.
- For values 2,4,6: mean = 4, variance = [(2−4)2+(4−4)2+(6−4)2]/3 = (4+0+4)/3 = 8/3, standard deviation = sqrt(8/3).
- Grouped variance: Use class marks and frequencies to compute Σfi mi and Σfi mi2 then calculate x̄ and variance.
- Population variance σ2 = Σ(xi − μ)2 / N
- Sample variance s2 = Σ(xi − x̄)2 / (n − 1)
- Computational formula: Σ(xi − x̄)2 = Σxi2 − n x̄2
- Standard deviation σ or s = sqrt(variance)
Coefficient of Variation and Relative Measures
Why Relative Measures Are Needed
Absolute measures of dispersion, such as standard deviation, depend on the unit and scale of measurement. When comparing variability across two series measured in different units or with different means, absolute measures can be misleading. The coefficient of variation (C.V.) is a dimensionless relative measure that expresses standard deviation as a proportion of the mean, enabling meaningful comparisons across series.
Definition and Calculation
Coefficient of variation is defined as C.V. = (Standard deviation / Mean) × 100%. Use the sample standard deviation with the sample mean when dealing with a sample. Because C.V. is expressed as a percentage, it indicates the size of typical fluctuations relative to the average level. It is especially useful in finance and economics where risk per unit of expected return or volatility relative to average price matters.
Practical Examples and Interpretation
For example, if commodity A has mean price Rs.50 with SD Rs.5, its C.V. is 10%; if commodity B has mean Rs.200 with SD Rs.30, its C.V. is 15%: although B has larger absolute variation, relative to its mean it is more volatile. C.V. helps investors compare riskiness across assets and policymakers compare stability of macro indicators across countries with different average levels.
Limitations and Cautions
C.V. is only meaningful for ratio-scale data where the mean is positive and has an absolute zero. It becomes unstable or meaningless when the mean is near zero or when negative means occur. Moreover, two series with similar C.V. may still differ in distribution shape: C.V. summarises only scale relative to mean, not skewness or kurtosis. Use C.V. alongside other descriptive statistics for fuller insight.
Other Relative Measures
Other relative measures include relative range (range / mean × 100%) and relative mean deviation (MD/mean × 100%). Choice of relative measure depends on what aspect of variability is most relevant and on data properties. Report the measure used and its interpretation clearly when making comparisons so that readers understand the basis for conclusions.
- If mean income = Rs.10000 and standard deviation = Rs.2500, then C.V. = (2500/10000)×100% = 25%.
- Comparing two commodities where Commodity A has mean price Rs.50 SD=5 (C.V.=10%) and Commodity B mean Rs.200 SD=30 (C.V.=15%) shows B is relatively more volatile.
- Coefficient of Variation C.V. = (Standard deviation / Mean) × 100%
Correlation: Concept, Types and Scatter Diagram
What is Correlation?
Correlation is a statistical measure that describes the degree and direction of association between two variables. It tells us whether, as one variable changes, the other tends to change in a consistent way. Correlation is central in economics because many relationships — such as income and consumption, advertising and sales, education and earnings — are studied in terms of how strongly and in what direction the variables move together.
Types of Correlation
Correlation can be positive (both variables increase together), negative (one increases while the other decreases) or zero (no consistent linear association). The strength of correlation ranges from weak to strong. Correlation can also be linear (points roughly lie along a straight line) or non-linear (curvilinear relationships where variables move together but not on a straight line). Recognising the nature of association is important before applying numerical measures like Pearson’s coefficient which assumes linearity.
Scatter Diagram: Visual Tool
A scatter diagram (scatter plot) displays paired observations (x, y) on a Cartesian plane and is the first step in correlation analysis. Each point represents a pair; the overall pattern indicates direction and strength of association. A tight cluster around an upward-sloping line indicates strong positive linear correlation; a scattered cloud with no apparent trend suggests weak or no correlation. Scatter plots also reveal outliers, clusters, and departures from linearity that affect numerical measures.
Interpreting Scatter Plots
Examine scatter plots for slope, tightness and shape. If the cloud of points slopes upward, correlation is positive; downward slope indicates negative correlation. The tighter the cloud around a line, the stronger the linear correlation. Curved shapes suggest non-linear association, for which rank-based methods or non-linear models may be appropriate. Outliers can distort measures of correlation, so investigate and consider whether to retain, exclude or transform outlying observations.
Cautions: Correlation vs Causation
Correlation indicates association but not causation. Two variables may correlate due to a common underlying factor or pure coincidence. For instance, ice-cream sales and drowning incidents may correlate seasonally without one causing the other. In economics, establishing causation requires controlled analysis, theory, temporal ordering and possibly experimental evidence or econometric methods beyond simple correlation. Correlation is a beginning point for deeper inquiry rather than a conclusion by itself.
- Scatter plot of household income (x) vs expenditure (y) showing a positive linear trend.
- Scatter diagram between advertising expenditure and sales for several months to visualise association.
Karl Pearson's Coefficient of Correlation
Definition and Purpose
Karl Pearson's coefficient of correlation (r) is the standard measure of the strength and direction of linear association between two quantitative variables. It summarises how closely observations cluster around a straight-line relationship. The coefficient ranges from −1 to +1: r = +1 indicates perfect positive linear relationship, r = −1 perfect negative linear relationship, and r = 0 indicates no linear relationship. Pearson’s r is widely used in economics to quantify associations such as price and demand movements, income and consumption levels, and investment and GDP growth.
Mathematical Expression
For a sample of n paired observations (xi, yi) with means x̄ and ȳ, Pearson’s r is r = Σ(xi − x̄)(yi − ȳ) / sqrt[Σ(xi − x̄)2 Σ(yi − ȳ)2]. This formula uses deviations from the mean and produces a dimensionless value. For practical computation the computational form r = [nΣxi yi − (Σxi)(Σyi)] / sqrt{[nΣxi2 − (Σxi)2][nΣyi2 − (Σyi)2]} is often used because it requires only sums of xi, yi, xi2, yi2 and xi yi, simplifying hand calculations.
Properties and Interpretation
Pearson’s r is symmetric: correlation between x and y equals correlation between y and x. Its magnitude indicates strength: values closer to 1 or −1 imply stronger linear association. As a rule of thumb, |r| < 0.3 suggests weak correlation, 0.3–0.6 moderate, 0.6–0.9 strong, though context matters. Always inspect the scatter plot since r measures only linear association; non-linear but strong relationships may yield low r. Outliers can distort r significantly, so analyze data visually and consider robust alternatives if needed.
Assumptions and Limitations
Pearson's r assumes variables are measured at least on interval scales and that the relationship is approximately linear. It is sensitive to the range of data: restricting the range of x or y reduces r even if the underlying relationship is strong. Also, correlation does not imply causation; a high r should lead to further investigation rather than immediate causal claims. For ordinal data or monotonic non-linear relationships, Spearman's rank correlation may be more appropriate.
Applications in Economics
Use Pearson’s r to quantify associations in empirical studies, to summarise relationships before regression analysis, and to compare strengths of linear relationships across datasets. Report r along with scatter plots and sample size, and discuss potential confounders or outliers that may influence the result.
- Compute r for paired data (x: years of schooling, y: income) using the computational formula and interpret sign and magnitude.
- Two small series x: 1,2,3 and y:2,4,6 yield r = 1 indicating perfect positive linear correlation.
- Pearson's r = [Σ(xi − x̄)(yi − ȳ)] / sqrt[Σ(xi − x̄)2 Σ(yi − ȳ)2]
- Computational form: r = [nΣxi yi − (Σxi)(Σyi)] / sqrt{[nΣxi2 − (Σxi)2][nΣyi2 − (Σyi)2]}
Spearman's Rank Correlation Coefficient
Purpose and Appropriateness
Spearman's rank correlation coefficient (rs) measures the strength and direction of a monotonic relationship between two variables based on ranks rather than original values. It is useful when data are ordinal, when the relationship is non-linear but monotonic, or when outliers make Pearson’s correlation unreliable. Because it uses ranks, Spearman’s rs is less sensitive to extreme values and does not require interval scaling.
Computation Steps
To compute rs, first assign ranks to the observations of each variable. If two or more observations have the same value (ties), assign them the average rank of their positions. For each paired observation, compute the difference di between the two ranks. Then use the formula rs = 1 − [6 Σdi2 / n(n2 − 1)], where n is the number of pairs. The value of rs lies between −1 and +1 with interpretations similar to Pearson’s r in terms of direction and strength of monotonic association.
Handling Ties and Small Samples
Ties complicate the calculation because the simple formula assumes distinct ranks. When ties are present, average ranks are used and the formula remains an approximation; for many ties more exact correction terms exist but they are not required at Class 11. In small samples, sample variability can make rs unstable, so interpret cautiously and consider complementing rank analysis with visual inspection of paired plots.
Interpretation and Examples
An rs close to +1 indicates that higher ranks of x correspond consistently to higher ranks of y (strong positive association), while rs near −1 indicates opposite ordering. For example, if students ranked by study hours and exam performance show similar orderings, rs will be high and positive, suggesting a monotonic relation. Spearman’s rank is particularly useful for survey data where responses are given on ordered scales (e.g., satisfaction from 1 to 5).
Comparison with Pearson's Coefficient
Spearman’s rs captures monotonic relationships that may be non-linear, whereas Pearson’s r captures linear relationships and is more efficient when linear assumptions hold. Use Spearman’s coefficient when data are ordinal, when measurement scales are unclear, or when robustness to outliers is required. Reporting both coefficients, along with scatter plots, gives a fuller picture of the association between variables.
- Two variables ranked 1–5; compute di for each pair and use rs = 1 − [6Σdi2 / n(n2 − 1)] to find association.
- Ranking students by marks in Economics and Mathematics and computing Spearman's rs to measure association.
- Spearman's rs = 1 − [6 Σdi2 / n(n2 − 1)]
Regression Analysis: Lines of Best Fit and Regression Equations
Purpose of Regression
Regression analysis estimates the relationship between a dependent variable y and an independent variable x. While correlation measures association, regression provides a predictive equation and quantifies how much y changes on average for a unit change in x. This is central in economics where one often wishes to predict outcomes (e.g., consumption) from causal or explanatory variables (e.g., income).
Regression Lines of y on x and x on y
For two variables there are two ordinary least squares (OLS) regression lines: the regression of y on x (used to predict y from x) and the regression of x on y (used to predict x from y). They are asymmetric: slopes differ except in perfect correlation. The regression line of y on x has form y − ȳ = byx (x − x̄) where byx = r (sy / sx). Similarly x − x̄ = bxy (y − ȳ) where bxy = r (sx / sy). These relationships link regression coefficients to Pearson’s correlation and standard deviations, providing intuitive understanding of slope.
Least Squares Principle
OLS chooses the line that minimises the sum of squared vertical deviations of observed y from predicted y. The intercept ensures the regression line passes through the centroid (x̄, ȳ). In practice students can compute slopes using r, sx and sy or directly from sums of cross-products. For grouped data class marks serve as xi and yi for approximate calculation.
Interpretation, Prediction and Cautions
Slope byx indicates average change in y for a one-unit increase in x; for example, a slope of 0.6 suggests y increases by 0.6 units per unit rise in x on average. Regression supports prediction within the data range, but extrapolation beyond observed values is risky. Regression identifies statistical association, and causal interpretation requires theory, time-ordering and control for omitted variables. Omitted variables, measurement error and reverse causality can bias estimates.
Goodness of Fit
Coefficient of determination r2 measures fraction of variance in y explained by x in a simple linear regression. An r2 close to 1 indicates that x explains most variability in y. However, acceptable r2 values depend on context: in social sciences lower r2 values may still be meaningful. Always accompany regression results with diagnostic checks such as scatter plots, residual patterns and discussion of assumptions.
- Given r=0.8, sx=10, sy=5, slope of y on x byx = r (sy/sx) = 0.8 × (5/10) = 0.4; regression line y − ȳ = 0.4(x − x̄).
- Predicting consumption for a given income using the regression equation obtained from sample data.
- Regression coefficient of y on x byx = r (sy / sx)
- Regression coefficient of x on y bxy = r (sx / sy)
- Regression equations: y − ȳ = byx (x − x̄) and x − x̄ = bxy (y − ȳ)
- Coefficient of determination r2 = (Pearson's r)2
Uses and Limitations of Statistics in Economics: Interpretation and Ethical Use
Practical Uses in Policy and Business
Statistics underpins decision-making in both public and private sectors. Governments use statistical indicators to estimate GDP, inflation, unemployment, poverty and inequality, which guide fiscal and monetary policy. Businesses rely on market surveys, demand forecasts and productivity measures for planning production, pricing and investment. Researchers use statistical summaries and tests to evaluate economic hypotheses and policy impacts. Statistics thus translates complex economic realities into manageable quantitative inputs for planning and evaluation.
Guidelines for Correct Interpretation
Interpreting statistics requires attention to data source, measurement definitions, collection methods and sample design. Always check whether data are from reliable agencies, whether definitions match across datasets, and whether time periods are comparable. When averages are reported, examine dispersion measures to understand whether the average is representative. Use graphical tools to detect patterns and outliers before relying on numerical summaries. Document assumptions and limitations when presenting results so users can assess credibility.
Common Misuses and Ethical Concerns
Statistics can be intentionally or unintentionally misused. Common abuses include selective reporting (cherry-picking favourable periods), misleading graphs (truncated axes, disproportionate scales), inappropriate averages (using arithmetic mean for highly skewed data), and confusing correlation with causation. Ethically, analysts should avoid manipulating visual presentation to mislead, must disclose methods and limitations, and should not claim causal effects without appropriate evidence. Transparency and full reporting build trust in statistical findings.
Critical Thinking for Students
Students should cultivate a questioning approach: ask about sample representativeness, examine raw data where possible, look for outliers, and test robustness of results to different measures. Understand that statistical uncertainty exists: sampling error, non-sampling error and model specification all affect conclusions. When reading media reports, seek the original source and evaluate whether conclusions are supported by data and methods.
Preparing for Advanced Analysis
This unit forms the foundation for advanced topics such as probability theory, sampling distributions, hypothesis testing and econometrics. Mastery of descriptive techniques, correlation and simple regression prepares students to undertake empirical research responsibly and to interpret economic statistics used in policy debates and academic literature. Ethical and careful use of statistics enhances the quality of economic decision-making and public discourse.
- A newspaper chart showing unemployment falling may hide that labour force also fell; check definitions and source.
- A company reporting average salary increase may omit that raises were concentrated at top; check dispersion measures.
Key Concepts
- Statistics
- A set of methods for collecting, organising, presenting and interpreting numerical data.
- Descriptive Statistics
- Methods that summarise and present data, such as averages and graphs.
- Inferential Statistics
- Techniques that draw conclusions about a population based on a sample using probability.
- Population
- The complete set of items or individuals under study.
- Sample
- A subset of the population selected for analysis.
- Frequency Distribution
- A tabular summary showing classes or values and their corresponding frequencies.
- Histogram
- A bar diagram representing frequency distribution of continuous data with adjacent bars touching.
- Mean
- Arithmetic average calculated as the sum of observations divided by their number.
- Median
- The middle value that divides an ordered data set into two equal parts.
- Mode
- The observation or value that occurs most frequently in a data set.
- Variance
- The average of squared deviations from the mean measuring dispersion.
- Standard Deviation
- The square root of variance, measuring average distance of observations from the mean.
- Coefficient of Variation
- Relative measure of dispersion equal to standard deviation divided by mean, expressed as percentage.
- Correlation
- A numerical measure of the degree and direction of association between two variables.
- Regression
- Method to estimate the average relationship between a dependent variable and one or more independent variables.
- Quartile Deviation
- Half the difference between the third and first quartiles, measuring spread of middle 50%.
- Ogive
- A cumulative frequency curve used to estimate medians and percentiles graphically.
- Spearman's Rank Correlation
- A correlation measure based on ranks, suitable for ordinal data or monotonic relationships.
Practice Questions
-
Define statistics and explain two uses of statistics in economics. / सांख्यिकी को परिभाषित करें और अर्थशास्त्र में सांख्यिकी के दो उपयोग बताइए।
Show answer
Statistics is a set of methods for collecting, organising, presenting and interpreting numerical data to make decisions. In economics, statistics is used to (1) estimate national income and construct price indices that guide macroeconomic policy, and (2) analyse relationships such as income and consumption to inform planning and forecasting. / सांख्यिकी संख्यात्मक डेटा को इकट्ठा, व्यवस्थित, प्रस्तुत और व्याख्यायित करने की विधियों का समूह है ताकि निर्णय लिए जा सकें। अर्थशास्त्र में, सांख्यिकी का उपयोग (1) राष्ट्रीय आय का अनुमान लगाने और मूल्य सूचकांक बनाने में किया जाता है जो प्रमुख नीतिगत निर्देश देता है, और (2) आय और उपभोग जैसे संबंधों का विश्लेषण कर योजना और पूर्वानुमान में मदद करने के लिए किया जाता है।
-
What is the difference between qualitative and quantitative data? Give one example of each. / गुणात्मक और मात्रात्मक डेटा में क्या अंतर है? प्रत्येक का एक उदाहरण दीजिए।
Show answer
Qualitative data describe categories or attributes (e.g., industry type: agriculture, manufacturing), while quantitative data are numerical measurements (e.g., monthly income in rupees). / गुणात्मक डेटा श्रेणियों या गुणों का वर्णन करते हैं (जैसे उद्योग प्रकार: कृषि, विनिर्माण), जबकि मात्रात्मक डेटा संख्यात्मक मापन होते हैं (जैसे मासिक आय रुपये में)।
-
Explain how to construct a grouped frequency distribution from continuous data. / सतत डेटा से समूहित आवृत्ति वितरण कैसे बनाते हैं, समझाइए।
Show answer
Steps: (1) Find range = max − min; (2) Choose number of classes (5–20 depending on sample size); (3) Compute class width ≈ range/number of classes and choose convenient width; (4) Determine class limits so classes are mutually exclusive and exhaustive; (5) Tally observations into classes and record frequencies; (6) Compute class marks and cumulative frequencies if required. State total frequency and class boundaries. / चरण: (1) रेंज निकालें = अधिकतम − न्यूनतम; (2) वर्गों की संख्या चुनें (नमूना आकार के अनुसार 5–20); (3) वर्ग चौड़ाई ≈ रेंज/वर्ग संख्या निकालें और सुविधाजनक चौड़ाई चुनें; (4) वर्ग सीमाएँ तय करें ताकि वे पारस्परिक रूप से विशेष और समेकित हों; (5) अवलोकनों को वर्गों में गिन कर आवृत्तियाँ लिखें; (6) आवश्यकता होने पर वर्ग चिह्न और संचयी आवृत्ति निकालें। कुल आवृत्ति और वर्ग सीमाएँ बताना न भूलें।
-
Calculate the arithmetic mean of the following discrete series: Value: 10, 20, 30, 40; Frequency: 2, 3, 4, 1. / निम्नलिखित विविक्त श्रेणी का अंकगणितीय माध्य निकालिए: मान:10,20,30,40; आवृत्ति:2,3,4,1।
Show answer
Compute Σfi xi = 10×2 + 20×3 + 30×4 + 40×1 = 20 + 60 + 120 + 40 = 240; Σfi = 2+3+4+1 = 10; Mean = 240 / 10 = 24. / Σfi xi = 10×2 + 20×3 + 30×4 + 40×1 = 20 + 60 + 120 + 40 = 240; Σfi = 10; माध्य = 240 / 10 = 24।
-
Given the grouped distribution with class intervals 0–9, 10–19, 20–29 and frequencies 5, 12, 8, find the median. / कक्षा अंतराल 0–9, 10–19, 20–29 और आवृत्तियाँ 5, 12, 8 वाली समूहित वितरण में मध्यमान (मेडियन) ज्ञात कीजिए।
Show answer
Total N = 5+12+8 = 25. N/2 = 12.5. Cumulative frequencies: up to 0–9 =5, up to 10–19 =17. Median class is 10–19. Using Median = L + [(N/2 − cf) / f] × h. Here L=10, cf=5 (cumulative before median class), f=12, h=10. Median = 10 + [(12.5 − 5)/12]×10 = 10 + (7.5/12)×10 = 10 + 6.25 = 16.25. / कुल N =25, N/2 =12.5। संचयी आवृत्ति: 0–9 तक=5, 10–19 तक=17। मध्यमान वर्ग 10–19 है। सूत्र Median = L + [(N/2 − cf) / f] × h. यहाँ L=10, cf=5, f=12, h=10। Median = 10 + [(12.5 − 5)/12]×10 = 10 + 6.25 = 16.25।
-
A dataset has mean 50 and standard deviation 5. What is the coefficient of variation? Explain what it means. / किसी डेटासेट का माध्य 50 और मानक विचलन 5 है। सहविचलन गुणांक क्या होगा? इसका अर्थ समझाइए।
Show answer
Coefficient of variation C.V. = (SD / Mean) × 100% = (5 / 50) × 100% = 10%. It means the standard deviation is 10% of the mean, indicating moderate relative variability; useful to compare variability with other series irrespective of units. / सहविचलन = (5/50)×100% = 10%। इसका अर्थ है मानक विचलन माध्य का 10% है, जो मध्यम सापेक्षिक अस्थिरता दिखाता है; यह अन्य श्रृंखलाओं के साथ तुलना के लिए उपयोगी है क्योंकि यह इकाइयों पर निर्भर नहीं है।
-
Compute Pearson's correlation coefficient for the pairs (x,y): (1,2), (2,3), (3,5). / जोड़े (x,y): (1,2), (2,3), (3,5) के लिए पियर्सन सहसम्बन्ध गुणांक निकालिए।
Show answer
Compute sums: Σx=1+2+3=6, Σy=2+3+5=10, Σxy=1×2+2×3+3×5=2+6+15=23, Σx2=1+4+9=14, Σy2=4+9+25=38, n=3. Use r = [nΣxy − (Σx)(Σy)] / sqrt{[nΣx2 − (Σx)2][nΣy2 − (Σy)2]} = [3×23 − 6×10] / sqrt{[3×14 − 36][3×38 − 100]} = [69 − 60] / sqrt{[42 − 36][114 − 100]} = 9 / sqrt{6×14} = 9 / sqrt{84} = 9 / 9.165 = 0.982. So r ≈ 0.982, indicating a very strong positive linear correlation. / Σx=6, Σy=10, Σxy=23, Σx2=14, Σy2=38, n=3। r = [3×23 − 6×10] / sqrt{[3×14 − 36][3×38 − 100]} = 9 / sqrt{84} ≈ 0.982। यह बहुत मजबूत धनात्मक रैखिक सहसम्बन्ध दर्शाता है।
-
Explain the difference between correlation and regression. / सहसम्बन्ध और प्रतिगमन में अंतर समझाइए।
Show answer
Correlation measures the strength and direction of association between two variables but does not distinguish dependent and independent variables; it is symmetric. Regression quantifies the expected change in a dependent variable for a unit change in an independent variable and provides equations for prediction; regression is asymmetric and distinguishes dependent/independent variables. Correlation does not imply causation whereas regression is used to model causal hypotheses (with care). / सहसम्बन्ध दो चर के बीच सम्बन्ध की दिशा और ताकत को मापता है लेकिन निर्भर और स्वतंत्र चर में भेद नहीं करता; यह सममित होता है। प्रतिगमन यह बताता है कि स्वतंत्र चर में एक इकाई परिवर्तन पर निर्भर चर में औसतन कितना परिवर्तन होगा और भविष्यवाणी के लिए समीकरण देता है; प्रतिगमन असममित होता है और निर्भर/स्वतंत्र चर अलग करता है। सहसम्बन्ध कारण बताता नहीं है जबकि प्रतिगमन को कारण संबंधी परिकल्पनाओं के मॉडल के रूप में उपयोग किया जा सकता है (सावधानी के साथ)।
-
Describe how an ogive can be used to estimate the 75th percentile (third quartile) of a grouped distribution. / किसी समूहित वितरण में ओगाइव का उपयोग करके 75वें पर्सेंटाइल (तीसरा चौथाईक) का अनुमान कैसे लगाया जाता है, वर्णन कीजिए।
Show answer
Construct a less-than ogive by plotting upper class boundaries against cumulative frequencies and join by a smooth line. Total N gives 75% frequency = 0.75N. Draw a horizontal line at cumulative frequency = 0.75N; where it meets the ogive, drop a vertical to the horizontal axis and read the corresponding variable value; this is the estimated 75th percentile (Q3). For more accuracy, ensure correct class boundaries and smooth joining. / ऊपरी वर्ग सीमाओं के विरुद्ध संचयी आवृत्तियाँ प्लॉट करके एक 'less-than' ओगाइव बनाइए और उसे समतल रेखा से जोड़िए। कुल N का 75% = 0.75N निकालेँ। संचयी आवृत्ति = 0.75N पर क्षैतिज रेखा खींचें; जहाँ यह ओगाइव से मिलती है वहाँ से एक लंबवत रेखा नीचे गिराकर क्षैतिज अक्ष पर वैरिएबल का मान पढ़ें; यही अनुमानित 75वाँ पर्सेंटाइल (Q3) होगा। अधिक सटीकता के लिए वर्ग सीमाएँ सही रखें और रेखा को चिकना बनाएं।
-
A regression line of y on x is given by y − 40 = 0.5(x − 20). If x = 30, find predicted y. / y पर x का प्रतिगमन रेखा y − 40 = 0.5(x − 20) दी गई है। यदि x = 30 हो तो पूर्वानुमानित y क्या होगा?
Show answer
Substitute x = 30: y − 40 = 0.5(30 − 20) = 0.5×10 = 5; so y = 45. The model predicts y to be 45 when x is 30. / x = 30 रखते हुए: y − 40 = 0.5(10) = 5; अतः y = 45। जब x = 30 हो, मॉडल y = 45 पूर्वानुमानित करता है।
-
Mention two limitations of statistics when applied to economic problems. / आर्थिक समस्याओं पर लागू होने पर सांख्यिकी की दो सीमाएँ बताइए।
Show answer
Two limitations: (1) Dependence on data quality — inaccurate, biased or incomplete data lead to misleading conclusions. (2) Correlation versus causation — statistical association may not indicate causal relationship; external factors or omitted variables can produce spurious correlations. / दो सीमाएँ: (1) डेटा की गुणवत्ता पर निर्भरता — गलत, पक्षपाती या अधूरी जानकारी भ्रामक निष्कर्ष देती है। (2) सहसम्बन्ध और कारणत्व का अंतर — सांख्यिकीय सम्बन्ध जरूरी नहीं कि कारण दिखाए; बाहरी तत्व या छोड़े गए चर बनावट में भ्रामक सहसम्बन्ध दे सकते हैं।
Related Laws & Principles
Explore allFoundational laws & principles connected to this chapter — tap to open in the Laws Explorer.