L
LLLOS.ai
Learn
L

Chapter 3 — Organisation Of Data

Class 11 · Economics

Overview

Chapter 3 — Organisation Of Data Master Diagram

Introduction: This chapter introduces Organisation of Data — the first step in statistical analysis. It explains how raw observations collected in economic surveys and experiments are transformed into a form that is easier to understand and interpret. The chapter differentiates types of data (qualitative vs quantitative; discrete vs continuous) and shows how classification and tabulation lead to meaningful summaries. Importance: Organising data is essential in economics because policy decisions, business plans and forecasts depend on clear, accurate presentation of information (e.g., prices, production, income, population). Correct organisation reduces complexity, removes ambiguity, prevents errors in later analysis (like calculating averages or measures of dispersion) and enables effective graphical presentation and comparison. Key themes: The chapter covers principles of classification and tabulation, construction of frequency distributions for ungrouped and grouped data, class intervals (class limits, class boundaries, class width), class marks (mid-points), and types of frequency (absolute, relative, cumulative). It also introduces graphical methods used to present frequency…

Learning Objectives

  • Define primary and secondary data with suitable examples
  • Distinguish between qualitative and quantitative data and between discrete and continuous variables
  • Explain the steps and rules used for classification and tabulation of data
  • Construct frequency distribution tables for ungrouped (discrete) data
  • Organise continuous data into grouped frequency distributions by choosing appropriate class intervals and class size
  • Calculate class limits, class boundaries and class marks (mid‑points) for grouped data
  • Compute absolute frequencies, relative (proportion) frequencies, percentage frequencies and cumulative frequencies
  • Draw histograms and frequency polygons for continuous grouped data

Topics in this chapter

7 topics · tap a topic title to jump straight to it.

📊1

Types of Data

📊 COMMERCE / ECONOMIC LAW

Types of Data

Key Point: Relative frequency = frequency / total observations

Overview: In organising data we classify it so appropriate methods of tabulation, presentation and analysis can be used. "Types of data" refers to different ways data can be grouped according to source, nature, continuity, time-dimension and measurement scale.

1. By Source

  • Primary data: Collected first-hand for a specific purpose (surveys, interviews, experiments). Advantage: tailored and current. Disadvantage: costly and time-consuming.
  • Secondary data: Already collected for another purpose (government publications, research papers, databases). Advantage: cheaper and quick. Disadvantage: may not fit exactly and may be outdated or biased.

2. By Nature

  • Qualitative (Categorical) data: Describes qualities or categories. Examples: gender, religion, colour, type of house. Usually summarized by counts or percentages.
  • Quantitative (Numerical) data: Numerical values that measure quantity. Examples: income, age, production. Can be further analysed by arithmetic operations.

3. By Continuity (for quantitative data)

  • Discrete data: Countable values (integers). Example: number of students in a class, number of factories.
  • Continuous data: Can take any value in an interval (measurements). Example: height, weight, distance, income measured precisely.

4. By Time-Dimension

  • Time-series data: Observations of a variable over time (monthly GDP, annual rainfall). Used to study trends, seasonal effects and cycles.
  • Cross-section data: Observations at a single point in time across units (income of households in 2024). Used to compare units at same time.
  • Pooled (Panel) data: Combination of time-series and cross-section (same households observed across several years). Useful for dynamic analysis.

5. By Measurement Scale

  • Nominal: Names or labels without order (e.g., blood group). Only mode and percentages are meaningful.
  • Ordinal: Categories with order but unequal intervals (e.g., education level: primary, secondary, tertiary). Median and percentiles meaningful.
  • Interval: Ordered with equal intervals but no true zero (e.g., Celsius temperature). Differences are meaningful, ratios are not.
  • Ratio: Like interval but with a true zero (e.g., income, weight). All arithmetic operations and ratios are meaningful.

Use in economics: Correct identification of data type determines choice of tables, graphs and statistical measures (mean, median, mode, dispersion measures, correlation, regression). For example, categorical data are shown with bar charts/pie charts; continuous quantitative data with histograms, frequency polygons and box plots; time-series with line graphs; cross-section with scatter plots.

📌 Examples
  • Primary data: A researcher conducts a household survey to collect monthly expenditure data directly from respondents.
  • Secondary data: Using RBI or NSSO published tables on employment or national income for analysis.
  • Qualitative: Classify households by type of fuel used for cooking (wood, LPG, kerosene).
  • Quantitative discrete: Number of factories in a district (0,1,2,...).
  • Quantitative continuous: Monthly income of workers measured to the rupee (can be treated as continuous in practice).
  • Time-series: GDP of India from 2010 to 2024 (annual observations).
🧮 Formulas
  1. \[Relative frequency = frequency / total observations\]
  2. \[Percentage = (frequency / total observations) × 100\]
  3. \[Cumulative frequency (up to a class) = sum of frequencies of that class and all previous classes\]
  4. \[Class mark (midpoint) = (lower limit + upper limit) / 2\]
  5. \[Class width = upper limit − lower limit\]
📊2

Sources and Methods of Data Collection

📊 COMMERCE / ECONOMIC LAW

Sources and Methods of Data Collection

Key Point: Sample mean (x̄) = (Σxi) / n — average of observations in the sample.

Overview

Data collection is the systematic process of gathering information for analysis. In economics (Class 11), we classify data by its source (where it comes from) and by the method used to collect it. Good data collection ensures accuracy, reliability and relevance for decision-making and statistical analysis.

Sources of Data

There are two broad categories:

  • Primary data: Data collected first-hand for a specific purpose. It is original and fresh.
  • Secondary data: Data already collected by someone else and available in published or unpublished form.

Primary sources (examples and use)

  • Surveys and questionnaires — ask respondents directly.
  • Interviews — structured, semi-structured, or unstructured.
  • Observation — direct (researcher observes behaviour) or indirect (use of cameras, records).
  • Experiments — controlled manipulation to observe effects.
  • Case studies — in-depth study of a single unit (e.g., a firm or household).

Secondary sources (examples and use)

  • Official publications (Census of India, NSSO/NSO reports, government statistics).
  • Administrative records (tax records, hospital registers, school enrolments).
  • Published sources (books, journals, newspapers, company annual reports).
  • Online databases, research repositories, and archival material.

Methods of Data Collection

Choice of method depends on objectives, budget, time, required accuracy and the nature of respondents.

Census versus Sampling

  • Census: Collect data from every unit in the population (e.g., decennial population census). Accurate but expensive/time-consuming.
  • Sampling: Collect data from a subset (sample) of the population and generalise results. More practical and cheaper; requires careful design to avoid sampling bias.

Common sampling designs

  • Probability sampling (each unit has known chance): simple random sampling, systematic sampling, stratified sampling, cluster sampling.
  • Non-probability sampling (selection not random): convenience sampling, judgmental (purposive) sampling, quota sampling, snowball sampling.

Primary data collection methods — brief descriptions

  • Questionnaire: Written set of questions—good for large samples. Questions can be closed or open-ended.
  • Interview: Oral questioning—face-to-face, telephone, or video. Can probe deeper but costlier.
  • Observation: Recording behaviour or events without asking. Useful when responses may be biased.
  • Schedule: Similar to questionnaire but filled by the investigator after interviewing respondent—used in large-scale studies (e.g., household surveys).
  • Experiment: Manipulating variables to test hypotheses—common in behavioural economics/lab studies.
  • Case study: Intensive examination of a single entity for in-depth understanding.

Quality, Errors and Ethics

  • Validity: Data measures what it is intended to measure.
  • Reliability: Data collection would yield similar results under consistent conditions.
  • Errors: sampling errors (due to sample selection) and non-sampling errors (measurement error, non-response, data processing mistakes).
  • Ethics: informed consent, confidentiality, no harm, honest reporting.

Steps in Data Collection (practical workflow)

  1. Define objective and population.
  2. Choose data source (primary/secondary) and method.
  3. Design instrument (questionnaire/schedule) and sample (if sampling).
  4. Pilot test instrument and revise.
  5. Collect data, supervise and record metadata (who, when, how).
  6. Process, clean and validate the data.

When to use which source/method?

  • Use primary data when specific, up-to-date, or otherwise unavailable information is needed (e.g., consumer preferences for a new product).
  • Use secondary data for background, trend analysis, or when reliable official data exist (e.g., GDP series, census counts).

Note: In practice researchers combine methods (mixed methods) to improve coverage and validity—e.g., use secondary data for context and primary surveys for current measurements.

📌 Examples
  • Census of India (every 10 years) — a complete enumeration (census) collecting population and housing data from every household.
  • National Sample Survey Office (NSSO/NSO) consumer expenditure survey — uses stratified multi-stage sampling and schedules filled by investigators.
  • Market research firm conducts an online questionnaire (primary data) to find consumer preference for a soft drink before launch.
  • A researcher uses hospital admission registers (secondary administrative records) to study disease trends over 5 years.
  • A sociologist conducts in-depth interviews and observational fieldwork (case studies + observation) to study slum livelihoods.
  • Exit polls during elections use sampling; results depend on sample design and response rate.
🧮 Formulas
  1. \[Sample mean (x̄) = (Σxi) / n — average of observations in the sample.\]
  2. \[Sample proportion (p̂) = x / n — proportion of sample with a characteristic (x = number of successes).\]
  3. \[Sample variance (s^2) = [Σ(xi - x̄)^2] / (n - 1) — measure of spread in the sample.\]
  4. \[Standard deviation (s) = sqrt(s^2).\]
  5. \[Response rate (%) = (Number of responses / Number of persons approached) × 100.\]
  6. \[Margin of error for proportion (approx.) E = Z * sqrt(p̂(1 - p̂) / n) — Z from standard normal for chosen confidence level (e.g., 1.96 for 95%).\]
📊3

Classification of Data

📊 COMMERCE / ECONOMIC LAW

Classification of Data

Key Point: Class mark (midpoint) m = (Lower class limit + Upper class limit) / 2

Meaning: Classification of data is the process of arranging raw data into groups or classes that have similar characteristics so that it becomes meaningful and easier to analyze. Proper classification helps in summarizing large volumes of data and in choosing appropriate methods of presentation and analysis.

Why classify? Raw data are often voluminous and unstructured. Classification simplifies data, highlights patterns, facilitates comparison and helps in drawing conclusions.

Main bases (types) of classification:

  • By source: Primary data (collected first-hand through surveys, experiments, observations) and Secondary data (collected earlier by someone else — censuses, reports, books, databases).
  • By nature: Qualitative (attribute/categorical data describing qualities — e.g., gender, occupation) and Quantitative (numerical data measuring quantity — e.g., income, marks).
  • By measurement/ordering:
    • Nominal: categories without order (e.g., religion, blood group).
    • Ordinal: categories with a meaningful order but not equal intervals (e.g., grades: A, B, C; satisfaction levels: low, medium, high).
    • Interval/Ratio: numerical scales with meaningful differences; ratio scale has a true zero (e.g., weight, income).
  • By continuity: Discrete data (take distinct separate values — counts such as number of children) and Continuous data (take any value in a range — measurements such as height, time).
  • By statistical unit: Individual data (values for each unit — marks of each student) and Aggregate data (summary measures for groups — total sales of a firm, average income of a region).
  • By time/space: Cross-sectional data (observations collected at the same time for different units — incomes of households in 2025) and Time-series data (observations of a variable taken over time — monthly CPI, yearly GDP).

How classification affects presentation: Choice of charts and tables depends on the type of data. Qualitative data are usually shown with bar charts or pie charts. Quantitative discrete data are shown with bar charts; quantitative continuous data are shown with histograms, frequency polygons or ogives. Time-series data are best shown with line graphs.

Important terms related to classification:

  • Class interval or class: a group into which continuous data are divided (e.g., 10–19, 20–29).
  • Class width: size of a class interval.
  • Class mark (midpoint): the central value of a class interval used for certain calculations.
  • Frequency: number of observations in a class or category.

Practical tips: Use appropriate categories (mutually exclusive and exhaustive). For continuous data choose suitable class width and number of classes so that the frequency distribution is neither too detailed nor too coarse.

📌 Examples
  • Primary vs Secondary: A student conducting a household income survey collects primary data; using government census tables is using secondary data.
  • Qualitative (Nominal): Blood groups of patients (A, B, AB, O).
  • Qualitative (Ordinal): Customer satisfaction ratings: dissatisfied, neutral, satisfied.
  • Quantitative Discrete: Number of children in families: 0,1,2,3,...
  • Quantitative Continuous: Heights of students measured in cm (e.g., 150.2 cm).
  • Cross-sectional: Household incomes of 100 families in 2024 (observed at one point in time).
🧮 Formulas
  1. \[Class mark (midpoint) m = (Lower class limit + Upper class limit) / 2\]
  2. \[Class width (continuous) = Upper class boundary − Lower class boundary (or for inclusive discrete classes = Upper limit − Lower limit + 1)\]
  3. \[Relative frequency = Frequency / Total number of observations\]
  4. \[Percentage frequency = (Frequency / Total) × 100\]
  5. \[Frequency density (for unequal class widths) = Frequency / Class width\]
  6. \[Cumulative frequency (≤ upper limit) = Sum of frequencies up to that class\]
📈4

Tabulation

📊 COMMERCE / ECONOMIC LAW

Tabulation

Key Point: Total frequency: n = Σ fi (sum of all class frequencies)

Definition: Tabulation is the process of arranging raw data in rows and columns (a table) according to a plan so that the information becomes clear, systematic and easily interpretable. In economics and statistics it is the first step in organising data for further analysis.

Purpose / Importance:

  • Summarises large volumes of data in compact form.
  • Makes comparison and pattern recognition easier.
  • Prepares data for calculation of statistical measures and for graphical presentation.

Components of a Table:

  • Title: concise description of what the table contains.
  • Headings: column headings that describe variables (e.g., Year, Population).
  • Stubs: row labels (e.g., States, Categories).
  • Body: the cells where data values are placed.
  • Footnote / Source: any clarifications or data sources.

Types of Tables:

  • Simple table: Lists observations (e.g., names and marks).
  • Frequency distribution table: Groups data into classes with frequencies.
  • One-way table: Frequency distribution of a single variable.
  • Two-way (contingency) table: Cross-classifies two variables (e.g., gender × employment status).

Rules / Guidelines for Constructing Tables:

  • Give a clear and complete title.
  • Arrange rows and columns logically (e.g., chronological order, increasing magnitude).
  • Use mutually exclusive and exhaustive classes when grouping continuous data.
  • Choose an appropriate number of classes (not too many, not too few).
  • Use consistent units and state them.
  • Place totals (marginal totals) where useful.
  • Make the table readable: align numbers, use adequate spacing and clear headings.

How Tabulation Helps Further Analysis: After tabulation you can compute frequencies, relative frequencies, percentages, cumulative frequencies and prepare graphs (histograms, pie charts, ogives) that summarise the data visually.

Limitations: A table can hide individual observations; grouping may cause loss of detail; poor design can mislead readers.

📌 Examples
  • Student marks: A table listing students' names, roll numbers and marks in five subjects; used to find pass percentage and class average.
  • Household income survey: A frequency distribution showing number of households in income classes (e.g., 0-20k, 20k-40k, etc.) to analyse income distribution.
  • Retail prices: A two-way table of product categories (rows) and months (columns) to compare price changes across time.
  • Rainfall data: Monthly rainfall in mm presented in a table to compute seasonal totals and prepare a line graph.
  • Employment by sector: A contingency table classifying workforce by sector (agriculture, industry, services) and gender to study labour patterns.
🧮 Formulas
  1. \[Total frequency: n = Σ fi (sum of all class frequencies)\]
  2. \[Relative frequency: ri = fi / n\]
  3. \[Percentage frequency: pi = (fi / n) × 100\]
  4. \[Cumulative frequency (up to class k): CFk = Σ (fi for classes ≤ k)\]
  5. \[Class width (for continuous data): h = (Upper limit - Lower limit) of a class (or approximate = range / number of classes)\]
  6. \[Class mark (mid-point): x_i = (Lower limit + Upper limit) / 2\]
📈5

Frequency Distribution and Grouping

📊 COMMERCE / ECONOMIC LAW

Frequency Distribution and Grouping

Key Point: Range = Maximum − Minimum

What is Frequency Distribution? A frequency distribution organises raw data into classes (intervals) and shows how often (frequency) values occur in each class. It simplifies large data sets and helps identify patterns (central tendency, dispersion, shape).

Why grouping? Raw data can be unwieldy. Grouping into classes reduces detail while preserving important structure so we can display data using histograms, frequency polygons and ogives.

Steps to construct a grouped frequency distribution

  • Collect and sort the raw data (optional).
  • Find the range: maximum − minimum.
  • Decide number of classes (k). Use Sturges' rule as a guideline: k ≈ 1 + 3.322 log10(n), then round.
  • Determine class width (h) = range / k, then round up to a convenient value.
  • Choose class limits (non-overlapping continuous intervals) so they cover the entire range.
  • Tally observations into classes to get frequency (f).
  • Calculate cumulative frequency, relative frequency, class midpoints, etc. for analysis and graphs.

Key terms

  • Class limits: lower and upper bounds of an interval.
  • Class boundaries: exact cut-points between classes (useful for continuous data).
  • Class mark (midpoint) = (lower limit + upper limit)/2.
  • Cumulative frequency: running total of frequencies up to a class.
  • Relative frequency = frequency / total observations.
  • Frequency density = frequency / class width (used when class widths differ).

Important notes on presentation: For continuous quantitative data, use a histogram (adjacent bars). For discrete or categorical data, use bar diagrams (bars separated). Use frequency polygon (join class midpoints) to compare distributions. Use ogive (cumulative frequency curve) to read medians and percentiles.

📌 Examples
  • Worked example (marks of 20 students): Data sorted: 37,45,45,47,49,54,56,59,60,61,66,67,72,73,78,80,82,85,88,91. Range = 91 − 37 = 54. Using Sturges' rule: k ≈ 1 + 3.322 log10(20) ≈ 6 classes. Class width ≈ 54/6 = 9 ⇒ choose 10. Classes and frequencies: 36–45: 3; 46–55: 3; 56–65: 4; 66–75: 4; 76–85: 4; 86–95: 2. Class midpoints: 40.5, 50.5, 60.5, 70.5, 80.5, 90.5. Cumulative frequencies: 3, 6, 10, 14, 18, 20. Relative frequencies: 0.15, 0.15, 0.20, 0.20, 0.20, 0.10. Percent frequencies: 15%, 15%, 20%, 20%, 20%, 10%.
  • Real-life example 1: Daily sales (in units) of a shop for 30 days can be grouped into classes (0–9, 10–19, 20–29, ...) to find how many days had low, medium or high sales and to build a histogram for visual analysis.
  • Real-life example 2: Age distribution of a town's population grouped into 0–9, 10–19, 20–29, ... helps planners see which age groups need schools, jobs or elderly care.
  • Real-life example 3: Household income grouped into income brackets to study inequality; class widths may differ, so use frequency density (f/width) when drawing histograms.
🧮 Formulas
  1. \[Range = Maximum − Minimum\]
  2. \[Sturges' rule (approx. number of classes): k ≈ 1 + 3.322 × log10(n)\]
  3. \[Class width (h) ≈ Range / k (round up to a convenient number)\]
  4. \[Class midpoint (class mark) xm = (Lower limit + Upper limit) / 2\]
  5. \[Cumulative frequency (CF) for class j = sum of frequencies of all classes up to j\]
  6. \[Relative frequency = f / n\]
📈6

Presentation of Data: Diagrams and Graphs

📊 COMMERCE / ECONOMIC LAW

Presentation of Data: Diagrams and Graphs

Key Point: Class width (h) = (Maximum value − Minimum value) / Number of classes (rounded as appropriate)

What it is: Presentation of data means arranging and representing numerical or categorical information in pictorial form so that patterns, trends and comparisons become easy to understand. Diagrams and graphs convert raw data (tables and frequency distributions) into visual forms.

Why it matters: Visual representations make complex data accessible, help comparison, highlight trends, and support decision-making in real life (business, economics, public policy, education).

Main types:

  • Qualitative / Categorical diagrams: Bar diagram, multiple bar diagram, component bar diagram, pictogram and pie chart — used for categorical comparisons (e.g., sectoral contribution to GDP, market share).
  • Quantitative / Frequency diagrams: Histogram, frequency polygon, ogive (cumulative frequency curve), and line graph — used for grouped numerical data and to show distributions or trends (e.g., marks distribution, income classes, time series).

How to choose: Use bar/pictogram for comparing categories, pie chart for showing percentage composition, line graph for time-series/trends, histogram/frequency polygon/ogive for continuous grouped data distribution, and pictorial/stacked/component bars when parts-of-whole or multiple comparisons are needed.

Basic construction principles:

  • Label axes and units clearly; give a descriptive title.
  • Choose an appropriate scale so the plot fills the space and is easy to read.
  • For grouped data, use class boundaries/widths correctly; for unequal class widths use frequency density in histograms.
  • Maintain proportionality — areas/heights/angles must reflect data values (e.g., pie-chart angles proportional to frequencies).

Advantages: Quick visual comparison, identification of central tendency and dispersion patterns, easy presentation to non-technical audiences.

Limitations: Can hide details or distort data if scales or areas are chosen improperly; not all kinds of analysis (e.g., exact values, hypothesis testing) can be done from a graph alone.

📌 Examples
  • Marks distribution of a classroom: Use a histogram or frequency polygon to show how many students scored in intervals (0–10, 11–20, ...).
  • Monthly household expenditure: Display as a pie chart to show percentage spent on food, rent, education, transport and savings.
  • Price trend of petrol over a year: Use a line graph (time on x-axis, price on y-axis) to show rise and fall across months.
  • Comparing production of different industries in a state: Use a bar diagram (each industry as a bar) or a stacked bar for contributions over years.
  • Population age distribution: Use a histogram (or age-sex pyramid) to show number of people in age groups; use ogive to find median age or percentiles.
🧮 Formulas
  1. \[Class width (h) = (Maximum value − Minimum value) / Number of classes (rounded as appropriate)\]
  2. \[Class midpoint (xi) = (Lower class boundary + Upper class boundary) / 2\]
  3. \[Relative frequency (ri) = fi / N (where fi = frequency of class i\]
    \[N = total frequency)\]
  4. \[Percentage of class = (fi / N) × 100\]
  5. \[Angle for pie chart (in degrees) = (fi / N) × 360\]
  6. \[Cumulative frequency (CFk) = Σ fi (sum of frequencies up to the k-th class)\]
📈7

Key Terms and Short Concepts

📊 COMMERCE / ECONOMIC LAW

Key Terms and Short Concepts

Key Point: Class mark (midpoint): x = (Lower limit + Upper limit) / 2

Overview
This topic collects short definitions and bite-sized concepts used in the chapter "Organisation of Data". It explains types of data, methods to organise data, basic components of frequency distributions and the common graphical forms used to present data.

Basic terms

  • Data — Raw facts and figures collected for analysis (numbers, attributes, observations).
  • Variable — A characteristic that can take different values (e.g., height, income).
  • Attribute — Qualitative characteristic (e.g., gender, colour).
  • Qualitative (categorical) — Non‑numeric data (names, categories).
  • Quantitative — Numeric data; can be discrete (countable values) or continuous (measurable on a continuum).
  • Univariate / Bivariate / Multivariate — Number of variables involved (one, two, many).

Sources of data

  • Primary data — Collected firsthand by the investigator (surveys, experiments, observations).
  • Secondary data — Collected earlier by someone else (census reports, published statistics).
  • Census — Complete enumeration of all units in a population.
  • Sample — Subset of the population used when census is impractical.

Sampling methods (short concept)

  • Random (simple random) — Every unit has equal chance.
  • Stratified — Population divided into strata; samples taken from each strata (improves representativeness).
  • Systematic — Every k-th unit chosen (after random start).
  • Cluster — Population divided into clusters; a few clusters chosen and all units in them surveyed.
  • Quota — Interviewers collect data until quotas for subgroups are met (non‑probability).

Frequency distribution — key components

  • Class interval — A range of values grouped together (e.g., 10–14).
  • Class limits — Lower and upper boundary values of a class (inclusive/exclusive depending on convention).
  • Class boundaries — Exact dividing lines between classes used to avoid gaps (useful for continuous data).
  • Class width (size) — Difference between upper and lower limits of a class.
  • Class mark (mid-point) — (Lower limit + Upper limit) / 2; used as representative value for the class.
  • Frequency (f) — Number of observations in a class.
  • Cumulative frequency (CF) — Running total of frequencies up to (≤) a class.
  • Relative frequency — f / N (proportion of total).

Tabulation principles (short)

  • Have a clear title and source.
  • Use appropriate class widths; classes should be mutually exclusive and exhaustive.
  • Arrange classes in ascending or descending order.
  • Include totals and cumulative columns if needed.

Graphical presentation — quick notes

  • Bar chart — For categorical (qualitative) data; bars separated.
  • Histogram — For continuous grouped data; contiguous bars with area proportional to frequency.
  • Frequency polygon — Join class‑marks at heights equal to frequencies by straight lines; useful for comparing distributions.
  • Ogive (cumulative frequency curve) — Plots class upper (or lower) boundaries against cumulative frequency; useful for finding medians/percentiles graphically.
  • Pie chart — Shows composition (percentages) of a whole; use for a few categories only.

Short concept tips

  • Continuous data needs contiguous classes (no gaps); discrete data can have single‑value classes.
  • Open-ended classes (like "60 and above") are allowed but require caution for mid‑point or mean estimation.
  • Choose class width so you have between about 5 and 15 classes (practical rule) depending on sample size.
📌 Examples
  • Census vs Sample: The decennial population census is an example of complete enumeration; an NSSO household expenditure survey uses sampling to estimate averages without surveying every household.
  • Stratified sampling: To estimate average marks in a school, divide students by class (grades) and take random samples from each class so each grade is proportionately represented.
  • Histogram and frequency polygon: Given heights of 200 students grouped into classes 140–144, 145–149, ..., draw contiguous bars for the histogram and plot mid‑points (142, 147, ...) joined by lines for the frequency polygon.
  • Ogive: To find the median age from grouped data of employees, plot cumulative frequency against upper class boundaries and read the value corresponding to N/2 on the vertical axis.
  • Pie chart: A company’s market share (A: 40%, B: 25%, C: 20%, D: 15%) is best shown by a pie chart to highlight percentage composition.
🧮 Formulas
  1. \[Class mark (midpoint): x = (Lower limit + Upper limit) / 2\]
  2. \[Class width (h): h = Upper limit − Lower limit (or next class lower − this class lower)\]
  3. \[Relative frequency: r = f / N (where f = frequency\]
    \[N = total observations)\]
  4. \[Percentage frequency: %f = (f / N) × 100\]
  5. \[Cumulative frequency (CF) for k-th class: CF_k = Σ_{i=1 to k} f_i\]
  6. \[Estimated mean (grouped data\]
    \[using class marks): x̄ = (Σ f_i x_i) / N (x_i = class mark)\]

Key Concepts

Data
Facts, numbers or information collected for analysis.
Population
The complete set of items or individuals under study.
Sample
A part or subset of the population selected for study.
Primary data
Data collected firsthand by the investigator for a specific purpose.
Secondary data
Data collected earlier by someone else and used by the investigator.
Qualitative data (Attributes)
Non-numeric information describing qualities or categories.
Quantitative data (Variables)
Numeric information that can be measured or counted.
Discrete data
Quantitative data that take distinct, separate values (often counts).
Continuous data
Quantitative data that can take any value within a range.
Univariate data
Data on a single characteristic or variable for each observation.
Bivariate data
Data on two variables observed simultaneously on the same units.
Time-series data
Data collected at successive points or periods of time.
Cross-sectional data
Data collected on many subjects at the same point in time.
Frequency distribution
A table showing classes or categories and their corresponding frequencies.
Classification
Arranging data into groups or classes based on common characteristics.
Tabulation
Systematic arrangement of data in rows and columns for clarity.
Cumulative frequency
Running total of frequencies up to a particular class or value.
Class interval
A range of values grouped together in a frequency distribution.
Histogram
A graphical representation of a frequency distribution using adjoining bars for classes.
Frequency polygon
A line graph formed by joining midpoints of class intervals plotted against frequencies.

Practice Questions

  1. What is meant by classification of data and why is it necessary? / आँकड़ों के वर्गीकरण से क्या अभिप्राय है और यह क्यों आवश्यक है?
    Show answer

    Classification is the process of arranging raw data into groups or classes having similar characteristics. It is necessary because raw data are voluminous and unstructured; classification simplifies them, highlights patterns, facilitates comparison and prepares data for presentation and analysis. / वर्गीकरण कच्चे आँकड़ों को समान विशेषताओं वाले समूहों या वर्गों में व्यवस्थित करने की प्रक्रिया है। यह आवश्यक है क्योंकि कच्चे आँकड़े बड़ी मात्रा में व असंरचित होते हैं; वर्गीकरण उन्हें सरल बनाता है, प्रतिमान उजागर करता है, तुलना सुगम करता है और प्रस्तुति व विश्लेषण हेतु तैयार करता है।

  2. Distinguish between discrete and continuous data with one example each. / असतत और सतत आँकड़ों में एक-एक उदाहरण सहित अंतर कीजिए।
    Show answer

    Discrete data take distinct, countable values (usually integers), e.g. the number of factories in a district. Continuous data can take any value within a range, e.g. the height of students measured in centimetres. / असतत आँकड़े पृथक, गणनीय मान (प्राय: पूर्णांक) लेते हैं, जैसे किसी ज़िले में कारखानों की संख्या। सतत आँकड़े किसी परास में कोई भी मान ले सकते हैं, जैसे सेंटीमीटर में मापी गई विद्यार्थियों की लंबाई।

  3. Name and explain any four components of a statistical table. / सांख्यिकीय सारणी के कोई चार घटकों का नाम देकर समझाइए।
    Show answer

    Title gives a concise description of the table's contents; column headings (captions) describe the variables in columns; stubs are the row labels on the left; and the body holds the actual data values in the cells. (A source/footnote may be added for clarifications.) / शीर्षक सारणी की विषयवस्तु का संक्षिप्त विवरण देता है; स्तंभ शीर्षक स्तंभों के चरों का वर्णन करते हैं; स्टब बायीं ओर पंक्ति लेबल होते हैं; और मुख्य भाग कोष्ठकों में वास्तविक आँकड़ा मान रखता है। (स्पष्टीकरण हेतु स्रोत/पादटिप्पणी जोड़ी जा सकती है।)

  4. For the class interval 20–29, find the class mark (mid-point) and class width. / वर्ग अंतराल 20–29 के लिए वर्ग चिह्न (मध्य-बिंदु) और वर्ग चौड़ाई ज्ञात कीजिए।
    Show answer

    Class mark = (Lower limit + Upper limit)/2 = (20 + 29)/2 = 24.5. For an inclusive class, class width = Upper limit − Lower limit + 1 = 29 − 20 + 1 = 10. / वर्ग चिह्न = (निम्न सीमा + उच्च सीमा)/2 = (20 + 29)/2 = 24.5। समावेशी वर्ग के लिए, वर्ग चौड़ाई = उच्च सीमा − निम्न सीमा + 1 = 29 − 20 + 1 = 10।

  5. Using Sturges' rule, determine the suggested number of classes for n = 100 observations. / स्टर्जेस नियम का उपयोग करते हुए n = 100 प्रेक्षणों के लिए सुझाई गई वर्ग संख्या ज्ञात कीजिए।
    Show answer

    Sturges' rule: k ≈ 1 + 3.322 log10(n) = 1 + 3.322 × log10(100) = 1 + 3.322 × 2 = 1 + 6.644 = 7.644 ≈ 8 classes. / स्टर्जेस नियम: k ≈ 1 + 3.322 log10(n) = 1 + 3.322 × log10(100) = 1 + 3.322 × 2 = 1 + 6.644 = 7.644 ≈ 8 वर्ग।

  6. Differentiate between a histogram and a bar diagram. / आयतचित्र और दंड आरेख में अंतर कीजिए।
    Show answer

    A histogram is used for continuous grouped data with contiguous bars (no gaps) where area represents frequency, drawn on class intervals. A bar diagram is used for qualitative/discrete data with separated bars of equal width whose heights represent frequencies or values. / आयतचित्र सतत समूहित आँकड़ों के लिए प्रयोग होता है जिसमें संलग्न दंड (बिना अंतराल) होते हैं जहाँ क्षेत्रफल बारंबारता दर्शाता है, वर्ग अंतरालों पर बनाया जाता है। दंड आरेख गुणात्मक/असतत आँकड़ों के लिए प्रयोग होता है जिसमें समान चौड़ाई वाले पृथक दंड होते हैं जिनकी ऊँचाई बारंबारता या मान दर्शाती है।

  7. In a frequency distribution the classes are 0–10, 10–20, 20–30 with frequencies 5, 8, 7. Find the cumulative frequency up to class 20–30 and the relative frequency of class 10–20. / एक बारंबारता वितरण में वर्ग 0–10, 10–20, 20–30 हैं जिनकी बारंबारताएँ 5, 8, 7 हैं। वर्ग 20–30 तक संचयी बारंबारता और वर्ग 10–20 की सापेक्ष बारंबारता ज्ञात कीजिए।
    Show answer

    Total N = 5+8+7 = 20. Cumulative frequency up to 20–30 = 5+8+7 = 20. Relative frequency of 10–20 = f/N = 8/20 = 0.40 (or 40%). / कुल N = 5+8+7 = 20। 20–30 तक संचयी बारंबारता = 5+8+7 = 20। 10–20 की सापेक्ष बारंबारता = f/N = 8/20 = 0.40 (या 40%)।

  8. What is an ogive and what is it used to find? / तोरण (ओजाइव) क्या है और इसका उपयोग किस चीज़ को ज्ञात करने में होता है?
    Show answer

    An ogive is a cumulative frequency curve plotted by taking class boundaries on the x-axis and cumulative frequencies on the y-axis. It is used to read the median, quartiles and percentiles of grouped data graphically. / तोरण एक संचयी बारंबारता वक्र है जिसे x-अक्ष पर वर्ग सीमाएँ और y-अक्ष पर संचयी बारंबारताएँ लेकर खींचा जाता है। इसका उपयोग समूहित आँकड़ों की मध्यिका, चतुर्थक व शतमक को आलेखीय रूप से पढ़ने में होता है।

Related Laws & Principles

Explore all

Foundational laws & principles connected to this chapter — tap to open in the Laws Explorer.

Loading related laws…
Sourced from 117 content files · LLOS Learn · browse all chapters