Overview
This chapter introduces the systematic process of collecting statistical data for economic analysis. It explains why reliable data are essential for understanding economic issues, making policies and conducting empirical research. The chapter distinguishes between primary and secondary data, and between census and sample methods, then presents common methods of collecting primary data (questionnaire, interview, observation, schedules). It outlines sampling techniques — probability methods (simple random, stratified, systematic, cluster) and non-probability methods (purposive, convenience, quota) — with their advantages and limitations. Practical aspects such as rules for designing questionnaires, choosing sample size, and common sources of error (sampling and non-sampling errors) are discussed, along with precautions to improve data quality. By the end, students will be able to choose appropriate data sources and collection methods, design basic instruments for primary data, and recognise and reduce errors in data collection for economic studies.
Learning Objectives
- Define primary and secondary data with suitable examples
- Explain the difference between census and sample survey and state their advantages and limitations
- Distinguish between probability and non-probability sampling methods with examples
- Classify common methods of data collection (survey, observation, experiment, and measurement) and indicate when each is appropriate
- Describe the characteristics of a good questionnaire and schedule
- Apply simple random, systematic, and stratified sampling techniques to select samples and compute sampling fraction
- Construct a valid questionnaire or interview schedule for a given survey topic
- Interpret sampling and non-sampling errors and suggest practical ways to minimize them
Topics in this chapter
14 topics · tap a topic title to jump straight to it.
Meaning and Importance of Data
Meaning and Importance of Data
Key Point: Frequency (f): count of observations in a class or category.
Meaning of Data: Data are raw facts, figures or observations collected for analysis and decision-making. In economics, data can be numerical (quantitative) — e.g., income, price, production — or non‑numerical (qualitative) — e.g., occupation, preference, type of enterprise. Data become information after processing (sorting, summarising, analysing).
Types (brief):
- Qualitative (categorical) vs Quantitative (discrete, continuous)
- Primary (collected first‑hand by surveys, experiments) vs Secondary (existing sources: reports, government records)
- Cross‑sectional (one time) vs Time‑series (over time)
Characteristics of good data: accuracy, relevance, completeness, consistency, timeliness, reliability and comparability.
Importance of Data in Economics: Data are the foundation of economic analysis and policy. Key uses include:
- Measurement and description — e.g., GDP, inflation, unemployment rates that describe economic performance.
- Planning and policy formulation — governments and firms use data to set targets and design interventions (budget, subsidies, monetary policy).
- Forecasting and trend analysis — time‑series data help predict future output, demand, prices.
- Evaluation and monitoring — assess effects of schemes, track progress, detect deviations.
- Research and hypothesis testing — data permit testing theories, estimating relationships (supply, demand elasticities).
- Resource allocation and decision‑making — firms use market data to price, produce, and allocate resources efficiently.
Limitations and cautions: Poor quality, bias in collection, inadequate coverage, measurement errors, outdated data and misinterpretation can lead to wrong conclusions. Always check source, method, and representativeness.
How economists turn data into information: collect (surveys, administrative records), tabulate (frequency distributions), summarise (averages, dispersion), visualise (charts), and analyse (correlations, regressions). Proper sampling, clear definitions and careful measurement are essential.
- Household income survey used to design targeted welfare programmes (primary, cross‑sectional).
- Consumer Price Index (CPI) — monthly price data used to measure inflation and adjust pensions/wages.
- National accounts (GDP) compiled from production, income and expenditure data for macroeconomic policy.
- Labour force survey reporting unemployment rate used by government to design employment schemes.
- Market research data on consumer preferences guiding a firm’s product launch and pricing.
- Crop yield records over years (time series) used by agricultural planners to predict food supply and procurement.
- \[Frequency (f): count of observations in a class or category.\]
- \[Relative frequency: rf = f / N (where N = total observations).\]
- \[Percentage: % = (f / N) × 100.\]
- \[Arithmetic mean (ungrouped): x̄ = Σx_i / N.\]
- \[Weighted mean: x̄_w = Σ(w_i · x_i) / Σw_i (useful for index numbers\]\[CPI).\]
- \[Median position (ungrouped): (N + 1) / 2th observation\]\[for grouped data: Median = L + [(N/2 − cfb) / fm] × h\]\[where L = lower boundary of median class\]\[cfb = cumulative frequency before median class\]\[fm = frequency of median class\]\[h = class width.\]
Types of Data
Types of Data
Key Point: Frequency (f_i): count of observations in class i
Overview
Data are facts or figures collected for analysis. In economics, data are classified in several ways depending on source, nature, measurement and time dimension. Understanding types of data helps choose appropriate collection methods, presentation and analysis tools.
Classification
- By Source
Primary data: Collected first-hand for a specific purpose (surveys, experiments, interviews).
Secondary data: Collected earlier by someone else for another purpose (government publications, research papers, databases). - By Nature
Qualitative (categorical): Non-numeric categories (gender, occupation, industry).
Quantitative: Numeric measurements. Can be discrete (countable values like number of firms) or continuous (measurable quantities like income, weight). - By Time Dimension
Cross-sectional data: Observations on many units at a single point in time (household incomes in 2025).
Time-series data: Observations on a single unit over several time periods (monthly CPI for 10 years).
Panel (longitudinal) data: Observations on multiple units over multiple time periods (same households surveyed every year). - By Coverage / Method
Census: Data collected from all units of the population.
Sample: Data collected from a subset (sample) of the population. - By Level of Aggregation
Individual (unit) data: Values for each unit (household income by household).
Aggregate data: Summarised values for groups or totals (national GDP).
Key characteristics & implications
- Qualitative data are best presented with frequency tables, bar charts or pie charts.
- Discrete quantitative data can be shown with bar charts or frequency tables.
- Continuous data are grouped into class intervals and shown with histograms, ogives or frequency polygons.
- Time-series require line graphs and special treatments (trend, seasonality).
- Source matters for reliability: primary data allow control over collection but cost more; secondary data are cheaper but may not fit the specific purpose.
- Primary data: Interviewing 200 households to record monthly expenditure on food.
- Secondary data: Using National Statistical Office reports for annual GDP figures.
- Qualitative data: Classifying workers by industry (agriculture, manufacturing, services).
- Discrete quantitative: Number of factories in a district (0,1,2,...).
- Continuous quantitative: Measuring annual income of households (in rupees).
- Cross-sectional: A survey measuring students' marks across schools in 2024.
- \[Frequency (f_i): count of observations in class i\]
- \[Relative frequency: r_i = f_i / N where N is total observations\]
- \[Percentage frequency: p_i = (f_i / N) × 100\]
- \[Cumulative frequency for class k: CF_k = sum of f_i up to class k\]
- \[Class width (for grouped continuous data): width = (range) / (number of classes) where range = max - min\]
- \[Arithmetic mean (sample): x̄ = (Σ x_i) / n where x_i are observations and n is sample size\]
Sources of Primary Data
Sources of Primary Data
Key Point: Sample mean: x̄ = Σxi / n (sum of observations divided by sample size).
Definition: Primary data are data collected firsthand by the investigator for a specific purpose or study. They are original, current and collected directly from sources.
Importance: Primary data are more reliable for the specific objective of the study because the investigator controls how, when and from whom the information is collected. They help answer precise research questions and allow control over measurement quality.
Main sources / methods of collecting primary data:
- Personal Interview (Face-to-face): Direct questioning of respondents by the investigator. Useful for detailed information and clarifying doubts. (Pros: high response quality; Cons: costly and time-consuming.)
- Telephone Interview: Respondents are interviewed over phone. Faster and cheaper than face-to-face but may have shorter responses and sampling bias.
- Questionnaire (Self-administered): A structured set of written questions given to respondents to fill. Good for large samples and standardization. (Watch for low response rates and misunderstanding of questions.)
- Schedule (Filled by Investigator): Investigator reads questions and records answers on a schedule — useful when respondents are illiterate or to ensure consistency.
- Observation Method: Investigator records behaviour, events or conditions directly (participant or non-participant observation). Best for actions that respondents may not report accurately. (Cons: observer bias, limited to observable behaviour.)
- Experiment: Controlled manipulation of one or more variables to observe effects (lab or field experiments). Common in behavioural and applied economic studies.
- Case Study: Intensive study of a single unit (individual, firm, village) to obtain in-depth primary information. Good for generating hypotheses and detailed insights.
- Focus Group / Group Interview: Guided discussion with a small group to explore attitudes and reasons behind choices. Useful in market research.
- Census and Sample Surveys: Census collects primary data from the entire population; sample survey collects from a representative subset. Choice depends on objectives, resources and required precision.
- Pilot Survey: A small-scale trial of the main survey to test instruments and procedures — a preparatory primary data source to improve design.
Quality issues & limitations: Primary data collection must guard against sampling bias, non-response bias, measurement error, interviewer bias and ethical issues (consent, confidentiality). Costs, time and feasibility also limit scope.
Practical steps when using primary data:
- Clearly define objectives and target population.
- Choose an appropriate method (questionnaire, observation, experiment, etc.).
- Design instruments (clear questions, pilot test).
- Decide sampling method and sample size.
- Collect data systematically and ethically.
- Clean, code and analyze data using appropriate statistical methods.
- National Population Census: Enumerators (investigators) collect basic demographic and socio-economic information directly from households ― a primary data collection on the entire population.
- Market research for a new soft drink: A company conducts a sample survey using questionnaires and taste tests (field experiment) to collect consumers' preferences and willingness to pay.
- Traffic study by a municipal body: Observers record vehicle counts and types at intersections (non-participant observation) to plan road improvements.
- School performance study: Investigators visit schools and interview teachers and students (schedules and interviews) to collect attendance and achievement data.
- Medical trial: Researchers conduct a randomized controlled trial (experiment) to test the effectiveness of a new drug; data on outcomes are primary.
- NGO household income survey: Field investigators administer questionnaires to collect up-to-date income, consumption and employment information from sampled households.
- \[Sample mean: x̄ = Σxi / n (sum of observations divided by sample size).\]
- \[Sample proportion: p̂ = x / n (x = number of successes\]\[n = sample size).\]
- \[Sample variance: s² = Σ(xi - x̄)² / (n - 1) (unbiased estimator for variance).\]
- \[Standard error of mean: SE(x̄) = s / √n (s = sample standard deviation).\]
- \[Required sample size for estimating a mean (approx.): n = (Z * σ / E)² where Z = z-value for confidence level, σ = estimated population sd\]\[E = desired margin of error.\]
- \[Required sample size for a proportion: n = (Z² * p * (1 - p)) / E² (use p = 0.5 if unknown for maximum variability).\]
Sources of Secondary Data
Sources of Secondary Data
Key Point: Percentage (share) = (Part / Whole) × 100
Definition: Secondary data are data that have already been collected, processed and published by someone else for some purpose other than the present research. They are reused by researchers, students, businesses and policymakers.
Why use secondary data? They save time and cost, provide access to large samples and long time series, and allow comparison across regions and periods. However, they may not fit the exact needs of a study and their quality must be checked.
Main sources of secondary data
- Government publications and agencies – Census of India, National Sample Survey/Office of the Registrar General, National Statistical Office (NSO), Reserve Bank of India (RBI), Ministry reports, Statistical Abstracts and Yearbooks. These provide authoritative demographic, economic and social statistics.
- International and intergovernmental organizations – World Bank, IMF, United Nations, WHO, UNESCO, ILO. Useful for cross-country comparisons and global indicators (GDP, health, education).
- Administrative and institutional records – School records, hospital patient registers, tax records, voter lists, police records. Regularly produced and useful for local-level analysis.
- Commercial and corporate sources – Company annual reports, financial statements, market research firms, trade associations and industry reports. Used for business studies and financial analysis.
- Published media and periodicals – Newspapers, magazines, academic journals, books, newsletters. Useful for current events, opinions and qualitative context.
- Online databases and digital repositories – Government open-data portals, institutional repositories, online statistical databases (e.g., data.gov.in, World Bank Data), scholarly archives.
- Research reports and working papers – University research, think-tank reports, consultancy studies and NGO publications containing survey results or analysis.
- Historical archives and libraries – Archival documents, past records, historical datasets that support longitudinal or historical studies.
Assessing secondary data quality
- Relevance: Does the data measure what you need (variables, period, geographic coverage)?
- Accuracy and reliability: Who collected it? What methods and sampling were used?
- Timeliness: Is the data up-to-date for your purpose?
- Consistency and comparability: Are definitions and units consistent across time/regions?
- Accessibility and permissions: Is the data publicly available or restricted?
Typical uses in Class 11 Economics – Use Census and NSS/NSO data for population, labour force and consumption patterns; RBI and Ministry of Finance data for macro indicators; company annual reports for business case studies; newspaper statistics for current economic events.
- Using Census data to find the rural–urban population ratio in a state for 2011 and 2021.
- Referencing RBI monthly bulletins to analyze trends in money supply (M1/M2) over five years.
- Using NSS consumption-expenditure reports to estimate average household expenditure on food items.
- Studying a company’s annual report to compute profitability ratios (e.g., net profit margin) for case study analysis.
- Using WHO and World Bank datasets to compare infant mortality rates across countries.
- Using school enrolment registers to study changes in drop-out rates at a local school over time.
- \[Percentage (share) = (Part / Whole) × 100\]
- \[Percentage change = [(New value - Old value) / Old value] × 100\]
- \[Simple growth rate (annual) = [(Value in year t - Value in year t-1) / Value in year t-1] × 100\]
- \[Compound Annual Growth Rate (CAGR) = [(Ending value / Beginning value)^(1 / n) - 1] × 100\]\[where n = number of years\]
- \[Arithmetic mean (for a data series) = (Σxi) / n\]
- \[Weighted mean = (Σwi·xi) / (Σwi)\]
Census vs Sample Survey
Census vs Sample Survey
Key Point: Population mean (census): μ = (Σ Xi) / N, where Xi are values for all N units.
Definition
A census is a complete enumeration in which information is collected from every unit of the population. A sample survey collects information from a subset (sample) of the population and uses it to make inferences about the whole population.
Key differences (concise)
- Coverage: Census = whole population; Sample survey = subset.
- Cost & Time: Census is usually costly and time-consuming; sample surveys are cheaper and quicker.
- Accuracy: Census can be more accurate if perfectly done (no sampling error) but may suffer from large non-sampling errors; sample surveys have sampling error but can be highly accurate with proper design.
- Feasibility: Census is feasible for small populations or when legal/constitutional requirements exist (e.g., national population census). Sample surveys are preferred for large or inaccessible populations, or when destructive testing is involved.
- Purpose: Census gives detailed disaggregated data; sample surveys are useful for estimates and hypothesis testing with limited resources.
When to prefer which?
- Use a census when the population is small, when exhaustive data are required, or when mandated (e.g., national census every 10 years).
- Use a sample survey when population is large, resources are limited, time is short, or when sampling is scientifically sufficient for the study objective.
Errors and quality issues
- Sampling error: Present only in sample surveys — difference between sample estimate and true population value due to chance.
- Non-sampling error: Includes measurement error, non-response, coverage error — can affect both census and surveys and often dominates total error.
Design steps for a sample survey (brief)
- Define the population and objectives.
- Construct a sampling frame (list of units).
- Choose a sampling method (simple random, stratified, systematic, cluster).
- Decide sample size and selection procedure.
- Collect data and compute estimates with measures of precision (standard errors, confidence intervals).
Practical notes
Well-designed sample surveys (like NSSO surveys, National Family Health Survey (NFHS), market research polls) provide reliable and timely information at much lower cost than a census. A census (e.g., Census of India) provides highly detailed baseline counts but is undertaken infrequently.
- Census of India: a nationwide complete enumeration of the population carried out every 10 years — collects demographic, social and housing data from every household.
- National Sample Survey (NSS): selects representative samples of households to estimate employment, consumption, and other socio-economic indicators — quicker and less costly than a census.
- School example: measuring the height of every student in a class (census) vs measuring heights of 20 randomly chosen students to estimate the class average (sample survey).
- Manufacturing quality control: testing every bulb for life (census) is destructive and costly, so manufacturers test a sample of bulbs to estimate defect rate (sample survey).
- Election opinion polls: use sample surveys (random or stratified samples) to estimate vote intentions instead of asking every voter.
- \[Population mean (census): μ = (Σ Xi) / N\]\[where Xi are values for all N units.\]
- \[Sample mean: x̄ = (Σ xi) / n\]\[where xi are values in the sample of size n.\]
- \[Population proportion (census): P = X / N (X = count of units with attribute).\]
- \[Sample proportion: p̂ = x / n (x = count in sample with attribute).\]
- \[Standard error of sample mean (known population SD σ): SE(x̄) = σ / √n.\]
- \[Estimated SE when population SD unknown: SE(x̄) ≈ s / √n\]\[where s is sample standard deviation.\]
Basic Sampling Concepts
Basic Sampling Concepts
Key Point: Sample mean: x̄ = (Σ xi) / n — average of sample observations.
What is sampling? Sampling is the process of selecting a part of a population to infer characteristics of the whole population. The selected part is called a sample and the whole group is called the population (or universe).
Key terms
- Population (N): Complete set of units of interest (e.g., all students in a school).
- Sample (n): A subset of the population chosen for study.
- Sampling unit: A single element or group considered for selection (e.g., one household).
- Sampling frame: A list or source from which the sample is drawn (e.g., voter list).
- Parameter: A numerical characteristic of the population (e.g., population mean μ).
- Statistic: A numerical characteristic computed from the sample (e.g., sample mean x̄).
- Sampling fraction: n/N, the portion of population included in the sample.
Why sample? Sampling is used because it is usually cheaper, faster and sometimes more practical than studying the entire population.
Types of sampling
- Probability (random) sampling — each unit has a known non-zero chance of selection. Main methods:
- Simple Random Sampling (SRS): Every possible sample of size n has equal probability. Can be with or without replacement.
- Systematic Sampling: Choose every kth unit from an ordered list after a random start (k = N/n).
- Stratified Sampling: Population divided into homogeneous strata (e.g., gender, region); take random samples from each stratum. Often yields more precise estimates when strata differ.
- Cluster Sampling: Population divided into clusters (e.g., villages), some clusters randomly selected and all or a sample of units within selected clusters surveyed. Useful for geographically spread populations.
- Non-probability sampling — selection not random; selection probability unknown. Main types:
- Convenience sampling: Select easily available units (e.g., shoppers at a mall).
- Judgment (purposive) sampling: Expert selects units believed to be typical.
- Quota sampling: Choose units to fill predefined quotas for certain groups (not random within quotas).
Errors in sampling
- Sampling error: Difference between sample estimate and true population value due to observing only part of population. It decreases as sample size increases (all else equal).
- Non-sampling errors: Errors not due to sampling — measurement error, non-response, response bias, coverage error (when sampling frame misses parts of population). These may be larger and harder to control than sampling error.
Steps in designing a sample survey
- Define the population and objectives.
- Choose the sampling frame.
- Select sampling method and determine sample size.
- Collect data and compute statistics.
- Estimate sampling error and draw conclusions.
When to prefer which method? Use SRS or stratified sampling for accuracy when a good frame exists. Use cluster sampling to save cost when population is widely spread. Use non-probability methods only for exploratory work or when random selection is impossible.
Advantages of sampling: cost-effective, faster, less resource intensive, practical for destructive testing. Limitations: potential for sampling error, bias if selection is not proper, results depend on sample design.
- Simple random sampling: From a school roll of 600 students, use random numbers to pick 60 students to estimate average study hours.
- Systematic sampling: From a list of 2,000 households, select every 20th household after a random start to survey electricity usage.
- Stratified sampling: To estimate average income in a city, divide population into income strata (low/middle/high) and randomly sample within each stratum proportional to its size.
- Cluster sampling: For a nationwide health survey, randomly select 50 villages (clusters) and survey all households in those villages (or a random sample within each selected village).
- Convenience sampling: Asking passengers at one railway station about travel habits — quick but not representative of all travellers.
- \[Sample mean: x̄ = (Σ xi) / n — average of sample observations.\]
- \[Sample proportion: p̂ = x / n — where x is number of successes in sample.\]
- \[Sample variance: s^2 = Σ(xi - x̄)^2 / (n - 1) — measure of spread in sample.\]
- \[Standard error of sample mean (approx.): SE(x̄) = σ / √n (if population SD σ known)\]\[otherwise use s / √n.\]
- \[Sampling fraction: f = n / N — proportion of population sampled.\]
- \[Number of possible samples in SRS without replacement: C(N\]\[n) = N! / (n!(N - n)!).\]
Probability (Random) Sampling Methods
Probability (Random) Sampling Methods
Key Point: Population size N, sample size n, finite population correction (FPC) = (1 - n/N).
Definition: Probability (random) sampling methods are techniques of selecting a sample from a population in such a way that every unit of the population has a known, non‑zero probability of being included. These methods allow the use of probability theory to make inferences about the population and to estimate sampling errors.
Key components: population, sampling frame (list of population units), sampling unit, sample size (n), and selection method. Probability sampling ensures representativeness and allows calculation of standard errors.
Major probability (random) sampling methods
- Simple Random Sampling (SRS): Every possible sample of size n from N has equal probability. Two variants: with replacement (SRSWR) and without replacement (SRSWOR). SRS is often implemented by random number tables or computer-generated random numbers applied to a sampling frame.
- Systematic Sampling: Select a random start between 1 and k, then pick every k-th unit (k = N/n). Simple to apply when units are listed in order. Works well when the list has no periodic pattern related to the study variable.
- Stratified Random Sampling: Population is divided into non-overlapping subgroups called strata (e.g., gender, region, stream). Random samples are drawn from each stratum. Can be proportional (sample size in stratum proportional to stratum size) or disproportional. Useful when strata are internally homogeneous but different from each other, reducing overall variance.
- Cluster (Random) Sampling: Population is divided into clusters (usually based on geography or natural groupings). A random sample of clusters is selected; then either all units in chosen clusters are surveyed (one-stage cluster sampling) or a random sample of units within chosen clusters is taken (two-stage). Clusters should ideally be mini-representations of the population; otherwise variance increases.
- Multistage Sampling: A combination of the above methods applied in stages (e.g., randomly select districts, then villages, then households). Common in large surveys like censuses and national sample surveys.
When to choose which method: SRS for small, well‑listed populations; systematic for simplicity and evenly spread samples; stratified when known subgroups differ strongly; cluster/multistage for large or geographically spread populations where listing all units is impractical.
Advantages of probability sampling: allows objective selection, supports estimation of sampling error, enables unbiased estimators.
Limitations: requires a sampling frame (which may be unavailable), sometimes costlier (especially SRS for large scattered populations), and cluster sampling can increase variance if clusters are homogeneous internally.
- Simple random sampling (SRSWOR): From a class of 60 students, use a random number table/app to pick 10 roll numbers and interview those students about study habits.
- Systematic sampling: To survey 200 houses on a long street of 2000 houses, choose every 10th house after a random start between 1 and 10.
- Stratified sampling: To estimate average marks in a school with 3 streams (Science, Commerce, Arts), divide students by stream and draw samples from each stream proportionally to the stream size.
- Cluster sampling: For a village health survey, randomly select 5 villages (clusters) out of 50 and interview all households in the selected villages (one-stage cluster sampling).
- Multistage sampling: For a national household survey, randomly select districts, then within selected districts pick villages, and within villages select households randomly.
- \[Population size N\]\[sample size n\]\[finite population correction (FPC) = (1 - n/N).\]
- \[Variance of sample mean under SRS without replacement (SRSWOR): Var(\bar{x}) = (1 - n/N) * (σ^2 / n)\]\[where σ^2 is population variance.\]
- \[Variance of sample mean under SRS with replacement (SRSWR): Var(\bar{x}) = σ^2 / n.\]
- \[Standard error (SE) of sample mean: SE(\bar{x}) = sqrt(Var(\bar{x})).\]
- \[For a proportion p̂ under SRSWOR: Var(p̂) = (1 - n/N) * p(1 - p) / n\]\[SE(p̂) = sqrt(Var(p̂)).\]
- \[Required sample size for estimating a mean (approximate\]\[large population): n ≈ (Z^2 * σ^2) / E^2\]\[where Z is z-score for desired confidence\]\[E is margin of error.\]
Non-probability (Non-random) Sampling Methods
Non-probability (Non-random) Sampling Methods
Key Point: Sample mean (used to summarize the collected sample): x̄ = (Σ xi) / n
Definition: Non-probability (non-random) sampling methods are techniques where members of the population do not have a known or equal chance of being selected. Selection depends on the researcher’s judgment, convenience, or respondents themselves rather than random mechanisms.
Key characteristics
- Selection is subjective or driven by accessibility rather than chance.
- Probabilities of inclusion are unknown, so sampling error and standard errors cannot be computed in the usual way.
- Quicker and less expensive than probability sampling, but results are less generalisable and more prone to bias.
Common types
- Convenience sampling: Choosing units that are easiest to reach (e.g., asking people on the street or students in a classroom). Useful for quick, exploratory studies.
- Purposive (Judgmental) sampling: Researcher selects units based on judgement about which units are most useful or representative for the study (e.g., selecting experienced teachers to evaluate a curriculum).
- Quota sampling: Population is divided into groups (quotas) and interviewer fills quotas non-randomly (e.g., ensure 40% young people, 60% older people but pick conveniently within groups).
- Snowball sampling: Existing study subjects recruit future subjects from among their acquaintances. Useful for hard-to-reach or hidden populations (e.g., drug users, migrant networks).
- Self-selection (Voluntary response): Individuals opt in to participate (e.g., online polls, call-in surveys). Often biased toward strong opinions.
Advantages
- Low cost and fast to implement.
- Useful for exploratory research, pilot studies, or when a sampling frame is unavailable.
- Practical for hard-to-reach populations (snowball, purposive).
Disadvantages
- High risk of selection bias and low external validity; results may not represent the population.
- No straightforward way to calculate sampling error or confidence intervals based on selection probabilities.
- May produce misleading estimates if treated as if they were random samples.
When to use
- Preliminary or exploratory studies where precision is not critical.
- When speed, cost, or lack of sampling frame makes probability sampling impractical.
- When studying rare or hidden populations where random sampling is impossible.
Notes on inference
Because inclusion probabilities are unknown, treat estimates from non-probability samples cautiously. If using results to suggest hypotheses or for descriptive summaries of the sample, they can be informative; for making population-wide inferences, probability sampling is preferable.
- Convenience: A researcher surveys students in their own college canteen to learn about daily snack preferences — quick but not representative of all students in the city.
- Purposive: To study the implementation of a new teaching method, the researcher interviews five highly experienced teachers chosen for their expertise.
- Quota: A market researcher ensures responses from exactly 50 men and 50 women by stopping collection once quotas are filled, but selects respondents conveniently within each quota.
- Snowball: To study support networks of recent immigrants, the researcher asks initial contacts to refer other immigrants they know, building the sample through referrals.
- Self-selection: An online retailer posts a customer satisfaction survey on its website and analyses responses from customers who choose to participate.
- \[Sample mean (used to summarize the collected sample): x̄ = (Σ xi) / n\]
- \[Sample proportion (for a characteristic in the sample): p̂ = x / n (x = number of sample units with the characteristic)\]
- \[Response rate (useful for surveys): Response rate (%) = (Number of completed responses / Number of eligible contacts) × 100\]
- \[Important note: Standard probability-based formulas for sampling error and confidence intervals (which require known selection probabilities) are not valid for non-probability samples without strong additional assumptions.\]
Designing Questionnaires and Schedules
Designing Questionnaires and Schedules
Key Point: Response rate (%) = (Number of questionnaires returned / Number of questionnaires distributed) × 100
Definition
A questionnaire is a set of written questions used to collect information directly from respondents. A schedule is similar but filled by the investigator (interviewer) after asking the respondent. Both are primary tools for primary data collection.
Purpose
To collect accurate, relevant and comparable information for research, surveys, evaluations or administrative records.
Types of questions
- Closed-ended: fixed choices (yes/no, multiple choice, rating scales). Easy to code and analyze.
- Open-ended: respondent writes answer in own words. Useful for detailed information but harder to analyze.
- Dichotomous: two alternatives (e.g., Yes/No).
- Likert scale: degree of agreement (e.g., Strongly agree → Strongly disagree).
- Ranking: respondents order items by preference.
Principles of a good questionnaire/schedule
- Clear objective: every question should serve the research objective.
- Simple language: use everyday words appropriate to respondents' education and culture.
- Avoid ambiguity: one idea per question; avoid double-barrelled questions.
- No leading or biased wording: avoid suggesting an answer.
- Mutually exclusive and exhaustive options: response categories should not overlap and should cover all likely answers (use 'Other, specify' where needed).
- Logical sequence and flow: start with easy, non-sensitive questions; group similar topics; use filter questions to direct respondents.
- Keep it short and relevant: long questionnaires reduce response rate and data quality.
- Pre-test (pilot): test with a small sample to find unclear questions, flow problems and timing issues.
- Clear instructions and layout: specify how to answer, mark multiple responses, skip patterns, and interviewer notes.
- Ensure confidentiality and consent: state purpose, use of data and assure privacy.
Design steps
- Define objectives and information needed.
- Decide question types (open/closed) and response formats.
- Create draft questions and order them logically (screening → core → background).
- Add instructions, coding spaces and skip patterns.
- Pre-test on a small group; revise based on feedback.
- Train interviewers (for schedules) and finalize the instrument.
- Collect data; conduct quality checks and code answers for analysis.
Common problems and how to avoid them
- Non-response: make questions brief, assure confidentiality, choose convenient timing.
- Misinterpretation: use simple wording, define technical terms, pre-test.
- Leading questions: rephrase to be neutral.
- Inadequate options: include 'Don't know' or 'Other' where appropriate.
Coding and tabulation
Before large-scale collection, assign numeric codes to likely responses (e.g., Male=1, Female=2; Yes=1, No=2). This speeds tabulation and analysis. For open-ended answers, develop a coding scheme after pre-testing.
Difference between Questionnaire and Schedule (short)
- Questionnaire: filled by respondent (self-administered).
- Schedule: filled by interviewer after asking questions (used when respondents are illiterate or when complex probing is needed).
Quality check
Check completeness, internal consistency (e.g., age vs. date of birth), logical skips, and unusual values immediately after collection so errors can be followed up.
Note: Designing good questionnaires and schedules is as much art as science: clarity, relevance and careful pre-testing are the keys to reliable primary data.
- School feedback form: A 10-question closed-ended questionnaire given to students to rate teaching quality (Likert scale), facilities, and suggestions (one open-ended question).
- Household expenditure schedule used by an investigator: interviewer reads questions on monthly food, fuel, education expenses and fills standardized categories—useful for consumption surveys.
- Customer satisfaction survey for a shop: short self-administered questionnaire with dichotomous 'Would you return? (Yes/No)', multiple-choice on reasons and a rating for overall experience.
- Public health camp schedule: interviewer records patient demographics, symptoms, diagnoses and prescribed treatment—structured to allow later tabulation.
- \[Response rate (%) = (Number of questionnaires returned / Number of questionnaires distributed) × 100\]
- \[Non-response rate (%) = 100 − Response rate (%)\]
- \[Sample size for estimating a proportion (large population): n = (Z^2 × p × q) / E^2\]\[where Z = z-score for confidence level (e.g., 1.96 for 95%)\]\[p = expected proportion\]\[q = 1−p\]\[E = allowable margin of error (decimal).\]
- \[Finite population correction (when population N is not large): n' = n / (1 + (n−1)/N)\]\[where n is sample size from formula above and N is population size.\]
- \[Sample size for estimating a mean: n = (Z × σ / E)^2\]\[where σ is estimated standard deviation and E is allowable error.\]
Interview and Observation Techniques
Interview and Observation Techniques
Key Point: Response rate (%) = (Number of completed interviews / Number of eligible units contacted) × 100
Definition & purpose: Interview and observation are primary data collection techniques. An interview is a systematic oral questionnaire (face-to-face, telephone, or online) to obtain facts, opinions or attitudes. Observation is watching subjects’ behaviour or events directly (recording what is seen) without asking questions. Both aim to gather reliable, valid data when surveys or experiments are not sufficient.
Interview techniques:
- Types: structured (scheduled/questionnaire), semi-structured, unstructured (open-ended); modes: personal/face-to-face, telephonic, online/virtual.
- When to use: when complex information, explanations, attitudes or sensitive pieces of information need probing; where non-verbal cues are important.
- Key steps: define objectives → design questions (clear, neutral, logical order) → choose mode and sample → train interviewers → pilot-test → conduct interview → record responses → code and clean data.
- Good practice: build rapport, use simple language, ask one idea per question, use probes and follow-ups, record answers accurately, ensure privacy and consent.
- Advantages: can clarify unclear answers, probe deeper, higher response quality for complex topics. Disadvantages: costly, time-consuming, interviewer bias, social desirability bias.
Observation techniques:
- Types: direct vs indirect; participant vs non‑participant; structured (checklist, coding scheme) vs unstructured (narrative notes); mechanical (cameras, counters) vs human observation.
- When to use: when actual behaviour is more reliable than reported behaviour (e.g., buying patterns, classroom behaviour, traffic flow).
- Key steps: set clear objectives → choose observation type and instruments (checklist, coding form) → train observers → pilot observation → record systematically (time sampling, event sampling) → ensure inter-observer reliability → code and analyse data.
- Advantages: actual behaviour recorded (not just reported), less reliance on memory, useful for non‑literate populations. Disadvantages: observer bias, Hawthorne effect (people change behaviour when observed), limited access to inner states or motives.
Validity, reliability & bias: Ensure validity (measuring what you intend) by careful question design and clear observation criteria. Improve reliability through standardized instruments, observer training, pilot tests, and measuring inter-observer agreement (e.g., Cohen’s kappa). Minimize biases: use neutral wording, conceal observer presence when ethical and feasible, rotate observers, and anonymize responses.
Ethics & practical points: obtain informed consent, protect confidentiality, avoid leading questions, respect cultural norms, and debrief where needed. Keep a field diary and backup recordings (with permission).
- Census enumerators conducting face-to-face interviews using a structured questionnaire to collect household information (education, occupation, household size).
- A market researcher observing shoppers in a supermarket (non-participant, structured checklist) to count how many pick up a new product and whether they read the label.
- A school inspector using participant observation in a classroom to record teaching methods and student engagement (narrative notes + checklist for structured items).
- A bank calling customers after service (telephonic interview) to measure satisfaction scores and probe reasons for low ratings.
- Traffic engineers using camera footage (mechanical observation) to count vehicles and measure peak-hour flow without interacting with drivers.
- \[Response rate (%) = (Number of completed interviews / Number of eligible units contacted) × 100\]
- \[Non-response rate (%) = 100 − Response rate (%)\]
- \[Standard error of a proportion: SE(p) = sqrt[p(1 − p) / n]\]\[where p is sample proportion and n is sample size\]
- \[Standard error of the mean: SE(\u03BC) = s / sqrt(n)\]\[where s is sample standard deviation\]
- \[Percent error (simple measure): Percent error = |Observed − True| / True × 100\]
- \[Cohen's kappa (inter-observer agreement): kappa = (P_o − P_e) / (1 − P_e)\]\[where P_o is observed agreement and P_e expected agreement by chance\]
Pilot Survey and Pre-testing
Pilot Survey and Pre-testing
Key Point: Response rate (%) = (Number of completed interviews / Number of eligible units contacted) * 100
Definition: A pilot survey is a small-scale trial run of the main survey carried out on a subset of the target population to test research procedures, logistics and the questionnaire. Pre-testing (or pretest) usually means testing the questionnaire and data collection instruments on a small number of respondents to identify problems with wording, understanding, order of questions, response categories and administration.
Purpose: Both help ensure that the main survey will collect valid, reliable and usable data. They reveal ambiguities, sensitive questions, skip-pattern errors, length issues, interviewer training gaps and timing problems before full implementation.
Key differences:
- Pilot survey: broader in scope; may test sampling procedures, field operations, data entry, and analysis routines as well as the questionnaire.
- Pre-testing: focused on the questionnaire instrument and how respondents interpret questions and response options.
Typical objectives:
- Check clarity and interpretation of questions.
- Estimate time per interview and fieldwork costs.
- Evaluate interviewer instructions and training needs.
- Test sampling and response rates and identify likely nonresponse issues.
- Detect coding and skip-pattern errors and technical issues in data capture.
Procedure / Steps:
- Design a reduced version of the main survey instrument for testing.
- Select a small but representative pilot sample (often 5-10% of planned sample or at least 30 respondents) or a convenience sample for pre-testing (10-30 respondents for questionnaire checks).
- Administer the instrument as in the main survey, including timing and interviewer instructions.
- Collect feedback from respondents and interviewers on comprehension, difficulty and discomfort.
- Analyze pilot data for missing values, unusual patterns, respondent times, and item nonresponse.
- Revise the instrument, instructions and procedures and, if necessary, run a second round of pre-testing.
Interpreting pilot results: Look for systematic misunderstandings, consistently skipped questions, very long/short completion times, high item nonresponse and inconsistent answers. Use these findings to refine wording, change response categories, reorder questions or adjust sampling procedures.
Advantages: reduces risk of large-scale errors, saves time and money, improves validity and reliability, and improves interviewer training. Limitations: pilot findings may not generalize if pilot sample is too small or unrepresentative; some rare problems appear only in large samples.
Relation to validity and reliability: Pre-testing and pilot surveys increase content validity (questions measure intended concepts) and reliability (consistent measurement) by identifying ambiguous items and sources of measurement error before the main survey.
- National household income survey: Before launching a 10,000-household survey, a pilot of 500 households is done to test sampling frames, survey duration, routing for employment and income questions, and likely response rates. Based on pilot findings the questionnaire is shortened and training is improved.
- School feedback questionnaire: A pre-test on 20 students reveals that a question about parental involvement is misunderstood. The wording is simplified and response categories are adjusted before distributing the questionnaire to all students.
- Market research for a new soft drink: A pilot survey of 100 consumers tests taste-rating questions and price acceptability. The pilot shows too many respondents choose 'neutral', so the scale is changed from a 3-point to a 5-point Likert scale.
- \[Response rate (%) = (Number of completed interviews / Number of eligible units contacted) * 100\]
- \[Estimated standard error for a proportion: SE = sqrt( p * (1 - p) / n )\]\[where p is the sample proportion and n is pilot sample size\]
- \[Using pilot estimate p to compute required sample size for proportion: n = (z^2 * p * (1 - p)) / e^2\]\[where z is the z-score for desired confidence level and e is the desired margin of error\]
- \[Required sample size for estimating a mean: n = (z * s / e)^2\]\[where s is standard deviation from pilot and e is margin of error\]
Errors and Biases in Data Collection
Errors and Biases in Data Collection
Key Point: Sampling error for a sample mean (standard error): SE = s / sqrt(n), where s = sample standard deviation, n = sample size.
Overview: Errors and biases in data collection are departures of collected data from the true values. They reduce the accuracy and/or validity of conclusions. Errors are broadly grouped into sampling and non-sampling errors; bias refers to systematic errors that push results in a particular direction.
1. Sampling error
- Occurs because a sample (not the whole population) is observed. It is random and decreases as sample size increases.
- Example effect: sample mean differs from population mean by chance.
2. Non-sampling errors
- Coverage error/Selection bias: some members of the population are left out of the sampling frame (e.g., no list of migrant workers). This leads to biased estimates.
- Non-response bias: selected people do not respond and their characteristics differ from respondents (e.g., wealthier households refuse surveys more often).
- Measurement error (response error): respondents give incorrect answers because of misunderstanding, recall problems, or social desirability (e.g., underreporting income or alcohol use).
- Interviewer bias: an interviewer’s tone, wording or body language influences responses.
- Processing and recording errors: data-entry mistakes, coding errors, or wrong aggregation.
- Non-sampling systematic errors (bias): mistakes in instruments, procedures or definitions that consistently push estimates up or down.
3. Distinguishing random error vs systematic error (bias): Random (sampling) error causes variability around the true value and can be reduced by increasing sample size. Systematic error (bias) shifts the entire estimate away from the true value and is not fixed by larger samples — it must be corrected in design or measurement.
4. Sources and prevention:
- Frame quality: use an up-to-date and complete sampling frame to reduce coverage error.
- Question design: use clear neutral wording to reduce measurement and interviewer bias.
- Training & supervision: train interviewers and monitor fieldwork to reduce interviewer and recording errors.
- Improve response rates: follow-ups, incentives, and mixed modes (face-to-face + phone) reduce non-response bias.
- Pre-testing and calibration: pilot surveys, standardised instruments, and calibration of measuring devices reduce systematic measurement error.
- Processing checks: data validation, double entry, and logical checks reduce processing errors.
5. Implications for analysis: Always report sampling error (confidence intervals) and discuss possible non-sampling biases. If bias is suspected, correct by weighting, imputation, or redesign of the survey when possible; otherwise interpret results cautiously.
Summary: A well-designed survey minimises both sampling error (by adequate sample size and design) and non-sampling error (by careful frame construction, instrument design, field procedures and processing). Recognising sources of bias is essential to making valid inferences from collected data.
- Census undercount: Homeless people and recent migrants are often missed in the sampling frame leading to underestimation of population — a coverage error.
- Telephone survey bias: Conducting a survey only by landline excludes households that use only mobile phones, causing selection bias.
- Social desirability in income reporting: Respondents underreport informal or illicit income in face-to-face household surveys (measurement/response bias).
- Recall bias in health surveys: Patients fail to remember past illnesses accurately, producing measurement error.
- Interviewer leading question: An interviewer’s phrasing like 'Most people pay their taxes on time, don’t they?' can push answers toward the expected response — interviewer bias.
- Non-response in online polls: Younger people may be overrepresented if older persons are less likely to respond online, causing non-response bias.
- \[Sampling error for a sample mean (standard error): SE = s / sqrt(n)\]\[where s = sample standard deviation\]\[n = sample size.\]
- \[Sampling error for a sample proportion: SE(p) = sqrt( p*(1 - p) / n )\]\[where p = sample proportion.\]
- \[Margin of error (approx. for 95% CI): ME = z * SE (z ≈ 1.96 for 95% confidence).\]
- \[Finite population correction (when n is a substantial fraction of population N): SE_corrected = (s / sqrt(n)) * sqrt((N - n) / (N - 1)).\]
- \[Sample size for desired margin of error E (proportion): n = (z^2 * p*(1-p)) / E^2 (use p=0.5 if p unknown to maximize n).\]
- \[Sample size for desired margin of error E (mean): n = (z^2 * s^2) / E^2.\]
Quality Control and Ethical Considerations
Quality Control and Ethical Considerations
Key Point: Response rate (%) = (Number of completed responses / Number of eligible units contacted) × 100
Quality Control
Quality control in data collection means ensuring that data are accurate, complete, consistent and reliable so they can support valid conclusions. It covers steps before, during and after fieldwork: careful questionnaire design, pilot testing, training of enumerators, supervision, spot-checks and re-interviews, systematic editing and coding, proper data entry and validation, and documentation of methods and errors.
Key technical checks include completeness (no missing required items), consistency (answers do not contradict each other), range checks (values fall within plausible limits), logical checks (cross-variable relationships hold), and outlier detection. Quality control also distinguishes between sampling errors (caused by using only a sample) and non-sampling errors (measurement error, non-response, processing mistakes) — many quality-control measures aim to reduce non-sampling errors.
Ethical Considerations
Ethical issues govern how data are collected, stored, analyzed and shared. Main principles are:
- Informed consent: respondents should know the purpose of the study and consent to participate.
- Confidentiality and privacy: personal data must be protected; identities should be anonymized unless consented otherwise.
- Voluntary participation: no undue pressure or coercion to participate.
- Do no harm: questions and use of data should not harm respondents socially, psychologically or economically.
- Honest reporting: avoid fabrication, falsification or selective reporting of data and methods.
- Respect for intellectual property: acknowledge sources; do not plagiarize data or ideas.
Ethics and quality are linked: poor ethics (e.g., falsified responses, deception) undermines data quality and trust. Good practice includes drafting a data-management plan, secure storage, limited access, clear consent forms, and transparent documentation of methods and limitations.
Practical steps to ensure both quality and ethics
- Design clear, non-leading questionnaires; pilot them and revise.
- Train enumerators on both technical procedures and ethical behaviour.
- Use validation rules in data entry (e.g., range and logical checks).
- Supervise fieldwork and perform random re-interviews (spot-checks).
- Anonymize personal identifiers before analysis and limit access to raw data.
- Report response rates, sampling design and known limitations in results.
- Census/Household survey: enumerators are trained, questionnaires are pilot-tested, supervisors do spot-checks and re-interviews to detect falsified interviews; personal identifiers are separated from survey data to protect privacy.
- School survey on family income for scholarships: students/parents give written consent; incomes are collected but stored anonymously and used only for eligibility decisions.
- Medical questionnaire: participants sign informed consent; sensitive health data are encrypted and only aggregated results are published to avoid identifying individuals.
- Market research: survey designers avoid leading questions that would bias product preferences; data cleaning removes inconsistent responses before analysis.
- Research misconduct example (hypothetical): fabricating favourable responses to meet targets leads to wrong policy decisions; audit trails and supervisor checks help prevent this.
- \[Response rate (%) = (Number of completed responses / Number of eligible units contacted) × 100\]
- \[Non-response rate (%) = 100 − Response rate (%)\]
- \[Error rate (%) = (Number of records with errors / Total number of records checked) × 100\]
- \[Standard error of sample mean: SE = s / √n (s = sample standard deviation\]\[n = sample size)\]
- \[Standard error of sample proportion: SE(p) = √[p(1 − p) / n]\]\[where p is sample proportion\]
Data Preparation (Brief)
Data Preparation (Brief)
Key Point: Range = Maximum value − Minimum value
What is Data Preparation? Data preparation is the process of converting raw data collected in surveys or experiments into a form suitable for analysis. For Class 11 Economics, this means editing, coding, classifying, tabulating and arranging data into frequency distributions that can be summarized and graphically presented.
Main steps in data preparation
- Editing: Check raw data for errors, omissions and inconsistencies. Correct obvious mistakes, decide how to treat missing values (ignore, impute by mean/median, or mark as separate category).
- Coding: Replace words or long values by short numeric codes (for example, occupational categories 1=Farmer, 2=Shopkeeper). For grouped numerical data, you may also use coding (u=(x−a)/h) to simplify calculations.
- Classification: Group data into classes (intervals) when there are many distinct values. Decide number of classes and class width.
- Tabulation: Prepare a frequency distribution table with columns such as class limits, class boundaries, class mark (mid-point), frequency (f), cumulative frequency (cf), relative frequency (f/N) and percentage frequency.
- Presentation: Summarize the frequency table using appropriate graphs (histogram, frequency polygon, ogive, bar chart, pie chart) and compute summary measures if needed.
Practical notes when classifying data
- Range = maximum value − minimum value. This helps set class width.
- Number of classes can be chosen by convenience. A common rule-of-thumb is Sturges' formula: k ≈ 1 + 3.322 log10 n (optional).
- Class width (h) ≈ Range / k. Round up to a convenient value (so classes cover the whole range without overlap).
- Class limits are written consistently (eg 10–19, 20–29). Class boundaries remove gaps (for discrete measured to integer precision, lower boundary = lower limit − 0.5, upper boundary = upper limit + 0.5).
- Class mark or midpoint (x¯m) = (lower limit + upper limit)/2. Used in drawing frequency polygon or computing grouped mean.
- For unequal class widths, use frequency density = frequency / class width on the vertical axis of a histogram.
Handling missing values and outliers (brief)
- Missing values: possible treatments include deletion, imputing by mean/median/mode, or using separate category depending on context and proportion missing.
- Outliers: verify whether they are data-entry errors. If genuine, consider winsorizing, trimming, or keeping them but report their effect on averages.
Why prepare data? Prepared data (cleaned and grouped) is easier to summarize, compare and visualize. It also reduces noise and helps compute averages, percentages and trends reliably for economic analysis.
- Marks of 30 students: Convert the list of raw marks into a frequency distribution using class intervals (e.g., 0-9, 10-19, …, 90-100), then compute frequency, relative frequency and draw a histogram.
- Heights of 25 students (in cm): Find range, choose about 5–7 class intervals (for instance 140–144, 145–149, …), compute class marks and prepare frequency table for a frequency polygon.
- Monthly household incomes in a survey: Group incomes into categories (e.g., 0–5000, 5001–10000, 10001–20000, >20000), code categories, compute percentage of households in each category and show results in a pie chart.
- Daily temperatures recorded for a month: Identify and correct any obvious recording errors, decide class intervals (e.g., 10–14°C, 15–19°C), then prepare an 'ogive' (cumulative frequency curve) to find the median temperature.
- \[Range = Maximum value − Minimum value\]
- \[Sturges' rule (rule of thumb for number of classes): k ≈ 1 + 3.322 log10 n\]
- \[Class width (h) ≈ Range / k (round up to convenient value)\]
- \[Class mark (midpoint) x_m = (Lower limit + Upper limit) / 2\]
- \[Relative frequency = f / N (where f is class frequency and N is total observations)\]
- \[Percentage frequency = (f / N) × 100\]
Key Concepts
- Population (Universe)
- The entire set of items or individuals about which information is required in a study.
- Sample
- A subset of the population selected for observation to make inferences about the whole population.
- Primary Data
- Data collected firsthand by the investigator for a specific purpose of the study.
- Secondary Data
- Data already collected by someone else for a different purpose and reused by the investigator.
- Census
- A complete enumeration where data are collected from every member of the population.
- Sample Survey
- A study where data are collected from a sample rather than the entire population.
- Sampling
- The process of selecting a sample from the population for data collection.
- Sampling Frame
- A list or representation of all elements in the population from which a sample is drawn.
- Questionnaire
- A structured set of written questions used to collect information from respondents.
- Interview Schedule
- A set of questions administered orally by an interviewer to respondents.
- Pilot Survey
- A small-scale preliminary study to test research instruments and procedures before the main survey.
- Simple Random Sampling
- A sampling method where every element of the population has an equal chance of being selected.
- Systematic Sampling
- Selecting every k-th element from an ordered list after a random start.
- Stratified Random Sampling
- Dividing the population into homogeneous subgroups (strata) and drawing random samples from each stratum.
- Cluster Sampling
- Dividing the population into clusters, randomly selecting some clusters, and surveying all or a sample of units within them.
- Sampling Error
- The difference between a sample statistic and the true population parameter caused by using a sample rather than the whole population.
- Non-sampling Error
- Errors not related to sampling, such as measurement errors, non-response, data processing mistakes, and biased questions.
- Parameter
- A numerical characteristic of a population, such as population mean or proportion.
- Statistic
- A numerical measure calculated from sample data used to estimate a population parameter.
- Frequency Distribution
- An arrangement of raw data in groups or classes showing the number of observations in each class.
Practice Questions
-
Differentiate between primary data and secondary data with one example each. / प्राथमिक आँकड़ों और द्वितीयक आँकड़ों में एक-एक उदाहरण सहित अंतर कीजिए।
Show answer
Primary data are collected first-hand by the investigator for a specific purpose, e.g. a household questionnaire on monthly expenditure. Secondary data are already collected and published by others, e.g. using Census of India population figures. / प्राथमिक आँकड़े अन्वेषक द्वारा किसी विशिष्ट उद्देश्य के लिए प्रत्यक्ष रूप से एकत्र किए जाते हैं, जैसे मासिक व्यय पर घरेलू प्रश्नावली। द्वितीयक आँकड़े पहले से ही दूसरों द्वारा एकत्र व प्रकाशित होते हैं, जैसे भारत की जनगणना के जनसंख्या आँकड़े उपयोग करना।
-
State two advantages and two limitations of a census compared with a sample survey. / प्रतिदर्श सर्वेक्षण की तुलना में जनगणना के दो लाभ और दो सीमाएँ बताइए।
Show answer
Advantages: a census covers the entire population giving complete, detailed disaggregated data and has no sampling error. Limitations: it is very costly and time-consuming, and may suffer large non-sampling (coverage/measurement) errors. / लाभ: जनगणना संपूर्ण समष्टि को सम्मिलित करती है जिससे पूर्ण, विस्तृत आँकड़े मिलते हैं और प्रतिदर्श त्रुटि नहीं होती। सीमाएँ: यह अत्यधिक खर्चीली व समय-साध्य है, तथा इसमें बड़ी गैर-प्रतिदर्श (आवरण/मापन) त्रुटियाँ हो सकती हैं।
-
A factory has 2,000 workers listed in order; an investigator wants a sample of 100. Explain how systematic sampling would be applied and find the interval k. / एक कारखाने में क्रमवार सूचीबद्ध 2,000 श्रमिक हैं; अन्वेषक 100 का प्रतिदर्श चाहता है। समझाइए कि क्रमबद्ध प्रतिदर्शण कैसे लागू होगा और अंतराल k ज्ञात कीजिए।
Show answer
Interval k = N/n = 2000/100 = 20. Choose a random start between 1 and 20, say 7, then select every 20th worker (7th, 27th, 47th, …) until 100 workers are chosen. / अंतराल k = N/n = 2000/100 = 20। 1 से 20 के बीच यादृच्छिक प्रारंभ चुनें, मान लीजिए 7, फिर हर 20वाँ श्रमिक (7वाँ, 27वाँ, 47वाँ, …) तब तक चुनें जब तक 100 श्रमिक न हो जाएँ।
-
Why is stratified sampling often more precise than simple random sampling? / स्तरित प्रतिदर्शण प्राय: सरल यादृच्छिक प्रतिदर्शण की तुलना में अधिक परिशुद्ध क्यों होता है?
Show answer
Stratified sampling divides the population into homogeneous strata (e.g. income groups) and samples within each, so every important subgroup is represented. This reduces overall variance and yields more precise estimates when strata differ markedly from one another. / स्तरित प्रतिदर्शण समष्टि को समरूप स्तरों (जैसे आय वर्ग) में बाँटकर प्रत्येक से प्रतिदर्श लेता है, अत: हर महत्वपूर्ण उपसमूह का प्रतिनिधित्व होता है। जब स्तर परस्पर बहुत भिन्न हों, तो यह कुल प्रसरण घटाकर अधिक परिशुद्ध आकलन देता है।
-
Distinguish a questionnaire from a schedule. / प्रश्नावली और अनुसूची में अंतर कीजिए।
Show answer
A questionnaire is a set of written questions filled in by the respondent themselves (self-administered). A schedule contains similar questions but is filled in by the investigator/enumerator after asking the respondent, useful when respondents are illiterate. / प्रश्नावली लिखित प्रश्नों का समूह है जिसे उत्तरदाता स्वयं भरता है (स्व-प्रशासित)। अनुसूची में समान प्रश्न होते हैं पर इसे अन्वेषक/प्रगणक उत्तरदाता से पूछकर भरता है, जो निरक्षर उत्तरदाताओं के लिए उपयोगी है।
-
Explain the difference between sampling error and non-sampling error, and state which one is reduced by increasing sample size. / प्रतिदर्श त्रुटि और गैर-प्रतिदर्श त्रुटि में अंतर समझाइए, और बताइए कि प्रतिदर्श आकार बढ़ाने से कौन-सी घटती है।
Show answer
Sampling error arises because only a part of the population is observed; it is random and decreases as sample size increases. Non-sampling error (measurement, non-response, coverage errors) can occur in both census and surveys and is not reduced merely by increasing sample size. / प्रतिदर्श त्रुटि इसलिए उत्पन्न होती है क्योंकि समष्टि का केवल एक भाग देखा जाता है; यह यादृच्छिक होती है और प्रतिदर्श आकार बढ़ने पर घटती है। गैर-प्रतिदर्श त्रुटि (मापन, अनुत्तर, आवरण त्रुटियाँ) जनगणना व सर्वेक्षण दोनों में हो सकती है और मात्र प्रतिदर्श आकार बढ़ाने से नहीं घटती।
-
What is the purpose of a pilot survey (pre-testing) before the main survey? / मुख्य सर्वेक्षण से पहले पायलट सर्वेक्षण (पूर्व-परीक्षण) का उद्देश्य क्या है?
Show answer
A pilot survey is a small-scale trial run that tests the questionnaire, procedures and timing to detect ambiguous questions, skip-pattern errors and likely non-response before full-scale collection. It reduces the risk of large costly errors and improves the validity and reliability of the main survey. / पायलट सर्वेक्षण एक छोटे पैमाने का परीक्षण है जो प्रश्नावली, प्रक्रियाओं व समय की जाँच कर अस्पष्ट प्रश्नों, क्रम-त्रुटियों व संभावित अनुत्तर का पता लगाता है। यह बड़ी खर्चीली त्रुटियों का जोखिम घटाता है और मुख्य सर्वेक्षण की वैधता व विश्वसनीयता बढ़ाता है।
-
Give two precautions to minimise non-response bias in a survey. / सर्वेक्षण में अनुत्तर अभिनति घटाने के लिए दो सावधानियाँ बताइए।
Show answer
Keep the questionnaire short and relevant and assure respondents of confidentiality so they are willing to reply; choose convenient timing and make follow-up contacts with non-respondents. These steps raise the response rate and reduce bias from differing characteristics of non-respondents. / प्रश्नावली संक्षिप्त व प्रासंगिक रखें और उत्तरदाताओं को गोपनीयता का आश्वासन दें ताकि वे उत्तर देने को तैयार हों; सुविधाजनक समय चुनें और अनुत्तरदाताओं से अनुवर्ती संपर्क करें। ये कदम प्रत्युत्तर दर बढ़ाते हैं और अनुत्तरदाताओं की भिन्न विशेषताओं से उत्पन्न अभिनति घटाते हैं।
Related Laws & Principles
Explore allFoundational laws & principles connected to this chapter — tap to open in the Laws Explorer.