Overview
You exceeded your current quota, please check your plan and billing details. For more information on this error, read the docs: https://platform.openai.com/docs/guides/error-codes/api-errors.
Topics in this chapter
2 topics · tap a topic title to jump straight to it.
Types of Data
Fig 1 — Educational Diagram: Types of Data
Types of Data
Key Point: Relative frequency = f / N (f = frequency of a class/category, N = total observations)
Overview: Data are facts or figures collected for analysis. In statistics (Class 11 Economics), data are classified by nature, source, measurement scale, arrangement, and number of variables. Understanding types of data helps choose collection methods and suitable analysis/graphs.
1. By Source
- Primary data: Collected first-hand for a specific purpose (surveys, questionnaires, experiments, observations). Example: household income collected via a field survey.
- Secondary data: Already collected by others and reused (census reports, government publications, research papers). Example: GDP time series from a government website.
2. By Nature / Measurement Scale
- Qualitative (Categorical): Non-numeric categories.
- Nominal: Categories without order (e.g., religion, gender, industry type).
- Ordinal: Categories with a meaningful order but not equal intervals (e.g., satisfaction: low, medium, high; education levels).
- Quantitative (Numerical): Numeric values.
- Discrete: Countable values (e.g., number of children, number of factories).
- Continuous: Measurable on a continuum with possible fractional values (e.g., income, weight, production output).
3. By Arrangement
- Cross-sectional data: Observations on many subjects at a single point in time (e.g., unemployment rates across states in 2024).
- Time-series data: Observations of a variable over several time periods for the same unit (e.g., quarterly GDP of India from 2010–2024).
4. By Number of Variables
- Univariate: Single variable (e.g., household income distribution).
- Bivariate: Two variables, often to study relationship (e.g., consumption vs income).
- Multivariate: More than two variables (e.g., income, education, and age together).
Why these distinctions matter: Choice of graphs, summary measures (mean, median, percentages), and statistical methods depend on data type. For example, nominal data use mode and proportions; continuous data permit mean, variance, histograms and regression analyses.
- Primary data (survey): Asking 500 households in a town about monthly expenditure—primary, cross-sectional, quantitative (continuous).
- Secondary data (published): Using Reserve Bank reports of monthly inflation from 2010–2024—secondary, time-series, quantitative (continuous).
- Qualitative nominal: Classifying firms by sector (agriculture, manufacturing, services) and counting firms in each category.
- Qualitative ordinal: Consumer satisfaction ratings (poor, fair, good, excellent) from a product feedback form.
- Quantitative discrete: Number of employees in small firms (0,1,2,...).
- Quantitative continuous: Measuring daily electricity consumption (kWh) of households—values can be fractional.
- \[Relative frequency = f / N (f = frequency of a class/category\]\[N = total observations)\]
- \[Percentage (for category) = (f / N) × 100\]
- \[Cumulative frequency (CF) for class k = sum of frequencies up to class k\]
- \[Class width (h) for grouped data (equal classes) ≈ Range / number_of_classes\]\[where Range = maximum − minimum\]
- \[Class mark (mid-point) = (Lower class limit + Upper class limit) / 2\]
- \[Frequency density (for unequal class widths) = frequency / class width\]
Designing Questionnaire and Schedule
Fig 2 — Educational Diagram: Designing Questionnaire and Schedule
Designing Questionnaire and Schedule
Key Point: Response rate (%) = (Number of completed questionnaires / Number of questionnaires distributed) × 100
What they are: A questionnaire is a set of written questions given to respondents for self-completion or for an interviewer to read out. A schedule is a structured interview instrument where an enumerator records answers on behalf of the respondent. Both are tools for primary data collection and must be carefully designed to collect valid, reliable and relevant information.
Purpose and first steps: Start with a clear research objective (what you want to measure and why). Translate each objective into specific information needs. Decide whether you need qualitative open answers or quantitative, pre-coded answers suitable for statistical analysis.
Structure and sequence:
- Introduction: purpose, confidentiality, estimated time, consent.
- Screening questions: eligibility, skip patterns.
- Warm-up questions: easy, non-threatening items to build rapport.
- Core questions: main information related to objectives (grouped by theme).
- Background/demographic questions: age, gender, education, income — usually at the end.
- Closing: thank-you note, contact for follow-up.
Types of questions and when to use them:
- Closed (fixed-response): dichotomous (yes/no), multiple choice, ordinal scales (Likert). Use for easy coding and quantitative analysis.
- Open-ended: allow respondents to answer in their own words; useful for exploratory or detailed qualitative information but harder to code.
- Filter/skip questions: direct respondents to relevant sections (important for schedules used by interviewers).
Wording and design rules:
- Use simple, neutral language; avoid jargon and technical terms.
- Avoid leading, loaded, double-barreled or ambiguous questions (e.g., do not ask two things at once).
- Make response categories exhaustive and mutually exclusive.
- Keep questions short and focused on a single idea.
- Provide clear time references (e.g., in the last 30 days, during the past year).
- Use consistent scales (e.g., 1–5) and clearly label scale endpoints.
Format, layout and usability: Good spacing, clear numbering, logical flow, and instructions for interviewers (if any) reduce errors. For self-administered questionnaires use readable fonts, avoid clutter, and estimate completion time.
Pilot testing and revision: Always pilot the questionnaire/schedule on a small, similar sample. Check for comprehension problems, timing, skip-pattern errors, ambiguous options and coding issues. Revise accordingly.
Coding and data preparation: Pre-code closed answers where possible. Develop a coding manual for open responses. Plan variable names, value labels and missing-value codes before data entry to reduce post-collection work.
Administration considerations (questionnaire vs schedule):
- Questionnaire (self-administered): cheaper, less interviewer bias, but needs literate respondents and has lower control over completion.
- Schedule (interviewer-administered): better for complex/long interviews, illiterate respondents, higher response rates, but possible interviewer bias and higher cost.
Practical and ethical issues: Ensure respondent consent, anonymity/confidentiality, and data security. Keep instrument length appropriate and respect respondent time. Train interviewers on probing rules, neutrality and recording answers accurately.
Quality checks: Use back-checks, supervisor reviews, consistency checks and spot re-interviews to detect fabrication or errors. Monitor response and non-response patterns during fieldwork.
- Household expenditure survey: A schedule where an enumerator asks about food, rent, utilities, and pre-coded expenditure ranges for each item; includes screening (is household resident?) and ends with demographics.
- School feedback questionnaire: Self-administered form for students with Likert-scale items (1–5) on teacher clarity, facilities, safety and an open item for suggestions.
- Market research for a new beverage: Short self-completion questionnaire at supermarkets with multiple-choice on purchase frequency, preferred flavors, price sensitivity and a choice task.
- Health camp intake schedule: Interviewer-administered form to record patient symptoms, yes/no history of chronic conditions, and pre-coded medication categories — designed for quick completion and coding.
- NSSO-style schedule: A structured schedule used by enumerators to collect employment and consumption data with detailed instructions, skip patterns and pre-coded categories.
- \[Response rate (%) = (Number of completed questionnaires / Number of questionnaires distributed) × 100\]
- \[Non-response rate (%) = 100 - Response rate (%)\]
- \[Percentage = (Frequency of category / Total valid responses) × 100\]
- \[Mean (for numeric item) = Sum of all responses / Number of valid responses\]
- \[Proportion = Number in subgroup / Total number (useful for binary outcomes)\]
Practice Questions
-
Distinguish between primary and secondary data with one example each. / प्राथमिक और द्वितीयक आंकड़ों में अंतर एक-एक उदाहरण सहित स्पष्ट कीजिए।
Show answer
Primary data are collected first-hand for a specific purpose (e.g., a household income survey), whereas secondary data are already collected by others and reused (e.g., GDP series from a government website). / प्राथमिक आंकड़े किसी विशिष्ट उद्देश्य हेतु प्रथम-हस्त एकत्र किए जाते हैं (जैसे पारिवारिक आय सर्वेक्षण), जबकि द्वितीयक आंकड़े पहले से किसी और द्वारा एकत्रित और पुनः उपयोग किए जाते हैं (जैसे सरकारी वेबसाइट से GDP श्रृंखला)।
-
Differentiate between nominal and ordinal qualitative data with examples. / नाममात्र (nominal) और क्रमसूचक (ordinal) गुणात्मक आंकड़ों में उदाहरण सहित अंतर कीजिए।
Show answer
Nominal data are categories without order (e.g., religion, gender, industry type); ordinal data have a meaningful order but unequal intervals (e.g., satisfaction: low, medium, high). / नाममात्र आंकड़ों में बिना क्रम वाली श्रेणियाँ होती हैं (जैसे धर्म, लिंग, उद्योग प्रकार); क्रमसूचक आंकड़ों में अर्थपूर्ण क्रम होता है पर अंतराल असमान होते हैं (जैसे संतुष्टि: निम्न, मध्यम, उच्च)।
-
Define cross-sectional and time-series data. / प्रतिनिधि-काल (cross-sectional) और काल-श्रेणी (time-series) आंकड़ों को परिभाषित कीजिए।
Show answer
Cross-sectional data record many subjects at a single point in time (e.g., unemployment across states in 2024); time-series data record one variable over several periods (e.g., India's quarterly GDP 2010–2024). / प्रतिनिधि-काल आंकड़े एक ही समय बिंदु पर अनेक इकाइयों को दर्ज करते हैं (जैसे 2024 में राज्यों की बेरोजगारी); काल-श्रेणी आंकड़े एक चर को कई अवधियों में दर्ज करते हैं (जैसे भारत का त्रैमासिक GDP 2010–2024)।
-
In a survey of 500 firms, 120 are in the services sector. Find the relative frequency and percentage for services. / 500 फर्मों के सर्वेक्षण में 120 सेवा क्षेत्र में हैं। सेवा क्षेत्र की सापेक्ष आवृत्ति और प्रतिशत ज्ञात कीजिए।
Show answer
Relative frequency = f/N = 120/500 = 0.24; Percentage = 0.24 × 100 = 24%. / सापेक्ष आवृत्ति = f/N = 120/500 = 0.24; प्रतिशत = 0.24 × 100 = 24%।
-
Why is a histogram suitable for continuous data while a bar chart suits categorical data? / सतत आंकड़ों के लिए आयतचित्र (histogram) और श्रेणीगत आंकड़ों के लिए दंड आरेख (bar chart) क्यों उपयुक्त है?
Show answer
A histogram uses adjoining bars over class intervals to show the distribution of continuous values, whereas a bar chart uses separated bars to compare counts of distinct, non-continuous categories. / आयतचित्र वर्ग-अंतरालों पर जुड़ी हुई पट्टियों से सतत मानों का वितरण दिखाता है, जबकि दंड आरेख अलग-अलग श्रेणियों की गणना की तुलना हेतु पृथक पट्टियाँ उपयोग करता है।
-
State two essential rules for wording questions in a questionnaire. / प्रश्नावली में प्रश्नों के शब्दांकन के दो आवश्यक नियम बताइए।
Show answer
Use simple, neutral language avoiding jargon, and avoid leading, loaded or double-barreled questions; response categories must be exhaustive and mutually exclusive. / सरल, तटस्थ भाषा का प्रयोग करें और तकनीकी शब्दों से बचें; अग्रगामी, पूर्वाग्रही या दोहरे प्रश्नों से बचें; उत्तर श्रेणियाँ सम्पूर्ण और परस्पर अपवर्जी होनी चाहिए।
-
Distinguish between a questionnaire and a schedule as instruments of data collection. / आंकड़ा संग्रहण के साधनों के रूप में प्रश्नावली और अनुसूची (schedule) में अंतर कीजिए।
Show answer
A questionnaire is self-completed by the respondent (cheaper, less interviewer bias, needs literacy), while a schedule is filled by an enumerator who records answers (suits illiterate/complex cases, higher response but costlier with possible interviewer bias). / प्रश्नावली उत्तरदाता स्वयं भरता है (सस्ती, कम साक्षात्कारकर्ता पूर्वाग्रह, साक्षरता आवश्यक), जबकि अनुसूची गणक भरता है जो उत्तर दर्ज करता है (निरक्षर/जटिल मामलों हेतु उपयुक्त, अधिक प्रतिक्रिया पर महँगी एवं संभावित पूर्वाग्रह सहित)।
-
Out of 800 questionnaires distributed, 620 were completed. Calculate the response and non-response rates. / बाँटी गई 800 प्रश्नावलियों में से 620 पूर्ण हुईं। प्रतिक्रिया दर और अप्रतिक्रिया दर की गणना कीजिए।
Show answer
Response rate = (620/800) × 100 = 77.5%; Non-response rate = 100 − 77.5 = 22.5%. / प्रतिक्रिया दर = (620/800) × 100 = 77.5%; अप्रतिक्रिया दर = 100 − 77.5 = 22.5%।
Related Laws & Principles
Explore allFoundational laws & principles behind this chapter. Each one opens a full page — what it says, why it matters, five practice questions and the mistakes to avoid.