Overview
In Class 9 the student learned to present data in tables and graphs and to find the mean, median and mode of ungrouped data. This chapter extends those three measures of central tendency to grouped data, where the observations are arranged in class intervals with frequencies. The mean of grouped data is computed by three methods: the direct method, the assumed mean method and the step-deviation method, each a shortcut on the previous one. The mode of grouped data is found from the modal class by a formula that uses the frequencies of the modal class and its neighbours. The median is found by locating the median class from cumulative frequencies and applying the median formula. The chapter also introduces cumulative frequency distributions of the less-than and more-than types, draws their graphs, called ogives, and reads the median from the point where the two ogives cross. Along the way the student learns to convert inclusive classes to exclusive ones, to handle open-ended and unequal intervals, and to choose the measure that best describes a situation. Statistics of this kind is used in every survey, in school results, in agriculture and in economics, and the Telangana SSC examination asks a four-mark problem and a graph from this chapter almost every year.
Learning Objectives
- Explain what a measure of central tendency is and when the mean, median or mode is the appropriate one.
- Compute the mean of grouped data by the direct method using class marks.
- Compute the mean by the assumed mean method and by the step-deviation method and explain why all three give the same answer.
- Identify the modal class and calculate the mode of a grouped frequency distribution.
- Construct cumulative frequency tables of the less-than and more-than types.
- Locate the median class and calculate the median of grouped data.
- Draw less-than and more-than ogives and find the median graphically from their intersection.
- Convert inclusive class intervals to exclusive ones before applying the mode and median formulas.
- Verify the empirical relation 3 Median = Mode + 2 Mean and use it to estimate one measure from the other two.
Topics in this chapter
13 topics · tap a topic title to jump straight to it.
Grouped data and measures of central tendency
A set of observations such as the marks of 50 students or the daily wages of 100 workers is called data. When the number of observations is large, listing them one by one is unhelpful, so we arrange them in class intervals such as 0–10, 10–20, 20–30, and record how many observations fall in each; this count is the frequency of the class. The resulting table is a grouped frequency distribution. In an interval 10–20, the number 10 is the lower limit and 20 the upper limit; the class size or width is the difference, h = 20 − 10 = 10. The class mark or mid-value is the average of the two limits: x = (lower limit + upper limit)/2, so for 10–20 it is 15. In the exclusive form 10–20, 20–30 the upper limit of one class is the lower limit of the next and an observation equal to 20 is placed in 20–30.
A measure of central tendency is a single number that represents the whole data set, a typical or central value. Three are used: the mean, which is the sum of all observations divided by their number; the median, which is the middle value when the observations are arranged in order; and the mode, which is the value occurring most often. For ungrouped data these were found in Class 9: the mean of 2, 4, 4, 5, 10 is 25/5 = 5, the median is 4, and the mode is 4.
When data are grouped, the individual observations are no longer known — we know only that 7 students scored between 10 and 20, not their exact marks. Therefore the mean, median and mode of grouped data are estimates that assume the observations in a class are spread evenly across it, or concentrated at its class mark. The chapter shows how to compute these estimates.
Which measure to use depends on the purpose. The mean uses every observation and is best for quantities that can be added, such as total rainfall or total production; but it is pulled towards extreme values, so the mean income of a village with one very rich family is misleading. The median is unaffected by extremes and is used for incomes, house prices and marks where a typical value is wanted. The mode is used when the most common value matters, such as the most popular shoe size or shirt size in a shop, or the most common number of children in a family. For a perfectly symmetric distribution the three coincide; for a distribution with a long tail on one side they separate, and the empirical relation 3 Median = Mode + 2 Mean roughly connects them.
- For the class 20–30, lower limit 20, upper limit 30, class size 10, class mark 25.
- Ungrouped data 3, 5, 5, 8, 9: mean = 30/5 = 6, median = 5, mode = 5.
- Incomes (in thousands) 5, 6, 6, 7, 90: mean = 22.8, median = 6. The median describes the typical family better because the single large income inflates the mean.
- A shoe shop wants to stock the size most people buy: it needs the mode, not the mean, of the sizes sold.
- Class mark x = (lower limit + upper limit)/2.
- Class size h = upper limit − lower limit.
- Mean of ungrouped data = (sum of observations)/(number of observations).
Mean of grouped data: the direct method
Suppose observations x₁, x₂, …, xₙ occur with frequencies f₁, f₂, …, fₙ. The mean is the total of all observations divided by the number of observations: x̄ = (f₁x₁ + f₂x₂ + … + fₙxₙ)/(f₁ + f₂ + … + fₙ) = Σfᵢxᵢ / Σfᵢ. Here Σ (sigma) means 'the sum of'. For grouped data we do not know the individual values, so we assume that all the observations in a class are located at its class mark xᵢ. Multiplying each class mark by its frequency, adding, and dividing by the total frequency gives the direct method.
Procedure. (1) Write the class intervals and frequencies fᵢ. (2) Find the class mark xᵢ of each class. (3) Compute fᵢxᵢ for each class. (4) Add the fᵢ column to get Σfᵢ = N and the fᵢxᵢ column to get Σfᵢxᵢ. (5) Divide.
Worked example. The marks of 30 students are: 10–25: 2, 25–40: 3, 40–55: 7, 55–70: 6, 70–85: 6, 85–100: 6. Class marks are 17.5, 32.5, 47.5, 62.5, 77.5, 92.5. The products fᵢxᵢ are 35, 97.5, 332.5, 375, 465, 555, whose sum is 1860. N = 30. Mean = 1860/30 = 62.
Table form:
| Class | fᵢ | xᵢ | fᵢxᵢ |
| 10–25 | 2 | 17.5 | 35 |
| 25–40 | 3 | 32.5 | 97.5 |
| 40–55 | 7 | 47.5 | 332.5 |
| 55–70 | 6 | 62.5 | 375 |
| 70–85 | 6 | 77.5 | 465 |
| 85–100 | 6 | 92.5 | 555 |
| Total | 30 | 1860 |
The direct method is exact in idea but tedious when the class marks are large numbers, such as 2500 or 3500 for wages, because the products become large and errors creep in. The next two methods reduce the arithmetic without changing the answer. The mean should always come out within the range of the data; a mean of 62 for marks between 10 and 100 is reasonable, whereas a mean of 620 signals an error, usually in dividing by the wrong total.
The mean has the units of the data. If the classes are in rupees the mean is in rupees; if they are in kilograms the mean is in kilograms. State the unit in the answer.
- Marks 10–25 (2), 25–40 (3), 40–55 (7), 55–70 (6), 70–85 (6), 85–100 (6): Σfx = 1860, N = 30, mean = 62.
- Number of plants in 20 houses: 0–2 (1), 2–4 (2), 4–6 (1), 6–8 (5), 8–10 (6), 10–12 (2), 12–14 (3). Class marks 1, 3, 5, 7, 9, 11, 13; Σfx = 1 + 6 + 5 + 35 + 54 + 22 + 39 = 162; mean = 162/20 = 8.1 plants.
- Pocket money of 12 students: 5–10 (3), 10–15 (5), 15–20 (4). Class marks 7.5, 12.5, 17.5; Σfx = 22.5 + 62.5 + 70 = 155; mean = 155/12 = Rs 12.92.
- Mean x̄ = Σfᵢxᵢ / Σfᵢ (direct method).
- xᵢ = class mark of the i-th class; N = Σfᵢ = total frequency.
Mean of grouped data: the assumed mean method
When the class marks are large, the products fᵢxᵢ are large and awkward. The assumed mean method avoids them. Choose any convenient class mark, usually the one in the middle of the table, and call it the assumed mean a. For each class compute the deviation dᵢ = xᵢ − a. The deviations are small numbers, some negative, some positive. Then compute fᵢdᵢ, add to get Σfᵢdᵢ, and use
x̄ = a + Σfᵢdᵢ / Σfᵢ.
Why it works. Since xᵢ = a + dᵢ, Σfᵢxᵢ = Σfᵢ(a + dᵢ) = aΣfᵢ + Σfᵢdᵢ. Dividing by Σfᵢ gives x̄ = a + Σfᵢdᵢ/Σfᵢ. The formula is therefore exactly the direct method rearranged; the answer does not depend on which a is chosen, only the ease of arithmetic does. If a is chosen near the true mean, Σfᵢdᵢ is small and the correction term is small.
Worked example. Daily wages of 50 workers: 100–120: 12, 120–140: 14, 140–160: 8, 160–180: 6, 180–200: 10. Class marks 110, 130, 150, 170, 190. Take a = 150. Deviations dᵢ = −40, −20, 0, 20, 40. Products fᵢdᵢ = −480, −280, 0, 120, 400; Σfᵢdᵢ = −240. N = 50. Mean = 150 + (−240)/50 = 150 − 4.8 = Rs 145.20.
| Class | fᵢ | xᵢ | dᵢ = xᵢ − 150 | fᵢdᵢ |
| 100–120 | 12 | 110 | −40 | −480 |
| 120–140 | 14 | 130 | −20 | −280 |
| 140–160 | 8 | 150 | 0 | 0 |
| 160–180 | 6 | 170 | 20 | 120 |
| 180–200 | 10 | 190 | 40 | 400 |
| Total | 50 | −240 |
Check by the direct method: Σfᵢxᵢ = 1320 + 1820 + 1200 + 1020 + 1900 = 7260; 7260/50 = 145.2. Same answer.
Points to remember. The assumed mean must be one of the class marks so that the deviations are simple; any class mark works, but the middle one keeps the numbers smallest. Keep the signs of the deviations carefully; a common error is to add all fᵢdᵢ as positive numbers. The sum Σfᵢdᵢ may be negative, in which case the mean is less than a. The method is most useful when the class size is not uniform, because the step-deviation method of the next section needs equal class sizes.
- Wages 100–120 (12), 120–140 (14), 140–160 (8), 160–180 (6), 180–200 (10); a = 150; Σfd = −240; mean = 150 − 4.8 = Rs 145.20.
- Marks 0–10 (5), 10–20 (8), 20–30 (12), 30–40 (10), 40–50 (5); a = 25; d = −20, −10, 0, 10, 20; Σfd = −100 − 80 + 0 + 100 + 100 = 20; mean = 25 + 20/40 = 25.5.
- Heights (cm) 150–155 (4), 155–160 (10), 160–165 (14), 165–170 (8), 170–175 (4); a = 162.5; d = −10, −5, 0, 5, 10; Σfd = −40 − 50 + 0 + 40 + 40 = −10; mean = 162.5 − 10/40 = 162.25 cm.
- dᵢ = xᵢ − a, where a is the assumed mean (a chosen class mark).
- x̄ = a + Σfᵢdᵢ / Σfᵢ.
- Derivation: Σfᵢxᵢ = aΣfᵢ + Σfᵢdᵢ.
Mean of grouped data: the step-deviation method
When all the classes have the same size h, the deviations dᵢ = xᵢ − a are all multiples of h: 0, ±h, ±2h, ±3h, and so on. Dividing each by h gives small integers uᵢ = (xᵢ − a)/h, usually …, −2, −1, 0, 1, 2, …. This is the step-deviation method. Compute fᵢuᵢ, add, and use
x̄ = a + h × (Σfᵢuᵢ / Σfᵢ).
Why it works. dᵢ = h uᵢ, so Σfᵢdᵢ = hΣfᵢuᵢ, and substituting in the assumed mean formula gives x̄ = a + hΣfᵢuᵢ/Σfᵢ. Again the answer is the same as the direct method; only the arithmetic is simpler.
Worked example. The percentage of female teachers in 35 districts: 15–25: 6, 25–35: 11, 35–45: 7, 45–55: 4, 55–65: 4, 65–75: 2, 75–85: 1. Class marks 20, 30, 40, 50, 60, 70, 80. Take a = 50, h = 10. Then uᵢ = −3, −2, −1, 0, 1, 2, 3. Products fᵢuᵢ = −18, −22, −7, 0, 4, 4, 3; Σfᵢuᵢ = −36. N = 35. Mean = 50 + 10 × (−36/35) = 50 − 10.29 = 39.71 per cent.
| Class | fᵢ | xᵢ | uᵢ | fᵢuᵢ |
| 15–25 | 6 | 20 | −3 | −18 |
| 25–35 | 11 | 30 | −2 | −22 |
| 35–45 | 7 | 40 | −1 | −7 |
| 45–55 | 4 | 50 | 0 | 0 |
| 55–65 | 4 | 60 | 1 | 4 |
| 65–75 | 2 | 70 | 2 | 4 |
| 75–85 | 1 | 80 | 3 | 3 |
| Total | 35 | −36 |
Choosing the method. Use the direct method when the class marks and frequencies are small numbers. Use the assumed mean method when class marks are large or the class sizes are unequal. Use the step-deviation method when the class sizes are equal and the class marks are large; this is the method most often expected in the examination for a four-mark problem, and the question sometimes names the method to be used. Whatever the method, the final mean is identical, and the student may check the answer by a second method if time permits.
A caution: h must be the common class size; if the classes are 0–10, 10–20 then h = 10 even if the frequencies suggest otherwise, and if the classes are of different widths the step-deviation method should not be used, or uᵢ must be computed with a chosen h that divides all the deviations.
- Female teachers: a = 50, h = 10, Σfu = −36, N = 35; mean = 50 − 360/35 = 39.71.
- Wages 100–120 (12), 120–140 (14), 140–160 (8), 160–180 (6), 180–200 (10); a = 150, h = 20; u = −2, −1, 0, 1, 2; Σfu = −24 − 14 + 0 + 6 + 20 = −12; mean = 150 + 20 × (−12/50) = 145.2.
- Number of wickets by 45 bowlers: 20–60 (7), 60–100 (5), 100–150 (16), 150–250 (12), 250–350 (2), 350–450 (3). Unequal classes; the assumed mean method gives a = 200, d = −160, −120, −75, 0, 100, 200; Σfd = −1120 − 600 − 1200 + 0 + 200 + 600 = −2120; mean = 200 − 2120/45 = 152.89 wickets.
- uᵢ = (xᵢ − a)/h, where h is the common class size.
- x̄ = a + h × (Σfᵢuᵢ / Σfᵢ).
- All three methods give the same mean; they differ only in arithmetic.
Mode of grouped data
The mode of ungrouped data is the observation that occurs most often. In grouped data we cannot see individual values, so we first find the modal class, the class with the highest frequency, and then estimate where within it the mode lies, using the frequencies of the neighbouring classes. If the class before the modal class is heavier than the class after it, the mode is pulled towards the lower end of the modal class, and vice versa. The formula that does this is
Mode = l + [(f₁ − f₀) / (2f₁ − f₀ − f₂)] × h,
where l is the lower limit of the modal class, h its class size, f₁ the frequency of the modal class, f₀ the frequency of the class preceding it and f₂ the frequency of the class succeeding it.
Worked example. Ages of patients admitted to a hospital: 5–15: 6, 15–25: 11, 25–35: 21, 35–45: 23, 45–55: 14, 55–65: 5. The highest frequency is 23, so the modal class is 35–45; l = 35, h = 10, f₁ = 23, f₀ = 21, f₂ = 14. Mode = 35 + [(23 − 21)/(46 − 21 − 14)] × 10 = 35 + (2/11) × 10 = 35 + 1.82 = 36.82 years. The mode is close to the lower limit because the preceding class (21) is nearly as heavy as the modal class while the succeeding class (14) is much lighter.
Second example. Number of students per teacher in 35 states: 15–20: 3, 20–25: 8, 25–30: 9, 30–35: 10, 35–40: 3, 40–45: 0, 45–50: 0, 50–55: 2. Modal class 30–35, l = 30, h = 5, f₁ = 10, f₀ = 9, f₂ = 3. Mode = 30 + [(10 − 9)/(20 − 9 − 3)] × 5 = 30 + (1/8) × 5 = 30.625, about 30.6 students per teacher.
Points to note. (1) The classes must be exclusive and of equal size around the modal class; if they are inclusive, such as 1–10, 11–20, convert them first by subtracting 0.5 from each lower limit and adding 0.5 to each upper limit. (2) If the modal class is the first class, f₀ = 0; if it is the last class, f₂ = 0. (3) The denominator 2f₁ − f₀ − f₂ is always positive because f₁ is the largest frequency. (4) The mode always lies inside the modal class; if it does not, an arithmetic error has occurred. (5) If two classes share the highest frequency the data are bimodal and the formula is not applied in this course.
The mode answers questions such as 'what is the most common age of patients?', 'which size sells most?', 'what is the typical monthly consumption?'. It is the only measure of central tendency that can be used for non-numerical data such as the most popular colour.
- Ages 5–15 (6), 15–25 (11), 25–35 (21), 35–45 (23), 45–55 (14), 55–65 (5): modal class 35–45; mode = 35 + (2/11) × 10 = 36.82 years.
- Students per teacher: modal class 30–35, f₁ = 10, f₀ = 9, f₂ = 3; mode = 30 + (1/8) × 5 = 30.6.
- Monthly electricity consumption (units): 65–85 (4), 85–105 (5), 105–125 (13), 125–145 (20), 145–165 (14), 165–185 (8), 185–205 (4). Modal class 125–145; mode = 125 + [(20 − 13)/(40 − 13 − 14)] × 20 = 125 + (7/13) × 20 = 135.77 units.
- Runs scored by batsmen: 3000–4000 (4), 4000–5000 (18), 5000–6000 (9), 6000–7000 (7), 7000–8000 (6), 8000–9000 (3), 9000–10000 (1), 10000–11000 (1). Modal class 4000–5000; mode = 4000 + [(18 − 4)/(36 − 4 − 9)] × 1000 = 4000 + 14000/23 = 4608.7 runs.
- Mode = l + [(f₁ − f₀)/(2f₁ − f₀ − f₂)] × h.
- l = lower limit of modal class; h = class size; f₁ = frequency of modal class; f₀ = frequency of preceding class; f₂ = frequency of succeeding class.
- Modal class = class with the maximum frequency.
Cumulative frequency distributions
The cumulative frequency of a class is the total of its frequency and the frequencies of all the classes before it. A table listing cumulative frequencies is a cumulative frequency distribution. It answers questions of the type 'how many observations are less than 40?' and is the tool for finding the median.
There are two types. In the less-than type, the cumulative frequency is written against the upper limit of each class and tells how many observations are less than that limit. In the more-than type, it is written against the lower limit and tells how many observations are more than or equal to that limit.
Worked example. Marks of 53 students: 0–10: 5, 10–20: 3, 20–30: 4, 30–40: 3, 40–50: 3, 50–60: 4, 60–70: 7, 70–80: 9, 80–90: 7, 90–100: 8. Less-than table: less than 10: 5; less than 20: 5 + 3 = 8; less than 30: 12; less than 40: 15; less than 50: 18; less than 60: 22; less than 70: 29; less than 80: 38; less than 90: 45; less than 100: 53. The last cumulative frequency equals the total N = 53. More-than table: more than or equal to 0: 53; ≥ 10: 53 − 5 = 48; ≥ 20: 45; ≥ 30: 41; ≥ 40: 38; ≥ 50: 35; ≥ 60: 31; ≥ 70: 24; ≥ 80: 15; ≥ 90: 8. The first entry equals N and the entries decrease.
| Marks | f | Less than (upper limit) | cf | More than or equal to (lower limit) | cf |
| 0–10 | 5 | less than 10 | 5 | ≥ 0 | 53 |
| 10–20 | 3 | less than 20 | 8 | ≥ 10 | 48 |
| 20–30 | 4 | less than 30 | 12 | ≥ 20 | 45 |
| 30–40 | 3 | less than 40 | 15 | ≥ 30 | 41 |
| 40–50 | 3 | less than 50 | 18 | ≥ 40 | 38 |
| 50–60 | 4 | less than 60 | 22 | ≥ 50 | 35 |
| 60–70 | 7 | less than 70 | 29 | ≥ 60 | 31 |
| 70–80 | 9 | less than 80 | 38 | ≥ 70 | 24 |
| 80–90 | 7 | less than 90 | 45 | ≥ 80 | 15 |
| 90–100 | 8 | less than 100 | 53 | ≥ 90 | 8 |
Sometimes the data are given already in cumulative form, such as 'less than 20: 8, less than 30: 12', and the ordinary frequencies must be recovered by subtracting successive cumulative frequencies: 12 − 8 = 4 for the class 20–30. Recognising whether a table is a frequency table or a cumulative table is the first step in any problem: cumulative frequencies always increase (less-than) or decrease (more-than) along the table, whereas ordinary frequencies rise and fall.
Cumulative frequencies also give quick answers: from the table, 38 students scored less than 80 and 15 scored 80 or more; the number scoring between 40 and 70 is 29 − 15 = 14.
- Marks of 53 students: less-than cf 5, 8, 12, 15, 18, 22, 29, 38, 45, 53; more-than cf 53, 48, 45, 41, 38, 35, 31, 24, 15, 8.
- Given less than 10: 4, less than 20: 9, less than 30: 17, less than 40: 20, the frequencies are 4, 5, 8, 3.
- From a less-than table, the number of observations between 40 and 70 = cf(70) − cf(40) = 29 − 15 = 14.
- Less-than cumulative frequency of a class = sum of frequencies of that class and all classes before it.
- More-than cumulative frequency of a class = sum of frequencies of that class and all classes after it.
- Frequency of a class = (less-than cf of the class) − (less-than cf of the previous class).
Median of grouped data
The median is the value that divides the ordered data into two equal halves: half the observations are below it and half above. For n ungrouped observations arranged in order, the median is the middle one if n is odd and the average of the two middle ones if n is even. For grouped data we find the class containing the middle observation and then estimate the median within it, assuming the observations are spread evenly across the class.
Procedure. (1) Form the less-than cumulative frequency column. (2) Find N = Σf and compute N/2. (3) The median class is the class whose cumulative frequency is the first to be greater than or equal to N/2. (4) Apply
Median = l + [(N/2 − cf) / f] × h,
where l is the lower limit of the median class, N the total frequency, cf the cumulative frequency of the class preceding the median class, f the frequency of the median class and h its class size.
Worked example. The marks of 53 students given in the previous section have cumulative frequencies 5, 8, 12, 15, 18, 22, 29, 38, 45, 53. N = 53, N/2 = 26.5. The first cumulative frequency ≥ 26.5 is 29, belonging to the class 60–70. So l = 60, cf = 22 (the cumulative frequency of 50–60), f = 7, h = 10. Median = 60 + [(26.5 − 22)/7] × 10 = 60 + (4.5/7) × 10 = 60 + 6.43 = 66.43 marks. So about half the students scored below 66.4 and half above.
Second example. Heights of 51 girls: less than 140: 4, less than 145: 11, less than 150: 29, less than 155: 40, less than 160: 46, less than 165: 51. This is already cumulative. The frequencies are 4, 7, 18, 11, 6, 5 for the classes 135–140, 140–145, 145–150, 150–155, 155–160, 160–165. N/2 = 25.5; the first cf ≥ 25.5 is 29, class 145–150. l = 145, cf = 11, f = 18, h = 5. Median = 145 + [(25.5 − 11)/18] × 5 = 145 + (14.5/18) × 5 = 145 + 4.03 = 149.03 cm.
Points to note. Use N/2 for grouped data even when N is odd; do not use (N + 1)/2. The cf in the formula is that of the class before the median class, not of the median class itself. The median must lie inside the median class. Inclusive classes must be converted to exclusive ones first. If a frequency is missing and the median is given, the formula becomes an equation for the unknown frequency; this is a favourite four-mark problem.
Why the formula works. Below the median class there are cf observations; we need N/2 − cf more from the median class. Assuming its f observations are spread evenly over width h, each observation occupies h/f of the width, so N/2 − cf observations cover (N/2 − cf) × h/f, which is added to the lower limit.
- Marks of 53 students: N/2 = 26.5, median class 60–70, median = 60 + (4.5/7) × 10 = 66.43.
- Heights of 51 girls (cumulative data): median class 145–150, median = 145 + (14.5/18) × 5 = 149.03 cm.
- Life of 400 lamps (hours): 1500–2000 (14), 2000–2500 (56), 2500–3000 (60), 3000–3500 (86), 3500–4000 (74), 4000–4500 (62), 4500–5000 (48). cf: 14, 70, 130, 216, 290, 352, 400. N/2 = 200, median class 3000–3500, median = 3000 + [(200 − 130)/86] × 500 = 3000 + 406.98 = 3406.98 hours.
- Weight of 30 students (kg): 40–45 (2), 45–50 (3), 50–55 (8), 55–60 (6), 60–65 (6), 65–70 (3), 70–75 (2). cf: 2, 5, 13, 19, 25, 28, 30. N/2 = 15, median class 55–60, median = 55 + [(15 − 13)/6] × 5 = 55 + 1.67 = 56.67 kg.
- Median = l + [(N/2 − cf)/f] × h.
- l = lower limit of the median class; cf = cumulative frequency of the class before it; f = its frequency; h = its class size.
- Median class = the class whose less-than cumulative frequency first reaches or exceeds N/2.
Converting inclusive classes and handling special cases
The mode and median formulas assume exclusive (continuous) classes, in which the upper limit of one class is the lower limit of the next: 0–10, 10–20, 20–30. Data are often given in inclusive form, such as 1–10, 11–20, 21–30, with a gap of 1 between the upper limit of one class and the lower limit of the next. Before applying the formulas, convert: subtract half the gap from every lower limit and add half the gap to every upper limit. For a gap of 1 the classes become 0.5–10.5, 10.5–20.5, 20.5–30.5. The class size becomes 10, and the class marks are unchanged (5.5, 15.5, 25.5), so the mean is unaffected by the conversion; only the mode and median require it.
Worked example. Number of letters in 100 surnames: 1–4: 6, 4–7: 30, 7–10: 40, 10–13: 16, 13–16: 4, 16–19: 4. These are already exclusive with h = 3. cf: 6, 36, 76, 92, 96, 100. N/2 = 50; median class 7–10; median = 7 + [(50 − 36)/40] × 3 = 7 + 1.05 = 8.05 letters. Modal class 7–10; mode = 7 + [(40 − 30)/(80 − 30 − 16)] × 3 = 7 + (10/34) × 3 = 7.88 letters. Mean by step deviation with a = 11.5, h = 3: u = −3, −2, −1, 0, 1, 2; Σfu = −18 − 60 − 40 + 0 + 4 + 8 = −106; mean = 11.5 − 3 × 106/100 = 8.32 letters.
Inclusive example. Marks 11–20: 5, 21–30: 12, 31–40: 18, 41–50: 10, 51–60: 5. Convert to 10.5–20.5, 20.5–30.5, 30.5–40.5, 40.5–50.5, 50.5–60.5. N = 50, cf: 5, 17, 35, 45, 50; N/2 = 25; median class 30.5–40.5; median = 30.5 + [(25 − 17)/18] × 10 = 30.5 + 4.44 = 34.94. Mode = 30.5 + [(18 − 12)/(36 − 12 − 10)] × 10 = 30.5 + (6/14) × 10 = 34.79.
Unequal class sizes. The median formula still works with h equal to the size of the median class alone. The mode formula is used only when the classes around the modal class are of equal width. For the mean, use the direct or assumed mean method.
Open-ended classes such as 'below 10' or '60 and above' have no class mark, so the mean cannot be computed without an assumption; the median can, provided the median class is a closed class, which is why the median is preferred for such data (incomes, ages).
Missing frequency. If one frequency is unknown and the median or mean is given, put the unknown as x, express N and cf in terms of x and solve the resulting equation. For example, if the classes 0–10, 10–20, 20–30, 30–40, 40–50 have frequencies 5, x, 20, 15, 7 and the median is 24, then N = 47 + x, cf before 20–30 is 5 + x, and 24 = 20 + [((47 + x)/2 − 5 − x)/20] × 10 gives 4 = (37 − x)/4, so x = 21.
- Inclusive classes 1–10, 11–20, 21–30 become 0.5–10.5, 10.5–20.5, 20.5–30.5; class marks stay 5.5, 15.5, 25.5.
- Surnames data: median = 8.05 letters, mode = 7.88 letters, mean = 8.32 letters.
- Inclusive marks 11–20 (5), 21–30 (12), 31–40 (18), 41–50 (10), 51–60 (5): median = 34.94, mode = 34.79.
- Frequencies 5, x, 20, 15, 7 with median 24: x = 21.
- Exclusive lower limit = inclusive lower limit − ½ gap; exclusive upper limit = inclusive upper limit + ½ gap.
- The mean is unchanged by conversion; the mode and median formulas require exclusive classes.
- For a missing frequency x: write N and cf in terms of x, substitute in the median (or mean) formula and solve.
Graphical representation: the less-than ogive
A cumulative frequency distribution can be drawn as a curve called an ogive (pronounced 'o-jive'), named after the pointed arch of the same shape in architecture. The less-than ogive is drawn as follows. (1) Take the upper limits of the classes on the horizontal axis with a suitable scale. (2) Take the less-than cumulative frequencies on the vertical axis. (3) Plot the points (upper limit, cumulative frequency) for every class. (4) Also plot the point (lower limit of the first class, 0), since no observation is below the first lower limit. (5) Join the points by a free-hand smooth curve, not by straight segments.
The curve starts at zero, rises continuously and flattens at the top where it reaches N. It rises steeply where the frequencies are large and gently where they are small. From it one can read how many observations are below any value, not only the class limits.
Worked example. Daily income of 50 workers: less than 120: 12, less than 140: 26, less than 160: 34, less than 180: 40, less than 200: 50. Plot (100, 0), (120, 12), (140, 26), (160, 34), (180, 40), (200, 50) with income on the x-axis (2 cm = Rs 20) and workers on the y-axis (1 cm = 5 workers), and draw a smooth curve through them.
Reading the median from a less-than ogive. Locate N/2 on the vertical axis (here 25). Draw a horizontal line from it to meet the ogive, and from that point drop a perpendicular to the horizontal axis. The foot of the perpendicular is the median. In the example, 25 lies just below 26, so the median is a little less than 140, about Rs 138.6. By the formula: median class 120–140, l = 120, cf = 12, f = 14, h = 20; median = 120 + [(25 − 12)/14] × 20 = 120 + 18.57 = 138.57, confirming the graph.
The graphical method gives an approximate answer, depending on the accuracy of the drawing and the scale chosen. The examination accepts an answer within reasonable reading error when the graph is asked for, but expects the formula value when the question says 'find the median'. A good practice is to compute by the formula and mark the point on the graph as a check.
Scale and neatness. Choose the scale so that the graph fills the page; label both axes with the quantity and unit; mark the plotted points clearly; write 'less than ogive' as the title. Points must be plotted against upper limits, never against class marks — that is the most frequent error.
- Income of 50 workers: plot (100, 0), (120, 12), (140, 26), (160, 34), (180, 40), (200, 50); the horizontal line at 25 meets the ogive above about 138.5, the median.
- Marks of 53 students: plot (0, 0), (10, 5), (20, 8), (30, 12), (40, 15), (50, 18), (60, 22), (70, 29), (80, 38), (90, 45), (100, 53); from 26.5 on the y-axis the median reads about 66.
- Weights of 35 students: less than 38: 0, less than 40: 3, less than 42: 5, less than 44: 9, less than 46: 14, less than 48: 28, less than 50: 32, less than 52: 35. N/2 = 17.5 gives a median of about 46.5 kg; by formula, median class 46–48, median = 46 + [(17.5 − 14)/14] × 2 = 46.5 kg.
- Less-than ogive: plot (upper limit, less-than cumulative frequency) and (first lower limit, 0); join smoothly.
- Median from the less-than ogive: the x-coordinate of the point where the horizontal line through N/2 meets the curve.
The more-than ogive and the median from two ogives
The more-than ogive is drawn from the more-than cumulative frequency table. (1) Take the lower limits of the classes on the horizontal axis. (2) Plot the points (lower limit, more-than cumulative frequency) for every class; the first point is (lower limit of the first class, N). (3) Also plot (upper limit of the last class, 0). (4) Join by a smooth curve. The curve starts at N and falls to zero.
Worked example. Production yield of 100 farms (quintals per hectare): 50–55: 2, 55–60: 8, 60–65: 12, 65–70: 24, 70–75: 38, 75–80: 16. More-than table: ≥ 50: 100; ≥ 55: 98; ≥ 60: 90; ≥ 65: 78; ≥ 70: 54; ≥ 75: 16; and (80, 0). Plot these six points and the end point and draw the falling curve.
Median from the more-than ogive alone. Exactly as with the less-than ogive: draw the horizontal line through N/2 = 50 to meet the curve and drop a perpendicular; it lands at about 70.5. By the formula, less-than cf are 2, 10, 22, 46, 84, 100; median class 70–75, l = 70, cf = 46, f = 38, h = 5; median = 70 + [(50 − 46)/38] × 5 = 70 + 0.53 = 70.53 quintals per hectare.
Median from both ogives. If the less-than and more-than ogives are drawn on the same axes, one rising and one falling, they cross at exactly one point. The x-coordinate of the point of intersection is the median, because at that value the number of observations below it (read from the less-than curve) equals the number above it (read from the more-than curve), which is what the median means. The y-coordinate of the intersection is N/2. So drawing both curves and dropping a perpendicular from their crossing gives the median without needing to locate N/2 first.
For the marks of 53 students, the less-than points are (10, 5), (20, 8), …, (100, 53) and the more-than points are (0, 53), (10, 48), (20, 45), (30, 41), (40, 38), (50, 35), (60, 31), (70, 24), (80, 15), (90, 8), (100, 0). The two curves cross near (66.4, 26.5), giving median ≈ 66.4, agreeing with the formula value 66.43.
Examination pattern. The four-mark graph question typically gives a frequency table and asks: 'Change the distribution to a more-than type and draw its ogive', or 'Draw both ogives and hence find the median'. Marks are awarded for the cumulative table, the correct points, the smooth curve, the axes and labels, and the median. Draw with a sharp pencil on graph paper, take the same scale for both curves, mark the intersection, and state the median with its unit.
- Yield of 100 farms: more-than points (50, 100), (55, 98), (60, 90), (65, 78), (70, 54), (75, 16), (80, 0); median ≈ 70.5 from the graph and 70.53 by formula.
- Marks of 53 students: both ogives cross near (66.4, 26.5); median ≈ 66.4.
- Income of 50 workers: more-than points (100, 50), (120, 38), (140, 24), (160, 16), (180, 10), (200, 0); with the less-than ogive they cross at about (138.6, 25).
- More-than ogive: plot (lower limit, more-than cumulative frequency) and (last upper limit, 0); join smoothly.
- The two ogives intersect at (Median, N/2); the perpendicular from the intersection to the x-axis gives the median.
Relation among mean, median and mode
The three measures of central tendency describe the same data from different angles. The mean is the balancing point: it uses every value and the sum of deviations from it is zero. The median is the positional middle: half the data lie on either side. The mode is the peak: the value with the greatest concentration.
For a symmetric distribution, one whose histogram is the same on both sides of the centre, the three coincide: mean = median = mode. For a distribution with a tail stretching to the right (higher values), such as incomes, the mean is pulled to the right by the few large values, the mode stays at the peak on the left, and the median lies between: mode < median < mean. For a tail to the left, the order reverses: mean < median < mode. In every case the median sits between the other two.
Karl Pearson gave an empirical (observed, not proved) relation that holds approximately for moderately skewed distributions:
3 Median = Mode + 2 Mean, or equivalently Mode = 3 Median − 2 Mean.
It is used to estimate one measure when the other two are known, and to check computed values. For the surnames data of an earlier section, mean = 8.32, median = 8.05, mode = 7.88. Check: 3 × 8.05 = 24.15 and 7.88 + 2 × 8.32 = 24.52; close but not equal, as expected of an empirical rule. For the hospital ages, mean = 35.38 and mode = 36.82 give an estimated median = (36.82 + 2 × 35.38)/3 = 35.86.
Worked example. If the mean of a distribution is 42 and the median is 40, the estimated mode is 3 × 40 − 2 × 42 = 120 − 84 = 36. If the mode is 15 and the mean 18, the median ≈ (15 + 36)/3 = 17.
Choosing the measure. Consider a factory with 20 workers earning Rs 5000, 5 supervisors earning Rs 10000 and one owner earning Rs 1 lakh a month. Mean salary = (100000 + 50000 + 100000)/26 = Rs 9615; median = Rs 5000; mode = Rs 5000. A union would quote the median or mode as the typical wage; the owner might quote the mean. Neither is wrong, but each answers a different question, and a student should say which measure suits the purpose: mean for totals and averages of additive quantities, median for typical values in skewed data, mode for the most common category.
The examination asks one-mark items on this relation ('if mean = 25 and median = 24, find the mode'), and reasoning questions on which measure is appropriate. It also asks for the mean, median and mode of the same table, after which the relation serves as a check.
- Mean 42, median 40: mode ≈ 3 × 40 − 2 × 42 = 36.
- Mode 15, mean 18: median ≈ (15 + 36)/3 = 17.
- Surnames data: mean 8.32, median 8.05, mode 7.88; 3 × median = 24.15 ≈ mode + 2 × mean = 24.52.
- Factory salaries: mean Rs 9615, median Rs 5000, mode Rs 5000; the mean is inflated by one large salary.
- 3 Median = Mode + 2 Mean (empirical relation).
- Mode = 3 Median − 2 Mean; Median = (Mode + 2 Mean)/3.
- Symmetric distribution: mean = median = mode; right-skewed: mode < median < mean.
Worked examination problems on the mean
Problem 1. Find the mean number of days 40 students were absent: 0–6: 11, 6–10: 10, 10–14: 7, 14–20: 4, 20–28: 4, 28–38: 3, 38–40: 1. The classes are unequal, so use the assumed mean method. Class marks: 3, 8, 12, 17, 24, 33, 39. Take a = 17. Deviations: −14, −9, −5, 0, 7, 16, 22. Products fd: −154, −90, −35, 0, 28, 48, 22; Σfd = −181. Mean = 17 + (−181)/40 = 17 − 4.525 = 12.475, about 12.48 days.
Problem 2. The distribution of daily pocket allowance of children of a locality is 11–13: 7, 13–15: 6, 15–17: 9, 17–19: 13, 19–21: f, 21–23: 5, 23–25: 4, and the mean is Rs 18. Find f. Class marks 12, 14, 16, 18, 20, 22, 24. Take a = 18, h = 2; u = −3, −2, −1, 0, 1, 2, 3; fu = −21, −12, −9, 0, f, 10, 12; Σfu = f − 20. N = 44 + f. Mean = 18 + 2(f − 20)/(44 + f) = 18 gives f − 20 = 0, so f = 20.
Problem 3. Concentration of sulphur dioxide (ppm) in 30 localities: 0.00–0.04: 4, 0.04–0.08: 9, 0.08–0.12: 9, 0.12–0.16: 2, 0.16–0.20: 4, 0.20–0.24: 2. Class marks 0.02, 0.06, 0.10, 0.14, 0.18, 0.22. Direct method: fx = 0.08, 0.54, 0.90, 0.28, 0.72, 0.44; Σfx = 2.96; mean = 2.96/30 = 0.0987 ppm, about 0.099 ppm.
Problem 4. Literacy rate (per cent) of 35 cities: 45–55: 3, 55–65: 10, 65–75: 11, 75–85: 8, 85–95: 3. Step deviation with a = 70, h = 10: u = −2, −1, 0, 1, 2; fu = −6, −10, 0, 8, 6; Σfu = −2. Mean = 70 + 10 × (−2/35) = 70 − 0.57 = 69.43 per cent.
Problem 5. The mean of the following frequency table is 50 and the total frequency is 120: 0–20: 17, 20–40: f₁, 40–60: 32, 60–80: f₂, 80–100: 19. Find f₁ and f₂. From N: 68 + f₁ + f₂ = 120, so f₁ + f₂ = 52. Class marks 10, 30, 50, 70, 90; take a = 50, h = 20; u = −2, −1, 0, 1, 2; Σfu = −34 − f₁ + 0 + f₂ + 38 = 4 + f₂ − f₁. Mean = 50 + 20(4 + f₂ − f₁)/120 = 50 gives f₂ − f₁ = −4. Solving with f₁ + f₂ = 52: f₁ = 28, f₂ = 24.
Layout for the examination: draw the table with columns for class, f, x, d or u, and fd or fu; write the totals; state a and h; write the formula; substitute; give the answer with its unit. Two unknown frequencies need two equations, one from the total and one from the mean.
- Absent days (unequal classes): assumed mean a = 17, Σfd = −181, N = 40, mean = 12.48 days.
- Pocket allowance with unknown f and mean Rs 18: Σfu = f − 20 = 0, so f = 20.
- SO₂ concentration: Σfx = 2.96, mean = 0.099 ppm.
- Two unknown frequencies with N = 120 and mean 50: f₁ + f₂ = 52 and f₂ − f₁ = −4 give f₁ = 28, f₂ = 24.
- Direct: x̄ = Σfx/Σf. Assumed mean: x̄ = a + Σfd/Σf. Step deviation: x̄ = a + h Σfu/Σf.
- Unknown frequency: equate the computed mean to the given mean and solve; a second unknown needs the total frequency equation.
Worked examination problems on the median and mode
Problem 1. Monthly consumption of electricity of 68 consumers: 65–85: 4, 85–105: 5, 105–125: 13, 125–145: 20, 145–165: 14, 165–185: 8, 185–205: 4. Find the median, mode and mean and compare. Cumulative frequencies: 4, 9, 22, 42, 56, 64, 68. N/2 = 34; median class 125–145; l = 125, cf = 22, f = 20, h = 20. Median = 125 + [(34 − 22)/20] × 20 = 125 + 12 = 137 units. Modal class 125–145; f₁ = 20, f₀ = 13, f₂ = 14. Mode = 125 + [(20 − 13)/(40 − 13 − 14)] × 20 = 125 + 140/13 = 135.77 units. Mean by step deviation with a = 135, h = 20: u = −3, −2, −1, 0, 1, 2, 3; fu = −12, −10, −13, 0, 14, 16, 12; Σfu = 7; mean = 135 + 20 × 7/68 = 137.06 units. The three are close, so the distribution is nearly symmetric.
Problem 2. A life insurance agent found the following data for the ages of 100 policy holders: below 20: 2, below 25: 6, below 30: 24, below 35: 45, below 40: 78, below 45: 89, below 50: 92, below 55: 98, below 60: 100. Find the median age. Convert to classes 15–20: 2, 20–25: 4, 25–30: 18, 30–35: 21, 35–40: 33, 40–45: 11, 45–50: 3, 50–55: 6, 55–60: 2. N/2 = 50; the cf first exceeding 50 is 78, class 35–40; l = 35, cf = 45, f = 33, h = 5. Median = 35 + [(50 − 45)/33] × 5 = 35 + 0.76 = 35.76 years.
Problem 3. The median of the following data is 525; find x and y if the total frequency is 100: 0–100: 2, 100–200: 5, 200–300: x, 300–400: 12, 400–500: 17, 500–600: 20, 600–700: y, 700–800: 9, 800–900: 7, 900–1000: 4. From the total: 76 + x + y = 100, so x + y = 24. Since the median 525 lies in 500–600, that is the median class: l = 500, f = 20, h = 100, cf = 36 + x. 525 = 500 + [(50 − 36 − x)/20] × 100, so 25 = 5(14 − x), 14 − x = 5, x = 9, y = 15.
Problem 4. Ages of 80 patients: 5–15: 6, 15–25: 11, 25–35: 21, 35–45: 23, 45–55: 14, 55–65: 5. Mode = 36.82 years (modal class 35–45). Mean by step deviation, a = 30, h = 10: u = −2, −1, 0, 1, 2, 3; fu = −12, −11, 0, 23, 28, 15; Σfu = 43; mean = 30 + 430/80 = 35.375 years. The maximum number of patients are about 36.8 years old, while the average age is 35.4 years.
These problems are the model for the four-mark question: build the cumulative column, identify the class, name l, cf, f, h (or l, f₀, f₁, f₂, h), substitute and simplify, and interpret the answer in a sentence.
- Electricity consumption: median 137 units, mode 135.77 units, mean 137.06 units.
- Policy holders (below-type data converted): median class 35–40, median 35.76 years.
- Median 525 with unknowns x, y and N = 100: x + y = 24 and 14 − x = 5 give x = 9, y = 15.
- Patients: mode 36.82 years, mean 35.375 years.
- Median = l + [(N/2 − cf)/f] × h; Mode = l + [(f₁ − f₀)/(2f₁ − f₀ − f₂)] × h.
- When the median is given, its class is the median class; substitute and solve for the unknown frequency.
- 'Below' data are less-than cumulative frequencies; subtract successively to recover class frequencies.
Key Concepts
- Grouped frequency distribution
- A table that arranges data in class intervals and records the number of observations, the frequency, in each class.
- Class mark
- The mid-value of a class interval, equal to (lower limit + upper limit)/2, used to represent all observations of the class.
- Class size
- The width of a class interval, the difference between its upper and lower limits, denoted h.
- Measure of central tendency
- A single value, such as the mean, median or mode, that represents the centre of a data set.
- Mean (direct method)
- x̄ = Σfᵢxᵢ/Σfᵢ, the sum of frequency times class mark divided by the total frequency.
- Assumed mean method
- A method of finding the mean using deviations dᵢ = xᵢ − a from a chosen class mark a, with x̄ = a + Σfᵢdᵢ/Σfᵢ.
- Step-deviation method
- A method using uᵢ = (xᵢ − a)/h for equal class sizes, with x̄ = a + h(Σfᵢuᵢ/Σfᵢ).
- Modal class
- The class interval with the highest frequency in a grouped distribution.
- Mode of grouped data
- The estimated most frequent value, l + [(f₁ − f₀)/(2f₁ − f₀ − f₂)]h, computed from the modal class and its neighbours.
- Cumulative frequency
- The running total of frequencies up to a class (less-than type) or from a class onwards (more-than type).
- Median class
- The class whose less-than cumulative frequency is the first to reach or exceed N/2.
- Median of grouped data
- The value l + [(N/2 − cf)/f]h that divides the distribution into two equal halves.
- Ogive
- A smooth curve obtained by plotting cumulative frequencies against class limits; less-than ogives rise and more-than ogives fall.
- Less-than ogive
- The graph of less-than cumulative frequencies plotted against the upper limits of the classes.
- More-than ogive
- The graph of more-than cumulative frequencies plotted against the lower limits of the classes.
- Empirical relation
- The approximate relation 3 Median = Mode + 2 Mean connecting the three measures of central tendency.
- Exclusive class interval
- A class such as 10–20 whose upper limit is the lower limit of the next class; required by the median and mode formulas.
- Inclusive class interval
- A class such as 11–20 with a gap before the next class, which is converted to exclusive form by adjusting the limits by half the gap.
End-of-Chapter Trial Paper & Test Questions
Topic-wise questions to test your understanding of every concept in this chapter.
-
Find the mean of the following distribution by the step-deviation method: Class 10–25: 2, 25–40: 3, 40–55: 7, 55–70: 6, 70–85: 6, 85–100: 6. / निम्नलिखित बंटन का माध्य पग-विचलन विधि से ज्ञात कीजिए: वर्ग 10–25: 2, 25–40: 3, 40–55: 7, 55–70: 6, 70–85: 6, 85–100: 6।
Show answer
Class marks are 17.5, 32.5, 47.5, 62.5, 77.5, 92.5 and h = 15. Take a = 47.5. Then u = (x − 47.5)/15 = −2, −1, 0, 1, 2, 3 and fu = −4, −3, 0, 6, 12, 18, so Σfu = 29 and N = 30. Mean = a + h(Σfu/N) = 47.5 + 15 × 29/30 = 47.5 + 14.5 = 62. / वर्ग चिह्न 17.5, 32.5, 47.5, 62.5, 77.5, 92.5 हैं और h = 15। a = 47.5 लीजिए। तब u = (x − 47.5)/15 = −2, −1, 0, 1, 2, 3 और fu = −4, −3, 0, 6, 12, 18, अतः Σfu = 29 और N = 30। माध्य = a + h(Σfu/N) = 47.5 + 15 × 29/30 = 47.5 + 14.5 = 62।
-
The following table shows the ages of patients admitted in a hospital during a year: 5–15: 6, 15–25: 11, 25–35: 21, 35–45: 23, 45–55: 14, 55–65: 5. Find the mode and the mean of the data and interpret the two measures. / निम्न तालिका एक वर्ष में अस्पताल में भर्ती रोगियों की आयु दर्शाती है: 5–15: 6, 15–25: 11, 25–35: 21, 35–45: 23, 45–55: 14, 55–65: 5। आँकड़ों का बहुलक और माध्य ज्ञात कीजिए तथा दोनों मापों की व्याख्या कीजिए।
Show answer
Modal class is 35–45 (frequency 23). l = 35, h = 10, f₁ = 23, f₀ = 21, f₂ = 14. Mode = 35 + [(23 − 21)/(46 − 21 − 14)] × 10 = 35 + (2/11) × 10 = 36.82 years. For the mean, class marks 10, 20, 30, 40, 50, 60; a = 30, h = 10; u = −2, −1, 0, 1, 2, 3; fu = −12, −11, 0, 23, 28, 15; Σfu = 43; N = 80. Mean = 30 + 10 × 43/80 = 35.375 years. Interpretation: the largest number of patients are about 36.8 years old, while the average age of all patients is about 35.4 years. / बहुलक वर्ग 35–45 है (बारंबारता 23)। l = 35, h = 10, f₁ = 23, f₀ = 21, f₂ = 14। बहुलक = 35 + [(23 − 21)/(46 − 21 − 14)] × 10 = 35 + (2/11) × 10 = 36.82 वर्ष। माध्य के लिए वर्ग चिह्न 10, 20, 30, 40, 50, 60; a = 30, h = 10; u = −2, −1, 0, 1, 2, 3; fu = −12, −11, 0, 23, 28, 15; Σfu = 43; N = 80। माध्य = 30 + 10 × 43/80 = 35.375 वर्ष। व्याख्या: सबसे अधिक रोगी लगभग 36.8 वर्ष की आयु के हैं, जबकि सभी रोगियों की औसत आयु लगभग 35.4 वर्ष है।
-
The following distribution gives the monthly consumption of electricity of 68 consumers: 65–85: 4, 85–105: 5, 105–125: 13, 125–145: 20, 145–165: 14, 165–185: 8, 185–205: 4. Find the median of the data. / निम्न बंटन 68 उपभोक्ताओं की मासिक बिजली खपत दर्शाता है: 65–85: 4, 85–105: 5, 105–125: 13, 125–145: 20, 145–165: 14, 165–185: 8, 185–205: 4। आँकड़ों का माध्यक ज्ञात कीजिए।
Show answer
Cumulative frequencies: 4, 9, 22, 42, 56, 64, 68. N = 68, N/2 = 34. The first cumulative frequency greater than 34 is 42, so the median class is 125–145 with l = 125, cf = 22, f = 20, h = 20. Median = 125 + [(34 − 22)/20] × 20 = 125 + 12 = 137 units. So half the consumers use less than 137 units a month. / संचयी बारंबारताएँ: 4, 9, 22, 42, 56, 64, 68। N = 68, N/2 = 34। 34 से बड़ी पहली संचयी बारंबारता 42 है, अतः माध्यक वर्ग 125–145 है जिसमें l = 125, cf = 22, f = 20, h = 20। माध्यक = 125 + [(34 − 22)/20] × 20 = 125 + 12 = 137 यूनिट। अतः आधे उपभोक्ता महीने में 137 यूनिट से कम उपयोग करते हैं।
-
The median of the following data is 525. Find the values of x and y if the total frequency is 100: 0–100: 2, 100–200: 5, 200–300: x, 300–400: 12, 400–500: 17, 500–600: 20, 600–700: y, 700–800: 9, 800–900: 7, 900–1000: 4. / निम्न आँकड़ों का माध्यक 525 है। यदि कुल बारंबारता 100 है तो x और y के मान ज्ञात कीजिए: 0–100: 2, 100–200: 5, 200–300: x, 300–400: 12, 400–500: 17, 500–600: 20, 600–700: y, 700–800: 9, 800–900: 7, 900–1000: 4।
Show answer
Total: 2 + 5 + x + 12 + 17 + 20 + y + 9 + 7 + 4 = 76 + x + y = 100, so x + y = 24. Since the median 525 lies in 500–600, that is the median class: l = 500, h = 100, f = 20, and cf (of classes before it) = 2 + 5 + x + 12 + 17 = 36 + x. Median formula: 525 = 500 + [(50 − 36 − x)/20] × 100, so 25 = 5(14 − x), giving 14 − x = 5, x = 9. Then y = 24 − 9 = 15. / कुल: 2 + 5 + x + 12 + 17 + 20 + y + 9 + 7 + 4 = 76 + x + y = 100, अतः x + y = 24। चूँकि माध्यक 525, 500–600 में है, वही माध्यक वर्ग है: l = 500, h = 100, f = 20, और उससे पहले के वर्गों की cf = 2 + 5 + x + 12 + 17 = 36 + x। माध्यक सूत्र: 525 = 500 + [(50 − 36 − x)/20] × 100, अतः 25 = 5(14 − x), जिससे 14 − x = 5, x = 9। तब y = 24 − 9 = 15।
-
Explain the assumed mean method for finding the mean of grouped data and show why it gives the same result as the direct method. / वर्गीकृत आँकड़ों का माध्य ज्ञात करने की कल्पित माध्य विधि समझाइए और दिखाइए कि यह प्रत्यक्ष विधि के समान परिणाम क्यों देती है।
Show answer
In the assumed mean method we choose a convenient class mark a, compute the deviation dᵢ = xᵢ − a of each class mark from it, multiply by the frequencies to get fᵢdᵢ, and use x̄ = a + Σfᵢdᵢ/Σfᵢ. Since xᵢ = a + dᵢ, we have Σfᵢxᵢ = Σfᵢ(a + dᵢ) = aΣfᵢ + Σfᵢdᵢ; dividing by Σfᵢ gives Σfᵢxᵢ/Σfᵢ = a + Σfᵢdᵢ/Σfᵢ, which is exactly the direct-method mean. The method only makes the arithmetic easier because the deviations are smaller numbers than the class marks; the answer is independent of the choice of a. / कल्पित माध्य विधि में हम एक सुविधाजनक वर्ग चिह्न a चुनते हैं, प्रत्येक वर्ग चिह्न का उससे विचलन dᵢ = xᵢ − a निकालते हैं, बारंबारताओं से गुणा करके fᵢdᵢ पाते हैं, और x̄ = a + Σfᵢdᵢ/Σfᵢ का उपयोग करते हैं। चूँकि xᵢ = a + dᵢ, अतः Σfᵢxᵢ = Σfᵢ(a + dᵢ) = aΣfᵢ + Σfᵢdᵢ; Σfᵢ से भाग देने पर Σfᵢxᵢ/Σfᵢ = a + Σfᵢdᵢ/Σfᵢ, जो ठीक प्रत्यक्ष विधि का माध्य है। यह विधि केवल गणना सरल बनाती है क्योंकि विचलन वर्ग चिह्नों से छोटी संख्याएँ होती हैं; उत्तर a के चुनाव पर निर्भर नहीं करता।
-
The following table gives the daily income of 50 workers of a factory: 100–120: 12, 120–140: 14, 140–160: 8, 160–180: 6, 180–200: 10. Convert the distribution to a less-than type cumulative frequency distribution and draw its ogive. How is the median read from it? / निम्न तालिका एक कारखाने के 50 श्रमिकों की दैनिक आय दर्शाती है: 100–120: 12, 120–140: 14, 140–160: 8, 160–180: 6, 180–200: 10। बंटन को 'से कम' प्रकार के संचयी बारंबारता बंटन में बदलिए और उसका तोरण खींचिए। इससे माध्यक कैसे पढ़ा जाता है?
Show answer
Less-than table: less than 120: 12; less than 140: 26; less than 160: 34; less than 180: 40; less than 200: 50. Plot the points (100, 0), (120, 12), (140, 26), (160, 34), (180, 40), (200, 50) with income on the x-axis and cumulative frequency on the y-axis, and join them by a smooth rising curve; this is the less-than ogive. To read the median, mark N/2 = 25 on the y-axis, draw a horizontal line to the curve, and drop a perpendicular to the x-axis; it meets the axis at about Rs 138.6, which agrees with the formula value 120 + [(25 − 12)/14] × 20 = 138.57. / 'से कम' तालिका: 120 से कम: 12; 140 से कम: 26; 160 से कम: 34; 180 से कम: 40; 200 से कम: 50। x-अक्ष पर आय और y-अक्ष पर संचयी बारंबारता लेकर बिंदु (100, 0), (120, 12), (140, 26), (160, 34), (180, 40), (200, 50) अंकित कीजिए और उन्हें एक चिकने बढ़ते वक्र से जोड़िए; यही 'से कम' तोरण है। माध्यक पढ़ने के लिए y-अक्ष पर N/2 = 25 अंकित कीजिए, वक्र तक क्षैतिज रेखा खींचिए और x-अक्ष पर लंब डालिए; यह अक्ष से लगभग 138.6 रुपये पर मिलता है, जो सूत्र मान 120 + [(25 − 12)/14] × 20 = 138.57 से मेल खाता है।
-
The mean of a distribution is 42 and its median is 40. Using the empirical relation, estimate the mode. Which measure of central tendency is most suitable for reporting the typical income of families in a village, and why? / एक बंटन का माध्य 42 और माध्यक 40 है। आनुभविक संबंध से बहुलक का अनुमान लगाइए। किसी गाँव के परिवारों की प्रतिनिधि आय बताने के लिए केंद्रीय प्रवृत्ति का कौन-सा माप सबसे उपयुक्त है, और क्यों?
Show answer
The empirical relation is 3 Median = Mode + 2 Mean, so Mode = 3 × 40 − 2 × 42 = 120 − 84 = 36. For incomes the median is most suitable, because income data are usually skewed: a few very rich families pull the mean far above what most families earn, whereas the median, the middle value, is unaffected by such extreme values and represents the typical family. / आनुभविक संबंध 3 माध्यक = बहुलक + 2 माध्य है, अतः बहुलक = 3 × 40 − 2 × 42 = 120 − 84 = 36। आय के लिए माध्यक सबसे उपयुक्त है, क्योंकि आय के आँकड़े प्रायः विषम होते हैं: कुछ बहुत धनी परिवार माध्य को अधिकांश परिवारों की आय से बहुत ऊपर खींच लेते हैं, जबकि माध्यक, अर्थात मध्य मान, ऐसे चरम मानों से प्रभावित नहीं होता और प्रतिनिधि परिवार को दर्शाता है।
-
Find the mode of the following distribution of the number of students per teacher in 35 states: 15–20: 3, 20–25: 8, 25–30: 9, 30–35: 10, 35–40: 3, 40–45: 0, 45–50: 0, 50–55: 2. / 35 राज्यों में प्रति शिक्षक विद्यार्थियों की संख्या के निम्न बंटन का बहुलक ज्ञात कीजिए: 15–20: 3, 20–25: 8, 25–30: 9, 30–35: 10, 35–40: 3, 40–45: 0, 45–50: 0, 50–55: 2।
Show answer
The highest frequency is 10, so the modal class is 30–35. Here l = 30, h = 5, f₁ = 10, f₀ = 9 (class 25–30), f₂ = 3 (class 35–40). Mode = l + [(f₁ − f₀)/(2f₁ − f₀ − f₂)] × h = 30 + [(10 − 9)/(20 − 9 − 3)] × 5 = 30 + (1/8) × 5 = 30 + 0.625 = 30.6 students per teacher (approximately). / सबसे अधिक बारंबारता 10 है, अतः बहुलक वर्ग 30–35 है। यहाँ l = 30, h = 5, f₁ = 10, f₀ = 9 (वर्ग 25–30), f₂ = 3 (वर्ग 35–40)। बहुलक = l + [(f₁ − f₀)/(2f₁ − f₀ − f₂)] × h = 30 + [(10 − 9)/(20 − 9 − 3)] × 5 = 30 + (1/8) × 5 = 30 + 0.625 = लगभग 30.6 विद्यार्थी प्रति शिक्षक।
-
A life insurance agent found the following data for the distribution of ages of 100 policy holders: below 20: 2, below 25: 6, below 30: 24, below 35: 45, below 40: 78, below 45: 89, below 50: 92, below 55: 98, below 60: 100. Find the median age. / एक जीवन बीमा एजेंट ने 100 पॉलिसीधारकों की आयु का निम्न बंटन पाया: 20 से कम: 2, 25 से कम: 6, 30 से कम: 24, 35 से कम: 45, 40 से कम: 78, 45 से कम: 89, 50 से कम: 92, 55 से कम: 98, 60 से कम: 100। माध्यक आयु ज्ञात कीजिए।
Show answer
The data are less-than cumulative frequencies. Class frequencies: 15–20: 2, 20–25: 4, 25–30: 18, 30–35: 21, 35–40: 33, 40–45: 11, 45–50: 3, 50–55: 6, 55–60: 2. N = 100, N/2 = 50. The first cumulative frequency exceeding 50 is 78, so the median class is 35–40 with l = 35, cf = 45, f = 33, h = 5. Median = 35 + [(50 − 45)/33] × 5 = 35 + 25/33 = 35 + 0.76 = 35.76 years. / आँकड़े 'से कम' संचयी बारंबारताएँ हैं। वर्ग बारंबारताएँ: 15–20: 2, 20–25: 4, 25–30: 18, 30–35: 21, 35–40: 33, 40–45: 11, 45–50: 3, 50–55: 6, 55–60: 2। N = 100, N/2 = 50। 50 से बड़ी पहली संचयी बारंबारता 78 है, अतः माध्यक वर्ग 35–40 है जिसमें l = 35, cf = 45, f = 33, h = 5। माध्यक = 35 + [(50 − 45)/33] × 5 = 35 + 25/33 = 35 + 0.76 = 35.76 वर्ष।
-
The distribution of daily pocket allowance of children of a locality is 11–13: 7, 13–15: 6, 15–17: 9, 17–19: 13, 19–21: f, 21–23: 5, 23–25: 4. The mean pocket allowance is Rs 18. Find the missing frequency f. / एक मोहल्ले के बच्चों के दैनिक जेबखर्च का बंटन 11–13: 7, 13–15: 6, 15–17: 9, 17–19: 13, 19–21: f, 21–23: 5, 23–25: 4 है। माध्य जेबखर्च 18 रुपये है। लुप्त बारंबारता f ज्ञात कीजिए।
Show answer
Class marks are 12, 14, 16, 18, 20, 22, 24; take a = 18, h = 2, so u = −3, −2, −1, 0, 1, 2, 3 and fu = −21, −12, −9, 0, f, 10, 12, giving Σfu = f − 20. N = 44 + f. Mean = a + h(Σfu/N) = 18 + 2(f − 20)/(44 + f). Since the mean is 18, 2(f − 20)/(44 + f) = 0, so f − 20 = 0 and f = 20. / वर्ग चिह्न 12, 14, 16, 18, 20, 22, 24 हैं; a = 18, h = 2 लीजिए, अतः u = −3, −2, −1, 0, 1, 2, 3 और fu = −21, −12, −9, 0, f, 10, 12, जिससे Σfu = f − 20। N = 44 + f। माध्य = a + h(Σfu/N) = 18 + 2(f − 20)/(44 + f)। चूँकि माध्य 18 है, 2(f − 20)/(44 + f) = 0, अतः f − 20 = 0 और f = 20।
-
Explain how the median of grouped data can be obtained from the point of intersection of the less-than and more-than ogives. / समझाइए कि 'से कम' और 'से अधिक' तोरणों के प्रतिच्छेद बिंदु से वर्गीकृत आँकड़ों का माध्यक कैसे प्राप्त किया जा सकता है।
Show answer
Draw both ogives on the same axes with the same scale: the less-than ogive rising from (first lower limit, 0) to (last upper limit, N) and the more-than ogive falling from (first lower limit, N) to (last upper limit, 0). They cross at one point. At the x-value of that point, the number of observations below it (from the less-than curve) equals the number above it (from the more-than curve), so each is N/2; this is exactly the definition of the median. Hence drop a perpendicular from the point of intersection to the x-axis; the foot of the perpendicular gives the median, and the y-coordinate of the intersection is N/2. / दोनों तोरणों को एक ही अक्षों और एक ही पैमाने पर खींचिए: 'से कम' तोरण (पहली निचली सीमा, 0) से (अंतिम ऊपरी सीमा, N) तक बढ़ता हुआ और 'से अधिक' तोरण (पहली निचली सीमा, N) से (अंतिम ऊपरी सीमा, 0) तक घटता हुआ। वे एक बिंदु पर कटते हैं। उस बिंदु के x-मान पर उससे नीचे के प्रेक्षणों की संख्या ('से कम' वक्र से) उससे ऊपर के प्रेक्षणों की संख्या ('से अधिक' वक्र से) के बराबर होती है, अतः प्रत्येक N/2 है; यही माध्यक की परिभाषा है। अतः प्रतिच्छेद बिंदु से x-अक्ष पर लंब डालिए; लंब का पाद माध्यक देता है, और प्रतिच्छेद बिंदु का y-निर्देशांक N/2 है।
Related Laws & Principles
Explore allFoundational laws & principles behind this chapter. Each one opens a full page — what it says, why it matters, five practice questions and the mistakes to avoid.