ਪੰਜਾਬੀਯੂਨੀpunjabiuni
Introductory Statistics

Facts About the F Distribution

੧੧੪ ਪੈਰੇ · 114 paragraphs

ਮਸ਼ੀਨੀ ਅਨੁਵਾਦ · ਬਿਨਾਂ ਜਾਂਚਇਹ ਮਸ਼ੀਨੀ ਅਨੁਵਾਦ ਹੈ ਅਤੇ ਅਜੇ ਮਨੁੱਖੀ ਸਮੀਖਿਆ ਨਹੀਂ ਹੋਈ। ਇਸਨੂੰ ਅੰਤਿਮ, ਪ੍ਰਮਾਣਿਤ ਅਨੁਵਾਦ ਦੀ ਬਜਾਏ ਕੰਮ ਅਧੀਨ ਖਰੜਾ ਸਮਝ ਕੇ ਪੜ੍ਹੋ।Machine-translated, not yet reviewed by a human. Read it as a working draft, not a settled translation — Sikhi.io (Punjabi Classics Pipeline) · google/gemini-2.5-flash-lite.

Here are some facts about the F distribution.

The curve is not symmetrical but skewed to the right.

There is a different curve for each set of dfs.

The F statistic is greater than or equal to zero.

As the degrees of freedom for the numerator and for the denominator get larger, the curve approximates the normal.

Other uses for the F distribution include comparing two variances and two-way Analysis of Variance. Two-Way Analysis is beyond the scope of this chapter.

Example 13.2

Problem

Let’s return to the slicing tomato exercise in Example 13.1. The means of the tomato yields under the five mulching conditions are represented by μ1, μ2, μ3, μ4, μ5. We will conduct a hypothesis test to determine if all means are the same or at least one is different. Using a significance level of 5%, test the null hypothesis that there is no difference in mean yields among the five groups against the alternative hypothesis that at least one mean is different from the rest.

Solution

The null and alternative hypotheses are:

H0: μ1 = μ2 = μ3 = μ4 = μ5

Ha: μi ≠ μj some i ≠ j

The one-way ANOVA results are shown in Table 13.5

row: Source of Variation | Sum of Squares (SS) | Degrees of Freedom (df) | Mean Square (MS) | F

row: Factor (Between) | 36,648,561 | 5 – 1 = 4 | 36,648,561 4 = 9,162,140 36,648,561 4 = 9,162,140 | 9,162,140 2,044,672.6 = 4.4810 9,162,140 2,044,672.6 = 4.4810

row: Error (Within) | 20,446,726 | 15 – 5 = 10 | 20,446,726 10 = 2,044,672.6 20,446,726 10 = 2,044,672.6

row: Total | 57,095,287 | 15 – 1 = 14

Distribution for the test: F4,10

df(num) = 5 – 1 = 4

df(denom) = 15 – 5 = 10

Test statistic: F = 4.4810

Probability Statement: p-value = P(F > 4.481) = 0.0248.

Compare α and the p-value: α = 0.05, p-value = 0.0248

Make a decision: Since α > p-value, we reject H0.

Conclusion: At the 5% significance level, we have reasonably strong evidence that differences in mean yields for slicing tomato plants grown under different mulching conditions are unlikely to be due to chance alone. We may conclude that at least some of mulches led to different mean yields.

Using the TI-83, 83+, 84, 84+ Calculator

To find these results on the calculator:

Press STAT. Press 1:EDIT. Put the data into the lists L1, L2, L3, L4, L5.

Press STAT, and arrow over to TESTS, and arrow down to ANOVA. Press ENTER, and then enter L1, L2, L3, L4, L5). Press ENTER. You will see that the values in the foregoing ANOVA table are easily produced by the calculator, including the test statistic and the p-value of the test.

The calculator displays: F = 4.4810 p = 0.0248 (p-value) Factor df = 4 SS = 36648560.9 MS = 9162140.23 Error df = 10 SS = 20446726 MS = 2044672.6

Try It 13.2

There are multiple variants of the virus that causes COVID-19. The length of hospital stays for patients afflicted with various strains of COVID-19 is shown in Table 13.6.

row: Delta Strain | Omicron Strain | Alpha Strain | Gamma Strain | Beta Strain

row: 13.9 | 11.7 | 18.2 | 16.9 | 9.3

row: 14.9 | 15.1 | 14.6 | 12.8 | 15.8

row: 16.8 | 9.9 | 10.1 | 11.2 | 16.4

Test whether the mean length of hospital stay is the same or different for the various strains of COVID-19. Construct the ANOVA table, find the p-value, and state your conclusion. Use a 5% significance level.

Example 13.3

Four sororities took a random sample of sisters regarding their grade means for the past term. The results are shown in Table 13.7.

row: Sorority 1 | Sorority 2 | Sorority 3 | Sorority 4

row: 2.17 | 2.63 | 2.63 | 3.79

row: 1.85 | 1.77 | 3.78 | 3.45

row: 2.83 | 3.25 | 4.00 | 3.08

row: 1.69 | 1.86 | 2.55 | 2.26

row: 3.33 | 2.21 | 2.45 | 3.18

Problem

Using a significance level of 1%, is there a difference in mean grades among the sororities?

1. ਹੱਲ

2. ਮੰਨ ਲਓ ਕਿ ਸੋਰੋਰਿਟੀਜ਼ ਦੇ ਆਬਾਦੀ ਦੇ ਮਾਧਿਅਮ μ1, μ2, μ3, μ4 ਹਨ। ਯਾਦ ਰੱਖੋ ਕਿ ਸਿਫਰ ਪਰਿਕਲਪਨਾ ਦਾ ਦਾਅਵਾ ਹੈ ਕਿ ਸੋਰੋਰਿਟੀ ਸਮੂਹ ਇੱਕੋ ਆਮ ਵੰਡ ਤੋਂ ਹਨ। ਵਿਕਲਪਕ ਪਰਿਕਲਪਨਾ ਕਹਿੰਦੀ ਹੈ ਕਿ ਘੱਟੋ-ਘੱਟ ਦੋ ਸੋਰੋਰਿਟੀ ਸਮੂਹ ਵੱਖ-ਵੱਖ ਆਮ ਵੰਡਾਂ ਵਾਲੀਆਂ ਆਬਾਦੀਆਂ ਤੋਂ ਆਉਂਦੇ ਹਨ। ਨੋਟ ਕਰੋ ਕਿ ਚਾਰ ਨਮੂਨਾ ਆਕਾਰ ਹਰੇਕ ਪੰਜ ਹਨ।

NOTE

3. ਇਹ ਇੱਕ ਸੰਤੁਲਿਤ ਡਿਜ਼ਾਈਨ ਦੀ ਉਦਾਹਰਨ ਹੈ, ਕਿਉਂਕਿ ਹਰੇਕ ਕਾਰਕ (ਭਾਵ, ਸੋਰੋਰਿਟੀ) ਵਿੱਚ ਨਿਰੀਖਣਾਂ ਦੀ ਸਮਾਨ ਗਿਣਤੀ ਹੁੰਦੀ ਹੈ।

4. H0: μ1 = μ2 = μ3 = μ4

5. Ha: ਮਾਧਿਅਮ μ1, μ2, μ3, μ4 ਸਾਰੇ ਬਰਾਬਰ ਨਹੀਂ ਹਨ।

6. ਟੈਸਟ ਲਈ ਵੰਡ: F3,16

7. ਜਿੱਥੇ k = 4 ਸਮੂਹ ਅਤੇ n = ਕੁੱਲ 20 ਨਮੂਨੇ ਹਨ

8. df(ਅੰਸ਼)= k – 1 = 4 – 1 = 3

9. df(ਹਰ)= n – k = 20 – 4 = 16

10. ਟੈਸਟ ਅੰਕੜੇ ਦੀ ਗਣਨਾ ਕਰੋ: F = 2.23

11. ਗ੍ਰਾਫ:

12. ਸੰਭਾਵਨਾ ਕਥਨ: p-ਮੁੱਲ = P(F > 2.23) = 0.1241

13. α ਅਤੇ p-ਮੁੱਲ ਦੀ ਤੁਲਨਾ ਕਰੋ: α = 0.01 p-ਮੁੱਲ = 0.1241 α < p-ਮੁੱਲ

14. ਫੈਸਲਾ ਲਓ: ਕਿਉਂਕਿ α < p-ਮੁੱਲ, ਤੁਸੀਂ H0 ਨੂੰ ਰੱਦ ਨਹੀਂ ਕਰ ਸਕਦੇ।

15. ਸਿੱਟਾ: ਸੋਰੋਰਿਟੀਜ਼ ਲਈ ਔਸਤ ਗ੍ਰੇਡਾਂ ਵਿੱਚ ਅੰਤਰ ਹੋਣ ਦਾ ਸਿੱਟਾ ਕੱਢਣ ਲਈ ਕੋਈ ਪੂਰਨ ਸਬੂਤ ਨਹੀਂ ਹੈ।

16. TI-83, 83+, 84, 84+ ਕੈਲਕੁਲੇਟਰ ਦੀ ਵਰਤੋਂ ਕਰਨਾ

17. ਡਾਟਾ ਨੂੰ ਸੂਚੀਆਂ L1, L2, L3, ਅਤੇ L4 ਵਿੱਚ ਪਾਓ। STAT ਦਬਾਓ ਅਤੇ TESTS ਵੱਲ ਐਰੋ ਕਰੋ। F:ANOVA ਤੱਕ ਹੇਠਾਂ ਐਰੋ ਕਰੋ। ENTER ਦਬਾਓ ਅਤੇ (L1,L2,L3,L4) ਦਰਜ ਕਰੋ।

18. ਕੈਲਕੁਲੇਟਰ F ਅੰਕੜੇ, p-ਮੁੱਲ ਅਤੇ ਇੱਕ-ਤਰੀਕੇ ਵਾਲੇ ANOVA ਟੇਬਲ ਲਈ ਮੁੱਲ ਪ੍ਰਦਰਸ਼ਿਤ ਕਰਦਾ ਹੈ: F = 2.2303 p = 0.1241 (p-ਮੁੱਲ) ਕਾਰਕ df = 3 SS = 2.88732 MS = 0.96244 ਤਰੁੱਟੀ df = 16 SS = 6.9044 MS = 0.431525

19. ਇਸਨੂੰ ਅਜ਼ਮਾਓ 13.3

20. ਚਾਰ ਖੇਡ ਟੀਮਾਂ ਨੇ ਪਿਛਲੇ ਸਾਲ ਦੇ ਜੀਪੀਏ ਬਾਰੇ ਖਿਡਾਰੀਆਂ ਦਾ ਇੱਕ ਬੇਤਰਤੀਬ ਨਮੂਨਾ ਲਿਆ। ਨਤੀਜੇ ਟੇਬਲ 13.8 ਵਿੱਚ ਦਿਖਾਏ ਗਏ ਹਨ।

21. ਕਤਾਰ: ਬਾਸਕਟਬਾਲ | ਬੇਸਬਾਲ | ਹਾਕੀ | ਲਾਕਰੋਸ

22. ਕਤਾਰ: 3.6 | 2.1 | 4.0 | 2.0

23. ਕਤਾਰ: 2.9 | 2.6 | 2.0 | 3.6

24. ਕਤਾਰ: 2.5 | 3.9 | 2.6 | 3.9

row: 3.3 | 3.1 | 3.2 | 2.7

row: 3.8 | 3.4 | 3.2 | 2.5

Use a significance level of 5%, and determine if there is a difference in GPA among the teams.

Example 13.4

A fourth grade class is studying the environment. One of the assignments is to grow bean plants in different soils. Tommy chose to grow his bean plants in soil found outside his classroom mixed with dryer lint. Tara chose to grow her bean plants in potting soil bought at the local nursery. Nick chose to grow his bean plants in soil from his mother's garden. No chemicals were used on the plants, only water. They were grown inside the classroom next to a large window. Each child grew five plants. At the end of the growing period, each plant was measured, producing the data (in inches) in Table 13.9.

row: Tommy's Plants | Tara's Plants | Nick's Plants

row: 24 | 25 | 23

row: 21 | 31 | 27

row: 23 | 23 | 22

row: 30 | 20 | 30

row: 23 | 28 | 20

Problem

Does it appear that the three media in which the bean plants were grown produce the same mean height? Test at a 3% level of significance.

Solution

This time, we will perform the calculations that lead to the F' statistic. Notice that each group has the same number of plants, so we will use the formula F' = n⋅ s x ¯ 2 s 2 pooled n⋅ s x ¯ 2 s 2 pooled .

First, calculate the sample mean and sample variance of each group.

row: Tommy's Plants | Tara's Plants | Nick's Plants

row: Sample Mean | 24.2 | 25.4 | 24.4

row: Sample Variance | 11.7 | 18.3 | 16.3

Next, calculate the variance of the three group means (Calculate the variance of 24.2, 25.4, and 24.4). Variance of the group means = 0.413 = s x ¯ 2 s x ¯ 2

Then MSbetween = n s x ¯ 2 n s x ¯ 2 = (5)(0.413) where n = 5 is the sample size (number of plants each child grew).

Calculate the mean of the three sample variances (Calculate the mean of 11.7, 18.3, and 16.3). Mean of the sample variances = 15.433 = s2 pooled

Then MSwithin = s2pooled = 15.433.

The F statistic (or F ratio) is F= M S between M S within = n s x ¯ 2 s 2 pooled = (5)(0.413) 15.433 =0.134 F= M S between M S within = n s x ¯ 2 s 2 pooled = (5)(0.413) 15.433 =0.134

The dfs for the numerator = the number of groups – 1 = 3 – 1 = 2.

The dfs for the denominator = the total number of samples – the number of groups = 15 – 3 = 12

The distribution for the test is F2,12 and the F statistic is F = 0.134

The p-value is P(F > 0.134) = 0.8759.

Decision: Since α = 0.03 and the p-value = 0.8759, do not reject H0. (Why?)

Conclusion: With a 3% level of significance, from the sample data, the evidence is not sufficient to conclude that the mean heights of the bean plants are different.

Using the TI-83, 83+, 84, 84+ Calculator

To calculate the p-value:

*Press 2nd DISTR

*Arrow down to Fcdf(and press ENTER.

*Enter 0.134, E99, 2, 12)

*Press ENTER

The p-value is 0.8759.

Try It 13.4

Another fourth grader also grew bean plants, but this time in a jelly-like mass. The heights were (in inches) 24, 28, 25, 30, and 32. Do a one-way ANOVA test on the four groups. Are the heights of the bean plants different? Use the same method as shown in Example 13.4.

Collaborative Exercise

From the class, create four groups of the same size as follows: men under 22, men at least 22, women under 22, women at least 22. Have each member of each group record the number of states in the United States they have visited. Run an ANOVA test to determine if the average number of states visited in the four groups are the same. Test at a 1% level of significance. Use one of the solution sheets in Appendix E Solution Sheets.