ਪੰਜਾਬੀਯੂਨੀpunjabiuni
Introductory Statistics

The Regression Equation

੧੨੮ ਪੈਰੇ · 128 paragraphs

ਮਸ਼ੀਨੀ ਅਨੁਵਾਦ · ਬਿਨਾਂ ਜਾਂਚਇਹ ਮਸ਼ੀਨੀ ਅਨੁਵਾਦ ਹੈ ਅਤੇ ਅਜੇ ਮਨੁੱਖੀ ਸਮੀਖਿਆ ਨਹੀਂ ਹੋਈ। ਇਸਨੂੰ ਅੰਤਿਮ, ਪ੍ਰਮਾਣਿਤ ਅਨੁਵਾਦ ਦੀ ਬਜਾਏ ਕੰਮ ਅਧੀਨ ਖਰੜਾ ਸਮਝ ਕੇ ਪੜ੍ਹੋ।Machine-translated, not yet reviewed by a human. Read it as a working draft, not a settled translation — Sikhi.io (Punjabi Classics Pipeline) · google/gemini-2.5-flash-lite.

ਅੰਕੜੇ ਸ਼ਾਇਦ ਹੀ ਕਦੇ ਸਿੱਧੀ ਰੇਖਾ ਵਿੱਚ ਬਿਲਕੁਲ ਫਿੱਟ ਹੁੰਦੇ ਹਨ। ਆਮ ਤੌਰ 'ਤੇ, ਤੁਹਾਨੂੰ ਮੋਟੀਆਂ-ਮੋਟੀਆਂ ਭਵਿੱਖਬਾਣੀਆਂ ਨਾਲ ਸੰਤੁਸ਼ਟ ਹੋਣਾ ਪੈਂਦਾ ਹੈ। ਆਮ ਤੌਰ 'ਤੇ, ਤੁਹਾਡੇ ਕੋਲ ਅੰਕੜਿਆਂ ਦਾ ਇੱਕ ਸਮੂਹ ਹੁੰਦਾ ਹੈ ਜਿਸਦਾ ਸਕੈਟਰ ਪਲਾਟ ਇੱਕ ਸਿੱਧੀ ਰੇਖਾ ਵਿੱਚ "ਫਿੱਟ" ਹੁੰਦਾ ਦਿਖਾਈ ਦਿੰਦਾ ਹੈ। ਇਸਨੂੰ ਸਰਵੋਤਮ ਫਿੱਟ ਦੀ ਰੇਖਾ ਜਾਂ ਘੱਟ-ਵਰਗ ਰੇਖਾ ਕਿਹਾ ਜਾਂਦਾ ਹੈ।

ਸਹਿਯੋਗੀ ਕਸਰਤ

ਜੇਕਰ ਤੁਸੀਂ ਕਿਸੇ ਵਿਅਕਤੀ ਦੀ ਛੋਟੀ ਉਂਗਲ (ਸਭ ਤੋਂ ਛੋਟੀ) ਦੀ ਲੰਬਾਈ ਜਾਣਦੇ ਹੋ, ਤਾਂ ਕੀ ਤੁਹਾਨੂੰ ਲੱਗਦਾ ਹੈ ਕਿ ਤੁਸੀਂ ਉਸ ਵਿਅਕਤੀ ਦੀ ਉਚਾਈ ਦੀ ਭਵਿੱਖਬਾਣੀ ਕਰ ਸਕਦੇ ਹੋ? ਆਪਣੇ ਜਮਾਤ ਦੇ ਵਿਦਿਆਰਥੀਆਂ ਤੋਂ ਅੰਕੜੇ ਇਕੱਠੇ ਕਰੋ (ਛੋਟੀ ਉਂਗਲ ਦੀ ਲੰਬਾਈ, ਇੰਚਾਂ ਵਿੱਚ)। ਸੁਤੰਤਰ ਵੇਰੀਏਬਲ, x, ਛੋਟੀ ਉਂਗਲ ਦੀ ਲੰਬਾਈ ਹੈ ਅਤੇ ਨਿਰਭਰ ਵੇਰੀਏਬਲ, y, ਉਚਾਈ ਹੈ। ਹਰੇਕ ਅੰਕੜੇ ਦੇ ਸਮੂਹ ਲਈ, ਗ੍ਰਾਫ ਪੇਪਰ 'ਤੇ ਬਿੰਦੂਆਂ ਨੂੰ ਪਲਾਟ ਕਰੋ। ਆਪਣੇ ਗ੍ਰਾਫ ਨੂੰ ਕਾਫ਼ੀ ਵੱਡਾ ਬਣਾਓ ਅਤੇ ਇੱਕ ਰੂਲਰ ਦੀ ਵਰਤੋਂ ਕਰੋ। ਫਿਰ "ਅੱਖਾਂ ਨਾਲ" ਇੱਕ ਰੇਖਾ ਖਿੱਚੋ ਜੋ ਅੰਕੜਿਆਂ ਵਿੱਚ "ਫਿੱਟ" ਹੁੰਦੀ ਦਿਖਾਈ ਦਿੰਦੀ ਹੈ। ਆਪਣੀ ਰੇਖਾ ਲਈ, ਦੋ ਸੁਵਿਧਾਜਨਕ ਬਿੰਦੂ ਚੁਣੋ ਅਤੇ ਰੇਖਾ ਦੇ ਢਲਾਨ ਨੂੰ ਲੱਭਣ ਲਈ ਉਹਨਾਂ ਦੀ ਵਰਤੋਂ ਕਰੋ। ਆਪਣੀ ਰੇਖਾ ਨੂੰ y-ਐਕਸਿਸ ਨੂੰ ਕੱਟਣ ਲਈ ਵਧਾ ਕੇ ਰੇਖਾ ਦਾ y-ਇੰਟਰਸੈਪਟ ਲੱਭੋ। ਢਲਾਨਾਂ ਅਤੇ y-ਇੰਟਰਸੈਪਟਾਂ ਦੀ ਵਰਤੋਂ ਕਰਦੇ ਹੋਏ, ਆਪਣੀ "ਸਰਵੋਤਮ ਫਿੱਟ" ਦਾ ਸਮੀਕਰਨ ਲਿਖੋ। ਕੀ ਤੁਹਾਨੂੰ ਲੱਗਦਾ ਹੈ ਕਿ ਹਰ ਕੋਈ ਇੱਕੋ ਸਮੀਕਰਨ ਰੱਖੇਗਾ? ਕਿਉਂ ਜਾਂ ਕਿਉਂ ਨਹੀਂ? ਤੁਹਾਡੇ ਸਮੀਕਰਨ ਅਨੁਸਾਰ, 2.5 ਇੰਚ ਦੀ ਛੋਟੀ ਉਂਗਲ ਦੀ ਲੰਬਾਈ ਲਈ ਅਨੁਮਾਨਿਤ ਉਚਾਈ ਕੀ ਹੈ?

ਉਦਾਹਰਨ 12.6

11 ਅੰਕੜਿਆਂ ਦੇ ਵਿਦਿਆਰਥੀਆਂ ਦੇ ਇੱਕ ਬੇਤਰਤੀਬ ਨਮੂਨੇ ਨੇ ਹੇਠਾਂ ਦਿੱਤੇ ਅੰਕੜੇ ਪੈਦਾ ਕੀਤੇ, ਜਿੱਥੇ x 80 ਵਿੱਚੋਂ ਤੀਜੇ ਇਮਤਿਹਾਨ ਦਾ ਸਕੋਰ ਹੈ, ਅਤੇ y 200 ਵਿੱਚੋਂ ਅੰਤਿਮ ਇਮਤਿਹਾਨ ਦਾ ਸਕੋਰ ਹੈ। ਕੀ ਤੁਸੀਂ ਇੱਕ ਬੇਤਰਤੀਬ ਵਿਦਿਆਰਥੀ ਦੇ ਅੰਤਿਮ ਇਮਤਿਹਾਨ ਦੇ ਸਕੋਰ ਦੀ ਭਵਿੱਖਬਾਣੀ ਕਰ ਸਕਦੇ ਹੋ ਜੇਕਰ ਤੁਸੀਂ ਤੀਜੇ ਇਮਤਿਹਾਨ ਦਾ ਸਕੋਰ ਜਾਣਦੇ ਹੋ?

ਸਤਰ: x (ਤੀਜਾ ਇਮਤਿਹਾਨ ਸਕੋਰ) | y (ਅੰਤਿਮ ਇਮਤਿਹਾਨ ਸਕੋਰ)

ਸਤਰ: 65 | 175

ਸਤਰ: 67 | 133

ਸਤਰ: 71 | 185

ਸਤਰ: 71 | 163

ਸਤਰ: 66 | 126

ਸਤਰ: 75 | 198

ਸਤਰ: 67 | 153

ਸਤਰ: 70 | 163

ਸਤਰ: 71 | 159

ਸਤਰ: 69 | 151

ਸਤਰ: 69 | 159

ਸਾਰਣੀ 12.3: ਤੀਜੇ ਇਮਤਿਹਾਨ ਦੇ ਸਕੋਰਾਂ ਦੇ ਆਧਾਰ 'ਤੇ ਅੰਤਿਮ ਇਮਤਿਹਾਨ ਦੇ ਸਕੋਰਾਂ ਨੂੰ ਦਰਸਾਉਂਦੀ ਸਾਰਣੀ।

ਚਿੱਤਰ 12.9: ਤੀਜੇ ਇਮਤਿਹਾਨ ਦੇ ਸਕੋਰਾਂ ਦੇ ਆਧਾਰ 'ਤੇ ਅੰਤਿਮ ਇਮਤਿਹਾਨ ਦੇ ਸਕੋਰਾਂ ਨੂੰ ਦਰਸਾਉਂਦਾ ਸਕੈਟਰ ਪਲਾਟ।

ਇਸਨੂੰ ਅਜ਼ਮਾਓ 12.6

SCUBA ਡਾਈਵਰਾਂ ਕੋਲ ਵੱਧ ਤੋਂ ਵੱਧ ਡਾਈਵ ਸਮਾਂ ਹੁੰਦਾ ਹੈ ਜਿਸਨੂੰ ਉਹ ਵੱਖ-ਵੱਖ ਡੂੰਘਾਈਆਂ 'ਤੇ ਜਾਣ ਵੇਲੇ ਪਾਰ ਨਹੀਂ ਕਰ ਸਕਦੇ। ਸਾਰਣੀ 12.4 ਵਿੱਚ ਦਿੱਤੇ ਅੰਕੜੇ ਵੱਖ-ਵੱਖ ਡੂੰਘਾਈਆਂ ਨੂੰ ਮਿੰਟਾਂ ਵਿੱਚ ਵੱਧ ਤੋਂ ਵੱਧ ਡਾਈਵ ਸਮੇਂ ਨਾਲ ਦਰਸਾਉਂਦੇ ਹਨ। ਘੱਟੋ-ਘੱਟ ਵਰਗ ਰੀਗਰੈਸ਼ਨ ਲਾਈਨ ਲੱਭਣ ਲਈ ਆਪਣੇ ਕੈਲਕੂਲੇਟਰ ਦੀ ਵਰਤੋਂ ਕਰੋ ਅਤੇ 110 ਫੁੱਟ ਲਈ ਵੱਧ ਤੋਂ ਵੱਧ ਡਾਈਵ ਸਮੇਂ ਦੀ ਭਵਿੱਖਬਾਣੀ ਕਰੋ।

ਸਤਰ: X (ਡੂੰਘਾਈ ਫੁੱਟ ਵਿੱਚ) | Y (ਵੱਧ ਤੋਂ ਵੱਧ ਡਾਈਵ ਸਮਾਂ)

ਸਤਰ: 50 | 80

ਸਤਰ: 60 | 55

row: 70 | 45

row: 80 | 35

row: 90 | 25

row: 100 | 22

The third exam score, x, is the independent variable and the final exam score, y, is the dependent variable. We will plot a regression line that best "fits" the data. If each of you were to fit a line "by eye," you would draw different lines. We can use what is called a least-squares regression line to obtain the best fit line.

Consider the following diagram. Each point of data is of the the form (x, y) and each point of the line of best fit using least-squares linear regression has the form (x, ŷ).

The ŷ is read "y hat" and is the estimated value of y. It is the value of y obtained using the regression line. It is not generally equal to y from data.

The term y0 – ŷ0 = ε0 is called the "error" or residual. It is not an error in the sense of a mistake. The absolute value of a residual measures the vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line.

If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates that actual data value for y.

In the diagram in Figure 12.10, y0 – ŷ0 = ε0 is the residual for the point shown. Here the point lies above the line and the residual is positive.

ε = the Greek letter epsilon

For each data point, you can calculate the residuals or errors, yi - ŷi = εi for i = 1, 2, 3, ..., 11.

Each |ε| is a vertical distance.

For the example about the third exam scores and the final exam scores for the 11 statistics students, there are 11 data points. Therefore, there are 11 ε values. If you square each ε and add, you get

This is called the Sum of Squared Errors (SSE).

Using calculus, you can determine the values of a and b that make the SSE a minimum. When you make the SSE a minimum, you have determined the points that are on the line of best fit. It turns out that the line of best fit has the equation:

where a= y ¯ −b x ¯ a= y ¯ −b x ¯ and b= Σ(x− x ¯ )(y− y ¯ ) Σ (x− x ¯ ) 2 b= Σ(x− x ¯ )(y− y ¯ ) Σ (x− x ¯ ) 2.

The sample means of the x values and the y values are x ¯ x ¯ and y ¯ y ¯, respectively. The best fit line always passes through the point ( x ¯ , y ¯ ) ( x ¯ , y ¯ ).

The slope b can be written as b=r( s y s x ) b=r( s y s x ) where sy = the standard deviation of the y values and sx = the standard deviation of the x values. r is the correlation coefficient, which is discussed in the next section.

Residuals Plots

A residuals plot can be used to help determine if a set of (x, y) data is linearly correlated. For each data point used to create the correlation line, a residual y - ŷ can be calculated, where y is the observed value of the response variable and ŷ is the value predicted by the correlation line. The difference between these values is called the residual. A residuals plot shows the explanatory variable x on the horizontal axis and the residual for that value on the vertical axis. The residuals plot is often shown together with a scatter plot of the data. While a scatter plot of the data should resemble a straight line, a residuals plot should appear random, with no pattern and no outliers. It should also show constant error variance, meaning the residuals should not consistently increase (or decrease) as the explanatory variable x increases.

A residuals plot can be created using StatCrunch or a TI calculator. The plot should appear random. A box plot of the residuals is also helpful to verify that there are no outliers in the data. By observing the scatter plot of the data, the residuals plot, and the box plot of residuals, together with the linear correlation coefficient, we can usually determine if it is reasonable to conclude that the data are linearly correlated.

EXAMPLE:

A shop owner uses a straight-line regression to estimate the number of ice cream cones that would be sold in a day based on the temperature at noon. The owner has data for a 2-year period and chose nine days at random. A scatter plot of the data is shown, together with a residuals plot.

row: Temperature ° F | Ice cream cones sold

70 | 105

85 | 240

65 | 49

72 | 147

80 | 231

61 | 38

75 | 193

78 | 196

68 | 89

10. Table 12.5: ਨੌਂ ਰੈਂਡਮ ਦਿਨਾਂ 'ਤੇ ਵੇਚੀਆਂ ਗਈਆਂ ਆਈਸ ਕਰੀਮ ਕੋਨਾਂ ਦੀ ਗਿਣਤੀ ਅਤੇ ਉਨ੍ਹਾਂ ਦਿਨਾਂ ਦੁਪਹਿਰ ਵੇਲੇ ਦਾ ਤਾਪਮਾਨ ਦਰਸਾਉਣ ਵਾਲੀ ਸਾਰਣੀ।

11. Figure 12.11: ਡਾਟਾ ਦਾ ਸਕੈਟਰ ਪਲਾਟ ਇੱਕ ਸਿੱਧੀ ਰੇਖਾ ਵਰਗਾ ਦਿਖਾਈ ਦਿੰਦਾ ਹੈ।

12. Figure 12.12: ਰੈਜ਼ੀਡਿਊਅਲ ਪਲਾਟ ਬੇਤਰਤੀਬ ਜਾਪਦਾ ਹੈ।

13. ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਲਈ ਘੱਟੋ-ਘੱਟ ਵਰਗ ਮਾਪਦੰਡ

14. ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ ਨੂੰ ਫਿੱਟ ਕਰਨ ਦੀ ਪ੍ਰਕਿਰਿਆ ਨੂੰ ਰੇਖੀ ਰਿਗਰੈਸ਼ਨ ਕਿਹਾ ਜਾਂਦਾ ਹੈ। ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ ਲੱਭਣ ਪਿੱਛੇ ਦਾ ਵਿਚਾਰ ਇਸ ਧਾਰਨਾ 'ਤੇ ਅਧਾਰਤ ਹੈ ਕਿ ਡਾਟਾ ਇੱਕ ਸਿੱਧੀ ਰੇਖਾ ਦੇ ਆਲੇ-ਦੁਆਲੇ ਖਿੰਡਿਆ ਹੋਇਆ ਹੈ। ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ ਲਈ ਮਾਪਦੰਡ ਇਹ ਹੈ ਕਿ ਵਰਗੀਕ੍ਰਿਤ ਗਲਤੀਆਂ (SSE) ਦਾ ਜੋੜ ਘੱਟ ਕੀਤਾ ਜਾਂਦਾ ਹੈ, ਅਰਥਾਤ, ਜਿੰਨਾ ਸੰਭਵ ਹੋ ਸਕੇ ਛੋਟਾ ਕੀਤਾ ਜਾਂਦਾ ਹੈ। ਤੁਸੀਂ ਜੋ ਵੀ ਹੋਰ ਰੇਖਾ ਚੁਣਦੇ ਹੋ, ਉਸ ਵਿੱਚ ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ ਨਾਲੋਂ ਉੱਚਾ SSE ਹੋਵੇਗਾ। ਇਸ ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ ਨੂੰ ਘੱਟੋ-ਘੱਟ ਵਰਗ ਰਿਗਰੈਸ਼ਨ ਰੇਖਾ ਕਿਹਾ ਜਾਂਦਾ ਹੈ।

NOTE

15. ਕੰਪਿਊਟਰ ਸਪ੍ਰੈਡਸ਼ੀਟ, ਅੰਕੜਾ ਸੌਫਟਵੇਅਰ, ਅਤੇ ਬਹੁਤ ਸਾਰੇ ਕੈਲਕੂਲੇਟਰ ਤੇਜ਼ੀ ਨਾਲ ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ ਦੀ ਗਣਨਾ ਕਰ ਸਕਦੇ ਹਨ ਅਤੇ ਗ੍ਰਾਫ ਬਣਾ ਸਕਦੇ ਹਨ। ਜੇ ਹੱਥੀਂ ਗਣਨਾ ਕੀਤੀ ਜਾਵੇ ਤਾਂ ਇਹ ਬਹੁਤ ਮਿਹਨਤ ਵਾਲੀ ਹੁੰਦੀ ਹੈ। ਇਸ ਭਾਗ ਦੇ ਅੰਤ ਵਿੱਚ TI-83, TI-83+, ਅਤੇ TI-84+ ਕੈਲਕੂਲੇਟਰਾਂ ਦੀ ਵਰਤੋਂ ਕਰਕੇ ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ ਲੱਭਣ ਅਤੇ ਸਕੈਟਰਪਲਾਟ ਬਣਾਉਣ ਲਈ ਹਦਾਇਤਾਂ ਦਿੱਤੀਆਂ ਗਈਆਂ ਹਨ।

16. ਤੀਜਾ ਇਮਤਿਹਾਨ ਬਨਾਮ ਅੰਤਿਮ ਇਮਤਿਹਾਨ ਉਦਾਹਰਨ: ਤੀਜਾ-ਇਮਤਿਹਾਨ/ਅੰਤਿਮ-ਇਮਤਿਹਾਨ ਉਦਾਹਰਨ ਲਈ ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ ਦਾ ਗ੍ਰਾਫ ਇਸ ਤਰ੍ਹਾਂ ਹੈ:

17. ਤੀਜਾ-ਇਮਤਿਹਾਨ/ਅੰਤਿਮ-ਇਮਤਿਹਾਨ ਉਦਾਹਰਨ ਲਈ ਘੱਟੋ-ਘੱਟ ਵਰਗ ਰਿਗਰੈਸ਼ਨ ਰੇਖਾ (ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ) ਦਾ ਸਮੀਕਰਨ ਹੈ:

REMINDER

18. ਯਾਦ ਰੱਖੋ, ਪਹਿਲਾਂ ਸਕੈਟਰ ਡਾਇਗਰਾਮ ਪਲਾਟ ਕਰਨਾ ਹਮੇਸ਼ਾ ਮਹੱਤਵਪੂਰਨ ਹੁੰਦਾ ਹੈ। ਜੇ ਸਕੈਟਰ ਪਲਾਟ ਇਹ ਦਰਸਾਉਂਦਾ ਹੈ ਕਿ ਵੇਰੀਏਬਲਜ਼ ਵਿਚਕਾਰ ਇੱਕ ਰੇਖੀ ਸਬੰਧ ਹੈ, ਤਾਂ ਸੈਂਪਲ ਡਾਟਾ ਵਿੱਚ x-ਮੁੱਲਾਂ ਦੇ ਡੋਮੇਨ ਦੇ ਅੰਦਰ y ਲਈ ਭਵਿੱਖਬਾਣੀ ਕਰਨ ਲਈ ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ ਦੀ ਵਰਤੋਂ ਕਰਨਾ ਵਾਜਬ ਹੈ, ਪਰ ਜ਼ਰੂਰੀ ਨਹੀਂ ਕਿ ਉਸ ਡੋਮੇਨ ਤੋਂ ਬਾਹਰ x-ਮੁੱਲਾਂ ਲਈ। ਤੁਸੀਂ ਤੀਜੇ ਇਮਤਿਹਾਨ ਵਿੱਚ 73 ਗ੍ਰੇਡ ਪ੍ਰਾਪਤ ਕਰਨ ਵਾਲੇ ਵਿਦਿਆਰਥੀ ਲਈ ਅੰਤਿਮ ਇਮਤਿਹਾਨ ਦੇ ਸਕੋਰ ਦੀ ਭਵਿੱਖਬਾਣੀ ਕਰਨ ਲਈ ਰੇਖਾ ਦੀ ਵਰਤੋਂ ਕਰ ਸਕਦੇ ਹੋ। ਤੁਹਾਨੂੰ ਤੀਜੇ ਇਮਤਿਹਾਨ ਵਿੱਚ 50 ਗ੍ਰੇਡ ਪ੍ਰਾਪਤ ਕਰਨ ਵਾਲੇ ਵਿਦਿਆਰਥੀ ਲਈ ਅੰਤਿਮ ਇਮਤਿਹਾਨ ਦੇ ਸਕੋਰ ਦੀ ਭਵਿੱਖਬਾਣੀ ਕਰਨ ਲਈ ਰੇਖਾ ਦੀ ਵਰਤੋਂ ਨਹੀਂ ਕਰਨੀ ਚਾਹੀਦੀ, ਕਿਉਂਕਿ 50 ਸੈਂਪਲ ਡਾਟਾ ਵਿੱਚ x-ਮੁੱਲਾਂ ਦੇ ਡੋਮੇਨ ਵਿੱਚ ਨਹੀਂ ਹੈ, ਜੋ ਕਿ 65 ਅਤੇ 75 ਦੇ ਵਿਚਕਾਰ ਹਨ।

UNDERSTANDING SLOPE

19. ਰੇਖਾ ਦਾ ਢਲਾਣ, b, ਦੱਸਦਾ ਹੈ ਕਿ ਵੇਰੀਏਬਲਜ਼ ਵਿੱਚ ਬਦਲਾਅ ਕਿਵੇਂ ਸਬੰਧਤ ਹਨ। ਡਾਟਾ ਦੁਆਰਾ ਦਰਸਾਈ ਗਈ ਸਥਿਤੀ ਦੇ ਸੰਦਰਭ ਵਿੱਚ ਰੇਖਾ ਦੇ ਢਲਾਣ ਦੀ ਵਿਆਖਿਆ ਕਰਨਾ ਮਹੱਤਵਪੂਰਨ ਹੈ। ਤੁਹਾਨੂੰ ਸਾਦੇ ਅੰਗਰੇਜ਼ੀ ਵਿੱਚ ਢਲਾਣ ਦੀ ਵਿਆਖਿਆ ਕਰਨ ਵਾਲਾ ਇੱਕ ਵਾਕ ਲਿਖਣ ਦੇ ਯੋਗ ਹੋਣਾ ਚਾਹੀਦਾ ਹੈ।

20. ਢਲਾਣ ਦੀ ਵਿਆਖਿਆ: ਸਭ ਤੋਂ ਵਧੀਆ ਫਿੱਟ ਰੇਖਾ ਦਾ ਢਲਾਣ ਸਾਨੂੰ ਦੱਸਦਾ ਹੈ ਕਿ ਸੁਤੰਤਰ (x) ਵੇਰੀਏਬਲ ਵਿੱਚ ਹਰ ਇੱਕ ਇਕਾਈ ਵਾਧੇ ਲਈ, ਨਿਰਭਰ ਵੇਰੀਏਬਲ (y) ਔਸਤਨ ਕਿਵੇਂ ਬਦਲਦਾ ਹੈ।

21. ਤੀਜਾ ਇਮਤਿਹਾਨ ਬਨਾਮ ਅੰਤਿਮ ਇਮਤਿਹਾਨ ਉਦਾਹਰਨ ਢਲਾਣ: ਰੇਖਾ ਦਾ ਢਲਾਣ b = 4.83 ਹੈ। ਵਿਆਖਿਆ: ਤੀਜੇ ਇਮਤਿਹਾਨ ਦੇ ਸਕੋਰ ਵਿੱਚ ਇੱਕ-ਅੰਕ ਦੇ ਵਾਧੇ ਲਈ, ਅੰਤਿਮ ਇਮਤਿਹਾਨ ਦਾ ਸਕੋਰ ਔਸਤਨ 4.83 ਅੰਕ ਵਧਦਾ ਹੈ।

22. TI-83, 83+, 84, 84+ ਕੈਲਕੂਲੇਟਰ ਦੀ ਵਰਤੋਂ ਕਰਨਾ

23. ਰੇਖੀ ਰਿਗਰੈਸ਼ਨ ਟੀ ਟੈਸਟ ਦੀ ਵਰਤੋਂ ਕਰਨਾ: LinRegTTest

24. STAT ਲਿਸਟ ਐਡੀਟਰ ਵਿੱਚ, ਲਿਸਟ L1 ਵਿੱਚ X ਡਾਟਾ ਅਤੇ ਲਿਸਟ L2 ਵਿੱਚ Y ਡਾਟਾ ਦਾਖਲ ਕਰੋ, ਇਸ ਤਰ੍ਹਾਂ ਜੋੜਿਆ ਗਿਆ ਹੈ ਕਿ ਸੰਬੰਧਿਤ (x,y) ਮੁੱਲ ਸੂਚੀਆਂ ਵਿੱਚ ਇੱਕ ਦੂਜੇ ਦੇ ਅੱਗੇ ਹਨ। (ਜੇਕਰ ਮੁੱਲਾਂ ਦੀ ਕੋਈ ਖਾਸ ਜੋੜੀ ਦੁਹਰਾਈ ਜਾਂਦੀ ਹੈ, ਤਾਂ ਇਸਨੂੰ ਜਿੰਨੀ ਵਾਰ ਡਾਟਾ ਵਿੱਚ ਦਿਖਾਈ ਦਿੰਦੀ ਹੈ, ਓਨੀ ਵਾਰ ਦਾਖਲ ਕਰੋ।)

On the STAT TESTS menu, scroll down with the cursor to select the LinRegTTest. (Be careful to select LinRegTTest, as some calculators may also have a different item called LinRegTInt.)

On the LinRegTTest input screen enter: Xlist: L1 ; Ylist: L2 ; Freq: 1

On the next line, at the prompt β or ρ, highlight "≠ 0" and press ENTER

Leave the line for "RegEq:" blank

Highlight Calculate and press ENTER.

The output screen contains a lot of information. For now we will focus on a few items from the output, and will return later to the other items. The second line says y = a + bx. Scroll down to find the values a = –173.513, and b = 4.8273; the equation of the best fit line is ŷ = –173.51 + 4.83x The two items at the bottom are r2 = 0.43969 and r = 0.663. For now, just note where to find these values; we will discuss them in the next two sections.

Graphing the Scatterplot and Regression Line

We are assuming your X data is already entered in list L1 and your Y data is in list L2

Press 2nd STATPLOT ENTER to use Plot 1

On the input screen for PLOT 1, highlight On, and press ENTER

For TYPE: highlight the very first icon which is the scatterplot and press ENTER

Indicate Xlist: L1 and Ylist: L2

For Mark: it does not matter which symbol you highlight.

Press the ZOOM key and then the number 9 (for menu item "ZoomStat") ; the calculator will fit the window to the data

To graph the best-fit line, press the "Y=" key and type the equation –173.5 + 4.83X into equation Y1. (The X key is immediately left of the STAT key). Press ZOOM 9 again to graph it.

Optional: If you want to change the viewing window, press the WINDOW key. Enter your desired window using Xmin, Xmax, Ymin, Ymax

NOTE

Another way to graph the line after you create a scatter plot is to use LinRegTTest.

Make sure you have done the scatter plot. Check it on your screen.

Go to LinRegTTest and enter the lists.

At RegEq: press VARS and arrow over to Y-VARS. Press 1 for 1:Function. Press 1 for 1:Y1. Then arrow down to Calculate and do the calculation for the line of best fit.

Press Y = (you will see the regression equation).

Press GRAPH. The line will be drawn."

The Correlation Coefficient r

Besides looking at the scatter plot and seeing that a line seems reasonable, how can you tell if the line is a good predictor? Use the correlation coefficient as another indicator (besides the scatterplot) of the strength of the relationship between x and y.

ਕਾਰਲ ਪੀਅਰਸਨ ਦੁਆਰਾ 1900 ਦੇ ਦਹਾਕੇ ਦੇ ਸ਼ੁਰੂ ਵਿੱਚ ਵਿਕਸਤ ਕੀਤਾ ਗਿਆ ਸਹਿ-ਸੰਬੰਧ ਗੁਣਾਂਕ, r, ਸੰਖਿਆਤਮਕ ਹੈ ਅਤੇ ਸੁਤੰਤਰ ਵੇਰੀਏਬਲ x ਅਤੇ ਨਿਰਭਰ ਵੇਰੀਏਬਲ y ਵਿਚਕਾਰ ਰੇਖੀ ਸਬੰਧ ਦੀ ਤਾਕਤ ਅਤੇ ਦਿਸ਼ਾ ਦਾ ਮਾਪ ਪ੍ਰਦਾਨ ਕਰਦਾ ਹੈ।

ਸਹਿ-ਸੰਬੰਧ ਗੁਣਾਂਕ ਦੀ ਗਣਨਾ ਇਸ ਤਰ੍ਹਾਂ ਕੀਤੀ ਜਾਂਦੀ ਹੈ

ਜਿੱਥੇ n = ਡਾਟਾ ਪੁਆਇੰਟਾਂ ਦੀ ਸੰਖਿਆ ਹੈ।

ਜੇਕਰ ਤੁਸੀਂ x ਅਤੇ y ਵਿਚਕਾਰ ਰੇਖੀ ਸਬੰਧ ਦਾ ਸ਼ੱਕ ਕਰਦੇ ਹੋ, ਤਾਂ r ਮਾਪ ਸਕਦਾ ਹੈ ਕਿ ਰੇਖੀ ਸਬੰਧ ਕਿੰਨਾ ਮਜ਼ਬੂਤ ਹੈ।

r ਦਾ ਮੁੱਲ ਸਾਨੂੰ ਕੀ ਦੱਸਦਾ ਹੈ:

r ਦਾ ਮੁੱਲ ਹਮੇਸ਼ਾ –1 ਅਤੇ +1 ਦੇ ਵਿਚਕਾਰ ਹੁੰਦਾ ਹੈ: –1 ≤ r ≤ 1।

r ਦਾ ਆਕਾਰ x ਅਤੇ y ਵਿਚਕਾਰ ਰੇਖੀ ਸਬੰਧ ਦੀ ਤਾਕਤ ਨੂੰ ਦਰਸਾਉਂਦਾ ਹੈ। r ਦੇ –1 ਜਾਂ +1 ਦੇ ਨੇੜੇ ਦੇ ਮੁੱਲ x ਅਤੇ y ਵਿਚਕਾਰ ਇੱਕ ਮਜ਼ਬੂਤ ਰੇਖੀ ਸਬੰਧ ਨੂੰ ਦਰਸਾਉਂਦੇ ਹਨ।

ਜੇਕਰ r = 0 ਹੈ, ਤਾਂ ਸੰਭਵ ਤੌਰ 'ਤੇ ਕੋਈ ਰੇਖੀ ਸਹਿ-ਸੰਬੰਧ ਨਹੀਂ ਹੈ। ਹਾਲਾਂਕਿ, ਸਕੈਟਰਪਲਾਟ ਨੂੰ ਦੇਖਣਾ ਮਹੱਤਵਪੂਰਨ ਹੈ, ਕਿਉਂਕਿ ਕਰਵਡ ਜਾਂ ਹਰੀਜ਼ੋਂਟਲ ਪੈਟਰਨ ਵਾਲੇ ਡਾਟਾ ਦਾ ਸਹਿ-ਸੰਬੰਧ 0 ਹੋ ਸਕਦਾ ਹੈ।

ਜੇਕਰ r = 1 ਹੈ, ਤਾਂ ਸੰਪੂਰਨ ਸਕਾਰਾਤਮਕ ਸਹਿ-ਸੰਬੰਧ ਹੈ। ਜੇਕਰ r = –1 ਹੈ, ਤਾਂ ਸੰਪੂਰਨ ਨਕਾਰਾਤਮਕ ਸਹਿ-ਸੰਬੰਧ ਹੈ। ਇਹਨਾਂ ਦੋਵਾਂ ਮਾਮਲਿਆਂ ਵਿੱਚ, ਸਾਰੇ ਅਸਲ ਡਾਟਾ ਪੁਆਇੰਟ ਇੱਕ ਸਿੱਧੀ ਲਾਈਨ 'ਤੇ ਪੈਂਦੇ ਹਨ। ਬੇਸ਼ੱਕ, ਅਸਲ ਦੁਨੀਆਂ ਵਿੱਚ, ਇਹ ਆਮ ਤੌਰ 'ਤੇ ਨਹੀਂ ਵਾਪਰੇਗਾ।

r ਦਾ ਚਿੰਨ੍ਹ ਸਾਨੂੰ ਕੀ ਦੱਸਦਾ ਹੈ

r ਦਾ ਸਕਾਰਾਤਮਕ ਮੁੱਲ ਦਾ ਮਤਲਬ ਹੈ ਕਿ ਜਦੋਂ x ਵਧਦਾ ਹੈ, ਤਾਂ y ਵਧਣ ਲੱਗਦਾ ਹੈ ਅਤੇ ਜਦੋਂ x ਘਟਦਾ ਹੈ, ਤਾਂ y ਘਟਣ ਲੱਗਦਾ ਹੈ (ਸਕਾਰਾਤਮਕ ਸਹਿ-ਸੰਬੰਧ)।

r ਦਾ ਨਕਾਰਾਤਮਕ ਮੁੱਲ ਦਾ ਮਤਲਬ ਹੈ ਕਿ ਜਦੋਂ x ਵਧਦਾ ਹੈ, ਤਾਂ y ਘਟਣ ਲੱਗਦਾ ਹੈ ਅਤੇ ਜਦੋਂ x ਘਟਦਾ ਹੈ, ਤਾਂ y ਵਧਣ ਲੱਗਦਾ ਹੈ (ਨਕਾਰਾਤਮਕ ਸਹਿ-ਸੰਬੰਧ)।

r ਦਾ ਚਿੰਨ੍ਹ ਬੈਸਟ-ਫਿਟ ਲਾਈਨ ਦੇ ਢਲਾਣ, b, ਦੇ ਚਿੰਨ੍ਹ ਦੇ ਸਮਾਨ ਹੈ।

NOTE

ਚਿੱਤਰ 12.15: (a) ਇੱਕ ਸਕੈਟਰ ਪਲਾਟ ਜੋ ਸਕਾਰਾਤਮਕ ਸਹਿ-ਸੰਬੰਧ ਵਾਲਾ ਡਾਟਾ ਦਿਖਾਉਂਦਾ ਹੈ। 0 < r < 1 (b) ਇੱਕ ਸਕੈਟਰ ਪਲਾਟ ਜੋ ਨਕਾਰਾਤਮਕ ਸਹਿ-ਸੰਬੰਧ ਵਾਲਾ ਡਾਟਾ ਦਿਖਾਉਂਦਾ ਹੈ। –1 < r < 0 (c) ਇੱਕ ਸਕੈਟਰ ਪਲਾਟ ਜੋ ਜ਼ੀਰੋ ਸਹਿ-ਸੰਬੰਧ ਵਾਲਾ ਡਾਟਾ ਦਿਖਾਉਂਦਾ ਹੈ। r = 0

r ਲਈ ਫਾਰਮੂਲਾ ਭਿਆਨਕ ਲੱਗਦਾ ਹੈ। ਹਾਲਾਂਕਿ, ਕੰਪਿਊਟਰ ਸਪ੍ਰੈਡਸ਼ੀਟ, ਸਟੈਟਿਸਟੀਕਲ ਸੌਫਟਵੇਅਰ, ਅਤੇ ਕਈ ਕੈਲਕੂਲੇਟਰ ਜਲਦੀ ਨਾਲ r ਦੀ ਗਣਨਾ ਕਰ ਸਕਦੇ ਹਨ। ਸਹਿ-ਸੰਬੰਧ ਗੁਣਾਂਕ r TI-83, TI-83+, ਜਾਂ TI-84+ ਕੈਲਕੂਲੇਟਰ 'ਤੇ LinRegTTest ਲਈ ਆਉਟਪੁੱਟ ਸਕ੍ਰੀਨਾਂ ਵਿੱਚ ਸਭ ਤੋਂ ਹੇਠਾਂ ਵਾਲਾ ਆਈਟਮ ਹੈ (ਪਿਛਲੇ ਭਾਗ ਵਿੱਚ ਹਦਾਇਤਾਂ ਲਈ ਦੇਖੋ)।

ਨਿਰਧਾਰਨ ਦਾ ਗੁਣਾਂਕ

ਵੇਰੀਏਬਲ r2 ਨੂੰ ਨਿਰਧਾਰਨ ਦਾ ਗੁਣਾਂਕ ਕਿਹਾ ਜਾਂਦਾ ਹੈ ਅਤੇ ਇਹ ਸਹਿ-ਸੰਬੰਧ ਗੁਣਾਂਕ ਦਾ ਵਰਗ ਹੈ, ਪਰ ਆਮ ਤੌਰ 'ਤੇ ਦਸ਼ਮਲਵ ਰੂਪ ਦੀ ਬਜਾਏ ਪ੍ਰਤੀਸ਼ਤ ਵਜੋਂ ਦੱਸਿਆ ਜਾਂਦਾ ਹੈ। ਇਸਦਾ ਡਾਟਾ ਦੇ ਸੰਦਰਭ ਵਿੱਚ ਇੱਕ ਵਿਆਖਿਆ ਹੈ:

r2, ਜਦੋਂ ਪ੍ਰਤੀਸ਼ਤ ਵਜੋਂ ਪ੍ਰਗਟ ਕੀਤਾ ਜਾਂਦਾ ਹੈ, ਤਾਂ ਨਿਰਭਰ (ਅਨੁਮਾਨਿਤ) ਵੇਰੀਏਬਲ y ਵਿੱਚ ਪਰਿਵਰਤਨ ਦਾ ਪ੍ਰਤੀਸ਼ਤ ਦਰਸਾਉਂਦਾ ਹੈ ਜੋ ਸੁਤੰਤਰ (ਵਿਆਖਿਆਤਮਕ) ਵੇਰੀਏਬਲ x ਵਿੱਚ ਪਰਿਵਰਤਨ ਦੁਆਰਾ ਰਿਗਰੈਸ਼ਨ (ਬੈਸਟ-ਫਿਟ) ਲਾਈਨ ਦੀ ਵਰਤੋਂ ਕਰਕੇ ਸਮਝਾਇਆ ਜਾ ਸਕਦਾ ਹੈ।

1 – r2, ਜਦੋਂ ਪ੍ਰਤੀਸ਼ਤ ਵਜੋਂ ਪ੍ਰਗਟ ਕੀਤਾ ਜਾਂਦਾ ਹੈ, ਤਾਂ y ਵਿੱਚ ਪਰਿਵਰਤਨ ਦਾ ਪ੍ਰਤੀਸ਼ਤ ਦਰਸਾਉਂਦਾ ਹੈ ਜੋ ਰਿਗਰੈਸ਼ਨ ਲਾਈਨ ਦੀ ਵਰਤੋਂ ਕਰਕੇ x ਵਿੱਚ ਪਰਿਵਰਤਨ ਦੁਆਰਾ ਸਮਝਾਇਆ ਨਹੀਂ ਗਿਆ ਹੈ। ਇਸਨੂੰ ਰਿਗਰੈਸ਼ਨ ਲਾਈਨ ਦੇ ਆਲੇ-ਦੁਆਲੇ ਦੇਖੇ ਗਏ ਡਾਟਾ ਪੁਆਇੰਟਾਂ ਦੇ ਖਿੰਡਾਅ ਵਜੋਂ ਦੇਖਿਆ ਜਾ ਸਕਦਾ ਹੈ।

ਪਿਛਲੇ ਭਾਗ ਵਿੱਚ ਪੇਸ਼ ਕੀਤੇ ਗਏ ਤੀਜੇ ਪ੍ਰੀਖਿਆ/ਅੰਤਮ ਪ੍ਰੀਖਿਆ ਉਦਾਹਰਨ 'ਤੇ ਵਿਚਾਰ ਕਰੋ

ਬੈਸਟ ਫਿਟ ਲਾਈਨ ਹੈ: ŷ = –173.51 + 4.83x

ਸਹਿ-ਸੰਬੰਧ ਗੁਣਾਂਕ r = 0.6631 ਹੈ

ਨਿਰਧਾਰਨ ਦਾ ਗੁਣਾਂਕ r2 = 0.66312 = 0.4397 ਹੈ

ਇਸ ਉਦਾਹਰਨ ਦੇ ਸੰਦਰਭ ਵਿੱਚ r2 ਦੀ ਵਿਆਖਿਆ:

ਅੰਤਿਮ ਪ੍ਰੀਖਿਆ ਦੇ ਅੰਕਾਂ ਵਿੱਚ ਲਗਭਗ 44% ਪਰਿਵਰਤਨ (0.4397 ਲਗਭਗ 0.44 ਹੈ) ਤੀਜੀ ਪ੍ਰੀਖਿਆ ਦੇ ਅੰਕਾਂ ਵਿੱਚ ਪਰਿਵਰਤਨ ਦੁਆਰਾ, ਸਭ ਤੋਂ ਢੁਕਵੀਂ ਰੇਖਾ-ਗਣਨਾ ਰੇਖਾ ਦੀ ਵਰਤੋਂ ਕਰਕੇ, ਸਮਝਾਇਆ ਜਾ ਸਕਦਾ ਹੈ।

ਇਸ ਲਈ, ਅੰਤਿਮ ਪ੍ਰੀਖਿਆ ਦੇ ਅੰਕਾਂ ਵਿੱਚ ਲਗਭਗ 56% ਪਰਿਵਰਤਨ (1 – 0.44 = 0.56) ਤੀਜੀ ਪ੍ਰੀਖਿਆ ਦੇ ਅੰਕਾਂ ਵਿੱਚ ਪਰਿਵਰਤਨ ਦੁਆਰਾ, ਸਭ ਤੋਂ ਢੁਕਵੀਂ ਰੇਖਾ-ਗਣਨਾ ਰੇਖਾ ਦੀ ਵਰਤੋਂ ਕਰਕੇ, ਸਮਝਾਇਆ ਨਹੀਂ ਜਾ ਸਕਦਾ। (ਇਹ ਰੇਖਾ ਦੇ ਆਲੇ-ਦੁਆਲੇ ਬਿੰਦੂਆਂ ਦੇ ਖਿੰਡਾਅ ਵਜੋਂ ਦੇਖਿਆ ਜਾਂਦਾ ਹੈ।)