We studied  about the estimation of population parameter from the sample statistics in the previous chapter  point estimate and interval estimate. We also studied about how point and interval estimates behaved at different confidential interval such as 95% and 99.7%. Also noted  how the population parameter lies in between mean ± z*SE of means distribution. Now it becomes necessary to test if our estimation of population parameter is to be accepted or rejected based on hypothesis testing. Hypothesis testing enables you to take decision about credibility/accuracy of the estimated population parameter.

  1. We assume certain value to the population parameter.
  2. To Test the validity of our assumption we collect samples from the population and find sample statistics
  3. Now we find the difference between the assumed value of population parameter and the actual value of sample statistic.
  4. Now we judge whether the difference is significant or not.
  5. The smaller the difference the greater the chance of concluding the fact that the assumed value of population parameter is correct. Hypothesis can be accepted.
  6. Instead, we may say that there is no sufficient statistical evidence to reject our hypothesis. The large the difference the smaller the chance of accepting the assumption made. So, we reject the hypothesis we made

Activities Involved:

  1. Assumption of two types of hypotheses. A) Null Hypothesis (H0)  and B) Alternative Hypothesis (H1)
  2. Applying different types of test statistics.
  3. Identification of test for a given problem.
  4. Identification of type of errors we encounter

What is Hypothesis?

Nothing but an assumption.

Example: Sterlite Copper’s Industrial Unit of Tuticorin is reported to have dumped copper slag in the river. Because of that the ground water has been polluted and contaminated. People started protesting against the pollution that the unit caused. people in that area started feeling ill.

 Assumption: People in that area are suffering from illness due to the pollution caused by that unit.

Question here is whether this is to be accepted or rejected by Pollution Control Board of India .

Activity: Small samples are to be collected and tested to prove that the hypothesis is true or wrong. Conducting Survey to get the feeling of entire Tuticorin people is impossible. So, we conduct a random sample survey. Using sample statistics, we may arrive at the conclusion whether the Hypothesis is true or wrong. Hypothesis testing is about making inference of the population parameter using sample statistics.

Types of Hypotheses:

Assumptions we make:

Example: Parameter – mean (population mu)

Null Hypothesis:

Before entering into sampling techniques, we assume the value of population parameter (Hypothesis value of parameter).Null Hypothesis is denoted by H0. Say the population mean income  is Rs.6000. H0: µ=6000

Alternative Hypothesis:

  1. If Null Hypothesis is not true what course of action we have  to do and to conclude.
  2. For this alternative hypothesis are to be made.

Another example: Proportion (population – p)

Suppose if we need to test the success rate of a particular treatment against COVID-19 we make a null hypothesis for the success rate ’p’(for the test value of 0.99 ) as

Level of Significance:

  1. Under Testing of Hypothesis, the level of significance plays a vital role in determining the  decision whether to accept or reject Null Hypothesis.
  2. The question whether the difference between sample statistic and the assumed value of population parameter(mean, variance, proportion)  is significant or not determines our acceptance  or rejection of Null hypothesis.
  3. Under estimation the confidence level  points out the percentage of sample statistic that falls within the confidence limits.
  4. Under Testing of  Hypothesis, the confidence level points the percentage of sample statistic that falls out-side the confidence limits
  5. Remember, even if the sample statistic falls within the non-shaded region, it does not prove that our Null Hypothesis is true/correct. It does not provide any statistical evidence to  reject our Null Hypothesis. Instead of saying we do not reject the hypothesis we say we accept. This is the standard practice.
  6. Actual acceptance is impossible  as we  cannot find the  population parameter in practice.
  7. The acceptance  and rejection region of  sample are shown below.

Selection of Significance Level:

You can have any number of selection levels. However, in  practice the  following two significance levels are used.

  1. 5% Level of Significance: Under this the possibility of rejecting Null Hypothesis is higher when Null Hypothesis is true.
  2. 1% Level of Significance: We will never accept the Null Hypothesis when Null Hypothesis is not true under 1% Level  

Possible situations Under Testing of Hypothesis:

  1. Null Hypothesis is true and test result also recommends us to accept  null hypothesis. This means we have taken a correct decision.
  2. Null hypothesis is false but test  result recommends  us to accept the null hypothesis. It  is a wrong decision. We should have rejected Null Hypothesis. But there is no statistical evidence to reject the null. So, we accept. Here TYPE II error creeps in. This is called producers risk and it denoted by β. 1−β is called the power of the test.

Reject on Two Conditions

  1. Null Hypothesis is false and test result also recommends us to reject null hypothesis. This means we have taken a correct decision.
  2. Null hypothesis is True but test result recommends us to reject  the null hypothesis. It is a wrong decision. We should have accepted the Null Hypothesis. But there is no statistical evidence to accept the null. So, we reject. Here TYPE I error creeps in. This is called consumers risk and it denoted by α

Errors that creep in:

Selection of appropriate Sampling Distribution:

We use of two types of distributions for testing purpose:

  1. Normal Distribution – Z-Distribution
  2. Students t-Distribution

Selection is based on two factors

  1. a) Size of the sample
  2. b) Standard Deviation of Population

 Rules to be applied:

Example:

Selection of Distribution and Table for the above problem:

  • Apply Normal Distribution
  • use z-tables to calculate area value of sample mean

Example 2:

  • the sample size is less than 30
  • The standard deviation of population is known(10)

Selection of Distribution and tables for calculation of sample mean

  • apply t-Distribution
  • Use t-tables to calculate area value of sample mean

What type of distribution is to be used?