Pages

Tuesday, July 9, 2019

What is Hypothesis Testing?

Hypothesis Testing



         We have considered many problems of the point estimate and interval estimate involving unknown parameters such as means and proportions. In this chapter, we must deal with a decision-making process in which we must come to some conclusion about unknown population parameters the so-called testing of hypothesis. A hypothesis is widely used in science and research. In this subject, a hypothesis is used to mean a statistical hypothesis. A statement, an assertion, conjecture, or an assumption concerning one or more population. 

          In hypothesis testing, the researcher must define the population under study, state the particular hypotheses to be investigated, giving the significance level, select a sample statistic, perform the required test and make conclusions. There are two specific statistical tests for hypothesis testing on means: the z test and the t-test. 

          Hypothesis testing is a technique for determining whether a research hypothesis is justified in the light of observed data and is considered the primary tool for making decisions based on statistical analysis.

Basic Concepts of Types of Hypothesis

Null Hypothesis

    
          This is a statement that specifies some aspects of a population that is known or assume to be true. It is always stated as an equality so as to specify an exact value of the parameter. This means that the null hypothesis statement is describe of “no significant difference, no significant effect, no significant relationship, no significant association” on whatever phenomena the researcher wants to test. Usually, the null hypothesis is denoted by Ho. 

Alternative Hypothesis


          This is called the research hypothesis, it is a statement that contradicts the null hypothesis, that is, a statement of “significance”. Usually, the alternative hypothesis is denoted by Ha or .


Type of Tailed Test

One-Tailed test or Directional test

          A one-tailed test indicates that the null hypothesis should be rejected when the computed test value is in the critical region on one side of the parameter. A one-tailed test is either right-tailed when the inequality in the alternative hypothesis is greater than (>) or left-tailed when the direction is less than (<).

Two-Tailed test of Non-directional test

          In this test, the null hypothesis should be rejected when the test value is in either of the two critical regions or the rejection region. 

Critical Region or Rejection Region


          This is a portion of the sampling distribution of the test statistic which comprises the set of all values of the test statistic that would justify rejecting the null hypothesis in favor of its alternative. 

Critical Value


          This is a value selected from a table for the appropriate test statistic. This determines or divides the acceptance and the rejection region. 

Test Statistic


This is a value obtained after computed from sample data.

Level of Significance

          This is a probability value which specifies the risk of incorrectly rejecting the null hypothesis when it is true. The level of significance is predetermined or set by the researcher beforehand. This is symbolized by α (a Greek letter alpha).



Types of Error in Hypothesis testing


          Since the decision on the null hypothesis is based on sample data, there is a probability of drawing incorrect conclusions from the evidence available. Researchers may commit these types of errors in the decision of rejecting or accepting the null hypothesis. These are the Type I error and Type II error. 

Type I Error


This is the error committed in rejecting the null hypothesis Ho when, in fact, it is true.

Type II Error


This is the error committed in accepting the null hypothesis Ho when, in fact, it is falls.



          The classical approach to hypothesis testing uses the critical value criterion for significance. These critical values are also called tabular values since they are obtained from statistical tables of various probability distribution such as the Table of t-distribution and the z – distribution to name a few.

Steps in Classical Hypothesis Testing


Step 1. State the null hypothesis, Ho and alternative hypothesis Ha.

Step 2. Set the level of significance of the test, α.

Step 3. Select an appropriate test statistic and establish the critical region.

Step 4. Compute the value of the test statistic.

Step 5. Make a decision concerning the null hypothesis.

Step 6. Draw a conclusion concerning the population.


Video: hypothesis testing

https://www.youtube.com - hypothesis testing

For more details and illustration about hypothesis testing go to my next discussion on the illustrations on hypothesis testing.

Areas Under the Normal Distribution Curve

Figure 1. Normal Distribution Table

Finding Areas Under the Normal Distribution Curve


          In finding areas under the normal curve, used table in Figure 1. This table provides areas under the standard normal curve below a specified value z. The first column in table of normal curve refers to the values of the standard score or z-score ranging from 0.0 to +3.0 standard deviations from the mean. The first-row heading specifies the value of z-score in two decimal places while the entries in the table are the areas below a given value of z, in four-decimal-place accuracy.

Example 1: Determine the area under the standard normal curve of z = 2.49.

Solution: First go down the left-hand column, labeled z to “2.4”. Then go cross that row until the column 0.09. That intersection is the area under the standard normal curve of z = 2.49, that is, 0.4936.

Example 2: Determine the area under the standard normal curve of z = -1.38.

Solution: Apply the property of the normal curve that is, the normal curve is symmetrical about the mean. Thus, the area under the normal curve of z = -1.38 is just the same in finding areas under the normal curve of z = 1.38. First go down the left-hand column, labeled z to “1.3”. Then go cross that row until the column 0.08. The intersection is the area under the standard normal curve of z = -1.38, that is, 0.4162.

Example 3: Find the area under the standard normal curve to the right of z = 0.25.

Solution: First go down the left-hand column, labeled z to “0.2”. Then go cross that row until the column 0.05. That intersection is the area under the standard normal curve from 0 to z = 0.25 which is 0.0987, thus, the area under the standard normal curve to the right of z = 0.25 is the difference between 0.5 and obtained area or 0.5 – 0.0987 = 0.4013.

Example 4: Find the area under the standard normal curve to the left of z = 1.73.

Solution: First go down the left-hand column, labeled z to “1.7”. Then go cross that row until the column 0.03. That intersection is the area under the standard normal curve from 0 to z = 1.73 which is 0.4582, thus, the area under the standard normal curve to the left of z = 1.73 is the sum between 0.5 and obtained area or 0.5 + 0.4582 = 0.9582.

Example 5: Find the area under the standard normal curve to the left of z = -1.08.

Solution: First go down the left-hand column, labeled z to “1.0”. Then go cross that row until the column 0.08. That intersection is the area under the standard normal curve from 0 to z = -1.08 which is 0.3599, thus, the area under the standard normal curve to the left of z = -1.08 is the difference between 0.5 and obtained area or 0.5 – 0.3599 = 0.1401.

Example 6: Find the area under the standard normal curve above z = -2.46.

Solution: First go down the left-hand column, labeled z to “2.4”. Then go cross that row until the column 0.06. That intersection is the area under the standard normal curve from 0 to z = -2.46 which is 0.4931, thus, the area under the standard normal curve to the left of z = -2.46 is the sum between 0.5 and obtained area or 0.5 + 0.4931 = 0.9931.

Example 7: Find the area under the standard normal curve between z = 0.82 and z = 2.72.

Solution: The area between 0 and 0.82 is 0.2939 and the area between 0 and 2.72 is 0.4967. Thus, the required area between the z values 0.82 and 2.72 is the difference between these two areas or 0.4967 – 0.2939 = 0.2028.

Example 8: Determine the area under the standard normal curve between z = -1.57 and z = 1.63.

Solution: The area between -1.57 and 0 is 0.4418 and the area between 0 and 1.63 is 0.4484. Thus, the required area between the z values -1.57 and 1.63 is the sum between these two areas or 0.4484 – 0.4418 = 0.8902.

Standardizing a Normal Curve

          Most of the time, the probabilities you are interested in will involve a random variable let say x, a normal random variable with mean μ and standard deviation σ. You must then standardize variable of interest, writing it as the equivalent form in terms of z-score the standard normal random variable. Once this is done, the probability of interest is the area that you find using the standard normal probability distribution. The formula to convert raw score into standard score is:




Example 1: For a continuous random variable that has a normal distribution with a mean of 20 and a standard deviation of 4, find the area under the normal curve from 19 and 26.


Now, the area between -0.25 and 0 is 0.0987 and the area between 0 and 1.5 is 0.3531. Therefore, the area under the normal curve between 19 and 26 is sum between these two areas or 0.3531 + 0.0987 = 0.4518.


Example 3: A brisk walk at 4 miles per hour burns an average of 300 calories per hour. If the standard deviation of the distribution is 8 calories, find the probability that a person who walks one hour at the rate of 4 miles per hour will burn the following calories. Assume the variable to be randomly normally distributed.

a. more than 280 calories.

b. less than 294 calories.

c. between 278 and 318 calories.















The area from -0.75 and 0 is 0.2734. Thus, the probability that a person who walks one hour at the rate of 4 miles per hour will burn less than 294 calories is the difference between 0.5 and the area or 0.5 – 0.2734 = 0.2266. Therefore, there is a 22.66% probability that a person who walks one hour at the rate of 4 miles per hour will burn less than 294 calories.
















The area between -2.75 and 0 is 0.4970 and the area between 0 and 2.25 is 0.4878., the probability is the sum between these two areas or 0.4878 + 0.4970 = 0.9848. Therefore, around 98.48% probability that a person who walks one hour at the rate of 4 miles per hour will burn between 278 and 318 calories respectively.

Fill in the details below with your comments or suggestions. The author is delighted to get your feedback. Statistics geeks, have fun reading.

What is Normal Distribution

The Normal Distribution

          Random variables can either be discrete or continuous. Continuous random variables can assume all values between any two given values of the variables. Many continuous random variables have distributions that are bell-shaped and are called approximately normally distributed variables. 

         The normal distribution, sometimes referred to as the Z distribution or Gaussian distribution, is considered the most important continuous probability distributions in statistics. This was named after a German mathematician who extensively studied the normal distribution. The graph of a normal distribution is called the normal curve. Variables that are normally distributed will have a mean and a standard deviation that may take on any value, but they share the following properties of the normal curve:

Figure 1. Normal Distribution Curve


1. The three measures of central tendency are located at the center of the normal curve. 

2. The normal distribution curve is bell-shaped. 

3. The normal distribution is unimodal. 

4. The curve is symmetrical about the center. 

5. The curve is continuous. 

6. The curve is asymptotic with respect to the horizontal. 

7. The total area above the horizontal axis under the normal distribution curve is equal to 1 or 100%. 

8. The points of inflection of the curve occur at points plus or minus one standard deviation unit above or below the mean. 

9. Roughly 68% of the area of the curve falls within the limits plus or minus one standard deviation unit from the mean.

Friday, July 5, 2019

Numerical Methods of Summarizing Data

Percentages & Proportions

These are the most commonly used numerical measures for summarizing data. The formulas are given below.





Where;

f – the frequency or the number of cases in any category

N- the total number of cases in all categories.

Ratio


Used to compare categories in terms of relative frequency. The formula is given below;



Where;

- the number of cases in the first category

- the number of cases in the second category

Rates


Rates are defined as the number of actual occurrences of some phenomenon divided by the number of possible occurrences per some unit of time. The formula is given below;


Rate of change or percentage change


Is useful for comparing the actual change between time periods.

Where;

– frequency at the new/current time period

– frequency at the old/previous time period

Collection, Summarization and Presentation of Data

Introduction


Data can be collected in different ways. It can be obtained from original data or from previous students.

Methods of Collecting Data

The methods of collecting data are:


1. Direct or interview method. This is a personal communication with the individual you want to interview.

2. Indirect or questionnaires method. This is done by sending questionnaires to the person from whom you would like to get the information.

3. Registration method. Utilizing existing records from various agencies.

4. Observation method. This can be done directly or indirectly.

5. Experiment method. This is done with the participation of a certain researcher. In other words, there is a human intervention occur during the process of data collection.

Two Documented Sources of Data


1. Primary Data. Data documented by the primary source.

2. Secondary Data. Data documented by a secondary source.



Sampling Techniques


The study of the entire population of interest in some situations is impractical or even impossible to include the entire population. Thus, take a part of a population, the so-called sample.


Sampling


Sampling is the process of selecting the sample or the study units from a previously defined population.

Sampling Error


Sampling error is the difference or deviation of the sample from the population with respect to the characteristics of interest in the study.

Sampling Frame


The list of units from which the sample were drawn in any sampling procedure.

Sampling Procedure


Sampling procedure refers to the manner in which the members of the population are selected as part of the sample. These are classified into probability or random sampling and non-probability sampling procedures.

Sample Size Determination


An important aspect of the sampling design is the sample size. The number of members that you include in the study must not be too small in order to come up with reliable estimates. According to some researchers and statisticians suggest, Slovin’s formula is an alternative approach to computing the sample size. The formula is given below;



Where;

n – the sample size N – the population size e – the desired margin of error

Probability Sampling Procedures


These comprise all sampling methods done when there is a sampling frame which ensures that all the probable sampled have an equal chance or probability of being selected for the study.

Types of Probability Sampling Procedures


Simple Random Sampling


This is the basic method on which all other methods of probability sampling are built. Each member of the population has an equal and independent chance of being selected.

Systematic Sampling


This is done by selecting a sampling interval k and using the sampling frame, the researcher selects every kth member of the population beginning at some random point and cycling through the list.

Stratified Sampling


The members of the population are classified into non-overlapping groups or strata on the basis of characteristics to be properly represented in the sample.

Cluster Sampling


This is usually used in studies of huge populations where the sampling frame may be too large to study or too time-consuming that is better to divide them first into clusters or groups and randomly select a sample cluster of choice.

Multistage Sampling


This is usually done in big community-based studies in which selection of the sampling unit is done by stages.

Non-probability Sampling Procedures


These are methods that do not include random sampling at some stage in the process. Further, these are applicable when there is no sampling frame available.

Types of Non-Probability Sampling Procedures


Convenience Sampling


In this method, the sample consists of elements that are most accessible or easiest to contact.

Judgement or Purposive Sampling


In this method, the researcher chooses a sample that agrees with his/her subjective judgment of a representative sample.

Quota Sampling


Is the non-probability sampling wherein the researcher just sets a quota or number of sampling units to be included in each grouping but uses convenience sampling to select the units within each grouping.

Snowball Sampling


Also called chain referral and referential sampling. This is used to find members of a group not otherwise visibly identified.

Thursday, July 4, 2019

Estimation of Parameters

Learning Objectives:

Given the learning materials and activities of this chapter, they will be able to:
Ø  Distinguish classical methods from Bayesian methods of estimates.
Ø  Calculate the standard error of a sample.
Ø  Calculate the margin of error for interval estimate.
Ø  Construct interval estimates of the population mean given a specified level of confidence.
Ø  Construct interval estimates of the population proportion with the specified confidence level.

Introduction

         There are two major areas in statistical inference, the estimation of parameters and hypothesis testing. Estimation is the process of estimating the value of a parameter from information obtained in a sample. An important aspect of estimation is the size of the sample. An estimator is a formula, the function or procedure used in estimating a population parameter. There are two methods of estimating population parameters, such as:

a.                       Classical method is based strictly on information obtained from a random sample selected from the population.

b.                 Bayesian method utilizes prior subjective knowledge about the probability distribution of the unknown parameters in conjunction with the information provided from the sample data.

In this text, we shall utilize the classical method to estimate unknown population parameters such as the mean, proportion and the variance by computing statistics from a random sample and applying the theory of sampling distributions. There are two ways in the classical method of estimation, namely: point estimate and interval estimate.

Point estimate
            Point estimate consists of a single value used to estimate a population parameter. For most parts, the point estimate will be different from the population mean due to sampling error. There is now way of knowing how close the point estimate is to the population parameter. For this reason, statisticians prefer another type of estimate.

Interval estimate
            Is an interval or a range of values used to estimate the parameter. In an interval estimate, the parameter is specified as being between two values. A degree of confidence can be assigned before an interval estimate is made. The confidence level is the probability that the interval estimate will contain the true population mean or population proportion.

        Three common confidence level are 90%, 95% and 99% confidence intervals. The table below summarizes the values of the standard deviates and the margin of error for the most commonly used confidence level.


       A term level of significance is defined as the probability of erroneously concluding that a confidence interval generated will contain the parameter. As observed in the table, the greater the level of confidence, the larger the z values, the larger the margin of error and of course the wider the confidence interval.

        An interval estimate is constructed by subtracting and adding the margin of error to a point estimator. The length of the confidence interval is determined by the sample size, the standard deviation, and the desired confidence level.

Estimating Means

         Estimation (or estimating) is the process of finding an estimate, or approximation, which is a value that is usable for some purpose even if input data may be incomplete, uncertain, or unstable. The value is nonetheless usable because it is derived from the best information available. Typically, estimation involves "using the value of a statistic derived from a sample to estimate the value of a corresponding population parameter". The sample provides information that can be projected, through various formal or informal processes, to determine a range most likely to describe the missing information. An estimate that turns out to be incorrect will be an overestimate if the estimate exceeded the actual result, and an underestimate if the estimate fell short of the actual result. 

Estimating Means Large Sample and the Standard Deviation is known

The Central Limit Theorem says that, for large samples (samples of size n ≥ 30), when viewed as a random variable the sample mean is normally distributed with mean and standard deviation. The Empirical Rule says that we must go about two standard deviations from the mean to capture 95% of the values of sample mean generated by sample after sample.

Confidence interval for means >=30 and the standard deviation is known the formula is



Example 1: A study of 40 bowlers showed that their average score was 186. The standard deviation of the population is 6.
a.       Find the 95% confidence interval of the mean score for all bowlers.
b.      Find the 99% confidence interval of the mean score of a sample of 100 bowlers instead of a sample of 40.


          Thus, it can be 95% confident that the true mean score of bowlers is between 184.14 and 187.86. This means that 95% of the time, the population mean score of bowlers will be roughly between 184 and 188. 


          Thus, the 99% confidence interval for the population mean score is ranging from 184.542 to 187.548. This means that we can be 99% confident that the population mean score is roughly between 185 to 188.

Confidence interval for means < 30 and the standard deviation is unknown (small sample)

         When the population standard deviation is unknown and the sample size is less than 30, the standard deviation from the sample can be used in place of the population standard deviation. In this case, the t-distribution is used to determine the confidence interval and the random variable is approximately normally distributed. The formula is:


          To determine the value of t critical locates the critical value from the table in t distribution with the corresponding degrees of freedom. The degrees of freedom are the values that are free to vary after a sample statistic has been computed. The degrees of freedom for the confidence interval for the mean is n – 1. Also, note that the sample standard deviation is used instead of the population standard deviation. 

Example 2: A sample of 20 tuna showed that they swim an average of 8.6 miles per hour. The standard deviation for the sample was 1.6. Find the 95% confidence interval of the true mean.

Solution: Given n = 20, the sample mean of 8.6 miles per hour, and the sample standard deviation s = 1.6 miles per hour. The degrees of freedom is n – 1 = 20 – 1, using the t-distribution table yielded a critical value of t =2.093. Hence, 


          Thus, the 95% confidence interval for the population mean time is ranging from 7.851 to 9.349 miles per hour. This means that we can be 95% confident that the true mean time of tuna can swim roughly between 8 and 9 miles per hour. 

When to use the z and t distribution:
-          If the population standard deviation is known and sample size is large, use z-test.
-          If the population standard deviation is unknown and sample size is large, use z-test.
-          If the population standard deviation is unknown and sample size is small, use t-test.

Estimating Proportions

       When the variable of interest is qualitative and are summarized in terms of frequencies, confidence intervals for estimating proportion may be constructed. 

          To construct confidence interval for estimating a population proportion based on a proportion obtain from a random sample, similar procedure used to estimate population mean. 


          Example 3: A local polling organization reports that based on a local-wide survey of 500 respondents, 43% of the vote will be in favor of the administration governatorial candidate in the May 2016 elections. Construct the 95% confidence interval for the proportion indicating preference for the administration candidate.
         Thus, the 95% confidence interval is from 38.7% to 47.3%. This means that with a sample of 500, the poll has a margin of error of ±4.3% and the pollster can be 95% confident that the administration candidate will obtain roughly between 39% and 47% of the votes.


Powerpoint presentation: Estimation of Parameters


 Click Here: https://www.scribd.com 



Note: For Comments, Questions, and Suggestions feel free to contact at enomaratas@jrmsu.edu.ph or ednielmaratas@gmail.com. You can also post at the comment section below.

Population and Sample, Parameter and Statistic, Descriptive and Inferential

What is Population? Sample?


In statistics, researchers, and educators commonly use the terms population and sample (Alferez & Duro, 2006). The population is often too large for us to examine each of its members.

           The population is the entire collection of all elements/experimental units under consideration in a statistical inquiry or to be studied. Sample a part of the population. If the sample is to provide information about the entire population, it must be representative of that group in some way. In actuality, unless a sample is picked at random, it cannot be expected to be representative of a population. This is because any nonrandom rule for selecting a sample will almost always produce one that is skewed toward some data values over others.

        For example, if we wish to determine the average income of households in Zamboanga del Norte, then the population of interest is the collection of all households in Zamboanga del Norte. However due to some constraints, the budget, the time, the manpower for instance, and then we would have to redefine the interest. This time we can delimit the scope of the study we utilized sampling size to include only the collection of all households in Dapitan and Dipolog City, considering this is done thru sampling techniques.

What is Parameter? Statistic?
          Parameter refers to any numerical value describing a characteristic of a population and is usually denoted by some Greek letters such as population standard deviation, σ and population mean, μ. 
         While the term statistic is a numerical measurement describing some characteristics of a sample. The symbols x and s are statistics which is unbiased of the parameter μ and σ.


Two Major Areas of Statistics

Descriptive Statistics
Is concerned with the methods for collecting, organizing, and describing a set of data so as to yield meaningful information (Walpole, 2000). Construction of tables, charts, and graphs and the computation of descriptive statistical measures also fall in this area.

Inferential Statistics
Inferential Statistics is also called Inductive Statistics or Statistical Inference. Comprises those procedures for drawing inferences or making generalizations about characteristics of a population-based on partial and incomplete information obtained sample data to infer to populations.

SEE YOU ON THE NEXT TOPIC:


Introduction to Statistics

Types of Variables

What are the Types of Variables?


Variable refers to a characteristics or phenomena that changes or varies over time for different individual or objects under consideration (Mendenhall, et.al., 2012).

Data
Data are the values that the variables can assume (Reston, 2004).

Experimental unit
Experimental unit is the individual or object on which a variable is measured.

Types of variable according to a functional relationship

Independent variable. This is sometimes termed as a predictor variable if the object is to predict the value of one variable on the basis of the other.

Dependent variable. This is sometimes called the criterion variable and whose value is predicted. For example, academic achievement is dependent on Intelligent Quotient, study habits, interests, attitudes and many more. Hence, IQ, study habits, interests, attitudes are independent variables. On the other hand, academic achievement is the dependent variable.

Types of variable according to the attribute of objects they classify

         Qualitative Variables. These are words or codes that represent a class or category. Further, produce data that can be categorized according to similarities or differences in kind. Also known as a categorical variable. 

Here are some examples: gender, taste ranking, religious affiliation, academic achievement, marital status, type of high school attended and many more.

Quantitative Variables. These are variables that classify objects or represent an amount or a count. This is a variable often represented by an arbitrary letter, let say x, produce numerical data. 

Here are some examples: Height, student enrolment, class size, family size, test scores, entrance test results, crime rate, salary, number of passengers, a volume of orange juice, etc.

Types of variable according to the continuity of values

Discrete variable. This refers to variables that can be obtained, can assume only a finite or through a countable number of values.

Examples are: Number of family members, number of new car sales, number of defective bulbs, faculty size, hospital staff size, number of students enrolled in Statistics course, number of bedrooms in a house, etc.

Continuous variable. Variables that can assume many values corresponding to the points on a line interval. 
Here are some examples: Crime rates, cell density, rainfall, temperature, air pressure, weight, height, study hours, time, salary, distance traveled, etc. 

Types of variable according to the Scale of Measurements

Nominal scale is often referred to as a categorical scale. This only satisfies the identity property of measurement. 

For example, gender, Religion, and political affiliation.

Ordinal scale has an ordered relationship to every other value on the scale. 

Example: Academic achievement, taste ranking, honors received, educational qualification etc.

           Interval scale has equal units of measurement, thus making it possible to interpret not only the order of scale scores but also the distance between them. It has the properties of identity, magnitude, and equal intervals. 
     
            Example: test scores, height (in cm), Intelligent Quotient (IQ) and many more.

Ratio scale has the property allows one to make statements of equality of intervals. This is the highest level of measurement which includes the inherent zero starting point. 

Examples: number of children in a family, student enrolment income and many more.
                                                                                                  




Why Study Statistics?

Study Statistics?

As early as in the Old Testament, statistical methods occur particularly censuses of population and wealth were taken by the Pharaohs and the ancient Hebrews long before Christ was born. With the turn of 20th century, the term came to be applied to a study of scientific methods of dealing with quantitative data with applications to practically all fields of study.

There are reasons why the scope of statistics and the need to study statistics have grown enormously in the last few decades.


First, Knowledge of basic statistics is essential for people going into research in various fields of human endeavor.

Second, numerical information is everywhere. 

Third, statistical techniques are used to make decisions that affect our daily lives.
Fourth, a person with an understanding of statistics is between able to decide whether his or her professional colleagues use their statistics to illuminate or merely to support their personal biases; that is, it helps one to decide whether the claims are valid or not.

Fifth, knowledge of statistics is essential for persons who wish to keep their education up-to-date.

Finally, an understanding of statistics can help anyone discriminate between facts and fancy in everyday life – in reading newspapers and watching television, and in making daily comparison and evaluations.


Generally, no matter what your future line of work, you will make decisions that involve data. That is, an understanding of statistical methods will help you make decisions more effectively.


What is Statistics? Statisticians?

Statistics?


How do we define the word statistics? The word statistics has more than one meaning. We encounter it frequently in our everyday language. The origin of the sources we shall use in the definition may be traced to the following authors, of which have very little in common meaning about statistics.  

Reston, (2004), the word statistics simply refers to a mass of data or a collection of facts and figures. The grades of students in the registrar’s office, entrance examination results of incoming freshmen in certain universities or colleges, the average temperature and average rainfall per month, weekly sales of a company A are such statistics in nature.

Freund, et, al. (1986) and Nocon, et, al. (2000), defines statistics as a branch of mathematics that deals with the theory and methods of collection, analysis, interpretation, and presentation of data. As such, the data may be categorical or numerical in nature.


Walpole, (2000) and Lind, et, al. (2000) defines as a science that is concerned with the concepts and techniques employed in the collection, presentation, analysis and interpretation of data to assist in making more effective decisions.


Statistics is the branch of science that deals with the collection, presentation, summarization, analyzation, and interpretation of data.


Statisticians

Are professionals who are trained to provide crucial guidance in determining what information is reliable and which predictions can be trusted through the application of statistical methods. Furthermore, it helps determine the sampling and data collection methods, monitors the execution of the study and the processing of data, and advice on the strengths and limitations of the results (www.amstat.org). The word statistician is used to refer to collect and analyze data, and then calculate results using a specific design. One who draws conclusions and make decisions in the face of uncertainty. They are present behind the scenes in every field of scientific endeavor. The tasks of statisticians include: Agricultural and Animal Sciences, Medical Sciences, Business areas, Engineering, Education, Economics, Mathematics, Banking, Accounting, and Auditing, Natural and Social Sciences etc..