os2_lab_04A.pdf - Lab 4A Foundations for statistical...

This preview shows page 1 - 3 out of 5 pages.

Lab 4A: Foundations for statistical inference - Sampling distributions In this lab, we investigate the ways in which the statistics from a random sample of data can serve as point estimates for population parameters. We’re interested in formulating a sampling distribution of our estimate in order to learn about the properties of the estimate, such as its distribution. The data We consider real estate data from the city of Ames, Iowa. The details of every real estate transaction in Ames is recorded by the City Assessor’s office. Our particular focus for this lab will be all residential home sales in Ames between 2006 and 2010. This collection represents our population of interest. In this lab we would like to learn about these home sales by taking smaller samples from the full population. Let’s load the data. download.file ( "" , destfile = "ames.RData" ) load ( "ames.RData" ) We see that there are quite a few variables in the data set, enough to do a very in-depth analysis. For this lab, we’ll restrict our attention to just two of the variables: the above ground living area of the house in square feet ( Gr.Liv.Area ) and the sale price ( SalePrice ). To save some effort throughout the lab, create two variables with short names that represent these two variables. area <- ames $ Gr.Liv.Area price <- ames $ SalePrice Let’s look at the distribution of area in our population of home sales by calculating a few summary statistics and making a histogram. summary (area) hist (area) Exercise 1 Describe this population distribution. The unknown sampling distribution In this lab we have access to the entire population, but this is rarely the case in real life. Gathering information on an entire population is often extremely costly or impossible. Because of this, we often take a sample of the population and use that to understand the properties of the population. If we were interested in estimating the mean living area in Ames based on a sample, we can use the following command to survey the population. This is a product of OpenIntro that is released under a Creative Commons Attribution-ShareAlike 3.0 Unported ( http: //creativecommons.org/licenses/by-sa/3.0/ ). This lab was written for OpenIntro by Andrew Bray and Mine C ¸ etinkaya-Rundel. 1
Image of page 1

Subscribe to view the full document.

samp1 <- sample (area, 50) This command collects a simple random sample of size 50 from the vector area , which is assigned to samp1 . This is like going into the City Assessor’s database and pulling up the files on 50 random home sales. Working with these 50 files would be considerably simpler than working with all 2930 home sales.
Image of page 2
Image of page 3
  • Fall '15

What students are saying

  • Left Quote Icon

    As a current student on this bumpy collegiate pathway, I stumbled upon Course Hero, where I can find study resources for nearly all my courses, get online help from tutors 24/7, and even share my old projects, papers, and lecture notes with other students.

    Student Picture

    Kiran Temple University Fox School of Business ‘17, Course Hero Intern

  • Left Quote Icon

    I cannot even describe how much Course Hero helped me this summer. It’s truly become something I can always rely on and help me. In the end, I was not only able to survive summer classes, but I was able to thrive thanks to Course Hero.

    Student Picture

    Dana University of Pennsylvania ‘17, Course Hero Intern

  • Left Quote Icon

    The ability to access any university’s resources through Course Hero proved invaluable in my case. I was behind on Tulane coursework and actually used UCLA’s materials to help me move forward and get everything together on time.

    Student Picture

    Jill Tulane University ‘16, Course Hero Intern