Group Members

Toan Vo, Christine Hou, LanFei Ma, Kaiden Schears

Introduction

Our mission is to increase student SAT scores by looking at factors like attendance and what section students focus on when studying for their SATs. We use two datasets in our analysis: The first one includes the mean SAT scores by section of each school in New York (2010); The second one is the enrollment and attendance history of the schools in New York. Our questions of interest include the following:

Through regression analysis and hypothesis testing we will provide statistical insight into these questions. We believe that students SAT scores would benefit from increasing their attendance in school as well as the knowledge that reading and writing are potentially linked thus improving one may have a doubling effect in overall score.

Background

Data

  • SAT College Board 2010 Results from NY schools 1
    • contains the means of the each section of the SAT by school.
    • College Board collected SAT mean scores of college-bound seniors, who finish SAT questionnaire and self-report the graduation status from high school in the year 2010.
    • 460 observations
    • 6 variables: DBN, School Name, Critical Reading Mean, Number of Test Takers, Mathematics Mean, Writing Mean
  • Historic Daily Attendance by School 2
    • contains Daily Attendance by School and by Year.
    • Collected by NYC Department of Education
    • 831,800 observations
    • 7 variables: School, Date, SchoolYear, Enrolled, Present, Absent, Released

Pre-processing Data:

The attendance data was filtered to only include data from 2010 to match with 2010 SAT data. The average attendance by school was calculated by taking the mean of the portion of students present divided by total students enrolled and then multiplied by 100. The attendance data and SAT data were joined and all columns with NA were dropped to get working Dataset with variables as described in the data summary.

Data Summary

  • Dataset used for analysis (variable description):
    • DBN: District Number - unique identifier first two numbers represent the district and following characters are a unique identifier for the school
    • num_takers: Number of students that took the SAT by school
    • reading_mean: mean reading score by school
    • writing_mean: mean writing score by school
    • math_mean: mean math score by school
    • overall_score: calculated column by summing up scores of Reading, Math and Writing for each school
    • avg_attendance: Average percentage of attendance found from pre-processing.
    • n: number of days attendance was taken from a school (number of rows associated with a school in filtered attendance data)
    • 377 observations
avg_reading avg_math avg_writing avg_overall avg_attendance avg_samples min_samples max_samples
404.5995 413.7905 398.0637 1216.454 83.10759 165.2069 70 179

Here is a data summary of the combined data to get a sense of important variables. Average SAT scores across the dataset for each section and overall are displayed. Additionally, the average attendance is displayed along with the average number of attendance samples of across all schools as well as the min and max of attendance samples. From this summary it is clear that not all schools reported attendance on the same days (there is a range of sample size 70-179) which may be something to consider as we do our analysis of the data.

The remaining portion of this report will be used to answer our two questions. In doing so we will use a regression model to answer how attendance has an impact on SAT scores. Additionally, we will use hypothesis testing with normal distributions to answer if writing and reading mean are the same or different.

Regression Analysis

Using the regression model, we can see whether average attendance in the school has positive or negative impact on the overall SAT mean score. We will visualize the correlation and residuals. From the distirbution of residual plot, we can see whether the model has outliers, and whether the model is reasonable. With the specific coefficient and intercept from regression model, we can have a specific estimated regression model such that we can make a prediction on how the 100% attendance in one school can help on the overall mean SAT score. Given an specific x variable, we will also check the confidence interval on it and display this in the plot.

Question 1: Does school attendance have an impact on the total SAT mean score?

In this question, we are interested in the correlation between total SAT mean score (reading_mean + writing_mean + math_mean) and the school attendance. The linear regression model is

\[ Y_{i} = \beta_0 + \beta_1x_i +\epsilon_i \]

  • \(Y_i\) is the total SAT mean score (response variable)
  • \(x_i\) is the school attendance (explanatory variable)
  • \(\beta_0\) is the intercept
  • \(\beta_1\) is the slope
  • \(\epsilon_i\) is the random error

Therefore, the estimated model in this question is

\[ \hat{Y_{i}} = \hat{\beta_0} + \hat{\beta_1}x_i \] where \(\hat{Y_{i}}\)is the predicted response corresponding to \(x_i\), \(\hat{\beta_0}\) is the estimated intercept, and \(\hat{\beta_1}\) is the estimated slope.

Using linear regression analysis, we have the intercept and slope value:

   (Intercept) avg_attendance 
      336.6020        10.5869 

So the predicted linear regression model with specific values is:

\[ \hat{Y_{i}} = 336.6020 + 10.5869 * x_i \]

With the predicted model, when the school attendance is 100%, the overall mean SAT score is:

[1] 1395.292

The 95% confidence interval prediction interval for \(x = 100\) is:

Confidence Interval and Hypothesis Testing Analysis

Question 2: Is there any difference in mean SAT Reading and Writing scores?

Some pre-analysis and observations:

We first did some observations on the mean SAT reading and writing scores, the density plot shows that the both two variable reading_mean and writing_mean follow an similar distribution, so we decide to research the differences in mean SAT reading and Writing scores

From the density plot, we found that there are differences in mean SAT reading and writing scores, but the differences are not significant.

Then we calculate mean and standard deviation of reading_mean and wriring_mean, and visualize the differences between those two variables using density plot.

mean_reading mean_writing sd_reading sd_writing
404.5995 398.0637 57.32686 58.29073

By calculating the mean and standard deviation of reading_mean and writing_mean, we found that there are differences in the mean and standard deviation

This plot shows there are differences between the mean reading and mean writing score since not all the differences value distribute at 0

Confidence Interval

After analyzing the difference plot, we seek to find a confidence interval for the mean difference between mean sat reading scores and sat writing scores.

Model

  • The population is all mean writing score and mean reading score measurement differences, in this case, n = 377

  • Let \(\Delta\) be the mean difference, reading_mean - writing_mean, in the population.

  • We have data pairs \((x_i, y_i)\) for \(i = 1,\ldots,377\).

  • Model the differences, \(d_i = x_i - y_i\).

\[ D_i \sim F(\Delta, \sigma), \quad \text{for $i = 1, \ldots, n$} \]

  • \(F\) is a generic distribution for the population of differences
  • \(\Delta\) is the mean of this distribution, in this case
  • \(\sigma\) is the standard deviation

Manual Calculations

diff_mu diff_sd n
6.535809 11.91635 377
  • The sample size is \(n = 377\)

  • The sample mean is \(\bar{d} = 6.54\)

  • The sample standard deviation is \(s = 11.9\)

    low high
    5.329049 7.742569

Interpretations

  • A typical difference between the sat reading and writing is about \(6.54\).
  • From above, the standard deviation among differences is about 11.9.
  • We are confident that the mean difference in measurements between the sat reading and writing score is quite small.

We are 95% confident that the mean difference between SAT reading scores and writing scores is \(5.33\) to \(7.74\) points.

t.test() confirmation

## 
##  Paired t-test
## 
## data:  comb_data$reading_mean and comb_data$writing_mean
## t = 10.649, df = 376, p-value < 2.2e-16
## alternative hypothesis: true mean difference is not equal to 0
## 95 percent confidence interval:
##  5.329049 7.742569
## sample estimates:
## mean difference 
##        6.535809

Hypothesis Test

After calculating the confidence interval, we tried to figure out whether there are systematic differences in SAT reading and writing scores

  • Same populations and samples and model as above

  • Hypotheses

\(H_0: \Delta = 0\)
\(H_a: \Delta \neq 0\)

  • Test Statistic

\[ T = \frac{ \bar{d} - 0 }{s/\sqrt{n}} \]

We calculated the the test statistics to be

## [1] 10.64944
  • Sampling distribution is t with 376 degrees of freedom

  • Calculate p-value

## [1] 2.538402e-23

There is very strong evidence that there are differences in mean sat writing scores and reading scores (p=2.538402e-23, two-sided paired t-test).

Discussion

Data

One way that data affects the analysis is that the data is not perfect. Many schools had a different number of days that they recorded attendance which might impact the accuracy of the data. Additionally the number of test takers was different for each school which impacts the mean scores for each school and could affect the analysis. For example if there is a low number of test takers and one student scores really high or low they could skew the mean. Additionally, the data was counted only if a student self-reported as a graduating senior, so the actual number of test takers could be much higher or lower. For the future it might be better to look into this and potential exclude schools with a low number of test takers to remove skewed data. However, we recognize excluding data might introduce other problems.

Regression

Even though the analysis goes through, there are some inferences we can get from the above linear regression analysis:

  • The general pattern between school attendance and total SAT mean score is increasing, and we can say that if higher percentage of students goes to school, the total SAT mean score can be higher.

  • Based on the scatter plot , we can find that the distribution of points is not very linear. The residual plot displays the U-shaped curve, indicating the non-linear relationship between school attendance percentage and the total SAT mean score. Therefore, to have better interpretations, linear regression model is not the best fit, and it’s better to try the regression model with more degrees. However, based on the course contents, we can’t choose more complicated regression analysis. Therefore, when we try to predict the overall SAT mean score when the school attendance is 100%, the result does not quite match what we can see from the scatter plot, because the actual total SAT mean score is much higher when school attendance is lower than 100%. The main reason is because we used the linear relationship but not quadratic.

  • In the future, if we step more into regression analysis, we can try complicated regression analysis to obtain more accurate predicted model on the association between school attendance percentage and total SAT mean score.

Hypothesis Testing

  • Broader interpretations: From both the plots, and calculations, it is clear that there is differences between mean reading scores and writing scores.

  • Potential short comings: The choice of test statistics may have problems since we use the mean of mean writing and reading scores. The result can be biased, the differences of mean writing and reading scores for individual samples may vary a lot.

  • Additional work: Explore test statistic to calculate P-value. Also explore more about what can be thought as significant difference in real world application rather than in statistical meaning. Perform a regression to further analyze relationship of reading and writing mean to see if they are correlated.

Conclusions

Using regression, a positive trend line can be seen between the predicted variable “Total SAT Mean Score” and “Attendance”. This is possible evidence for a higher attendance leading to higher mean SAT scores. However,The residuals plot suggests that the relationship is not linear and appears to be quadratic. Regarding our hypothesis testing: There is a significant difference between reading score mean and writing score mean (p=2.538402e-23, two-sided paired t-test). This evidence suggests that increasing reading scores won’t increase your writing by the same amount. However, the Reading and writing mean seem to be correlated but further analysis needs to confirm this suspicion.

References


  1. https://data.cityofnewyork.us/Education/SAT-College-Board-2010-School-Level-Results/zt9s-n5aj↩︎

  2. https://data.cityofnewyork.us/Education/2009-2012-Historical-Daily-Attendance-By-School/wpqj-3buw↩︎