This app allows students to explore how we model relationships between variables using data from the 2018-19 English Premier League seasons. Specifically, the app looks at the relationshiop between Expected Goals (xG) and possession across three tabs. Each tab allows the student to 'play' interactively with models of the data.
Look at a single team's data per game, or switch to “All Teams” to see all teams at once. The student should select different teams and 'say what they see'. What patterns do they notice? There is also an option to try a summary line by clicking a button. This produces a line with a randomly sampled intercept and slope. The student can comment on whether they think each line is a good or bad representation of the pattern. They can also select an option to show the line of best fit and the curve of best fit and comment on which best represents the pattern they see.
This tab shows the aggregated data for all teams (one data point per team). Students can adjust the intercept and slope sliders to find the line that they feel best represents the data. They can compare their line to the OLS line of best fit. This tool can be used to explore fit of a model to the data and to understand what the slope and intercept of a line represent.
This third tab allows students to pick any two teams and see how the relationship between xG and possession changes for different teams. This can be used to explore differences in relationships between variables and statistical moderation. For which teams is the relationship between xG and possesion strongest or weakest?
Click the info icon in the top-right corner of each tab for more detail on how to use it.
Click the app title above at any time to come back to this page.
Start by picking a team and 'say what you see'. Now change the team. What pattern do you notice now? Is there a consistent pattern across teams? What about if you select all teams? What do you notice about xG as possession increases?
Then click the Try a line button. This produces a line to summarise the pattern. Try different lines. Do you think they are a good or bad representation of the dots? Is each line better or worse than the previous line?
Finally, select the option to show the trend or curve. This will show the numerical line/curve of 'best fit'. Does the line or curve better represents the pattern of dots for a particular team?
This tab shows the aggregated data for all teams (one data point per team). There are two sliders that affect the line on the plot. Use the Average xG slider to change the 'typical' amount of xG seen in the data, and use the Change in xG slider to adjust the rate at which xG changes as possession increases. How do the sliders affect the line? How do they relate to the intercept and slope of the line?
Use the sliders to find the line that you feel best represents the data.
Now click “Show best line” to overlay the numerical 'line of best fit'. How close is your line to the 'best fit'? Is 'best fit' the same as 'good fit'?
Pick any two teams and look at the relationship between xG and possession for the two teams. Is it the same? What do you see? What does this mean for the different teams? For which teams is the relationship between xG and possesion strongest or weakest?