Chapter 33 Final Chapter
We are beginning to approach the end of this book but much remains for those who want to learn more.
33.1 Common Misconceptions
Analysis with mathematics and statistics often becomes abstract and it is then easy to confuse concepts. Here follow some reminders based on commonly occurring misconceptions.
Mathematics \(\neq\) reality. In various examples we have used mathematics to describe theoretical relations and connections. Theoretical reasoning can sound so convincing that it is taken as strong arguments for how the world works. It is then important to remind oneself that mathematics and logic are not proof of how the world works, other than in very limited reasoning. Humans are complex, changeable and can function in many different ways in different situations.
Correlation is not causality. Causal relationships are something we read into what we observe. It cannot be proven mathematically. Even in a controlled experiment there can be room for interpretation and thereby more or less erroneous conclusions regarding causal relationships. The risks for erroneous interpretations also exist when we work with quasi-experiments.
Mean and variance. Let us assume that we have found a positive correlation between two variables, for example that higher income covaries with greater happiness. Covariation for groups does not necessarily mean that we can make statements about individual observations, or persons, with any great certainty. If the variance in the variables is sufficiently large in relation to the correlation then there will also be many low-income earners who are happier than many high-income earners.
There is always a risk that we are wrong. Suppose we estimate a correlation between the variables \(Y\) and \(X\), where our statistical tests indicate that our null hypothesis (no correlation) is false at significance \(\alpha = 0.01\), which feels safe. But there is still 1% probability that we reject the null hypothesis erroneously, a type 1 error. Furthermore, the power of our statistical test can be low. Other studies using other data or other methods could give other results that indicate that our calculations and/or our results and/or our conclusions were erroneous in some way, or that they maybe only apply to our specific case or some other problem.
Other factors can be important. Results that indicate that \(X\) has a causal effect on \(Y\) are not proof that some other phenomenon does not have an effect on \(Y\). If we for example show that medicine A affects a disease symptom it does not mean that there does not exist some other medicine B that can also affect the same symptom.
Variables do not need to be normally distributed. For a regression model \(Y = a + bX + U\) we often assume that the error term \(U\) follows a normal distribution. But the variables \(Y\) and \(X\) do not need to follow a normal distribution.
Results that are not statistically significant may still be interesting. When we work with regression analysis it is easy to believe that regression results without statistically significant results are uninteresting. But many times it is the statistically non-significant results that are the important ones, or at least as important. Say for example that there exists a widespread erroneous notion in society that \(X\) affects \(Y\). In that case a large amount of statistically non-significant results (that \(X\) does not seem to have any effect on \(Y\)) are very interesting.
33.2 Things That Perhaps Should Have Been Included in This Book
Here are some tips for those who want to read more.
Differential and difference equations are mathematical functions that describe how a variable can be a function of the change of the same variable. This can among other things be used to describe development over time. This type of mathematics can among other things be used to describe theoretical processes from one point in time to another. In empirical analysis, difference equations are used to study data that is sorted over time, within what is called time series econometrics.
Geometry is a part of mathematics that deals with distances and spaces, for example shape and size of different objects in a space. A particular part of geometry is trigonometry where we study angles and dimensions of triangles. In upper secondary school mathematics courses the trigonometric functions \(\cos()\), \(\sin()\) and \(\tan()\) are introduced. Within social science we can among other things use geometry to study cyclical phenomena, such as the economy’s ups and downs. Geometry is also used within statistics and can therefore be valuable for those who want to deepen themselves within this.
Time series and panel data. In the examples we have primarily used cross-sectional data, observations that belong to one and the same point in time, for example how different countries a specific year. Data that is instead organized over time is called time series or time series data, for example GDP in a region or country over a collection of years. Data that consists of several time series for different groups is called panel data, for example GDP for several countries and years in the same table. Statistical analysis of time series and panel data brings particular challenges and sometimes requires other estimators than the least squares method. There is specific literature for this and at universities there are specific courses within time series econometrics.
Other statistical schools. The statistical method described in this book is based on what is called frequentist statistics, or classical statistics. There are also other approaches to how we should approach probability. Another popular approach is what is called Bayesian statistics, named after the mathematician Thomas Bayes (1702–1761). Bayesian statistics relates differently to knowledge and probability and emphasizes more subjective valuations.
Text analysis. In earlier sections we went through an example of how we can use correlation analysis for unstructured data, such as calculating to what extent the characters in Star Wars appear in the same films. The mathematics and statistics are basically the same regardless of what type of information we study. But different types of data sometimes require other methods and often encounter other types of challenges.
33.3 The Last Word
University departments differ in what they consider good science and which skills they value. In practice, most analytical and research work takes place outside academia. If you intend to work in analysis, a broad range of skills will serve you well: wide knowledge of the world, technical proficiency, and the ability to communicate clearly — in writing, through visuals, and in front of an audience. The ability to find and evaluate information, ideally in more than one language, is especially worth developing.
Finally, take a moment to appreciate what you have accomplished. There is always more to learn, and the work never truly ends. But you have now made it to the end of this book. Well done!