Theories

Scholarship works with theories. What these look like exactly differs considerably between disciplines. While the natural sciences often work with mathematical models, that is, formulas that describe the relationship between variables explicitly and unambiguously and allow predictions, the social sciences often work with verbal theories in the style of “X and Y are positively related” or “the higher X, the higher Y,” and the traditional humanities work, for example, with verbal explanations. Verbal theories have the advantage of tending to be easy to understand and broadly applicable, but the terms they use are often subject to individual, cultural, or temporal influences, and discussants risk talking past one another in scholarly discourse.

For formal theories, all variables involved are precisely defined, and such theories often have a strongly restricted scope of validity (e.g., many physical laws hold only under tightly controlled conditions, such as in a vacuum, at a specific temperature, and so on). A concern raised in the context of the replication crisis is that theories are not clear enough to predict when replications will succeed, and that this is one of the causes of low replication rates (Buzbas & Devezer, 2023; P. Smaldino, 2019). A theory about the consequences of identifying with gender roles, for example, must account for changes in gender roles and their particularities across different countries. It is hardly surprising that one and the same experiment on this topic yields different results in the USA in 1980 than in Germany in 2020. What is problematic, however, is that – even though such qualifications seem sensible and necessary for many social-science theories – statements to this effect are rarely made.

Verbal theories are not inherently less scientific: within their respective fields, scientific theories always stand out from everyday explanations through their particularly high degree of systematicity (Hoyningen-Huene & Kincaid, 2023). However, fields that place value on predicting events cannot do without formal theories (Muthukrishna & Henrich, 2019). It should be emphasized that certain disciplines place no value on prediction (e.g., history, or fields that proceed primarily hermeneutically). Fields such as psychology, quantitative sociology, and parts of the humanities (“digital humanities”) are currently moving closer to formal models – in social psychology, there was already a call to formalize theories once before, during a crisis in the 1960s (Lakens, 2023). Because theories lacking objectivity are rarely used by different researchers and, due to their flexible interpretation, are difficult to refute, an enormous quantity of useless theories has emerged there (Ferguson & Heene, 2012). Among these are also mutually contradictory theories: for instance, Banker et al. (2017) argued that “ego depletion,” that is, the depletion of self-control resources, causes people to rely more on cues from other people (p. 2), whereas Francis et al. (2018) conjectured the opposite – that depletion prevents cues from being processed at all. Both provided data supporting their respective theories, yet a follow-up investigation found that both were probably wrong (Röseler et al., 2020).

Robinaugh et al. (2021) discuss examples of the conversion of verbal theories into formal ones. This process results in new, more specific predictions that can be derived. When a theory makes more precise predictions and the set of possible events that would contradict the theory grows, this represents an increase in empirical content (Glöckner & Betsch, 2011; Popper, 1959/2008).

Empirical Content and Strong Inference

Theories can differ in their empirical content. Concretely, this refers to how specific their predictions are. The more possible observations would refute a theory, the higher its empirical content.

Let us take the case where our theory allows us to make predictions about what kind of car will drive along a particular street at a particular time. The figure shows all possible cars. For simplicity, our example world contains only nine different cars, which differ in terms of the features color (green, black, blue), rear spoiler (with, without), and wheel color (gray, yellow).

  • The purple theory states: The observed car has gray wheels. Without a theory, all cars would be equally likely to us; the purple theory “forbids” the car from having yellow wheels. It rules out 3/9 of the cars.

  • The red theory states: The observed car is blue. The probability of refuting it would be higher in our sample world, namely 6/9. Because the red theory is, so to speak, a riskier bet a priori – that is, without further prior knowledge – it has higher empirical content.

  • The orange theory has the highest possible empirical content: The observed car is green, has no rear spoiler, and has gray wheels. It rules out all but one case (8/9).

The example with the nine possible car types is, of course, highly simplified. In certain fields, however, researchers occasionally manage to reduce the results of experiments to a few possible outcomes and thereby weigh theories against one another. Platt (1964) calls this the method of strong inference and argues that fields proceeding this way experience rapid progress. Building on this, P. E. Smaldino (2017) calls for more theories or models and argues that researchers should always offer several explanations simultaneously. This can have the advantage that researchers do not commit to a single possibility and that theories are not treated as someone’s property. As long as a theory can be clearly attributed to one person, there is a risk that criticism of the theory will be confused with criticism of the person.

Visualization of theories with different empirical content.

Deduction and Induction

Methods are being reformed, and scientists discuss how science works, how it should proceed, and which methods are sensible or nonsensical. As becomes clear from the hermeneutic circle, one path to knowledge consists of combining a set of observations into a regularity or law (induction), while another consists of deriving predictions about observations not yet made from a law or theory (deduction). This distinction is repeatedly neglected or obscured in scientific discourse. For example, a debate in consumer psychology revolved for years around which path was better, even though both paths are equally legitimate and complement one another (Calder et al., 1981). Something similar applies to conflicts between qualitative and quantitative approaches, which, formally considered, tend to proceed inductively or deductively, respectively (Borgstede & Scholz, 2021). In replication research, the inductive side has traditionally received more attention (Hüffmeier et al., 2016; Yamashita & Neiriz, 2024): every difference between a replication study and the original study is cited as a possible cause for the failure of the replication attempt, in order to preserve the trustworthiness of the original findings (Baumeister & Vohs, 2016). This overlooks the fact that minor differences between the original and replication study (e.g., the measures used, the average age of participants, the language of the instructions) are not captured by theories – and should therefore, according to the theories themselves, be irrelevant – and that a failed replication clearly reveals the limits of the theory, from which recommendations for modifying the theory can be derived (Cesario, 2014; Dijksterhuis, 2014). An overview of the approaches can be found in the following table.

Characteristics of inductive and deductive approaches; taken and adapted from an unpublished manuscript by Röseler & Leder.
Facet Deductive Approach (Theory-driven) Inductive Approach (Phenomenon-driven)
Generalizability lies in… the theory: it is maximally general a priori (e.g., it holds for all people until demonstrated otherwise). the data: only diverse observations across different contexts allow the assumption that the phenomenon is universally valid.
Change in generalizability Generality decreases with more observations. Generality increases with more observations (provided they are confirmatory in nature).
Type of test Predictions of the theory are primarily subjected to attempts at refutation. Repeated observations confirm the original individual case.
Choice of study setting Student samples from a single country or laboratory studies are unproblematic. The context of the study should reflect the target conditions (e.g., when applying the findings in practice) as closely as possible.

Auxiliary Hypotheses

Replication failures can be explained via the following paths:

  1. Type I error in the original study: The original finding was merely a chance finding or arose through scientific misconduct (see the chapter “Researchers’ Degrees of Freedom”).

  2. Type II error in the replication study: The original study was correct; the replication study made an error (e.g., too small a sample, poor calibration of instruments, or scientific misconduct).

  3. Boundary condition of the phenomenon: Both studies are trustworthy. The replication study differs in a way that matters for the theory (e.g., the replication study was conducted with people from a different country, and the theory only holds for people from the “original country”).

Option 3 is constructive and accepts both individual findings as robust. This requires a theoretically relevant difference between the original and replication study, which, given the infinite number of possible important factors, applies in most cases (Smedslund, 2015). This path can then be used to modify the theory or to formulate an additional theory that must likewise be taken into account for the context of the study. Things become difficult when researchers conduct a replication to the best of their knowledge, it “fails” (that is, it does not demonstrate what it was meant to demonstrate), and other researchers criticize the replication for having done something “wrong.” After Hagger et al. (2016), in consultation with Roy Baumeister, tested his ego depletion theory with a large-scale study, Baumeister & Vohs (2016) criticized that it should have been expected from the outset that the study would not work, and described the study as misguided. Vohs, who was a co-author of the critique, conducted another large-scale replication study a few years later. Although this time the researchers were able to follow their own advice, they again failed to find the expected effect (Vohs et al., 2021).

Further Information

  • Ramminger (2023) discusses a philosophical perspective on the relationship between theory, measurement, and replication.

  • Yarkoni (2019) argues that replication problems originate in the generalization of results to theories.

  • In a talk, Fanelli discusses the complexity of research as a reason for replication failures and proposes a theory for measuring complexity (Fanelli et al., 2022). A video of a talk is available online: https://www.youtube.com/watch?v=CEAV7420jBk.


References

Banker, S., Ainsworth, S. E., Baumeister, R. F., Ariely, D., & Vohs, K. D. (2017). The sticky anchor hypothesis: Ego depletion increases susceptibility to situational cues. Journal of Behavioral Decision Making, 87(1), 23. https://doi.org/10.1002/bdm.2022
Baumeister, R. F., & Vohs, K. D. (2016). Misguided effort with elusive implications. Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 11(4), 574–575. https://doi.org/10.1177/1745691616652878
Borgstede, M., & Scholz, M. (2021). Quantitative and qualitative approaches to generalization and replication-a representationalist view. Frontiers in Psychology, 12, 605191. https://doi.org/10.3389/fpsyg.2021.605191
Buzbas, E. O., & Devezer, B. (2023). Tension between theory and practice of replication. Journal of Trial & Error, 4(1).
Calder, B. J., Phillips, L. W., & Tybout, A. M. (1981). Designing research for application. Journal of Consumer Research, 8(2), 197. https://doi.org/10.1086/208856
Cesario, J. (2014). Priming, replication, and the hardest science. Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 9(1), 40–48. https://doi.org/10.1177/1745691613513470
Dijksterhuis, A. (2014). Welcome back theory! Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 9(1), 72–75. https://doi.org/10.1177/1745691613513472
Fanelli, D., Tan, P. B., Amaral, O. B., & Neves, K. (2022). A metric of knowledge as information compression reflects reproducibility predictions in biomedical experiments. https://doi.org/10.31222/osf.io/5r36g
Ferguson, C. J., & Heene, M. (2012). A vast graveyard of undead theories: Publication bias and psychological science’s aversion to the null. Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 7(6), 555–561. https://doi.org/10.1177/1745691612459059
Francis, Z., Milyavskaya, M., Lin, H., & Inzlicht, M. (2018). Development of a within-subject, repeated-measures ego-depletion paradigm. Social Psychology, 49(5), 271–286. https://doi.org/10.1027/1864-9335/a000348
Glöckner, A., & Betsch, T. (2011). The empirical content of theories in judgment and decision making: Shortcomings and remedies. Judgment and Decision Making, 6(8), 711–721. https://doi.org/10.1017/s1930297500004149
Hagger, M. S., Chatzisarantis, N. L. D., Alberts, H., Anggono, C. O., Batailler, C., Birt, A. R., Brand, R., Brandt, M. J., Brewer, G., Bruyneel, S., Calvillo, D. P., Campbell, W. K., Cannon, P. R., Carlucci, M., Carruth, N. P., Cheung, T., Crowell, A., Ridder, D. T. D. de, Dewitte, S., … Zwienenberg, M. (2016). A multilab preregistered replication of the ego-depletion effect. Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 11(4), 546–573. https://doi.org/10.1177/1745691616652873
Hoyningen-Huene, P., & Kincaid, H. (2023). What makes economics special: Orientational paradigms. J. Econ. Methodol., 1–15. https://doi.org/10.1080/1350178x.2023.2192231
Hüffmeier, J., Mazei, J., & Schultze, T. (2016). Reconceptualizing replication as a sequence of different studies: A replication typology. Journal of Experimental Social Psychology, 66, 81–92. https://doi.org/10.1016/j.jesp.2015.09.009
Lakens, D. (2023). Concerns about replicability, theorizing, applicability, generalizability, and methodology across two crises in social psychology. https://doi.org/10.31234/osf.io/dtvs7
Muthukrishna, M., & Henrich, J. (2019). A problem in theory. Nature Human Behaviour, 349(Suppl 1), aac4716. https://doi.org/10.1038/s41562-018-0522-1
Platt, J. R. (1964). Strong inference: Certain systematic methods of scientific thinking may produce much more rapid progress than others. Science, 146(3642), 347–353.
Popper, K. R. (1959/2008). The logic of scientific discovery (Repr. 2008). Routledge Classics; Routledge. https://doi.org/10.4324/9780203994627
Ramminger, J. J. (2023). Vermessen? Zur möglichkeit philosophischer beiträge für den diskurs der quantitativen psychologie. Cultura & Psyché, 4(2), 215–224. https://doi.org/10.1007/s43638-023-00081-3
Robinaugh, D. J., Haslbeck, J. M. B., Ryan, O., Fried, E. I., & Waldorp, L. J. (2021). Invisible hands and fine calipers: A call to use formal theory as a toolkit for theory construction. Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 16(4), 725–743. https://doi.org/10.1177/1745691620974697
Röseler, L., Schütz, A., Baumeister, R. F., & Starker, U. (2020). Does ego depletion reduce judgment adjustment for both internally and externally generated anchors? Journal of Experimental Social Psychology, 87, 103942. https://doi.org/10.1016/j.jesp.2019.103942
Smaldino, P. (2019). Better methods can’t make up for mediocre theory. Nature, 575(7781), 9. https://doi.org/10.1038/d41586-019-03350-5
Smaldino, P. E. (2017). Models are stupid, and we need more of them. In R. R. Vallacher (Ed.), Computational social psychology (pp. 311–331). Routledge. https://doi.org/10.4324/9781315173726-14
Smedslund, J. (2015). Why psychology cannot be an empirical science. Integrative Psychological & Behavioral Science. https://doi.org/10.1007/s12124-015-9339-x
Vohs, K. D., Schmeichel, B. J., Lohmann, S., Gronau, Q. F., Finley, A. J., Ainsworth, S. E., Alquist, J. L., Baker, M. D., Brizi, A., & Bunyi, A. (2021). A multisite preregistered paradigmatic test of the ego-depletion effect. Psychological Science, 32(10), 1566–1581. https://doi.org/10.31234/osf.io/e497p
Yamashita, T., & Neiriz, R. (2024). Why replicate? Systematic review of calls for replication in language teaching. Research Methods in Applied Linguistics, 3(1), 100091. https://doi.org/10.1016/j.rmal.2023.100091
Yarkoni, T. (2019). The generalizability crisis. https://doi.org/10.31234/osf.io/jqw35