Stocktaking

A similar crisis of confidence occurred in social psychology in the 1960s (Lakens, 2023). A crucial difference between the old crisis and the current one is the stocktaking that has taken place: parallel to the peculiar findings on Bem’s precognition studies, psychologists around Brian Nosek formed an international network and investigated the replicability of 100 studies from prominent psychology journals (Open Science Collaboration, 2015). They found that only 39 of the 100 original findings could be replicated. For all other studies, the replication results differed from the original results. Many further large-scale projects followed, all with similar results: replication rates fell far below what had been hoped for.

A critical look at the Open Science Collaboration, 2015

Although this “Reproducibility Project: Psychology” shook the entire field to its core and paved the way for a paradigm shift, some researchers also point to negative effects on subsequent replication research. By publishing 100 studies simultaneously through a group of more than 100 researchers, the project set unrealistic standards for replication research. At the same time, quality control was less rigorous, since the individual studies could not all be reviewed to the extent that would have been the case for a traditional publication, e.g. (Röseler et al., 2022). Some equally ambitious projects have since been published, such as the Many Labs studies, e.g. (Klein et al., 2014; Klein et al., 2018), or attempts in which independent groups tested and replicated the same hypotheses (Landy et al., 2020). Such projects are often limited to studies that can be replicated within an online survey. Formats such as longitudinal studies or behavioral observations are underrepresented due to the greater practical difficulty involved.

Numerous consortia followed. Some projects focused on individual phenomena. For example, 17 research groups joined forces to replicate the facial feedback finding (Strack et al., 1988) in (Wagenmakers et al., 2016). In this experiment, participants hold a pen in their teeth, and depending on the pen’s orientation, they either tense the muscles used for smiling or do not. In the “smiling” condition, participants subsequently rated comics as funnier. The replication failed. In 2022, a further study with more than 3,000 participants from 19 countries was published – this time with direct involvement from Fritz Strack, who had conducted the original study (Coles et al., 2022). Again, it was shown that the position of a pen in the mouth has no effect on the evaluation of stimuli. Beyond social-psychological findings, researchers also focused on areas such as research with infants (Byers-Heinlein et al., 2020) or studies from specific journals (Camerer et al., 2018).

Defining replicability

Determining what could and could not be replicated is only possible once replication success has been defined. In common usage among researchers, “was replicated” means that a study investigating the same research question arrived at results similar to those of an original study. Replication failures are described as “could not be replicated.” Subtly different from this, “was not replicated” can mean that no replication attempts exist, or that they could not even be conducted because the original study failed to adequately describe key details (Errington et al., 2021). For a science that is over 100 years old, it seems surprising that there is still no clear definition of key concepts surrounding replication, let alone that replicating studies has become routine. While different fields have settled on divergent taxonomies – that is, models for classifying different types of replication – terms related to replication are used in this book as described in the table below. Depending on whether the same or different data and the same or different analyses are used, we speak of reproducibility, replicability, robustness, or generalizability.

Replication taxonomy following the Turing Way (The Turing Way Community & Scriberia, 2024).
Data
same different
Analysis same reproducible replicable
different robust generalizable
Statistical comparison of original and replication findings

Every measurement carries some degree of imprecision. In the social sciences, measuring properties or phenomena such as decision-making heuristics is extremely difficult. When original results are compared with replication results, both results carry this imprecision. This makes it difficult to say whether differences arose from random fluctuation or from problems with the original study. There are numerous statistical methods for comparing the two results. They typically differ in how much they account for the imprecision of the original study. Because of older methodological standards, original results tend to be extremely imprecise, and the more heavily they are weighted, the more favorable the comparison turns out – that is, the higher the estimated replication rate. Depending on the method used, replication rates can then range between 40% and 80%. A comparison of criteria against the FORRT replication database is available online.

Further reading

For a more systematic taxonomy of replication types, grounded in information science, see Plesser (2018). LeBel et al. (LeBel et al., 2019) have proposed a taxonomy for replication study outcomes based on statistical methods. The closeness of replications is discussed philosophically, for example, by Choi (2023) and Leonelli (2023).

Replication taxonomy.
Distinguishing criterion Categories
Outcome of a replication study

Successful

Failed

Unclear or mixed

Closeness of a replication study to the original study (following LeBel et al. [2019] and Hüffmeier et al. [2016])

Direct replication

(same experimenters,
same experimental materials,
new participants)

Close replication
(different experimenters,
experimental materials as similar as possible,
new participants)

Conceptual or constructive replication
(different experimenters,
different experimental materials,
new participants)

Goal of the replication

Reproduction
Arriving at the same results using the same data and the same code

Replication
Arriving at the same results using different data

“One swallow does not make a summer”

Whether a scientific finding “holds water” – that is, whether it has a valid claim to truth – depends, in replication research, on many factors beyond the way it was originally established. What were the results of the replication study? How many studies were conducted, and how varied were they? What exactly did the methods look like? What were the differences between the replications and the original study? While individual studies always yield some gain in knowledge, at the very least whether a particular method is feasible (Sikorski & Andreoletti, 2023), they can vary considerably depending on the field of research (Landy et al., 2020; McShane et al., 2022). To see the bigger picture, more is needed – for example, a statistical summary of many individual studies combined into one overall study (a meta-analysis). An example using fictitious data is shown in the following figure.

If one looks at many studies that have examined the relationship between two things – for instance, income and educational attainment – the studies will differ in their details: what exactly was counted as income (net, gross, social benefits, family members’ income, values over a period of time or from a specific point in the past, etc.), or which people were surveyed (students, working professionals, whether respondents were paid, etc.). All these differences may affect the relationship in question, and even when they do not, relationships are often subject to fluctuations arising from the measurement methods themselves. In this example, the strength of the relationship can be reduced to a single number along with an associated precision. Here, the number is called the “correlation” and the precision the “error bar.” The forest plot (also called a blobbogram) displays possible correlations from various studies. Within a meta-analysis, study results can then be combined and differences examined.

Figure 1: Forest plot with simulated correlations from 15 fictitious studies.

Phenomenon-centered replication projects

In contrast to the broadly scoped Reproducibility Project: Psychology (Open Science Collaboration, 2015) and other attempts to estimate replication rates for entire fields (Brodeur et al., 2024; Camerer et al., 2016; Feldman, 2021), other efforts have focused on fundamental phenomena. In such cases, dozens of groups around the world have joined forces, agreed on a study design, and carried out the studies with an enormous number of participants. Most of these projects originate from psychology. While the effect sizes found in them – that is, roughly speaking, the magnitude of a relationship or finding – were in almost all cases far below those of previous studies (Kvarven et al., 2020), for the majority of the studies they were also null, meaning the phenomena were “not detectable” at all (Alogna et al., 2014; Bouwmeester et al., 2017; Cheung et al., 2016; Eerland et al., 2016; O’Donnell et al., 2018; Rife et al., 2024; Vaidis et al., 2024; Wagenmakers et al., 2016). For instance, it was shown with enormous precision that a story about a professor does not make participants perform better on a subsequent test of intelligence (O’Donnell et al., 2018).

Efficient use of resources?

How should resources be handled in replications? This question inevitably arises whenever many researchers join forces. Does each group create the study independently? Does everyone adhere to a jointly agreed protocol? Do they run the study sequentially, so as to learn from one another? In Registered Replication Reports, a study design is typically agreed upon in advance with other researchers (e.g. the authors of the original study). In other cases, a study design is developed collaboratively that should be ideal for testing the theory (the creative destruction approach, Tierney et al., 2020). Teams in different countries then translate the protocol and adhere to it closely during implementation. These protocols are sometimes not pre-tested (Buttliere, 2024), but they are often based on successful, well-known studies. This has the advantage that differences between groups cannot be attributed to differences in implementation, and that cultures can be compared (Kakinohana et al., 2022). A disadvantage, however, is that if the experiment already fails to work at one, two, or five sites, it becomes questionable whether the remaining 30 groups should even attempt it. In the words of Buttliere (2024): “Who gets better results? Thirty-nine people doing something for the first time, or one person doing something 39 times?”

Discipline-centered replication projects

Fewer than half of all psychological findings are replicable, then. Does this mean that all social-science textbooks across all disciplines are half wrong? The clear answer is no. The accurate answer is it depends.

Beyond psychology

The extent to which this depends on the discipline within the social sciences has so far been examined primarily in psychology. Current trends suggest that replication rates in personality psychology and cognitive psychology (Soto, 2019) are higher than in social psychology (Open Science Collaboration, 2015) or in marketing (Charlton, 2022). While hundreds of replication attempts for social-psychological studies have already been published, there are currently far fewer in other fields such as marketing – as of October 2022, only nine. Fields outside psychology are likewise affected by replication problems. Almost all disciplines are affected by problems of replicability, reproducibility, and traceability. New approaches to addressing these issues are being discussed in medicine, biology, chemistry, physics, history, political science, education, computer science, and many other fields.


References

Alogna, V. K., Attaya, M. K., Aucoin, P., Bahník, Š., Birch, S., Birt, A. R., Bornstein, B. H., Bouwmeester, S., Brandimonte, M. A., Brown, C., Buswell, K., Carlson, C., Carlson, M., Chu, S., Cislak, A., Colarusso, M., Colloff, M. F., Dellapaolera, K. S., Delvenne, J.-F., … Zwaan, R. A. (2014). Registered replication report: Schooler and engstler-schooler (1990). Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 9(5), 556–578. https://doi.org/10.1177/1745691614545653
Bouwmeester, S., Verkoeijen, P. P. J. L., Aczel, B., Barbosa, F., Bègue, L., Brañas-Garza, P., Chmura, T. G. H., Cornelissen, G., Døssing, F. S., Espín, A. M., Evans, A. M., Ferreira-Santos, F., Fiedler, S., Flegr, J., Ghaffari, M., Glöckner, A., Goeschl, T., Guo, L., Hauser, O. P., … Wollbrant, C. E. (2017). Registered replication report: Rand, greene, and nowak (2012). Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 12(3), 527–542. https://doi.org/10.1177/1745691617693624
Brodeur, A., Mikola, D., & Cook, N. (2024). Mass reproducibility and replicability: A new hope. https://doi.org/10.2139/ssrn.4790780
Buttliere, B. (2024). Was this registered report pilot tested? Examination of vaidis, sleegers, van leeuwen, DeMarree, ... & priolo, d. (2024). In PsyArXiv. https://doi.org/10.31234/osf.io/c6r8x
Byers-Heinlein, K., Bergmann, C., Davies, C., Frank, M. C., Hamlin, J. K., Kline, M., Kominsky, J. F., Kosie, J. E., Lew-Williams, C., & Liu, L. (2020). Building a collaborative psychological science: Lessons learned from ManyBabies 1. Canadian Psychology/Psychologie Canadienne, 61(4), 349. https://doi.org/10.31234/osf.io/dmhk2
Camerer, C. F., Dreber, A., Forsell, E., Ho, T.-H., Huber, J., Johannesson, M., Kirchler, M., Almenberg, J., Altmejd, A., Chan, T., Heikensten, E., Holzmeister, F., Imai, T., Isaksson, S., Nave, G., Pfeiffer, T., Razen, M., & Wu, H. (2016). Evaluating replicability of laboratory experiments in economics. Science (New York, N.Y.), 351(6280), 1433–1436. https://doi.org/10.1126/science.aaf0918
Camerer, C. F., Dreber, A., Holzmeister, F., Ho, T.-H., Huber, J., Johannesson, M., Kirchler, M., Nave, G., Nosek, B. A., Pfeiffer, T., Altmejd, A., Buttrick, N., Chan, T., Chen, Y., Forsell, E., Gampa, A., Heikensten, E., Hummer, L., Imai, T., … Wu, H. (2018). Evaluating the replicability of social science experiments in nature and science between 2010 and 2015. Nature Human Behaviour, 2(9), 637–644. https://doi.org/10.1038/s41562-018-0399-z
Charlton, A. (2022). Replications of marketing studies. https://openmkt.org/research/replications-of-marketing-studies/
Cheung, I., Campbell, L., LeBel, E. P., Ackerman, R. A., Aykutog˘lu, B., Bahník, Š., Bowen, J. D., Bredow, C. A., Bromberg, C., Caprariello, P. A., Carcedo, R. J., Carson, K. J., Cobb, R. J., Collins, N. L., Corretti, C. A., DiDonato, T. E., Ellithorpe, C., Fernández-Rouco, N., Fuglestad, P. T., … Yong, J. C. (2016). Registered replication report: Study 1 from finkel, rusbult, kumashiro, & hannon (2002). Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 11(5), 750–764. https://doi.org/10.1177/1745691616664694
Choi, H. H. (2023). In defense of the resampling account of replication. Journal of Theoretical and Philosophical Psychology, 43(4), 249–251. https://doi.org/10.1037/teo0000224
Coles, N. A., March, D. S., Marmolejo-Ramos, F., Larsen, J. T., Arinze, N. C., Ndukaihe, I. L. G., Willis, M. L., Foroni, F., Reggev, N., Mokady, A., Forscher, P. S., Hunter, J. F., Kaminski, G., Yüvrük, E., Kapucu, A., Nagy, T., Hajdu, N., Tejada, J., Freitag, R. M. K., … Liuzza, M. T. (2022). A multi-lab test of the facial feedback hypothesis by the many smiles collaboration. Nature Human Behaviour. https://doi.org/10.1038/s41562-022-01458-9
Eerland, A., Sherrill, A. M., Magliano, J. P., Zwaan, R. A., Arnal, J. D., Aucoin, P., Berger, S. A., Birt, A. R., Capezza, N., Carlucci, M., Crocker, C., Ferretti, T. R., Kibbe, M. R., Knepp, M. M., Kurby, C. A., Melcher, J. M., Michael, S. W., Poirier, C., & Prenoveau, J. M. (2016). Registered replication report: Hart & albarracín (2011). Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 11(1), 158–171. https://doi.org/10.1177/1745691615605826
Errington, T. M., Mathur, M., Soderberg, C. K., Denis, A., Perfito, N., Iorns, E., & Nosek, B. A. (2021). Investigating the replicability of preclinical cancer biology. eLIFE, 10. https://doi.org/10.7554/eLife.71601
Feldman, G. (2021). Replications and extensions of classic findings in judgment and decision making. https://doi.org/10.17605/OSF.IO/5Z4A8
Kakinohana, R. K., Pilati, R., & Klein, R. A. (2022). Does anchoring vary across cultures? Expanding the many labs analysis. European Journal of Social Psychology. https://doi.org/10.1002/ejsp.2924
Klein, R. A., Ratliff, K. A., Vianello, M., Adams, R. B., Bahník, Š., Bernstein, M. J., Bocian, K., Brandt, M. J., Brooks, B., Brumbaugh, C. C., Cemalcilar, Z., Chandler, J., Cheong, W., Davis, W. E., Devos, T., Eisner, M., Frankowska, N., Furrow, D., Galliani, E. M., … Nosek, B. A. (2014). Investigating variation in replicability. Social Psychology, 45(3), 142–152. https://doi.org/10.1027/1864-9335/a000178
Klein, R. A., Vianello, M., Hasselman, F., Adams, B. G., Adams, R. B., Alper, S., Aveyard, M., Axt, J., Babalola, M. T., Bahník, Š., Berkics, M., Bernstein, M. J., Berry, D. R., Bialobrzeska, O., Bocian, K., Brandt, M., Busching, R., Cai, H., Cambier, F., … Nosek, B. A. (2018). Many labs 2: Investigating variation in replicability across sample and setting. https://doi.org/10.31234/osf.io/9654g
Kvarven, A., Strømland, E., & Johannesson, M. (2020). Comparing meta-analyses and preregistered multiple-laboratory replication projects. Nature Human Behaviour, 4(4), 423–434. https://doi.org/10.1038/s41562-019-0787-z
Lakens, D. (2023). Concerns about replicability, theorizing, applicability, generalizability, and methodology across two crises in social psychology. https://doi.org/10.31234/osf.io/dtvs7
Landy, J. F., Jia, M. L., Ding, I. L., Viganola, D., Tierney, W., Dreber, A., Johannesson, M., Pfeiffer, T., Ebersole, C. R., Gronau, Q. F., Ly, A., van den Bergh, D., Marsman, M., Derks, K., Wagenmakers, E.-J., Proctor, A., Bartels, D. M., Bauman, C. W., Brady, W. J., … Uhlmann, E. L. (2020). Crowdsourcing hypothesis tests: Making transparent how design choices shape research results. Psychological Bulletin, 146(5), 451–479. https://doi.org/10.1037/bul0000220
LeBel, E. P., Vanpaemel, W., Cheung, I., & Campbell, L. (2019). A brief guide to evaluate replications. Meta-Psychology, 3. https://doi.org/10.15626/MP.2018.843
Leonelli, S. (2023). Philosophy of open science. Cambridge University Press. https://doi.org/10.1017/9781009416368
McShane, B. B., Böckenholt, U., & Hansen, K. T. (2022). Variation and covariation in large-scale replication projects: An evaluation of replicability. J. Am. Stat. Assoc., 117(540), 1605–1621. https://doi.org/10.1080/01621459.2022.2054816
O’Donnell, M., Nelson, L. D., Ackermann, E., Aczel, B., Akhtar, A., Aldrovandi, S., Alshaif, N., Andringa, R., Aveyard, M., Babincak, P., Balatekin, N., Baldwin, S. A., Banik, G., Baskin, E., Bell, R., Białobrzeska, O., Birt, A. R., Boot, W. R., Braithwaite, S. R., … Zrubka, M. (2018). Registered replication report: Dijksterhuis and van knippenberg (1998). Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 13(2), 268–294. https://doi.org/10.1177/1745691618755704
Open Science Collaboration. (2015). Psychology: Estimating the reproducibility of psychological science. Science (New York, N.Y.), 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
Plesser, H. E. (2018). Reproducibility vs. Replicability: A brief history of a confused terminology. Frontiers in Neuroinformatics, 11. https://doi.org/10.3389/fninf.2017.00076
Rife, S., Lambert, Q., Calin-Jageman, R., Matus, A., Banik, G., Barberia, I., Beaudry, J., Bernauer, H., Calvillo, D., & Chopik, W. (2024). Registered replication report: Study 3 from trafimow and hughes (2012). Advances in Methods and Practices in Psychological Science. https://doi.org/10.31234/osf.io/esu9z_v2
Röseler, L., Gendlina, T., Krapp, J., Labusch, N., & Schütz, A. (2022). Successes and failures of replications: A meta-analysis of independent replication studies based on the OSF registries. https://doi.org/10.31222/osf.io/8psw2
Sikorski, M., & Andreoletti, M. (2023). Epistemic functions of replicability in experimental sciences: Defending the orthodox view. Found. Sci. https://doi.org/10.1007/s10699-023-09901-4
Soto, C. J. (2019). How replicable are links between personality traits and consequential life outcomes? The life outcomes of personality replication project. Psychological Science, 30(5), 711–727. https://doi.org/10.1177/0956797619831612
Strack, F., Martin, L. L., & Stepper, S. (1988). Inhibiting and facilitating conditions of the human smile: A nonobtrusive test of the facial feedback hypothesis. Journal of Personality and Social Psychology, 54(5), 768–777. https://doi.org/10.1037/0022-3514.54.5.768
The Turing Way Community, & Scriberia. (2024). Illustrations from the turing way: Shared under CC-BY 4.0 for reuse [figure]. Zenodo. https://doi.org/10.5281/zenodo.13882307
Tierney, W., Hardy III, J. H., Ebersole, C. R., Leavitt, K., Viganola, D., Clemente, E. G., Gordon, M., Dreber, A., Johannesson, M., & Pfeiffer, T. (2020). Creative destruction in science. Organizational Behavior and Human Decision Processes, 161, 291–309.
Vaidis, D. C., Sleegers, W. W., Van Leeuwen, F., DeMarree, K. G., Sætrevik, B., Ross, R. M., Schmidt, K., Protzko, J., Morvinski, C., & Ghasemi, O. (2024). A multilab replication of the induced-compliance paradigm of cognitive dissonance. Advances in Methods and Practices in Psychological Science, 7(1), 25152459231213375.
Wagenmakers, E.-J., Beek, T., Dijkhoff, L., Gronau, Q. F., Acosta, A., Adams, R. B., Albohn, D. N., Allard, E. S., Benning, S. D., Blouin-Hudon, E.-M., Bulnes, L. C., Caldwell, T. L., Calin-Jageman, R. J., Capaldi, C. A., Carfagno, N. S., Chasten, K. T., Cleeremans, A., Connell, L., DeCicco, J. M., … Zwaan, R. A. (2016). Registered replication report: Strack, martin, & stepper (1988). Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 11(6), 917–928. https://doi.org/10.1177/1745691616674458