Methods

In science there is no one method (Feyerabend, 1975/2002); rather, methods are developed for problems, problems are solved, and methods are refined or abandoned. Methods are thus like tools, and not everything can be assembled with a single screwdriver. Open science reforms bring countless methodological innovations, improvements, and proposals that researchers often find overwhelming: advancing research projects, teaching seminars and lectures, raising external funding, and now also open science? While many problems are attributed to inadequate methods training (Lakens, 2021b), which researchers then have to make up for on their own time, some methods actually make the work easier. For doctoral students in biological psychology, for example, the guide ARIADNE was developed (https://igor-biodgps.github.io/ARIADNE/graph/graph.html). This chapter offers an overview of methodological developments and debates in the social sciences and in disciplines that primarily work with statistical methods.

Meta-Analyses

Meta-analyses are studies of studies. Researchers typically extract results from already published studies, write to other researchers in a field to ask about unpublished studies, and use statistical methods to analyze the similarities and differences between the results. Fletcher (2022) argues that only meta-analyses can demonstrate the generalizability of (statistical) phenomena. In an ideal world, everyone could carry on as before and meta-analytic models would simply correct for the problems. Given the devastating extent of publication bias, however, this is currently not possible. How meta-analyses can nevertheless be informative is discussed in general terms by Carlsson et al. (2024), and I list specific problems and solutions below.

Assessing Publication Bias and P-Hacking

For meta-analyses, the rule is: garbage in, garbage out. Anyone who combines many poorly conducted studies in a meta-analysis ends up with a poor summary. This is, for instance, what happened when Hagger et al. (2010) found a substantial effect for a model of willpower in their meta-analysis, yet subsequent large-scale replication attempts and analyses all failed to find an effect of comparable size (Dang et al., 2020; Friese et al., 2018; Hagger et al., 2016; Vohs et al., 2021). What can still be done — and should be done in every meta-analysis — is an assessment of data quality, for example the extent of publication bias. There are methods that check whether unpublished studies exist, and methods that correct for potentially missing studies. Some of these only work with more than 200 studies, while others can already be applied to a dozen studies.

Funnel Plot

One of the oldest methods is the funnel plot (Light & Pillemer, 1984). It plots the precision and effect size of individual studies in a single diagram. Ideally, the points should form the shape of a funnel: the more precise a study is (for example, due to a large sample), the closer the association it measures should lie to the true mean. Less precise studies deviate unsystematically, sometimes above and sometimes below. Because non-significant results are rarely published, funnel plots almost never actually show a funnel shape — instead, the non-significant results are simply missing.

The figure below shows a funnel plot for a study on the relationship between math anxiety and math performance. The pattern is nearly symmetrical, indicating only weak publication bias. The small associations at the bottom edge are skewed slightly to the right, and the precise effects at the top are not all the same size but vary considerably.

Figure 1: Funnel plot for the relationship between math anxiety and math performance.

P-Curve

In p-hacking, data are analyzed in multiple ways and only the analysis yielding a low, and therefore significant, p-value is reported. Results thus become significant not because the hypotheses are correct, but because the data were analyzed until they became significant. In the chapter on p-hacking we saw how p-values are distributed depending on whether the hypothesis is correct or not. The p-curve (Simonsohn et al., 2014b, 2014a; Simonsohn et al., 2015) makes use of exactly this fact. The p-values from a set of studies are plotted in a diagram and their distribution is examined. If there is no p-hacking, the values are either uniformly distributed (all p-values occur equally often) or cluster near 0 (smaller p-values occur more often). P-hacking, however, causes the values to cluster near the 5% threshold, since the data need not be “hacked” any further than that. The method gained particular prominence when Simmons & Simonsohn (2017) applied it to one of the most famous phenomena in psychology, power posing, and found that p-hacking had likely occurred there. Critics later pointed out other ways a suspicious p-curve can arise even without p-hacking, and the method is now rarely used (Erdfelder & Heck, 2019).

Z-Curve

Instead of using p-values, meta-analytic findings can also be converted into so-called z-values. These are normally distributed, and additional algorithms can be used to estimate, based on observed effects, how many further effects there ought to be. This method, called the z-curve (Bartoš & Schimmack, 2022), can thus correct for the file-drawer effect and for p-hacking. Its output also includes an estimate of what the replication rate would be if all the analyzed studies were conducted again. According to recent studies, these estimates work quite well, although they require a large amount of data (Röseler, 2023; Sotola, 2023; Sotola & Credé, 2022).

Sensitivity Analysis

An approach that works not only for meta-analyses but for almost all statistical analyses is the so-called sensitivity or robustness analysis. Different analytical pathways are run through and the extent to which they affect the results is examined. In meta-analyses, for example, many possible procedures can be computed at once. Such a “shotgun” approach was coined by Kepes et al. (2017) and has since been adopted in other studies (Körner et al., 2022). Carter et al. (2019) provide an overview of various procedures and the conditions under which they are suitable for particular data.

Forensic Meta-Science

Various fields have developed methods for identifying implausible data patterns. Such techniques check, for instance, whether reported values are even possible: if 10 people each have a value of either 0 or 1, the mean of those 10 values cannot be 0.15, but only 0, 0.1, 0.2, 0.3, … 1.0. A collection of guides for checking such problems is COSIG (Collection of Open Science Integrity Guides) (Richardson, 2025).

Rules of Thumb for Assessing Individual Articles

A meta-analysis is laborious and can take several years. Even researchers who lack the funds for student assistants to help code and check hundreds or thousands of studies have little chance of conducting a proper analysis on their own. The following rules of thumb — and they are not meant to be more than that — offer shortcuts for assessing scientific quality.

Many Significant Studies

Since a crisis of confidence in social psychology in the 1960s (Lakens, 2023), many journals have required multiple studies per research article. As a result, resources have been invested in several smaller studies rather than one solid study. Most of these “multi-study papers” then report exclusively significant results across as many as 10 studies. While many studies with uniformly significant results may look impressive at first glance, closer inspection raises suspicion: individual studies typically have an 80-95% probability of yielding a significant result in the central analysis. This probability (statistical power) decreases when several studies are run in sequence. It is comparable to a marksman who hits a glass bottle with a rifle 99% of the time. The probability that he hits a single bottle with one shot is thus 99%. The probability that he hits 50 out of 50 bottles with 50 shots is lower, namely 99%^50 (to the power of fifty) = 60.5%. Scientific studies undergo a similar “power deflation.” The probability of conducting 4 significant studies, each with 80% power, is 40.96%. Actually publishing exactly such a set of studies is then extremely unlikely (Lakens & Etz, 2017).

Effect Sizes “Just Barely Significant”

Following the logic of the p-curve, it is unlikely that p-values fall between 1% and 5%. Due to p-hacking, however, this happens frequently. A p-value close to 5% also comes with a confidence interval for the effect size that is close to 0, e.g., (Jané et al., 2024). Suppose someone conducts two studies on a topic and both have p-values near 5% with roughly equal numbers of participants; the question then arises why the sample size was not increased for the later study — given a result that was only just significant, it is clear that the researcher was “lucky,” since statistical power was not particularly high. When planning the sample for the next study, one should therefore build on the first study and adjust the plan accordingly, e.g., (Lakens, 2021a).

Figure 2: Width of confidence intervals as a function of correlation size for N = 250.

Checking for Reproducibility

If a finding is tested again using the same data and, ideally, the same program or analysis code, to check whether the same numbers come out — not merely whether the hypothesis is confirmed again — this is called a reproduction of the results. Unlike a replication, no new data are collected. That results are reproducible ought to be the absolute minimum standard for scientific reports, yet it is far from being met. Reproduction studies are still rare. The average success rate across fields (economics, education, biomedicine, health sciences, geosciences) is around 50% (C. Chang & Li, 2022; Cobey et al., 2023; Koukouraki & Kray, 2023), and even the success rate for simple analyses using programs that explicitly output reproducibility protocols is extremely low (Thibault et al., 2024).

Terminological Confusion

While the term replication in economics refers both to testing an existing study with new data and to re-testing with the same data, psychology uses the term reproduction for the latter. In biological reproduction research, the term “reproducibility” is used instead to avoid ambiguity. In yet other cases, such as Open Science Collaboration (2015), replications (new data) are referred to as “reproducibility,” while repeated tests with the same data are called “computational reproducibility.” Finally, in some areas the boundaries blur — for example, when a replication of the PISA study findings uses partly the same and partly new data, or when the data are computer-generated and the same program can use a pseudo-random number generator to produce different data with the same underlying structure.

For a few years now, the journal Meta-Psychology has been one of the first in psychology to conduct reproducibility checks for all published articles. These are carried out by researchers voluntarily or as part of their work for the journal. While this practice has already been called for at other journals (Lindsay, 2023), it is still the exception. Reproduction checks from all sorts of disciplines can be published at Rescience (http://rescience.github.io). Independent of journals, the Codecheck community offers to check the code of any research article to verify that it runs and that the results can be recomputed (Nüst & Eglen, 2021). Articles with verified code can then reference the CODECHECK report, which is uploaded as a public article on Zenodo.org.

For 2024, the Institute for Replication announced that it would reproduce studies from the journal Nature Human Behaviour (“Promoting Reproduction and Replication at Scale,” 2024). Nature Human Behaviour is one of the most prestigious journals in research on human behavior — though prestige should not be equated with scientific quality. It is managed by the Springer publishing group and charges the highest article processing fee, around €9,000 per article. The strategic decision to focus on articles from that journal has the advantage that the people conducting the reproducibility checks may be able to publish them there, and that the reproductions receive considerable attention. Given the quality standards such journals claim for themselves, and the fact that free journals like Meta-Psychology can carry out the procedure without external help from the Institute for Replication, we again see the familiar pattern of publishers exploiting their prestige to extract free, profit-generating labor from the scientific community. In the end, it is once again not the journal itself that contributes to scientific quality assurance, but the Institute for Replication.

A shortcut for checking correctness, used by many journals, is the program statcheck. It automatically detects classic statistical tests and checks, based on the reported values, whether they are internally consistent. Hartgerink (2016) checked results from over 50,000 articles with the program and had the articles commented on via PubPeer. Because the algorithm — as openly acknowledged in the comments — occasionally flags values as erroneous incorrectly, and because the authors of the articles were not warned about the comments beforehand, the DGPs (German Psychological Society) condemned the practice (German-language statement). The responses from the Statcheck group and from Chris Hartgerink are no longer available.

Reproduction at the Push of a Button

Push-button replications refer to results that can be recomputed by any researcher with little effort — at the push of a button, as it were. While social science journals increasingly require that data and analysis code be published in a form that allows results to be recomputed, the journal Image Processing Online (IPOL, https://www.ipol.im) embodies the ideal of this approach: for every article published there, a demo is available in which, after selecting an image, the algorithm published in the article is run live.

Large-Scale Reproduction Projects

Various research disciplines have launched large-scale projects to estimate reproducibility for entire disciplines. A pioneer in this field was Höffler’s ReplicationWiki (https://replication.uni-goettingen.de). Subsequent projects such as the Replication Network (https://replicationnetwork.com) drew heavily on the data compiled there. For economics, Brodeur, Mikola, et al. (2024) reported a reproducibility rate of 70%, and in management science 55% (Fišar et al., 2024). The Institute for Replication (I4R) also overlaps with the Social Science Reproduction Platform of the Berkeley Initiative for Transparency in the Social Sciences (BITSS; https://www.bitss.org/resources/social-science-reproduction-platform). While I4R published a database of all results in 2024, the BITSS platform has already been available for some time. Alongside these predominantly economics- and political-science-oriented initiatives, the interdisciplinary Multi100 project investigates the robustness of 100 results from various fields (https://osf.io/q5h2c). A completed project originating in psychology was a university’s offer to check researchers’ work for reproducibility (Baker et al., 2023). In the geosciences, CODECHECK (https://codecheck.org.uk) offers a community in which researchers certify the one-time reproducibility of code. A journal that publishes reproducibility checks is Rescience C (https://rescience.github.io).

Open Code

Publicly available data and code are necessary for reproduction and robustness checks. Journals face a trade-off here between making submission harder — and thus making themselves less attractive — by imposing higher requirements, and promoting scientific quality assurance. A similar problem exists among the operators of panels in which large surveys or performance tests are regularly conducted, such as the international PISA study or the German Socio-Economic Panel (SOEP), a long-running household panel run by DIW Berlin. In published analyses of SOEP data, code is shared in only 20% of cases (German-language source: https://www.wifa.uni-leipzig.de/fileadmin/Fakultät_Wifa/Institut_für_Theoretische_Volkswirtschaftslehre/Professur_Makroökonomik/Economics_Research_Seminar/ERS-Paper_Marcus.pdf).

Robustness Analyses

Similar to sensitivity or robustness analyses in meta-analyses, individual studies can also explore further paths in the “garden of forking paths.” As a reminder: the path from data to results is long and involves many different decisions. To show that the result does not actually depend on these decisions, one can demonstrate what the results look like if different decisions had been made. The most extreme form of such robustness analyses is the multiverse analysis (see, e.g., Mazei et al., 2025 for recommendations). Here, an attempt is made to make all possible decisions simultaneously. The resulting set of results is then analyzed or presented in some form (e.g., averaged) (Jacobsen et al., 2024). Another option is the multi-analyst study. This concerns the dependence of results on the decisions of different researchers, with many people analyzing the data independently of one another. In the end, it is checked how strongly the results agree across researchers.

The figure below shows the results of different analysis methods applied to a fixed dataset (fictional data). Different types of correlations, different sample sizes, and different hypotheses were used. The result changes slightly each time, so that the value ranges between 0.20 and 0.35, but the positive (and significant) correlation is preserved throughout.

Figure 3: Results of different analysis methods for a fixed, fictional dataset.

Reproducible Manuscripts

In the online version of this book, it is possible to display the code behind a figure before the figure itself. With this code, the figure can easily be reconstructed or reproduced. Research articles, too, can be written in this way. Text, programming code, and results in the form of numbers, tables, and diagrams are all written within a single program, sparing researchers the need to copy or retype numbers. Recomputing results is also made much easier. These programming environments are based on the so-called Markdown language, and countless other programming languages can be embedded within such documents.

While researchers often lack the expertise or time to make their manuscripts reproducible, pilot projects already exist (Baker et al., 2023), as do journals that support researchers in doing so (Carlsson et al., 2017). In combination with multiverse analyses, it is also possible, for research articles published online, to write the text so that it reacts interactively to alternative analytical decisions. Readers can thus make decisions within the manuscript and directly see how the results change.

Statistics

Probably every research discipline that uses statistics has been affected by the replication crisis. Very much in the spirit of “post hoc ergo propter hoc” (after this, therefore because of this), this fact is often interpreted to mean that the use of statistical methods is the cause of the replication problems. While counterarguments hold that the methods are simply being used incorrectly (Lakens, 2021c), some researchers also propose changes or alternatives. A group of 72 psychologists, for instance, called for lowering the significance threshold for new findings from 5% to 0.5% (Benjamin et al., 2018), making p-hacking harder. Others propose banning null hypothesis significance testing (NHST) altogether and using other methods instead: Wagenmakers (2007) advocates for Bayesian statistics, and the journal Basic and Applied Social Psychology has banned the use of significance tests entirely — a move that may actually have increased the problem of false-positive findings (Fricker Jr et al., 2019).

Open Data and Open Materials

Because of frequent fixed-term contracts and repeated moves between universities, but also because of discontinued doctorates or people leaving academia due to the Wissenschaftszeitvertragsgesetz (Germany’s fixed-term academic contract law), projects must often be handed over to other researchers. If the research materials and data are not documented and prepared properly, time and effort are lost in the process. In extreme cases, animals were raised and operated on in a laboratory, and the investigation cannot be continued. Open data and materials are meant to prevent this, to enable collaborative work, and to make errors correctable. In more extreme cases, researchers attempt to publish articles based on fabricated data. Only where openly accessible data exist can such fraud be detected (Carlisle, 2021).

Numerous studies have already shown that data are frequently not shared even upon request, and that this has not changed in the wake of the replication crisis (Vanpaemel et al., 2015). Moreover, whether data are shared says nothing about whether they contain errors (Claesen et al., 2023). More and more journals require the publication of data (e.g., https://topfactor.org/journals?factor=Data+Transparency), funding bodies require data management plans, tools for automated data preparation are under development (https://leibniz-psychology.org/das-institut/drittmittelprojekte/datawiz-ii), and numerous research data repositories — websites where data can be uploaded and found — have emerged.

The most important prerequisites for researchers to be able to share data are the consent of participants (where applicable), anonymization where necessary (e.g., for data on health or political views), and holding the rights to the data. Participant consent is routinely obtained before a study begins; anonymization takes place either during data collection or afterward (e.g., for qualitative data using AMNESIA); and the rights are usually only lacking when the data were collected on behalf of a company. The most difficult part is probably anonymization. Campbell et al. (2023), for example, report how they anonymized accounts from survivors of sexual assault using a multi-stage process in which names, dates, locations, trauma histories, and other sensitive information were redacted.

Where Are Research Data Uploaded?

Depending on the field and institution, research data are archived in different places. Universities often have their own services, but field-specific repositories are usually better for discoverability: psychologists frequently use the Open Science Framework (OSF), PsychArchives from the Leibniz Institute for Psychology, or Researchbox.org. For the social sciences, the Leibniz Institute for the Social Sciences offers various resources via GESIS. In chemistry, LISTER is software under development that semi-automatically describes data based on electronic lab notebooks. Re3data provides an overview of research data repositories (e.g., by field).

Repositories for research data.
Field Repository
Interdisciplinary, mainly psychology osf.io
Social sciences data.gesis.org
Political science and social sciences icpsr.umich.edu
Life sciences Pangaea.de
Arts and humanities de.Dariah.eu
Linguistics Clarin.eu
Biology gfbio.org
Materials science nomad-lab.eu
Qualitative data qdr.syr.edu
Interdisciplinary openbis.ch
Interdisciplinary about.coscine.de
Interdisciplinary frdr-dfdr.ca/repo
Where Are Research Data Published?

Repositories allow research data to be published with a mouse click. If additional features such as interactive analyses or peer review are desired, researchers use dedicated tools and specialized journals. Psychological datasets, for example, can be published in the Journal of Open Psychology Data, the R package PsyMetaData (Rodriguez & Williams, 2022) contains data from psychological meta-analyses, and the journal Inggrid publishes data from engineering research.

On dedicated websites, researchers can access chat log data voluntarily shared and anonymized by individuals at MOCODA, download data on human cooperation at CODA, or analyze estimation judgments at OpAQ.

Criteria for Preparing Research Data

Merely uploading research data to a website is not enough to make research more transparent. Typically, the data are linked in the research article in which they were used, and a codebook is provided that explains what the various values mean.

The table below shows an excerpt from a fictional dataset with three variables. Real datasets usually contain many more variables (e.g., the Abitur grade broken down by subject, demographic data such as age and gender, date of the survey, and sometimes variables with cryptic names like “V1_Z01” or “V0815”). Here, the three variables (i.e., columns) are “ID,” a sequential number identifying each participant; “IQ,” the measured IQ score from a particular intelligence test; and “Abitur grade,” the grade from participants’ German school-leaving exam, which participants reported themselves. (German school-leaving grades run from 1.0, the best, to 6.0, the worst — the inverse of a US-style GPA, where a higher number is better.) The codebook contains this information.

Example of a dataset.
ID IQ Abitur grade
1 103 2.6
2 86 2.4
3 112 1.8

FAIR and CARE

The FAIR principles for research data were developed on an interdisciplinary basis. They recommend that data be archived so as to be findable (F), accessible (A), interoperable (I) across different computer systems, and reusable (R). De Waard (2016) structures these requirements as a pyramid, with storage as the foundation, sharing above that, and quality assurance at the top. According to an EU report, the annual cost of data not complying with the FAIR principles amounts to €10.2 billion. Discipline-specific templates for making shared data FAIR are currently being developed by the Center for Open Science (www.cos.io/blog/cedar-embeddable-editor).

The CARE principles go a step further. They were designed for data collection involving Indigenous peoples, and originate from the Indigenous data sovereignty movement (e.g., the Global Indigenous Data Alliance). Building on FAIR, they call for collective benefit (C) of the data — for example, usability by society. Local examples from Germany include flood hazard maps (German-language) and the city of Münster’s “Cool City Map” (German-language), which lets residents mark places that stay cool on hot days. The people represented in the data must be given authority to control (A) how the data are used — that is, they should have a say in how the data appear. To preserve their self-determination, data should also be shared responsibly (R), and their rights and well-being should be central to the research (Ethics). When non-scientists actively participate in data collection or preparation, this is referred to as citizen science. For example, people can send their collected love letters to the German-language Liebesbriefarchiv (Love Letter Archive), which allows the German language, social conventions, and cultural change to be studied more comprehensively than would be possible using only a handful of famous scholars, or people can operate space telescopes over the internet from their own home computer (http://www.aim-muenster.de). This does not mean that laypeople conduct research on the basis of which they overturn “classic” research findings (Levy, 2022), but rather that they gain insight into the research process and can contribute to knowledge generation under guidance.

Requesting Research Data from Governments

When governments commission companies to answer research questions, citizens can request the underlying data. In the United Kingdom, for example, there is the platform WhatDoTheyKnow; in Germany, the equivalent is FragDenStaat. In one study, Maier et al. (2024) used such data to check whether recommendations regarding open science practices had been taken into account.

Concerns About Open Data

Data Police

The discourse around open science is at times highly charged: researchers who ask for data, or use data to identify errors, are labeled data parasites, data police, or even data terrorists. The claim that “those who have nothing to hide have nothing to fear” has a dystopian undertone, and in politically charged times, and in the face of plagiarism hunters, researchers are understandably reluctant to make themselves fully transparent. In this context, however, it should be kept in mind that science is not a solitary hobby but a profession with social responsibility. Anyone hiding something here should rightly be suspected of not doing proper science. Fittingly, the outside wall of the University and State Library in Münster bears large red letters reading “Gehorche keinem” — “obey no one.” Form your own opinion instead — go to the researchers’ work and data and see the truth for yourself. When researchers do not share their data and publish articles behind paywalls, this constitutes an unnecessary obstacle to independent opinion-forming.

Data Theft

Sharing data often conflicts with the goal of maintaining a competitive advantage over other researchers. This means that researchers withhold their data, publish as many articles as possible based on it, and only share the data once there is “nothing left to extract” from it. The concern is that other researchers might be quicker to publish articles based on the data. In practice, however, the data end up not being published at all. It is important to understand, regarding this concern, that sharing data does not mean giving away data. When sharing, researchers can attach a license that, for example, specifies how the data must be cited. If other researchers fail to comply, they risk their careers. Furthermore, it is easier to prove who originally collected the data if that person published it early on.

Citation Counts

Open data is sometimes promoted with the argument that it leads to more citations. This motivates some scientists more than good scientific practice does, but is probably not actually true (Colavizza et al., 2020).

What Does Long-Term Archiving Mean?

Funding bodies sometimes require long-term archiving. Depending on the context, this means that data must be stored and retrievable for 20 or 50 years. A somewhat extreme variant of long-term archiving was carried out with data from Github.io: all data that had been uploaded there as of February 2, 2020 were saved and transported to a disused coal mine in Norway. There they are meant to remain retrievable for up to 1,000 years.

Increasing Replicability

The approaches discussed so far, such as meta-analyses or reproducibility checks, can often be applied to existing projects. The following proposals, however, are more difficult to apply retroactively — they are mainly suited to new research and aim to increase the replicability of newly published studies. This includes stricter methodological standards or replications carried out before publication (“internal replications”), as is standard practice, for instance, in genetic epidemiology. A positive side effect here is that research groups join forces for replication studies and exchange data (Royal Netherlands Academy of Arts and Sciences, 2018). As discussed in the book’s conclusion, it is difficult to make a general, cross-disciplinary statement about whether these proposals actually affect replicability, given how rare replication studies still are. Even in social psychology, replications are currently still the least-implemented measure among all open science recommendations (Glöckner et al., 2024). Either way, the value of measures that generally increase the transparency of research is clear with respect to replication difficulties and p-hacking.

Transparent Reports

Aczel et al. (2020) designed a Transparency Checklist that researchers can work through point by point to check whether their research report is transparent. In the online app (https://www.shinyapps.org/apps/TransparencyChecklist/), a report can then be generated and attached to the research article. The checklist is divided into the topics of preregistration, methods, results, and data/code/materials. Regarding results, for example, it asks whether the number of observations was reported for all groups. The checklist is currently available in about 30 languages. The checklist by Aczel et al. (2020) is, however, primarily suited to quantitative studies. For qualitative and mixed-methods studies, Symonds & Tang (2024) have developed an assessment scheme.

A further and shorter variant is the 21-word solution. Here, a proposed statement (Simmons et al., 2012) is included in the report, assuring readers that no studies (or results) were withheld. It is far less rigorous and comprehensive than the transparency checklist, but it lowers the barrier to engaging with the transparency of research reports.

For animal studies, the ARRIVE guidelines were developed. Building on these, the LAG-R standards aim to ensure clear reporting of the genetic make-up of the animals studied (Teboul et al., 2024).

Preregistration

Research is either confirmatory (i.e., researchers have thought through in advance which data they will collect, which analyses they will run, and what the possible results might be) or exploratory (i.e., researchers try to approach a topic without preconceptions, generating questions for future, possibly confirmatory, research). Whenever research is confirmatory, it should be preregistered. This means that all the thinking researchers should have done in advance should be written down before data collection begins. Preregistrations are meant to prevent researchers from ending up in a garden of forking paths — or rather, they help fix the path in advance, so that it is no longer possible afterward to change the path in order to produce the desired results.

Figure 4: Preregistration fixes the path through the garden of forking paths in advance.

A good preregistration is characterized by researchers stripping themselves of all degrees of freedom (Wicherts et al., 2016) and specifying in advance all the decisions they will make on the way from data to results. Preregistrations should also be well structured, so that other people can easily verify what was determined beforehand (Simmons et al., 2020). To record all necessary decisions (e.g., how to handle missing values or the planned number of participants) in a structured way, preregistration templates exist. These exist for all kinds of areas, such as social-psychological experiments (van ’t Veer & Giner-Sorolla, 2016) or replication studies (Brandt et al., 2014). A collection of more than 20 templates along with their intended uses is available here. Researchers with little experience in a given area are supported by such templates. For instance, preregistration templates for meta-analyses include overviews of literature databases and already incorporate gold standards for transparent reporting (PRISMA, David Moher et al., n.d.). Moreover, the preregistration materials can later be reused for the manuscript itself.

The method of preregistration developed within psychology, drawing on the registration of studies involving human participants in medicine and the registration of meta-analyses. Like some other models, it transfers to many other fields without major adaptation — for example, to paleontology (Drage & Wong Hearing, 2023).

Preregistration Should Always Include an Analysis Plan and Analysis Code

Regarding the spread of preregistration, there is an ongoing debate about whether it should be introduced in as simple and resource-efficient a form as possible, or whether extensive and carefully prepared preregistrations should be required. For instance, the widely used template from aspredicted.org consists of only 11 questions and does not allow analysis code or other files to be attached. Since the purpose of preregistration is to prevent p-hacking, compromises on the scope of preregistrations make little sense.

Analysis code is based on a particular program and corresponding programming language and carries out the preparation of the data (e.g., computing scores, excluding observations) as well as the analyses themselves. Traditionally, it is written after the data have been collected. This can render studies worthless, because it may only become apparent after data collection that the planned analyses are not feasible, and because studies may then be analyzed in whatever way produces the desired results (p-hacking). Akker et al. (2023) were unable to demonstrate, for psychology research articles, that preregistration protects against p-hacking, but they did not distinguish between preregistrations with and without an analysis plan. For Registered Reports, which typically require a fully worked-out analysis plan, Scheel et al. (2021) were able to show that the proportion of hypothesis-confirming results drops sharply. Using thousands of statistical tests from economics research articles, Brodeur, Cook, et al. (2024) showed that preregistration only protects against p-hacking when an analysis plan is present. An analysis plan differs from analysis code in that the latter is unambiguous. A plan might state, for example, “we will compute the correlation between variable X and variable Y and test whether it is significant.” This does not specify which significance level will be used, whether a one-tailed or two-tailed test will be performed, which type of correlation will be computed (Pearson, Spearman, Kendall), how missing values will be handled, and so on. The probability of obtaining a significant result when the true correlation is 0 then rises from the nominal significance level of, say, 5% to 10%. This cannot happen with programming code (e.g., “cor.test(x, y)”), since the default settings used when something is left unspecified are clearly defined there. Reviewers of research should therefore pay close attention to exactly what is included in a preregistration (Thibault et al., 2023).

Deviations from Preregistrations

“The more deliberately people proceed, the more effectively chance strikes them” (my translation), as Friedrich Dürrenmatt puts it in point 8 of his 21 Points to The Physicists (Dürrenmatt, 2012). The same holds for preregistrations: researchers cannot anticipate every eventuality. It must therefore be expected that, due to programming errors, reasoning errors, or previously unconsidered arguments, researchers will need to proceed differently than specified in the preregistration. Deviations from preregistrations are very common (Heirene et al., 2021) — for instance, the actual sample size differs from the target sample size in more than 50% of cases. Once again, transparency is what distinguishes the right way to handle this: every deviation should be clearly communicated, and it should be discussed why the deviation occurred and how it affects the results (e.g., by checking whether results change when the target sample rather than the full, larger-than-planned sample is used) (Heirene et al., 2021; see also Lakens, 2024). Currently, reviewers do not check preregistrations or deviations from them (Syed, 2023; TARG Meta-Research Group and Collaborators et al., 2021). Tools that do this automatically are, however, under development (e.g., RegCheck, https://regcheck.app).

Where Are Studies Preregistered?

A broad infrastructure already exists for preregistering studies. Researchers can, for example, preregister studies via the Open Science Framework (osf.io) and also upload data, materials, analyses, and manuscripts there, as well as later publish preprints. Projects and preregistrations can also remain private or anonymized, and preregistrations can be placed under a multi-year embargo — a period during which the preregistration is not yet public and can only be accessed via a special link. All public preregistrations are automatically indexed by Google Scholar and are thereafter findable, citable, and also visible in the OSF search engine. Within OSF, either open templates or specific existing templates, such as the one by Brandt et al. (2014), can be used for preregistration (further templates will be incorporated in the future).

Another preregistration provider is aspredicted.org, though no files can be attached there, and its 11 questions are primarily suited to classic psychological studies. Medical studies are registered — usually without the details that are important for genuine preregistration — via https://clinicaltrials.gov, and meta-analyses via PROSPERO (http://www.crd.york.ac.uk/PROSPERO/).

Preregistration Checklist

Before conducting the study

  • The methodology and analysis plan of the study have been described completely and in a structured way (ideally using a template)

  • The analysis script has been written and tested using test data or random numbers, and it works

  • The preregistration has been time-stamped and archived before the study (e.g., via OSF.io)

After conducting the study

  • A link to the preregistration is included in the manuscript (e.g., in the methods section or in a paragraph on the transparency and openness of the study, ideally paired with links to data and materials)

  • All deviations from the preregistration are listed and justified

What If the Desired Results Do Not Come Out?

A question researchers often ask me when discussing preregistration once again demonstrates the discrepancy between research that is good for science and research that is good for one’s career. A preregistration prevents results from being dressed up (p-hacking). This increases the risk of obtaining results that are hard to publish, for example because they are not groundbreaking or not as expected. While a preregistration thus distinguishes researchers who are genuinely interested in good science, it can currently — in various research disciplines — have negative career consequences for not embellishing one’s data.

Planning Sample Size

Another method, recommended mainly for confirmatory research using statistical methods, is planning the sample size. This planning can be based on available resources or other constraints, but in the social sciences and medicine it is often done via power analyses. Statistical power is the probability of detecting an association of a given size if such an association actually exists. The underlying logic is that an association should not go undetected simply because the study was not adequately designed to detect it.

Power analyses have been known for a long time, and it is regularly criticized, at least within psychology, that they are conducted too rarely (Cohen, 1988, 1992). While they are still often absent or insufficiently integrated into students’ methods training, there are now extensive guides (Lakens, 2021a), software tools (Champely, 2020; Faul et al., 2007; Faul et al., 2009; Zhang & Mai, 2022), and video tutorials. Methods have meanwhile also been developed for alternative significance tests [equivalence tests; Lakens (2017); Lakens et al. (2018)] and for replications (Simonsohn, 2015). A particular type of sample planning is based on Bayesian statistics and is recommended especially in ethically sensitive situations such as animal behavior research: Richter (2024) proposes computing the so-called Bayes factor — which, unlike p-values from significance tests, also converges when there is no true association — after every observation or trial, and concluding the study once a certain value has been exceeded. It is important to note that this approach is explicitly suited only to Bayesian analyses.

Effect Sizes Instead of P-Values

It is now clear that statistical methods are frequently misunderstood and misused (Gigerenzer, 2004; Lakens, 2021b; Perezgonzalez, 2014). The discussion is currently shifting from p-values (Uygun Tunç et al., 2023) toward effect sizes and the question of what role different values play. For example, Vohs et al. (2021) were able to demonstrate that the controversial effect of diminishing self-control in the ego depletion paradigm does exist (i.e., it is not equal to 0) — but in their study it was so small that a successful demonstration alone would require over 6,700 participants, more than any experiment ever conducted on the topic, and almost twice the 3,524 participants used by Vohs et al. (2021). It remains controversial whether researchers should focus on easily demonstrable phenomena, so-called “low-hanging fruit” (Baumeister, 2020), whether there is a threshold below which phenomena are practically undetectable and thus of no interest (Primbs et al., 2023), and what practical relevance such small effects have (Anvari et al., 2021).

Qualitative Research

In qualitative research, there is often no clear research question fixed in advance; instead, it is developed over the course of the research process. Alongside anonymization (Campbell et al., 2023), preregistration therefore also plays a special role. Nevertheless, there are typical procedures — for example, decisions about how to code responses — that can be preregistered using dedicated templates.

Open Source Software

Research without programs for literature search, reference management, data analysis, and academic writing is unthinkable in most disciplines. Even the systems used by academic journals to structure the submission process (submission, review, revision, publication) are based on software. This relevant area within software is called “research software engineering,” and the challenge is that researchers, despite having had little prior contact with writing software (e.g., analysis code), must acquire expertise in this area. This ranges from general recommendations on the “idioms” of programming languages (e.g., the Zen of Python) all the way to collaborative work using systems like GitHub or GitLab.

Open source means that a program’s source code is readable and can be freely reused. This makes it possible to trace exactly what the program does. While, for example, the statistical software IBM SPSS has no such visible code, in GNU R every step of a statistical calculation can be traced. Errors that can jeopardize entire branches of science (Soergel, 2014) are easier to detect in open source programs. In cryptography there is talk of Kerckhoffs’s principle, which states that even encryption methods are more secure when only the key — not the encryption method itself — is kept secret. For example, software for analyzing fMRI data (a type of “brain scan”) was not properly validated until 15 years after its introduction, and its quality turned out to be inadequate (Eklund et al., 2016). Furthermore, proprietary software — software “owned” by a company or an individual — carries the risk of lock-in: people come to rely on the program and align their workflows closely with it. At some point they become dependent on it, for example being able to manage their data only with that one program, or facing enormous costs if they were to switch (Brembs et al., 2023). A license that identifies software as open source is the GNU license. Widely used proprietary software outside of research includes Microsoft Windows and the Office suite, Adobe Acrobat for reading PDF files, or Zoom for video calls. While the German government of the 2021-2025 legislative period set itself the goal of putting software projects out to tender as open source projects, it failed to follow through on this.

In short: for research in which traceability plays a primary role, open source software should be used wherever possible. For some methods this is currently not possible; there, it is the task of university libraries, as custodians of knowledge (Quan, 2021), and of professional societies, as the disciplines’ points of contact, to provide or commission the necessary infrastructure.

A “Copy” of the Open Science Framework

One of the most valuable tools in the social sciences is the Open Science Framework (osf.io). It gives researchers the ability to preregister studies and publish data, research reports, and materials. The entire codebase is available online and was copied for another project (GakuNin RDM) and adapted to its needs. Licensing terms clearly permit the code to be copied and reused. This openness of knowledge makes it possible to create similar offerings without having to start from scratch.

Commercial and non-commercial software.
Infrastructure Commercial Software Non-Commercial Software
Literature database SCOPUS OpenAlex (Priem et al., 2022)
Reference management Citavi, Mendeley Zotero (Puckett, 2011)
Data collection Unipark, Millisecond Inquisit PsychoPy (Peirce et al., 2022)
Statistical data analysis IBM SPSS, Stata GNU R (R Core Team, 2018), PSPP (Yagnik, 2014), JASP (Love et al., 2019)
Review and publication Editorial Manager Open Journal System (Willinsky, 2005)

Slow Science

Many open science developments can be summarized under the heading of “slow science.” The traditional mass production of low-quality research articles stands in contrast to a more mindful engagement and thorough scrutiny. Hyman (2024), for example, criticizes the fact that many of the proposed solutions put researchers at a disadvantage in the competition to publish. He proposes instead to solve these problems using artificial intelligence. Publishers such as Elsevier are already using AI for peer review, even though models like ChatGPT 4.0 are not actually suited to this task (Thelwall, 2024).

Big Team Science

Large-scale research projects such as ManyBabies (Byers-Heinlein et al., 2020), replication projects (Open Science Collaboration, 2015), initiatives like FORRT (Azevedo et al., 2019), and software development efforts such as GNU R (R Core Team, 2018) have shown what large groups of researchers are capable of, and that many current problems cannot be solved by individual researchers alone. In physics, the record for the largest number of authors on a single study is held by a paper from CERN on the Higgs boson. About 15 of the 33 pages consist solely of the names of the people involved, with additional pages for their institutions (Aad et al., 2015). In this context, the concept of “authors” is shifting toward that of “contributors” (A. Holcombe, 2019).

To clearly state who did what, researchers can use the Contributor Roles Taxonomy (CRediT, A. O. Holcombe, 2019), which consists of standardized role descriptions. Apps can then be used to generate simple tables or lists showing the contributions of everyone involved (A. O. Holcombe et al., 2020).

Further Reading

References

Aad, G., Abbott, B., Abdallah, J., Abdinov, O., Aben, R., Abolins, M., AbouZeid, O., Abramowicz, H., Abreu, H., & Abreu, R. (2015). Combined measurement of the higgs boson mass in pp collisions at s= 7 and 8 TeV with the ATLAS and CMS experiments. Physical Review Letters, 114(19), 191803.
Aczel, B., Szaszi, B., Sarafoglou, A., Kekecs, Z., Kucharský, Š., Benjamin, D., Chambers, C. D., Fisher, A., Gelman, A., Gernsbacher, M. A., Ioannidis, J. P., Johnson, E., Jonas, K., Kousta, S., Lilienfeld, S. O., Lindsay, D. S., Morey, C. C., Munafò, M., Newell, B. R., … Wagenmakers, E.-J. (2020). A consensus-based transparency checklist. Nature Human Behaviour, 4(1), 4–6. https://doi.org/10.1038/s41562-019-0772-6
Adler, S. J., Röseler, L., & Schöniger, M. K. (2023). A toolbox to evaluate the trustworthiness of published findings. Journal of Business Research, 167, 114189. https://doi.org/10.1016/j.jbusres.2023.114189
Akker, O. R. van den, Assen, M. A. van, Bakker, M., Elsherif, M., Wong, T. K., & Wicherts, J. M. (2023). Preregistration in practice: A comparison of preregistered and non-preregistered studies in psychology. Behavior Research Methods, 1–10. https://doi.org/10.31222/osf.io/fhdbs
Anvari, F., Kievit, R., Lakens, D., Przybylski, A. K., Tiokhin, L., Wiernik, B. M., & Orben, A. (2021). Evaluating the practical relevance of observed effect sizes in psychological research. https://doi.org/10.31234/osf.io/g3vtr
Azevedo, F., Parsons, S., Micheli, L., Strand, J. F., Rinke, E. M., Guay, S., Elsherif, M. M., Quinn, K. A., Wagge, J. R., Steltenpohl, C. N., Kalandadze, T., Vasilev, M. R., Oliveira, C. M., Aczel, B., Miranda, J. F., Baker, B. J., Galang, C. M. O., Pennington, C. R., Marques, T., … FORRT. (2019). Introducing a framework for open and reproducible research training (FORRT). https://doi.org/10.2218/eorc.2022.6968
Baker, D., Berg, M., Hansford, K., Quinn, B., Segala, F. G., & English, E. (2023). ReproduceMe: Lessons from a pilot project on computational reproducibility. https://doi.org/10.31234/osf.io/k8d4u
Bartoš, F., & Schimmack, U. (2022). Z-curve 2.0: Estimating replication rates and discovery rates. Meta-Psychology, 6. https://doi.org/10.15626/MP.2021.2720
Baumeister, R. (2020). Do effect sizes in psychology laboratory experiments mean anything in reality? https://doi.org/10.31234/osf.io/mpw4t
Benjamin, D. J., Berger, J. O., Johannesson, M., Nosek, B. A., Wagenmakers, E.-J., Berk, R., Bollen, K. A., Brembs, B., Brown, L., Camerer, C., Cesarini, D., Chambers, C. D., Clyde, M., Cook, T. D., Boeck, P. de, Dienes, Z., Dreber, A., Easwaran, K., Efferson, C., … Johnson, V. E. (2018). Redefine statistical significance. Nature Human Behaviour, 2(1), 6–10. https://doi.org/10.1038/s41562-017-0189-z
Brandt, M. J., IJzerman, H., Dijksterhuis, A., Farach, F. J., Geller, J., Giner-Sorolla, R., Grange, J. A., Perugini, M., Spies, J. R., & van ’t Veer, A. (2014). The replication recipe: What makes for a convincing replication? Journal of Experimental Social Psychology, 50, 217–224. https://doi.org/10.1016/j.jesp.2013.10.005
Brembs, B., Huneman, P., Schönbrodt, F., Nilsonne, G., Susi, T., Siems, R., Perakakis, P., Trachana, V., Ma, L., & Rodriguez-Cuadrado, S. (2023). Replacing academic journals. Royal Society Open Science, 10(7). https://doi.org/10.1098/rsos.230206
Brodeur, A., Cook, N. M., Hartley, J. S., & Heyes, A. (2024). Do preregistration and preanalysis plans reduce p-hacking and publication bias? Evidence from 15,992 test statistics and suggestions for improvement. Journal of Political Economy Microeconomics, 2(3), 527–561. https://doi.org/10.1086/730455
Brodeur, A., Mikola, D., & Cook, N. (2024). Mass reproducibility and replicability: A new hope. https://doi.org/10.2139/ssrn.4790780
Byers-Heinlein, K., Bergmann, C., Davies, C., Frank, M. C., Hamlin, J. K., Kline, M., Kominsky, J. F., Kosie, J. E., Lew-Williams, C., & Liu, L. (2020). Building a collaborative psychological science: Lessons learned from ManyBabies 1. Canadian Psychology/Psychologie Canadienne, 61(4), 349. https://doi.org/10.31234/osf.io/dmhk2
C. Chang, A., & Li, P. (2022). Is economics research replicable? Sixty published papers from thirteen journals say “often not.” Crit. Fin. Rev., 11(1), 185–206. https://doi.org/10.1561/104.00000053
Campbell, R., Javorka, M., Engleton, J., Fishwick, K., Gregory, K., & Goodman-Williams, R. (2023). Open-science guidance for qualitative research: An empirically validated approach for de-identifying sensitive narrative data. Advances in Methods and Practices in Psychological Science, 6(4), Article 25152459231205832. https://doi.org/10.1177/25152459231205832
Carlisle, J. (2021). False individual patient data and zombie randomised controlled trials submitted to anaesthesia. Anaesthesia, 76(4), 472–479. https://doi.org/10.1111/anae.15263
Carlsson, R., Batinović, L., Hyltse, N., Kalmendal, A., Nordström, T., & Topor, M. (2024). A beginner’s guide to open and reproducible systematic reviews in psychology. https://doi.org/10.31234/osf.io/59vsw
Carlsson, R., Danielsson, H., Heene, M., Innes-Ker, Å., Lakens, D., Schimmack, U., Schönbrodt, F. D., van Asssen, M., & Weinstein, Y. (2017). Inaugural editorial of meta-psychology. Meta-Psychology, 1, a1001. https://doi.org/10.15626/MP2017.1001
Carter, E. C., Schönbrodt, F. D., Gervais, W. M., & Hilgard, J. (2019). Correcting for bias in psychology: A comparison of meta-analytic methods. Advances in Methods and Practices in Psychological Science, 2(2), 115–144. https://doi.org/10.1177/2515245919847196
Champely, S. (2020). Pwr: Basic functions for power analysis. R package version 1.3-0. https://CRAN.R-project.org/package=pwr
Claesen, A., Vanpaemel, W., Maerten, A.-S., Verliefde, T., Tuerlinckx, F., & Heyman, T. (2023). Data sharing upon request and statistical consistency errors in psychology: A replication of wicherts, bakker and molenaar (2011). Plos One, 18(4), e0284243. https://doi.org/10.1371/journal.pone.0284243
Cobey, K. D., Fehlmann, C. A., Christ Franco, M., Ayala, A. P., Sikora, L., Rice, D. B., Xu, C., Ioannidis, J. P., Lalu, M. M., & Ménard, A. (2023). Epidemiological characteristics and prevalence rates of research reproducibility across disciplines: A scoping review of articles published in 2018-2019. Elife, 12, e78518. https://doi.org/10.7554/elife.78518
Cohen, J. (1988). Statistical power analysis for the behavioral sciences. Routledge. https://doi.org/10.4324/9780203771587
Cohen, J. (1992). A power primer. Psychological Bulletin, 112(3), 409–410. https://doi.org/10.1037//0033-2909.112.3.409
Colavizza, G., Hrynaszkiewicz, I., Staden, I., Whitaker, K., & McGillivray, B. (2020). The citation advantage of linking publications to research data. PloS One, 15(4), e0230416. https://doi.org/10.1371/journal.pone.0230416
Dang, J., Barker, P., Baumert, A., Bentvelzen, M., Berkman, E., Buchholz, N., Buczny, J., Chen, Z., Cristofaro, V. de, Vries, L. de, Dewitte, S., Giacomantonio, M., Gong, R., Homan, M., Imhoff, R., Ismail, I., Jia, L., Kubiak, T., Lange, F., … Zinkernagel, A. (2020). A multilab replication of the ego depletion effect. Social Psychological and Personality Science, 194855061988770. https://doi.org/10.1177/1948550619887702
David Moher, Larissa Shamseer, Mike Clarke, Davina Ghersi, Alessandro Liberati, Mark Petticrew, Paul Shekelle, & Lesley A Stewart. (n.d.). Preferred reporting items for systematic review and meta-analysis protocols (PRISMA-p) 2015 statement. https://doi.org/10.1186/2046-4053-4-1
De Waard, A. (2016). Research data management at elsevier: Supporting networks of data and workflows. Information Services & Use, 36(1-2), 49–55. https://doi.org/10.3233/isu-160805
Drage, H., & Wong Hearing, T. (2023). Diamond open access with preregistration: A new publishing model for palaeontology. In EarthArXiv. https://doi.org/10.31223/x5kh30
Dürrenmatt, F. (2012). Die physiker: Eine komödie in zwei akten. Diogenes Verlag AG.
Eklund, A., Nichols, T. E., & Knutsson, H. (2016). Cluster failure: Why fMRI inferences for spatial extent have inflated false-positive rates. Proceedings of the National Academy of Sciences, 113(28), 7900–7905. https://doi.org/10.1073/pnas.1602413113
Erdfelder, E., & Heck, D. W. (2019). Detecting evidential value and p-hacking with the p-curve tool. Zeitschrift für Psychologie, 227(4), 249–260. https://doi.org/10.1027/2151-2604/a000383
Faul, F., Erdfelder, E., Buchner, A., & Lang, A.-G. (2009). Statistical power analyses using g*power 3.1: Tests for correlation and regression analyses. Behavior Research Methods, 41(4), 1149–1160. https://doi.org/10.3758/BRM.41.4.1149
Faul, F., Erdfelder, E., Lang, A.-G., & Buchner, A. (2007). G*power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior Research Methods, 39(2), 175–191. https://doi.org/10.3758/bf03193146
Feyerabend, P. K. (1975/2002). Against method (Reprinted der 3. ed. 1993). Verso.
Fišar, M., Greiner, B., Huber, C., Katok, E., Ozkes, A. I., & Management Science Reproducibility Collaboration. (2024). Reproducibility in management science. Management Science, 70(3), 1343–1356. https://doi.org/10.2139/ssrn.4620006
Fletcher, S. C. (2022). Replication is for meta-analysis. Philosophy of Science, 89(5), 960–969. https://doi.org/10.1017/psa.2022.38
Fricker Jr, R. D., Burke, K., Han, X., & Woodall, W. H. (2019). Assessing the statistical analyses used in basic and applied social psychology after their p-value ban. The American Statistician, 73(sup1), 374–384. https://doi.org/10.1080/00031305.2018.1537892
Friese, M., Loschelder, D. D., Gieseler, K., Frankenbach, J., & Inzlicht, M. (2018). Is ego depletion real? An analysis of arguments. Personality and Social Psychology Review : An Official Journal of the Society for Personality and Social Psychology, Inc. https://doi.org/10.1177/1088868318762183
Giehl, K., Mutsaerts, H.-J., Aarts, K., Barkhof, F., Caspers, S., Chetelat, G., Colin, M.-E., Düzel, E., Frisoni, G. B., & Ikram, M. A. (2024). Sharing brain imaging data in the open science era: How and why? The Lancet Digital Health, 6(7), e526–e535.
Gigerenzer, G. (2004). Mindless statistics. The Journal of Socio-Economics, 33(5), 587–606. https://doi.org/10.1016/j.socec.2004.09.033
Glöckner, A., Gollwitzer, M., Hahn, L., Lange, J., Sassenberg, K., & Unkelbach, C. (2024). Quality, replicability, and transparency in research in social psychology: Implementation of recommendations in germany. Social Psychology, 55(3), 134–147. https://doi.org/10.1027/1864-9335/a000548
Hagger, M. S., Chatzisarantis, N. L. D., Alberts, H., Anggono, C. O., Batailler, C., Birt, A. R., Brand, R., Brandt, M. J., Brewer, G., Bruyneel, S., Calvillo, D. P., Campbell, W. K., Cannon, P. R., Carlucci, M., Carruth, N. P., Cheung, T., Crowell, A., Ridder, D. T. D. de, Dewitte, S., … Zwienenberg, M. (2016). A multilab preregistered replication of the ego-depletion effect. Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 11(4), 546–573. https://doi.org/10.1177/1745691616652873
Hagger, M. S., Wood, C., Stiff, C., & Chatzisarantis, N. L. D. (2010). Ego depletion and the strength model of self-control: A meta-analysis. Psychological Bulletin, 136(4), 495–525. https://doi.org/10.1037/a0019486
Harrer, M., Cuijpers, P., Furukawa, T., & Ebert, D. (2021). Doing meta-analysis with r: A hands-on guide. Chapman; Hall/CRC.
Hartgerink, C. H. (2016). 688,112 statistical results: Content mining psychology articles for statistical test results. Data, 1(3), 14. https://doi.org/10.3390/data1030014
Heirene, R., LaPlante, D., Louderback, E. R., Keen, B., Bakker, M., Serafimovska, A., & Gainsbury, S. M. (2021). Preregistration specificity & adherence: A review of preregistered gambling studies & cross-disciplinary comparison. https://doi.org/10.31234/osf.io/nj4es
Holcombe, A. (2019). Farewell authors, hello contributors. Nature, 571(7763), 147–148. https://doi.org/10.1038/d41586-019-02084-8
Holcombe, A. O. (2019). Contributorship, not authorship: Use CRediT to indicate who did what. Publications, 7(3), 48. https://doi.org/10.3390/publications7030048
Holcombe, A. O., Kovacs, M., Aust, F., & Aczel, B. (2020). Documenting contributions to scholarly articles using CRediT and tenzing. PLoS One, 15(12), e0244611. https://doi.org/10.1371/journal.pone.0244611
Hyman, M. (2024). Freeing social and medical scientists from the replication crisis. Available at SSRN 4898637.
Jacobsen, N. S. J., Kristanto, D., Welp, S., Inceler, Y. C., & Debener, S. (2024). Preprocessing choices for P3 analyses with mobile EEG: A systematic literature review and interactive exploration. Psychophysiology, 62(1). https://doi.org/10.1111/psyp.14743
Jané, M. B., Xiao, Q., Yeung, S. K., Ben-Shachar, M. S., Caldwell, A. R., Cousineau, D., Dunleavy, D. J., Elsherif, M., Johnson, B. T., & Moreau, D. (2024). Guide to effect sizes and confidence intervals. https://doi.org/10.17605/OSF.IO/D8C4G
Kepes, S., Bushman, B. J., & Anderson, C. A. (2017). Violent video game effects remain a societal concern: Reply to hilgard, engelhardt, and rouder (2017). Psychological Bulletin, 143(7), 775–782. https://doi.org/10.1037/bul0000112
Kepes, S., Wang, W., & Cortina, J. M. (2023). Assessing publication bias: A 7-step user’s guide with best-practice recommendations. Journal of Business and Psychology, 38(5), 957–982. https://doi.org/10.1007/s10869-022-09840-0
Körner, R., Röseler, L., Schütz, A., & Bushman, B. J. (2022). Dominance and prestige: Meta-analytic review of experimentally induced body position effects on behavioral, self-report, and physiological dependent variables. Psychological Bulletin, 148(1-2), 67–85. https://doi.org/10.1037/bul0000356
Koukouraki, E., & Kray, C. (2023). Map reproducibility in geoscientific publications: An exploratory study. Schloss Dagstuhl - Leibniz-Zentrum für Informatik.
Lakens, D. (2017). Equivalence tests: A practical primer for t tests, correlations, and meta-analyses. Social Psychological and Personality Science, 8(4), 355–362. https://doi.org/10.1177/1948550617697177
Lakens, D. (2021a). Sample size justification. https://doi.org/10.31234/osf.io/9d3yf
Lakens, D. (2021b). The practical alternative to the p value is the correctly used p value. Perspectives on Psychological Science, 16(3), 639–648. https://doi.org/10.31234/osf.io/shm8v
Lakens, D. (2021c). The practical alternative to the p value is the correctly used p value. Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 16(3), 639–648. https://doi.org/10.1177/1745691620958012
Lakens, D. (2023). Concerns about replicability, theorizing, applicability, generalizability, and methodology across two crises in social psychology. https://doi.org/10.31234/osf.io/dtvs7
Lakens, D. (2024). When and how to deviate from a preregistration. Collabra: Psychology, 10(1). https://doi.org/10.31234/osf.io/ha29k
Lakens, D., & Etz, A. J. (2017). Too true to be bad: When sets of studies with significant and nonsignificant findings are probably true. Social Psychological and Personality Science, 8(8), 875–881. https://doi.org/10.1177/1948550617693058
Lakens, D., Scheel, A. M., & Isager, P. M. (2018). Equivalence testing for psychological research: A tutorial. Advances in Methods and Practices in Psychological Science, 1(2), 259–269. https://doi.org/10.1177/2515245918770963
Levy, N. (2022). Do your own research! Synthese, 200(5), 356. https://doi.org/10.1007/s11229-022-03793-w
Light, R. J., & Pillemer, D. B. (1984). Summing up: The science of reviewing research. Harvard University Press.
Lindsay, D. S. (2023). A plea to psychology professional societies that publish journals: Assess computational reproducibility. Meta-Psychology, 7. https://doi.org/10.15626/mp.2023.4020
Love, J., Selker, R., Marsman, M., Jamil, T., Dropmann, D., Verhagen, J., Ly, A., Gronau, Q. F., Šmı́ra, M., & Epskamp, S. (2019). JASP: Graphical statistical software for common statistical designs. Journal of Statistical Software, 88, 1–17. https://doi.org/10.18637/jss.v088.i02
Maier, M., Bartoš, F., Raihani, N., Shanks, D. R., Stanley, T., Wagenmakers, E.-J., & Harris, A. J. (2024). Exploring open science practices in behavioural public policy research. Royal Society Open Science, 11(2), 231486.
Mazei, J., Rudolph, C. W., Zacher, H., & Hüffmeier, J. (2025). Do not put all of your eggs in one basket: Multiverse analysis in applied psychology. Journal of Applied Psychology, 110(11), 1511–1537. https://doi.org/10.1037/apl0001291
Nüst, D., & Eglen, S. J. (2021). CODECHECK: An open science initiative for the independent execution of computations underlying research articles during peer review to improve reproducibility. F1000Res., 10, 253. https://doi.org/10.12688/f1000research.51738.2
Open Science Collaboration. (2015). Psychology: Estimating the reproducibility of psychological science. Science (New York, N.Y.), 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
Peirce, J., Hirst, R., & MacAskill, M. (2022). Building experiments in PsychoPy. Sage. https://doi.org/10.4135/9781036231378
Perezgonzalez, J. D. (2014). A reconceptualization of significance testing. Theory & Psychology, 24(6), 852–859. https://doi.org/10.1177/0959354314546157
Priem, J., Piwowar, H., & Orr, R. (2022). OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. arXiv Preprint arXiv:2205.01833.
Primbs, M. A., Pennington, C. R., Lakens, D., Silan, M. A. A., Lieck, D. S., Forscher, P. S., Buchanan, E. M., & Westwood, S. J. (2023). Are small effects the indispensable foundation for a cumulative psychological science? A reply to götz et al.(2022). Perspectives on Psychological Science, 18(2), 508–512. https://doi.org/10.31234/osf.io/6s8bj
Promoting reproduction and replication at scale. (2024). Nat. Hum. Behav., 8(1), 1. https://doi.org/10.1038/s41562-024-01818-7
Puckett, J. (2011). Zotero: A guide for librarians, researchers, and educators. Assoc of Cllge & Rsrch Libr.
Quan, J. (2021). Toward reproducibility: Academic libraries and open science (pp. 49–66). Cambridge University Press. https://doi.org/10.29085/9781783304615.004
R Core Team. (2018). R: A language and environment for statistical computing. R Foundation for Statistical Computing. https://www.R-project.org/
Richardson, R. (2025). The collection of open science integrity guides (COSIG): Expanding participation in post-publication peer review. https://doi.org/10.5281/zenodo.15564777
Richter, S. H. (2024). Challenging current scientific practice: How a shift in research methodology could reduce animal use. Lab Animal, 53(1), 9–12. https://doi.org/10.1038/s41684-023-01308-9
Rodriguez, J. E., & Williams, D. R. (2022). Psymetadata: An r package containing open datasets from meta-analyses in psychology. https://doi.org/10.31234/osf.io/myxsj
Röseler, L. (2023). Predicting replication rates with z-curve: A brief exploratory validation study using the replication database. https://doi.org/10.31222/osf.io/ewb2t
Royal Netherlands Academy of Arts and Sciences. (2018). Replication studies—improving reproducibility in the empirical sciences.
Scheel, A. M., Schijen, M. R., & Lakens, D. (2021). An excess of positive results: Comparing the standard psychology literature with registered reports. Advances in Methods and Practices in Psychological Science, 4(2), 25152459211007467. https://doi.org/10.31234/osf.io/p6e9c
Schönbrodt, F., Gollwitzer, M., & Abele-Brehm, A. (2017). Der umgang mit forschungsdaten im fach psychologie: Konkretisierung der DFG-leitlinien. Psychologische Rundschau. https://doi.org/10.1026/0033-3042/a000341
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2012). A 21 word solution. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.2160588
Simmons, J. P., Nelson, L., & Simonsohn, U. (2020). Pre–registration is a game changer. But, like random assignment, it is neither necessary nor sufficient for credible science. Journal of Consumer Psychology. https://doi.org/10.1002/jcpy.1207
Simmons, J. P., & Simonsohn, U. (2017). Power posing: P-curving the evidence. Psychological Science, 28(5), 687–693. https://doi.org/10.1177/0956797616658563
Simonsohn, U. (2015). Small telescopes: Detectability and the evaluation of replication results. Psychological Science, 26(5), 559–569. https://doi.org/10.1177/0956797614567341
Simonsohn, U., Nelson, L. D., & Simmons, J. P. (2014a). P-curve and effect size: Correcting for publication bias using only significant results. Perspectives on Psychological Science : A Journal of the Association for Psychological Science, 9(6), 666–681. https://doi.org/10.1177/1745691614553988
Simonsohn, U., Nelson, L. D., & Simmons, J. P. (2014b). P-curve: A key to the file-drawer. Journal of Experimental Psychology. General, 143(2), 534–547. https://doi.org/10.1037/a0033242
Simonsohn, U., Simmons, J. P., & Nelson, L. D. (2015). Better p-curves: Making p-curve analysis more robust to errors, fraud, and ambitious p-hacking, a reply to ulrich and miller (2015). Journal of Experimental Psychology. General, 144(6), 1146–1152. https://doi.org/10.1037/xge0000104
Soergel, D. A. (2014). Rampant software errors may undermine scientific results. F1000Research, 3. https://doi.org/10.12688/f1000research.5930.2
Sotola, L. K. (2023). How can i study from below, that which is above? Meta-Psychology, 7. https://doi.org/10.15626/MP.2022.3299
Sotola, L. K., & Credé, M. (2022). On the predicted replicability of two decades of experimental research on system justification: A z–curve analysis. European Journal of Social Psychology, 52(5-6), 895–909. https://doi.org/10.1002/ejsp.2858
Syed, M. (2023). Some data indicating that editors and reviewers do not check preregistrations during the review process. PsyArXiv Preprints. https://doi.org/10.31234/osf.io/nh7qw
Symonds, J. E., & Tang, X. (2024). Quality appraisal checklist for quantitative, qualitative, and mixed methods studies. https://doi.org/10.31234/osf.io/djmcq_v2
TARG Meta-Research Group and Collaborators, Thibault, R. T., Clark, R., Pedder, H., Akker, O. van den, Westwood, S., Thompson, J., & Munafo, M. (2021). Estimating the prevalence of discrepancies between study registrations and publications: A systematic review and meta-analyses. MedRxiv, 2021–2007.
Teboul, L., Amos-Landgraf, J., Benavides, F. J., Birling, M.-C., Brown, S. D., Bryda, E., Bunton-Stasyshyn, R., Chin, H.-J., Crispo, M., & Delerue, F. (2024). Improving laboratory animal genetic reporting: LAG-r guidelines. Nature Communications, 15(1), 5574.
Thelwall, M. (2024). Can ChatGPT evaluate research quality? Journal of Data and Information Science. https://doi.org/10.2478/jdis-2024-0013
Thibault, R. T., Pennington, C. R., & Munafò, M. R. (2023). Reflections on preregistration: Core criteria, badges, complementary workflows. Journal of Trial & Error, 2(1), 10–36850. https://doi.org/10.31234/osf.io/w6tj4
Thibault, R. T., Zavalis, E. A., Malicki, M., & Pedder, H. (2024). An evaluation of reproducibility and errors in published sample size calculations performed using g* power. medRxiv, 2024–2007. https://doi.org/10.1101/2024.07.15.24310458
Uygun Tunç, D., Tunç, M. N., & Lakens, D. (2023). The epistemic and pragmatic function of dichotomous claims based on statistical hypothesis tests. Theory & Psychology, 33(3), 403–423. https://doi.org/10.31234/osf.io/af9by
van ’t Veer, A. E., & Giner-Sorolla, R. (2016). Pre-registration in social psychology—a discussion and suggested template. Journal of Experimental Social Psychology, 67, 2–12. https://doi.org/10.1016/j.jesp.2016.03.004
Vanpaemel, W., Vermorgen, M., Deriemaecker, L., & Storms, G. (2015). Are we wasting a good crisis? The availability of psychological research data after the storm. Collabra, 1(1), 3. https://doi.org/10.1525/collabra.13
Vohs, K. D., Schmeichel, B. J., Lohmann, S., Gronau, Q. F., Finley, A. J., Ainsworth, S. E., Alquist, J. L., Baker, M. D., Brizi, A., & Bunyi, A. (2021). A multisite preregistered paradigmatic test of the ego-depletion effect. Psychological Science, 32(10), 1566–1581. https://doi.org/10.31234/osf.io/e497p
Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems ofp values. Psychonomic Bulletin & Review, 14(5), 779–804. https://doi.org/10.3758/BF03194105
Wicherts, J. M., Veldkamp, C. L. S., Augusteijn, H. E. M., Bakker, M., van Aert, R. C. M., & van Assen, M. A. L. M. (2016). Degrees of freedom in planning, running, analyzing, and reporting psychological studies: A checklist to avoid p-hacking. Frontiers in Psychology, 7, 1832. https://doi.org/10.3389/fpsyg.2016.01832
Willinsky, J. (2005). Open journal systems: An example of open source software for journal management and publishing. Library Hi Tech, 23(4), 504–519.
Wilson, G., Bryan, J., Cranston, K., Kitzes, J., Nederbragt, L., & Teal, T. K. (2017). Good enough practices in scientific computing. PLoS Computational Biology, 13(6), e1005510. https://doi.org/10.1371/journal.pcbi.1005510
Yagnik, J. (2014). PSPP a free and open source tool for data analysis.
Zhang, Z., & Mai, Y. (2022). WebPower: Basic and advanced statistical power analysis (r package version 0.7). https://CRAN.R-project.org/package=WebPower