Data Colada
by Uri Simonsohn, Leif Nelson, Joseph Simmons · Data Colada
Blog by the researchers behind p-curve and several high-profile data-fraud investigations, posting worked forensic analyses of published papers, p-hacking diagnostics, and critiques of meta-analytic and statistical practice. The only source here that shows bias detection being performed on real published papers rather than described. Each post is a case study in what a manipulated or p-hacked dataset actually looks like, which is knowledge no textbook transmits well.
More resources on Bias Mitigation
cos.io
Center for Open Science (cos.io) is a nonprofit that builds open-source tools and promotes reproducibility and transparency in research. The site explains and provides access to the Open Science Framework (OSF), preregistration, data sharing resources, and educational materials on replication and open science practices.
Improving your statistical inferences
This course aims to help you to draw better statistical inferences from empirical research. First, we will discuss how to correctly interpret p-values, effect sizes, confidence intervals, Bayes Factors, and likelihood ratios, and how these statistics answer different questions you might be interested in. Then, you will learn how to design experiments where the false positive rate is controlled, and how to decide upon the sample size for your study, for example in order to achieve high statistical power. Subsequently, you will learn how to interpret evidence in the scientific literature given widespread publication bias, for example by learning about p-curve analysis. Finally, we will talk about how to do philosophy of science, theory construction, and cumulative science, including how to perform replication studies, why and how to pre-register your experiment, and how to share your results following Open Science principles. In practical, hands on assignments, you will learn how to simulate t-tests to learn which p-values you can expect, calculate likelihood ratio's and get an introduction the binomial Bayesian statistics, and learn about the positive predictive value which expresses the probability published research findings are true. We will experience the problems with optional stopping and learn how to prevent these problems by using sequential analyses. You will calculate effect sizes, see how confidence intervals work through simulations, and practice doing a-priori power analyses. Finally, you will learn how to examine whether the null hypothesis is true using equivalence testing and Bayesian statistics, and how to pre-register a study, and share your data on the Open Science Framework. All videos now have Chinese subtitles. More than 30.000 learners have enrolled so far! If you enjoyed this course, I can recommend following it up with me new course "Improving Your Statistical Questions"
Jon Kabat-Zinn Guided Mindfulness Meditation
A 2007 Google talk in which Jon Kabat-Zinn explains what mindfulness is, the research behind MBSR, and how attention training affects stress and health, interspersed with short guided sitting practice and audience questions.
Calling Bullshit: Data Reasoning in a Digital World
University of Washington course with free video lectures, syllabus, and case studies on detecting misleading quantitative claims. Covers publication bias, selection effects, correlation versus causation, and deceptive data visualisation. Lectures 4 and 7 (statistical traps, publication bias) are taught with real published examples, and the material is a genuine credit-bearing UW course rather than repackaged popular science — which matters in a topic swamped by cognitive-bias self-help.
Causal Inference: What If
Graduate textbook, free as a PDF, treating confounding, selection bias, and measurement bias as structural problems represented with causal diagrams rather than as statistical afterthoughts. Part I requires no modelling background. Chapters 7-9 (Confounding, Selection Bias, Measurement Bias) are the deepest treatment available of why these biases arise and what identification assumptions actually neutralise them — mental models, not procedures.
The Preregistration Revolution
Separates prediction from postdiction and sets out how preregistering hypotheses and analysis plans constrains selective reporting. Addresses nine practical objections, including protocol deviations, exploratory work, and preregistration of analyses on existing data.