Reality Bending Lab › Publications › Choosing informative priors in Bayesian regression models: a simulation study and tutorial using Stan and R.
Choosing informative priors in Bayesian regression models: a simulation study and tutorial using Stan and R.
Bayesian regression is at its most useful when samples are small, which is exactly when the prior matters most and researchers are least sure how to set one. A simulation study maps how prior location and scale move the posterior at different sample sizes, and a case-control example (N=526) shows a frequentist odds ratio of 8.87 with an interval running to 165 becoming a stable 4.01 under a literature-based prior.
Abstract
Background Bayesian regression models provide a robust framework for complex data analysis, which is particularly advantageous in scenarios with small sample sizes, common in psychology or medical research. However, specifying appropriate prior distributions that incorporate existing knowledge to regularize model parameters remains a challenge for many researchers. This can lead to unstable or implausible estimates. This study aims to demonstrate the impact of different prior distributions on regression models and to provide a practical guide for choosing and justifying informative priors to produce more stable and credible results. Methods The study involved two parts. First, a simulation study was conducted to systematically assess the sensitivity of Bayesian linear regression models to prior specification. We systematically varied the sample size, prior location, and prior scale to observe their impact on posterior estimates for a known true effect size. Second, a case–control study using real-world patient data ( N = 526) demonstrated the practical application of choosing informative priors. Bayesian logistic regression models were used to analyze the relationship between severe dementia and fall incidence, comparing results from priors based on existing literature (“believer”), conservative priors (“agnostic”), and priors assuming an opposite effect (“skeptical”). Results The simulation study showed that strongly informative priors had a substantial influence on posterior estimates, particularly for smaller sample sizes. As the sample size increased, the influence of the data increased, and the estimates converged toward the true effect. In the case–control study, a standard frequentist logistic regression produced an odds ratio of 8.87 with a very wide and unstable confidence interval (1.66–165.19), likely due to data sparsity. In contrast, a Bayesian model using a moderately informative “believer” prior derived from existing research yielded a more stable and plausible odds ratio of 4.01 with a substantially narrower credible interval (1.99–8.78). Conclusion Careful and transparent specification of informative priors is a critical tool in Bayesian analysis, especially when data are sparse. By incorporating justified evidence-based assumptions, researchers can regularize models to prevent implausible outcomes and produce more stable, interpretable, and credible results. This approach enhances the robustness of statistical inference in fields where small sample sizes are a frequent challenge.

Cite
Lüdecke, D., Makowski, A. C., Klein, J., Ben-Shachar, M. S., & Makowski, D. (2026). Choosing informative priors in Bayesian regression models: a simulation study and tutorial using Stan and R.. Frontiers in Psychology. https://doi.org/10.3389/fpsyg.2026.1856582
See also
- bayestestR: Describing Effects and their Uncertainty, Existence and Significance within the Bayesian Framework
- Indices of Effect Existence and Significance in the Bayesian Framework
- This Is Not the ex-Gaussian Model You Are Looking For: On the Default Parameterization of Bayesian ex-Gaussian Models in brms