Skip to content

Latest commit

Β 

History

193 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Methods Lab beta 0.7

Statistical Methods, Made Tangible.

License: MIT Static Badge Static Badge Static Badge Static Badge

A suite of 46 interactive tools that build statistical intuition from the ground up β€” from OLS regression, ANOVA and multilevel models, through the counter-intuitive paradoxes that trip up even seasoned researchers, into modern causal inference, psychometrics (classical test theory, measurement models, reliability, IRT), and clinical diagnostics. Everything runs entirely in the browser. No code, no server, no data ever leaves your machine.

[ Learning Path ] β€’ [ Ecosystem ] β€’ [ Philosophy ] β€’ [ Usage ] β€’ [ Sister Project ]


Language note: The interface, inline explanations, and help panels are written in German, as the lab is used in teaching at the University of OsnabrΓΌck. The methods and code are, of course, language-independent.


πŸ”­ Scientific Philosophy

The Methods Lab rejects the "black-box" experience of statistical software, where a number appears and its meaning stays hidden. Instead, every tool makes a concept manipulable: move a slider, drag a data point, switch a scenario β€” and watch the consequence unfold in real time.

  • Regression is the spine. Almost every method in applied statistics is regression in disguise. The lab is deliberately ordered so that the foundations (Section 1) make the paradoxes (Section 3) and the causal methods (Section 4) feel obvious in hindsight rather than mysterious.
  • Paradoxes are not magic. Berkson, Lord, Simpson, regression to the mean β€” each becomes transparent once you see the underlying selection or conditioning structure. The lab shows the structure, not just the surprising result.
  • Individual cases, not just group means. From diagnostic intervals to the Jacobson-Truax method for reliable and clinically significant change, the lab takes seriously the question every practitioner faces: what does this mean for this one person?
  • Honest about assumptions. Tools surface what is usually swept under the rug β€” non-normal norm samples, measurement error, the arbitrariness of cut-offs, the gap between "reliable" and "meaningful."

πŸ—Ί The Learning Path

The lab is organized into six sections that build on one another. Work through them in order, or jump in wherever your current knowledge begins.

Section What you will learn Key Tools
1 Β· Regression & Relationships How linear relationships are estimated and decomposed β€” the foundation for everything else OLS Regression Β· Partial Correlation Β· Mediation Β· Moderation Β· ANOVA Β· Multilevel Models
2 Β· Inference & Planning What a p-value actually claims, how big a sample needs to be, what counts as a meaningful effect, and what missing data does Significance Tests Β· Effect Sizes Β· Power & Sample Size Β· Missing Data
3 Β· Statistical Paradoxes Why counter-intuitive results are, in fact, plain regression or sampling logic Regression to the Mean Β· Attenuation Β· Berkson Β· Lord Β· Simpson Β· Law of Small Numbers
4 Β· Causal Inference When association may be read as causation β€” and the toolkit of modern causal analysis Causal Foundations Β· RDD Β· Difference-in-Differences Β· Propensity Score Matching
5 Β· Test Theory & Measurement How psychological constructs are measured and how well Classical Test Theory Β· Measurement Models (CFA) Β· Factor Analysis Β· Reliability (Ξ± vs. Ο‰) Β· IRT Β· DIF
6 Β· Diagnostics & Test Quality How accurately an instrument detects what it should β€” at the level of the individual Sensitivity/Specificity Β· Diagnostic Validity Β· Taylor-Russell Β· ICC Β· Diagnostic Intervals Β· Norm-Score Distortion Β· Jacobson-Truax

πŸ›  The Ecosystem

Legend: βœ“ finished Β· β—· planned

1 Β· Regression & Relationships

The foundation: how to estimate linear relationships, what "controlling for" means, and how effects split into direct and indirect paths.

  • βœ“ OLS & Multiple Regression β€” Simple and multiple regression: coefficients, RΒ², assumptions, and typical violations. Interactive data injection.
  • βœ“ Partial Correlation β€” The correlation of two variables purged of a third. Suppressor effects, semi-partial correlation, path coefficients.
  • βœ“ Mediation Analysis β€” Hayes/PROCESS Model 4: DAG visualization, direct and indirect effect, Sobel/Aroian live, bootstrap on demand.
  • βœ“ Moderation Analysis β€” Hayes/PROCESS Model 1: simple slopes, Johnson-Neyman technique, floodlight plot. When does W moderate the effect of X on Y?
  • βœ“ ANOVA β€” One-way and two-way analysis of variance with interaction β€” the same dataset recomputed as a regression with dummy vs. effect coding. Same F, same p, different parametrization.
  • βœ“ Multilevel Models β€” Complete / no / partial pooling compared, shrinkage caterpillar plot, ICC slider, random slopes (τ₁, ρ), Simpson's paradox experienced live, plus a dedicated "Shrinkage Extreme" module showing when and why partial pooling improves out-of-sample prediction for small/unbalanced groups.

2 Β· Inference & Planning

How large must a sample be? What is a meaningful effect? And what happens when data are missing?

  • βœ“ Significance Tests β€” What a p-value actually claims: repeated draws under Hβ‚€ make the definition tangible, the "dance of the p-values" shows its instability. Five classic misinterpretations corrected.
  • βœ“ Effect Sizes β€” Cohen's d, r, Ξ·Β², f, OR/RR: conversions, visualizations, confidence intervals.
  • βœ“ Effect Size Calculator β€” Exact computation on your own data: CSV upload, d_z/d_rm/d_av, Student's t vs. Welch's t side by side, Hedges/Nakagawa/Bonett corrections, confidence intervals via Bonett, Rosenthal, and the noncentral t-distribution.
  • βœ“ Power & Sample Size β€” Four-curve family, Hβ‚€/H₁ distributions, nΓ—d heatmap, sensitivity curve, MDES, SESOI. Ξ±, Ξ², d, n fully interactive.
  • βœ“ Missing Data β€” MCAR, MAR, MNAR: understand the mechanisms, compare four imputation strategies, traffic-light grid for bias and efficiency.

3 Β· Statistical Paradoxes

Counter-intuitive effects that trap even experienced researchers. All of them rest on regression logic β€” which is why this section comes after the foundations.

  • βœ“ Regression to the Mean β€” Blood-pressure example, r-slider, bidirectional selection. Why "bad" extreme values look better the second time.
  • βœ“ Measurement-Error Attenuation β€” Intelligence & school achievement: disattenuation, learning cards on N-instability and reliability choice.
  • βœ“ Berkson's Paradox β€” Talent + effort β†’ success: DAG, diagonal selection boundary, collider bias. Negative correlation in selected samples.
  • βœ“ Lord's Paradox β€” ANCOVA vs. change scores: when do the two analyses lead to opposite conclusions? DAG, key formula, lege-artis decision.
  • βœ“ Simpson's Paradox β€” An aggregate (pooled) trend reverses on disaggregation: within- vs. pooled regression, confounding slider, reversal & amplification scenarios. Cross-link to the multilevel tool.
  • βœ“ Range Restriction β€” How selection on the predictor X attenuates the correlation (variance restriction). Scatter full vs. restricted, dispersion bars, the Thorndike r(u) relationship, Case-II correction. The slope stays, r shrinks. Cross-linked with Berkson and Taylor-Russell.
  • βœ“ Lindley's Paradox β€” p < .05 yet the Bayes factor supports Hβ‚€: the frequentist–Bayes conflict. Lindley mode, z-distribution Hβ‚€ vs. H₁, divergence over n. Synergy with the Bayes Thinking Lab.
  • βœ“ Will Rogers Phenomenon β€” When the Okies left Oklahoma for California, the average IQ of both states rose. Reclassification raises the mean of both groups at once, while the overall mean stays put β€” animated. Also: stage migration.
  • βœ“ Data Dredging & the Streetlight Paradox β€” Search a large dataset for significance and you will find spurious hits by chance. Multiple comparisons, P(β‰₯1)=1βˆ’(1βˆ’Ξ±)ᡐ, Bonferroni, HARKing headlines, plus stepwise regression (Freedman) β€” spuriously significant predictors, RΒ² inflation.
  • βœ“ Inspection Paradox β€” You land by chance in above-average-length intervals (buses, clinical cohorts, survival data). Length-biased sampling: the experienced gap = E[LΒ²]/E[L] β‰₯ E[L].
  • βœ“ Survivorship Bias β€” Wald's bombers: where to add armor? Mark the damaged zones, then reveal β€” the missing (shot-down) planes show where it actually matters. Selecting on the outcome distorts every conclusion drawn from survivors alone.
  • βœ“ Law of Small Numbers β€” The "record" municipalities (highest and lowest rates) are systematically the smallest β€” pure sampling variance, not a real effect. Funnel plot, ranked list, SE ∝ 1/√n. Related to Regression to the Mean.
  • βœ“ Friendship Paradox β€” Your friends have more friends than you do, on average β€” size-biased sampling over edges. Click people, make a prediction, reveal. Related to the Inspection Paradox.

4 Β· Causal Inference

Under what conditions may association be read as causation? The potential-outcomes framework, natural experiments, and matching.

  • βœ“ Causal Inference β€” Foundations β€” DAGs, potential outcomes, IPW, G-computation: five modules in a full-screen layout. Strong confounding, live ATE estimation.
  • βœ“ Regression Discontinuity (RDD) β€” Sharp RDD: scenarios A–D, polynomial-degree slider, fuzzy RDD, and polynomial pitfalls as expandable modules.
  • βœ“ Difference-in-Differences β€” Regression model, ATT, selection bias, naΓ―ve comparisons, live-updating arrows. The parallel-trends assumption made visual.
  • βœ“ Propensity Score Matching β€” Logistic PS estimation, greedy 1:1 matching, IPW (ATT), love plot, estimator comparison before/after matching.

5 Β· Test Theory & Measurement

How psychological constructs are measured β€” and how well a test does that. Ordered as a learning path: classical foundations β†’ latent measurement models β†’ reliability β†’ modern IRT.

  • βœ“ Classical Test Theory β€” Basics β€” True-score model X = T + E, reliability as a variance ratio, SEM, confidence band around a single test score, Spearman-Brown test-length prophecy. The anchor of the whole reliability theme.
  • βœ“ Measurement Models (CFA intro) β€” Parallel, tau-equivalent, congeneric: path diagram, implied covariance matrix, automatic model classification, Ο‰_total vs. Ξ± live, optional intercepts for essential equivalence. The bridge from factor analysis to reliability.
  • βœ“ Factor Analysis β€” PAF (genuine EFA), correct oblique solution (pattern matrix Ξ›), reliability panel (Ξ± / Ο‰t / Ο‰h Schmid-Leiman), three scenarios, biplot.
  • βœ“ Reliability: Ξ± vs. Ο‰ β€” Cronbach's Ξ± vs. McDonald's Ο‰_total / Ο‰_hierarchical: when Ξ± underestimates (congeneric) and when it feigns unidimensionality (multidimensional). Variance decomposition, split-half distribution, three scenarios.
  • βœ“ IRT β€” Dichotomous Models β€” 1PL/2PL/3PL/4PL + Rasch: ICC, TIF, Wright map, MLE person estimation. The Rasch-vs-1PL distinction made explicit.
  • βœ“ IRT β€” Ordinal Models β€” PCM, GPCM, GRM: CRF, ESC, item information for every item. Disordered-threshold warning, factor-analysis link in the help.
  • βœ“ Differential Item Functioning (DIF) β€” 2PL model, four items, Ξ”b/Ξ”a sliders, ICC comparison, difference curve, group distributions, Raju SA/UA, ETS A/B/C.
  • βœ“ Measurement Invariance β€” Configural / metric / scalar: the CFA twin of DIF. Set a true difference and a degree of non-invariance and see how much of a group difference is real vs. a measurement artifact.
  • β—· Generalizability Theory β€” G-theory as the multi-facet extension of CTT: variance components across raters, items, occasions; G- and D-studies.

6 Β· Diagnostics & Test Quality

How well does an instrument detect what it should β€” at the level of the individual patient?

  • βœ“ Sensitivity & Specificity β€” ROC curve, AUC, PPV/NPV as a function of prevalence: interactive cut-off, live 2Γ—2 table.
  • βœ“ Diagnostic Validity β€” Construct, criterion, and content validity: validity coefficients, Taylor-Russell tables, utility analysis.
  • βœ“ Taylor-Russell Tables β€” The selection utility of a test: success rate (PPV) among those selected, as a function of validity, selection ratio, and base rate. Interactive table plus nomogram with a target-cross reticle; bivariate-normal computation. Cross-linked with Range Restriction.
  • βœ“ Test Bias β€” Cleary model, Meade & Fetzer, adverse impact: three modules, four scenarios. Differential prediction vs. fairness in practice.
  • βœ“ Diagnostic Intervals β€” Confidence, prediction, and tolerance intervals compared: what each one says and which question it answers.
  • βœ“ Jacobson-Truax Analysis β€” Reliable Change Index & clinical significance: RCI band, cut-off criteria a/b/c/d, five-group classification (analogous to the R package JTRCI) in the pre-post plot. BDI-II and SCL-90 as worked examples.
  • βœ“ Norm-Score Distortion β€” What linear T-scores do to right-skewed distributions (SCL-90 analogy). Gamma, log-normal, exponential, Weibull, ex-Gaussian; empirical vs. T-norm percentile ranks; area-transformation (Lienert & Raatz) overlay.
  • βœ“ Profile Analysis β€” Profile comparison with Cattell rβ‚š, McCrae Iβ‚šβ‚/rβ‚šβ‚, ICC_de: elevation, scatter, and shape evaluated separately.
  • βœ“ ICC-Lab β€” Intraclass correlation: all six Shrout-&-Fleiss forms, rater count, absolute vs. consistency agreement.
  • βœ“ Inter-Rater Agreement β€” Cohen's ΞΊ, weighted ΞΊ, the kappa paradox, prevalence & bias: when ΞΊ misleads and what to report instead.

πŸš€ Getting Started

The Methods Lab is a serverless web application. No installation, no backend, and no data leaves your machine.

Option 1 β€” Use it online (recommended)

Open the hosted version directly in your browser:

🌐 www.methodslab.uni-osnabrueck.de

No setup required.

Option 2 β€” Run it locally

  1. Clone the repository:
    git clone https://github.com/raduesing/MethodsLab.git
  2. Open index.html in any modern web browser.
  3. Follow the learning path β€” or jump directly to the tool that matches your current need.

Tip: Every tool includes a built-in ? Help panel with worked examples, formulae, and references, plus inline explanations for each parameter and a Light/Dark theme toggle.


πŸ›  Built With

Vanilla JS Β· HTML5 Canvas Β· No framework Β· No build step Β· Runs entirely in the browser.


πŸ”— Sister Project β€” Bayes Thinking Lab

Frequentist methods are at home here β€” but for priors, posterior distributions, ROPE decisions, Bayes factors, and brms model building there is the Bayes Thinking Lab: 21 interactive tools that teach Bayesian thinking from the ground up. The two labs complement each other β€” Lindley's Paradox, for instance, needs both perspectives.

🌐 www.bayes-thinking-lab.uni-osnabrueck.de Β· πŸ’» github.com/raduesing/Bayes_Thinking_Lab


πŸŽ“ Citation

If you use the Methods Lab for research, teaching, or software development, please cite it as follows:

APA Style

DΓΌsing, R. (2026). Methods Lab: An interactive suite of statistical methods tools for teaching and research (Version 0.7 beta). GitHub. https://github.com/raduesing/MethodsLab

BibTeX

@software{Duesing_Methods_Lab_2026,
  author  = {DΓΌsing, Rainer},
  title   = {{Methods Lab: An interactive suite of statistical methods tools for teaching and research}},
  url     = {https://github.com/raduesing/MethodsLab},
  version = {0.7 beta},
  year    = {2026}
}

πŸ“¬ Contact

Methods Lab Β· University of OsnabrΓΌck Fachgebiet Forschungsmethodik, Diagnostik & Evaluation Dr. Rainer DΓΌsing Β· ResearchGate βœ‰οΈ wwwmelab@uni-osnabrueck.de


Built for teaching. Runs in any browser. No data ever leaves your machine.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages