Doing ML in a field that runs on R and SPSS
In short: psychiatry researchers trust R and SPSS because they've checked those tools for years. Python output earns the same trust when it reproduces a result they already know, comes back in the…
- published
- read time
- 5 min
- words
- 907
- lang
- en
- filed under
- Research
In short: psychiatry researchers trust R and SPSS because they've checked those tools for years. Python output earns the same trust when it reproduces a result they already know, comes back in the shape they already read, and lands in the tool they already use.
Most researchers in psychiatry and psychology analyse their data in R or SPSS. That's not inertia. Their statistics courses used those tools, their reviewers read that output, and the procedures have been checked by a lot of people over a long time.
When I joined the computational side of a multi-site research team, we wrote our machine learning toolbox in Python. We said so plainly in the README: Python was more convenient to build in and to run in different environments, at different sites, on different machines. That was true. It was also the start of a conversation, because a psychiatrist looking at a number from a tool they've never used has every right to ask whether it's correct.
Why the tools differ, honestly
Neither side is wrong. They are built for different jobs.
What each side is good at
R and SPSS
- Decades of checked statistical procedures
- Output that reviewers and journals expect
- Formula syntax that reads like the hypothesis
- Point and click in SPSS, no code needed
Python
- The machine learning libraries live here
- Same code on a laptop, a server or a cluster
- Easy to package as one tool with arguments
- Cross-validation and pipelines built in
The gap is not about which language is better. It is that a clinician reads a regression table the way an engineer reads a stack trace. They know where to look and what looks wrong. Hand them a confusion matrix and a dictionary of hyperparameters and they lose that instinct. Trust has to be rebuilt from something they recognise.
Start by reproducing something boring
Before showing a new model, show an old result. Pick an analysis the group already ran in R or SPSS, a logistic regression they've published or presented, and run it in Python on the same data. If the coefficients match, you've bought credibility for everything that comes after.
import pandas as pd
import statsmodels.formula.api as smf
df = pd.read_csv("cohort_extract.csv")
# The R version the team already trusts:
# glm(outcome ~ exposure + age + sex, family = binomial, data = df)
fit = smf.logit("outcome ~ exposure + age + C(sex)", data=df).fit()
print(fit.summary()) # same layout of estimates, SE, z and p
print(fit.conf_int()) # 95% confidence intervals
statsmodels uses the same formula style as R, so the line on screen reads like the line they would have typed. That matters more than it sounds. When the code looks like their hypothesis, they can review it.
Give the output back in their shape
A machine learning toolbox doesn't need to hand over a pickle file. It can write the same kind of table a researcher already reads: one row per variable, an estimate, an interval. It can write a plain CSV that opens in SPSS or R without any conversion.
Once a result is in their own tool, they can poke at it. Re-plot it. Run a quick test of their own against it. That checking is the trust. You can't shortcut it with a nicer slide.
For the models that aren't regressions, explanation tools like SHAP help. They show which inputs pushed a prediction up or down, which maps loosely onto the "which factors matter" question the team already asks. I say loosely on purpose. A SHAP value is not an effect size, and saying so out loud earns more trust than pretending it is.
Arguments, not edits
The toolbox was controlled by arguments, and the README pointed users to a table of the parameters they could change. Nobody had to open a Python file to run a different outcome or a different set of predictors. A simplified sketch of the idea:
import argparse
p = argparse.ArgumentParser(description="Run one analysis")
p.add_argument("--data", required=True)
p.add_argument("--outcome", required=True)
p.add_argument("--model", choices=["logit", "rf", "svm"], default="logit")
p.add_argument("--folds", type=int, default=5)
p.add_argument("--seed", type=int, default=42)
args = p.parse_args()
This does two things. It means a researcher who doesn't write Python can still run the analysis and change it. And it means every run can be described in one line, which can go into a methods section as is. "We ran the toolbox with these arguments" is a sentence a reviewer can check.
Say what it isn't
Our README carried a disclaimer: the toolbox does not offer a prognosis or a diagnosis. That line did real work. Clinicians are right to be wary of a model that sounds like it is making a clinical call. Being clear that this was a research tool, for finding patterns in cohort data, made the rest of the conversation easier.
The public code was also validated on an open-access dataset, because the real cohort data could not be shared. Anyone who doubted the pipeline could run it themselves on data they could see.
If you're bringing Python into a team that lives in R or SPSS, try this order. Reproduce one result they already published. Write your output as a CSV they can open. Put every option behind an argument. Then show them the new model.
related