You are currently viewing Interactions are still not important: not in medical data either

Interactions are still not important: not in medical data either

Yesterday I wrote about a study I did of the OKCupid dataset. The point was to show that interaction effects are almost always tiny and you should adjust your effect size priors downwards, far far downwards. Naturally, some people objected to the dataset being used, which is amusingly a kind of interaction claim (that is, that interaction effects are weaker in dating site data than other data). However, this is all fine because there’s lots of other large scale datasets one can use for the same research design. So I decided to do it again and for this I chose the American NHANES. These large studies are done every 2 years and measure a large number of mostly biological and medical properties (e.g. blood work) that aren’t impacted by self-report issues or being on a dating site. To boost the sample size, I used data from the last 2 waves (from 2017–2020 and 2021–2023, total n = 14,608 adults). From these I selected 139 variables measured for at least 8,000 people; each model uses the people with data on all three of its variables (median n = 11,024).

The analytic strategy was about the same but expanded a bit:

  1. We pick 3 random variables, set the first as the outcome, and the other 2 as the predictors, thus getting the general model Y ~ X1 + X2. This we test against the interaction model Y ~ X1 * X2. The variables are continuous variables, occasionally quasi-continuous, so we don’t have to use anything other than OLS
  2. In variants of the above, we swap X2 for sex (binary), since sex interactions are probably those that are most common/largest.
  3. We use LOOCV to assess how much variance the interactions add. This is probably better than trusting the analytic approximations (adjusted R²).

Alright, so did this dataset change the conclusions? Not really. We find just about the same as before:

Or if you prefer a table:

So in a sizable number, even a majority, of models the variance explained % gain was negative, which is to say that there was so little signal of an interaction that model overfitting caused the gain to be negative. Some somewhat sizable interactions were found in about 1% of the models. But if you look more closely at them, you can see the problem is not an interaction but that the main effects are misspecified as linear when they are really nonlinear:

In this case, the model misfit is also caused by a data dependency. Someone measured at age 20 cannot have their age of heaviest weight at more than 20. It’s a kind of duh finding and not what people have in mind when they claim interactions.

I thus redid the models adding splines for the main effects and still testing for linear interactions (Y ~ spline(X1) + spline(X2) + X1:X2). Allowing for nonlinear main effects further reduced the apparent effect sizes of the interactions:

But this still leaves the possibility that the interaction itself is nonlinear (Y ~ spline(X1) * spline(X2)). Here’s a clean case of that happening:

Men’s level of estradiol is basically constant throughout their life (flat line fit), while women’s follow the fertile window (requiring a spline), and thus the interaction also requires a spline. Thus, redoing the models once again with nonlinear interactions allowed:

Given the extreme model flexibility, it is very apt to overfit even on large datasets, and 87% of non-sex models were worse than without the interaction. After all of this, something like 0.6% of random trio models, and 2% of random trio models with sex added at least 1% variance explained. Most of these are still going to be marginal. Like this one:

Looking closely, you can see the lines diverge slightly at the high end. And this is a relatively large real interaction effect!

As such, the results from the prior study cannot be explained away as being due to the nature of the dataset used, unless of course you can come up with another ad hoc hypothesis for why it also fails in NHANES, and presumably why interactions would show up in other kinds of data. These 2 studies cannot rule out that you can find some kind of dataset where practically relevant interaction effects show up (say, gene expression levels or other low-level biological systems). For now, however, the burden of proof is on those thinking interactions are commonplace and practically meaningful. Meanwhile, whenever you read a new study claiming some interaction effect with p = 4% based on a sample of 100 students, you are right to be very skeptical indeed.