{"id":16041,"date":"2026-10-03T12:29:53","date_gmt":"2026-10-03T11:29:53","guid":{"rendered":"https:\/\/emilkirkegaard.dk\/en\/?p=16041"},"modified":"2026-10-03T12:39:37","modified_gmt":"2026-10-03T11:39:37","slug":"interactions-are-generally-not-important-in-social-science","status":"publish","type":"post","link":"https:\/\/emilkirkegaard.dk\/en\/2026\/10\/interactions-are-generally-not-important-in-social-science\/","title":{"rendered":"Interactions are generally not important in social science"},"content":{"rendered":"<p>Some years ago both <a href=\"https:\/\/rpubs.com\/EmilOWK\/interactions_social_science\">me<\/a> and <a href=\"https:\/\/rpubs.com\/Jonatan\/interactions\">Jonatan Pallesen<\/a> looked into the generality of statistical interaction effects. These look like this:<\/p>\n<p><a href=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_illustration.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-16042\" src=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_illustration.png\" alt=\"\" width=\"1800\" height=\"720\" srcset=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_illustration.png 1800w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_illustration-300x120.png 300w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_illustration-1024x410.png 1024w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_illustration-768x307.png 768w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_illustration-1536x614.png 1536w\" sizes=\"auto, (max-width: 1800px) 100vw, 1800px\" \/><\/a><\/p>\n<p>In real life, however, interactions are mostly not a real thing. This is despite their sometimes great intuitive appeal and well-known examples (e.g. mating success\/attractiveness as function of height and sex and their interaction). So I decided to hammer home this message so that there is something robust and convincing one can refer to to prove the point. For this purpose, I reached for my trusty <del>revolver<\/del> OKCupid dataset. It&#8217;s very large, and has an extreme variety of items. Yes, it is from a dating site, but every other well-known finding I&#8217;ve tried replicating in the dataset worked fine. The very fact that it works so well, as do other convenience samples, is evidence against the importance of interactions. So here&#8217;s what I did:<\/p>\n<ol>\n<li>Item-level data: pick a random ordinal variable as the outcome (Y), and 2 others as the predictors (X1, X2). Then fit the additive and interaction models: 1) Y ~ X1 + X2, vs. 2) Y ~ X1 * X2. Repeat 1000s of times. This approach needs ordinal regression since the outcomes are ordinal (all items have 2-4 answer options, which in the base case are ordinal or binary, sometimes nominal).<\/li>\n<li>Scale-level data: aggregate items into scales from <a href=\"https:\/\/www.emilkirkegaard.com\/p\/dimensions-of-the-dating-mind\">my recent 10-dimension model<\/a> using a mixture of item types (binary, ordinal, nominal), then repeat the same approach as above using scales instead of items.<\/li>\n<li>Scale-level data: aggregate items into scales that I had previously investigated in <a href=\"https:\/\/emilkirkegaard.dk\/en\/2025\/08\/do-you-like-dogs-cats-both-or-neither\/\">the cat-dog study<\/a>, then repeat the same approach as above. The scales here are quite different from the ones above as they were handpicked to measure particular dimensions of interest instead of &#8216;falling out of&#8217; the data.<\/li>\n<li>Variants of the above swapping X2 for sex, to specifically look at sex interactions, the most intuitively plausible subset.<\/li>\n<\/ol>\n<p>First we have to think of what we care about. The most obvious is to look at the p-values. This was the focus of the prior studies on interactions in this dataset. It is, however, a mistake I think. The dataset is so large that many terms reach p&lt;5% even with tiny effect sizes. What we really care about is not so much whether there is literally any signal larger than pure noise, but whether these signals are of practical concern. For this reason, here we will focus on changes in model explanatory power, that is, R\u00b2 values (<a href=\"https:\/\/emilkirkegaard.dk\/en\/2022\/10\/variance-explained-is-mostly-bad\/\">which can be deceptive<\/a> but bear with me).<\/p>\n<p>Alright, so what was found? It can be summarized in a few ways graphically:<\/p>\n<p><a href=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_raw_vs_adj.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-16048\" src=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_raw_vs_adj.png\" alt=\"\" width=\"1800\" height=\"720\" srcset=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_raw_vs_adj.png 1800w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_raw_vs_adj-300x120.png 300w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_raw_vs_adj-1024x410.png 1024w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_raw_vs_adj-768x307.png 768w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_raw_vs_adj-1536x614.png 1536w\" sizes=\"auto, (max-width: 1800px) 100vw, 1800px\" \/><\/a><\/p>\n<p>In general, random models do not explain a great amount of variance in the outcome, but they do explain more than 0. Everything, after all, is correlated if you look close enough. For the item models, the median additive model explained 0.67% and the interaction model explained 0.67% for a relative gain of 0%. For the 2 sets of scale-level analyses, the results were about the same. For the 10-dimensional scales, the median additive model explained 2.83%, and the interaction model explained 2.97%, or a relative increase of 4.8%. For the replication using the cat-dog set of scales (general intelligence, mental health, antisocial behavior, prudence, conservatism, extroversion, reading, enjoys discussion, religiousness, drugs, kink\/sex), the median additive model explained 5.61% and the median interaction model 5.64%, for a gain of 0.008%. However, even these values can be considered overestimates. The medians don&#8217;t come from the same models. If one compared within the models (trios), then the relative % gains are -0.3% for the items, 0.1% for the cat-dog scales, and 1.2% for the 10 dimensions. So a typical model that explains 2% of variance without interactions grows to something like 2.01%, 2.002%, 2.024%. Practically useless. For the models with sex, they did slightly better, but not much better in absolute terms. The table below summarizes the numerical results:<\/p>\n<p><a href=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_summary_table.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-16044\" src=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_summary_table.png\" alt=\"\" width=\"1600\" height=\"440\" srcset=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_summary_table.png 1600w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_summary_table-300x83.png 300w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_summary_table-1024x282.png 1024w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_summary_table-768x211.png 768w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_summary_table-1536x422.png 1536w\" sizes=\"auto, (max-width: 1600px) 100vw, 1600px\" \/><\/a><\/p>\n<p>Alright, but since we already looked through the data to find interactions, what were the largest ones that were found? For the item-level analysis, they look like this:<\/p>\n<p><a href=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-16043\" src=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest.png\" alt=\"\" width=\"2100\" height=\"1639\" srcset=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest.png 2100w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest-300x234.png 300w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest-1024x799.png 1024w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest-768x599.png 768w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest-1536x1199.png 1536w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest-2048x1598.png 2048w\" sizes=\"auto, (max-width: 2100px) 100vw, 2100px\" \/><\/a><\/p>\n<p>It is somewhat tricky to understand the plot since there are 3 moving parts. The lines are the 2 sexes in their appropriate colors. The interaction term is always sex, and 4 of the 4 involve changing your name upon marriage, traditional female behavior. Thus, conservative women will lean towards &#8220;yes&#8221; and conservative men towards &#8220;no&#8221;. Any question that relates to conservatism thus works to get this effect (male gallantry, flag burning, watching sports on TV, mate guarding). These examples are somewhat trivial, as when I found that the largest effect of your astrological sign was predicting which season you like the most (few people prefer winter, but those born in the winter prefer it slightly more than average). For the scales, the largest interactions were:<\/p>\n<p><a href=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest_scales.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-16045\" src=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest_scales.png\" alt=\"\" width=\"2000\" height=\"1680\" srcset=\"https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest_scales.png 2000w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest_scales-300x252.png 300w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest_scales-1024x860.png 1024w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest_scales-768x645.png 768w, https:\/\/emilkirkegaard.dk\/en\/wp-content\/uploads\/interactions_largest_scales-1536x1290.png 1536w\" sizes=\"auto, (max-width: 2000px) 100vw, 2000px\" \/><\/a><\/p>\n<p>These are basically the same thing 4 times. It&#8217;s a measure of conservatism, a measure of religiousness, and a measure of libertinism (sex or drugs). These effects are also relatively minor, the largest added an additional 1.25% variance but the main model already explained 25%, so the relative gain was only 5%.<\/p>\n<p>One limitation of the above is that the measures aren&#8217;t perfectly reliable. Mathematically, in practice interactions also suck because interaction terms have the reliability of the product of the terms, so if both main variables have a semi-respectable reliability of 0.80, the interaction term is only .80\u00b2=0.64. This severely reduces statistical power. One can solve this in theory using latent variable modeling. Sadly, the most commonly used framework for this in R (lavaan) does still not support interactions between latent terms, but there are now some other packages that allow for this (<a href=\"https:\/\/modsem.org\/\">modsem<\/a>). Just for curiosity, I refit the models above using this framework to see how much things change. This was one of those times when the last 5% of a study opened up a can of worms that would potentially take hours to resolve. The various versions of latent interaction models took a long time to fit (even for a single model), and\/or produced unstable and implausible results, apparently related to inability to handle ordinal indicators (as opposed to continuous ones as normally used). Failing this, I fell back on a simpler method and brought out the trusty <a href=\"https:\/\/cran.r-project.org\/web\/packages\/simex\/index.html\">simex package<\/a>. This is a clever method where instead of trying to model the error-free variables, you add <em>more<\/em> error to them and work backwards to 0 error. It only needs to know the amount of (random, classical) error, which is the same as the reliabilities. And it fit instantly. It tells us that in the above cases of the four models, coefficients are about 30% larger, so the variance they explain roughly 1.5-2x more variance, which means they are still quite small in relative terms, especially since the main effects also grow. Fitting the <strong>simex<\/strong> for all the scale-level models lets us generalize the findings further. These results showed that when taking into account measurement error, the interaction effects explained 1.3-2.5x as much variance, which in absolute terms was still too small to matter in most cases.<\/p>\n<p>What does this mean for research? It means you should be wary of findings claiming interaction effects. Indeed, it was reported in <a href=\"https:\/\/www.science.org\/doi\/10.1126\/science.aac4716\">one<\/a> <a href=\"https:\/\/academic.oup.com\/qje\/advance-article-abstract\/doi\/10.1093\/qje\/qjy029\/5195544?redirectedFrom=fulltext&amp;login=false\">two<\/a> of the large-scale replication studies that interaction effects replicated more poorly than average. In other words, since they are almost never large enough to matter, you should have a higher standard of evidence requirement to believe them. A p-value of &lt; 5% is certainly not enough, and isn&#8217;t even enough for regular main effects. The lack of interaction effects also implies, indirectly, that selective\/convenience samples are not likely to be severely biased. Why? Because whatever characteristics they are selected on are not likely to change the association you are interested in. The OKCupid dataset is a good demonstration of this. It replicated basically every normal finding I examined: 1) anti-social behavior and intelligence, 2) conservatism and pronatalism\/fertility, 3) mental health and non-heterosexuality\/trans, 4) astrology being bunk, 5) differences between dog and cat enthusiasts, and others. This is despite being an online convenience sample of strange people filling out 100s of questions on a dating site in semi-public.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Some years ago both me and Jonatan Pallesen looked into the generality of statistical interaction effects. These look like this: In real life, however, interactions are mostly not a real thing. This is despite their sometimes great intuitive appeal and well-known examples (e.g. mating success\/attractiveness as function of height and sex and their interaction). So [&hellip;]<\/p>\n","protected":false},"author":17,"featured_media":16043,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1839,1766],"tags":[2707,1866],"class_list":["post-16041","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-psychometics","category-math-science","tag-interactions","tag-okcupid","entry","has-media"],"_links":{"self":[{"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/posts\/16041","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/users\/17"}],"replies":[{"embeddable":true,"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/comments?post=16041"}],"version-history":[{"count":3,"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/posts\/16041\/revisions"}],"predecessor-version":[{"id":16049,"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/posts\/16041\/revisions\/16049"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/media\/16043"}],"wp:attachment":[{"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/media?parent=16041"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/categories?post=16041"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/emilkirkegaard.dk\/en\/wp-json\/wp\/v2\/tags?post=16041"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}