Engagement, User Satisfaction, and the Amplification of Divisive Content on Social Media
A pre-registered audit of Twitter's engagement-based timeline against a reverse-chronological baseline — measured on real users, not modelled. What the claims are, and what they rest on.
In Papers
Engagement, User Satisfaction, and the Amplification of Divisive Content on Social Media — Smitha Milli, Micah Carroll, Yike Wang, Sashrika Pandey, Sebastian Zhao and Anca D. Dragan, May 26, 2023.
Claim ledger
Assessments are the model’s knowledge, not verification.
- 1¶
The paper's headline finding: relative to a reverse-chronological feed, Twitter's engagement ranking selects more emotional and out-group hostile content, and that content worsens readers' view of their out-group.
“In a pre-registered algorithmic audit, we found that, relative to a reverse-chronological baseline, Twitter’s engagement-based ranking algorithm amplifies emotionally charged, out-group hostile content that users say makes them feel worse about their political out-group.”
assertion · unclear
consistent · high confidence — This is the paper's own abstract-level summary and matches the internal statistics reported later in the same document (partisanship +0.24SD, out-group animosity +0.24SD, out-group perception -0.17SD). It is consistent with the broader emerging literature (Brady et al.'s PRIME framework, Rathje et al. on out-group animosity and engagement) that engagement optimization favors divisive content, though this is the first paired audit I'm aware of that isolates the ranking algorithm's effect from follow-graph choice on Twitter specifically.
To check: Independent replication with a similar paired reverse-chronological/engagement-timeline audit design on Twitter/X or another platform.
- 2¶
Users rate the algorithm's political tweet selections lower than the chronological ones, which the authors read as the algorithm failing on stated preference.
“Furthermore, we find that users do not prefer the political tweets selected by the algorithm, suggesting that the engagement-based algorithm underperforms in satisfying users’ stated preferences.”
assertion · unclear
consistent · high confidence — Matches the paper's own reported stated-preference effect for political tweets (-0.18SD, p=0.005) shown later in the same document; internally consistent.
To check: The -0.18 SD political-tweet stated-preference effect in the paper's Table 15/S3.
- 3¶
A simulated stated-preference ranking lowers anger, partisanship and out-group hostility but may increase exposure to belief-confirming content.
“Finally, we explore the implications of an alternative approach that ranks content based on users’ stated preferences and find a reduction in angry, partisan, and out-group hostile content, but also a potential reinforcement of pro-attitudinal content.”
assertion · unclear
consistent · high confidence · novel — Matches the paper's own exploratory SP-timeline results (Figure 1, Figure 2) reported later; the in-group-reinforcement caveat is a genuinely interesting, less-obvious finding that qualifies the 'just use stated preferences' prescription common in prior CS literature.
To check: Figure 2's decomposition of animosity reduction by in-group/out-group target.
- 4¶
As of April 2023 Twitter's ranker combined predictions across ten distinct engagement behaviours.
“For example, in April 2023, Twitter’s ranking algorithm was based on predicting whether a user would engage with a particular tweet using ten different types of engagement [50] .”
quantity · unclear
consistent · medium confidence — Twitter open-sourced part of its recommendation pipeline ('the-algorithm' repo) in March 2023, including a 'heavy ranker' multi-task model predicting several engagement types (like, retweet, reply, profile click, dwell time, etc.). The general shape of this claim matches what I recall of that release, though I can't independently confirm the exact count of ten without the source repo.
To check: Twitter/X's open-sourced 'the-algorithm-ml' repository's recap/heavy-ranker label list, cited as ref [50].
- 5¶
Quoting Brady et al.: engagement optimisation systematically exploits human social-learning biases, because those biases predict attention.
“[13] suggest that, “content algorithms systematically exploit human social-learning biases because they are designed to optimize attentional capture and engagement time on the platform, and social-learning biases strongly predict what users will want to see.””
assertion · unclear
consistent · medium confidence — Accurately attributed to Brady, Jackson, Lindström & Crockett's 'Algorithm-mediated social learning in online social networks' (Trends in Cognitive Sciences, 2023), whose PRIME (prestige, in-group, moral, emotional) framework is a real and reasonably well-known theoretical contribution in this literature.
To check: The cited Brady et al. 2023 Trends in Cognitive Sciences paper's text.
- 6¶
The observed emotion–engagement correlation in prior work is confounded by users' own choice of accounts to follow.
“However, it is possible that those posts received more engagement, not due to the algorithm prioritizing them, but because users chose to follow accounts that post more emotional content.”
assertion · unclear
plausible · high confidence — This is a standard confounding concern in observational social-media research (correlation between engagement and emotional content doesn't disentangle algorithm effects from follow-graph composition); it's the well-recognized motivation for needing a paired/controlled audit design, which the paper goes on to provide.
To check: Comparison of observational studies (e.g., Brady et al. 2017 PNAS) against paired audits like this one, holding follow graph fixed.
- 7¶
Methodological premise: using each user's own reverse-chronological feed as control holds follow choices fixed.
“By comparing to the reverse-chronological baseline, we are able to understand the effects of the engagement-based timeline, beyond users’ own deliberate decisions of who to follow.”
assertion · unclear
consistent · high confidence — Accurately describes the study's own paired design (same user's follow graph feeds both timelines), which is a sound and standard causal-inference logic for isolating ranking effects from follow-choice effects.
To check: The study's methodology section describing simultaneous collection of engagement and reverse-chronological timelines for the same user.
- 8¶
Even a reverse-chronological control is insufficient, because the algorithm could be serving fine-grained preferences users cannot express through follows.
“The algorithm might be surfacing more emotional content because, within the accounts a user follows, the user may actually prefer more emotional posts, even if they lack a direct mechanism to filter content at such a granular level.”
assertion · unclear
plausible · high confidence — A reasonable and non-trivial methodological point—reverse-chronological control doesn't rule out fine-grained revealed preference within the follow graph—that correctly motivates the paper's added stated-preference survey layer.
To check: Whether the stated-preference results (Section 2.1) show the algorithm's selections are disfavored even within-follow-graph, which the paper reports for political tweets.
- 9¶
Internal randomised Twitter experiments show the engagement timeline raises time spent versus reverse-chronological.
“Users’ revealed preferences appear to favor the engagement-based timeline over the reverse-chronological one, as randomized experiments at Twitter have shown that it increases the amount of time users spend on the platform compared to the reverse-chronological timeline [37] .”
assertion · unclear
plausible · medium confidence — Cites Milli, Belli & Hardt 2022 ('Causal Inference Struggles with Agency on Online Platforms'), a real paper by an overlapping author about internal Twitter A/B testing; the general finding that engagement-optimized feeds increase time-on-platform relative to chronological feeds matches common industry knowledge about ranking system evaluation, though I cannot independently verify Twitter's specific internal experiment results.
To check: The cited Milli, Belli & Hardt (2022) FAccT paper's reported internal Twitter experiment results.
- 10¶
The authors assumed retention-optimised ranking would favour positive emotion in readers, since happy users return.
“The rationale for this difference is that platforms test ranking algorithms in A/B tests to identify those that optimize user retention [19] ; thus, we assumed that fostering positive emotions would encourage users to return to the platform.”
assertion · unclear
consistent · high confidence — This is the authors' own stated methodological assumption, directly reported in the source; the assumption that retention-optimizing platforms would favor positive-emotion content is a reasonable (if ultimately falsified, per the next claim) prior.
To check: Cross-reference with the paper's pre-registration (OSF) for the stated rationale behind Hypothesis 5.
- 11¶
Pre-registered Hypothesis 5 predicted readers would feel happier and less negative on the engagement-based timeline.
“Participants will feel happier and less angry, anxious, or sad reading tweets in their personalized timeline, compared to the reverse-chronological timeline.”
prediction · unclear
unverifiable · high confidence — This is a pre-registered hypothesis statement (a prediction as submitted to OSF before data collection), not a factual claim to check against my knowledge — its truth is precisely what the study's own subsequent results address.
To check: The paper's OSF pre-registration at osf.io/upw9a, which should list this as Hypothesis 5 prior to data collection.
- 12¶
Hypothesis 5 was the single pre-registered prediction the data contradicted; readers reported more of all four emotions, happiness included.
“(Ultimately this turns out to be the only hypothesis we did not find supporting evidence for — readers reported feeling higher levels of all four emotions on the engagement-based timeline.)”
assertion · unclear
consistent · high confidence — Matches the paper's own Table 15 results (all four reader-emotion ATEs — angry, sad, anxious, happy — are positive and significant), confirming this specific pre-registered hypothesis (H5) was falsified as claimed. Notably counterintuitive and worth flagging: engagement-optimization increased even happiness reports slightly, contradicting the simple 'engagement=positive-affect' retention story.
To check: Table 15's Reader Happy ATE (+0.119 SD, p<0.001) alongside the other three reader-emotion rows.
- 13¶
All 26 pre-registered outcomes that reach p<0.05 survive FDR correction at 0.01.
“In total, we tested 26 outcomes, and all results that are significant (at a $p$ -value threshold of $0.05$ ) remain significant at a false discovery rate (FDR) of $0.01$ .”
quantity · unclear
consistent · high confidence — Standard, appropriately rigorous multiple-testing correction (Benjamini-Krieger-Yekutieli two-stage FDR) matching the described methodology and the supplementary Table 15's adjusted p-values shown in the source.
To check: Table 15's 'Adjusted p-value' column against the raw p-value column.
- 14¶
Engagement ranking raised tweet partisanship and out-group animosity each by about a quarter of a standard deviation.
“Relative to the reverse-chronological baseline, we found that the engagement-based algorithm amplified tweets that exhibited greater partisanship ( $0.24$ SD, $p<0.001$ ) and expressed more out-group animosity ( $0.24$ SD, $p<0.001$ ).”
quantity · unclear
consistent · high confidence — Matches Table 15 exactly (Partisanship 0.244 SD, Out-group Animosity 0.236 SD, both p=0.0002). Internally consistent, a moderate-sized effect (~quarter SD) typical for content-ranking audits of this kind.
To check: Table 15's Partisanship and Out-group Animosity rows.
- 15¶
Engagement-selected tweets moved readers' feelings against their out-group and mildly toward their in-group.
“Furthermore, tweets from the engagement-based algorithm made users feel significantly worse about their political out-group ( $-0.17$ SD, $p<0.001$ ) and better about their in-group ( $0.08$ SD, $p=0.0014$ ).”
quantity · unclear
consistent · high confidence — Matches Table 15 (Out-group Perc. -0.171 SD, In-group Perc. 0.081 SD, p=0.0014). Note the in-group effect is notably smaller than the out-group effect, an asymmetry the paper itself flags and links to prior work on negativity bias in sharing (Yu, Wojcieszak & Casas).
To check: Table 15's In-group Perc. and Out-group Perc. (all users) rows.
- 16¶
Across all tweets, the algorithm amplified expressed anger most strongly, with smaller effects for sadness and anxiety.
“The engagement-based algorithm significantly amplified tweets that expressed negative emotions—anger ( $0.47$ SD, $p<0.001$ ), sadness ( $0.22$ SD, $p<0.001$ ), and anxiety ( $0.23$ SD, $p<0.001$ ).”
quantity · unclear
consistent · high confidence — Matches Table 15 (Author Angry 0.473, Sad 0.220, Anxious 0.232, all p=0.0002). Anger stands out as by far the largest effect among the negative emotions.
To check: Table 15's Author Angry/Sad/Anxious rows under 'Emotional effects (all tweets)'.
- 17¶
Restricted to political tweets, anger dominates the amplification effect, at 0.75 SD expressed and 0.37 SD felt.
“When considering only political tweets, we found that anger was by far the predominant emotion amplified by the engagement-based algorithm, both in terms of the emotions expressed by authors ( $0.75$ SD, $p<0.001$ ) and the emotions felt by readers ( $0.37$ SD, $p<0.001$ ).”
quantity · unclear
consistent · high confidence — Matches Table 15's political-tweet subgroup (Author Angry 0.754, Reader Angry 0.377), and correctly notes that sadness/anxiety/happiness effects were much weaker or non-significant for political tweets specifically in the same table.
To check: Table 15's political-tweet Author Angry and Reader Angry rows versus the other political emotion rows (mostly non-significant).
- 18¶
Across all tweets, stated preference for engagement-timeline tweets is only marginally higher.
“We found that overall, tweets shown by the engagement-based algorithm are rated slightly higher ( $0.06$ SD, $p=0.022$ ).”
quantity · unclear
consistent · high confidence — Matches Table 15's Reader Pref (all tweets) row (0.065 SD, p=0.0226), a small but statistically significant preference for the algorithmically-ranked tweets overall — which sets up the contrast with the negative political-tweet result.
To check: Table 15's Reader Pref (all tweets) row.
- 19¶
For political tweets specifically, users valued the algorithm's picks significantly less than the chronological ones.
“Interestingly, however, the political tweets recommended by the engagement-based algorithm led to significantly lower user value than the political tweets in the reverse-chronological timeline ( $-0.18$ SD, $p=0.005$ ).”
contrarian · unclear
consistent · high confidence · novel — Matches Table 15 (Reader Pref, political tweets: -0.180 SD, p=0.0054). This overall-positive/political-negative split is arguably the paper's most interesting and least-anticipated result — it's a genuine dissociation worth flagging as a distinct contribution beyond the more expected emotion-amplification findings.
To check: Table 15's Reader Pref (political tweets) row versus Reader Pref (all tweets) row.
- 20¶
The constructed stated-preference timeline lowered negative emotion and raised happiness relative to the engagement timeline, for both authors and readers.
“As shown in Figure 1, relative to the engagement timeline, the SP timeline reduced negativity (anger, sadness, anxiety) and increased happiness, both in terms of the emotions expressed by authors and the emotions felt by readers.”
assertion · unclear
consistent · medium confidence — Directly quoted from the source's discussion of Figure 1; I cannot independently verify Figure 1's underlying numbers since only the caption/description (not the raw exploratory-SP-timeline data table) is fully shown in the excerpt, but the description is internally coherent with the rest of the paper's narrative.
To check: Figure 1's full ATE comparison table for the exploratory stated-preference (SP) timeline (not fully reproduced in the excerpt).
- 21¶
The algorithm satisfies revealed preference (time spent) while failing stated preference, most sharply on political content.
“Thus, the engagement-based algorithm seems to be catering to users’ revealed preferences (in terms of engagement and usage patterns) but not to their stated preferences , particularly when it comes to political content.”
assertion · unclear
consistent · high confidence — A fair summary conclusion given the juxtaposition of the cited internal Twitter time-spent A/B results with the paper's own stated-preference survey findings; follows logically from claims 8, 17, and 18 as reported.
To check: Juxtaposition of ref [37] (Milli et al. 2022 time-spent results) against this paper's own Reader Pref ATEs.
- 22¶
Misalignment with stated preference rules out deliberate user choice as a full explanation of the algorithm's content effects.
“The fact that the engagement-based algorithm is not aligned with users’ stated preference for tweets also suggests that its effects cannot entirely be explained by users’ deliberative choices.”
assertion · unclear
plausible · medium confidence — The logic is reasonable but slightly stronger than the evidence strictly requires: a mismatch in stated preference for tweet content doesn't fully rule out that users' 'deliberative choices' about time allocation (revealed preference) could still be doing real explanatory work; this is an inferential leap the authors themselves seem aware needs the stated-preference measure as a proxy for deliberation, which is itself debatable.
To check: Whether stated 'Value' survey responses are a valid proxy for deliberative/System-2 preference versus another noisy signal.
- 23¶
The audit establishes amplification of emotional content over and above both follow choices and tweet-level stated preferences.
“Our study provides insight into this debate by providing evidence that the engagement-based algorithm amplifies emotionally charged content, beyond users’ choices of who to follow, and even beyond what users’ stated preferences at the tweet-level would indicate.”
assertion · unclear
consistent · high confidence — A fair synthesis of the paper's two-layer control design (reverse-chronological + stated-preference survey), directly supported by the emotion-amplification and stated-preference-mismatch results reported elsewhere in the source.
To check: Joint reading of Tables 15's emotion ATEs and Reader Pref ATEs.
- 24¶
Engagement ranking is more polarising than the user's own follow graph would produce.
“Our results suggest that the engagement-based ranking algorithm selects more polarizing content compared to what would be expected from users’ following choices alone.”
assertion · unclear
consistent · high confidence — Follows from the in-group/out-group perception effects (claim 14) and partisanship/animosity effects (claim 13), both measured relative to the same-user reverse-chronological baseline, which holds follow choices fixed by design.
To check: Table 15's political-effects block collectively.
- 25¶
Guess et al. (2023) found no significant effect of Facebook and Instagram ranking algorithms on affective polarisation over weeks to months.
“They found no significant impact on users’ affective polarization.”
assertion · unclear
consistent · medium confidence — Refers to the 2020 US Facebook/Instagram Election Study consortium papers published in Science (2023), including Guess et al.'s feed-algorithm experiment; my recollection is that these large field experiments found feed-algorithm manipulations (e.g., switching to chronological) had limited or null detectable effects on affective polarization/attitude measures over the study window, consistent with this characterization, though I can't recall every specific null result with full precision.
To check: Guess et al. 2023, 'How do social media feed algorithms affect attitudes and behavior in an election campaign?', Science 381(6656).
- 26¶
Null findings on polarisation miss equilibrium effects, since algorithms also change what content gets produced.
“Crucially, however, the scope of Guess et al. (2023)’s research does not encompass “general equilibrium” effects, such as the possibility for ranking algorithms to indirectly shape user attitudes by incentivizing and changing the type of content that users produce in the first place.”
assertion · unclear
plausible · medium confidence · novel — A reasonable methodological critique — the Guess et al. field experiments measured short/medium-term individual-level attitude change, not longer-run content-ecosystem shifts driven by creator incentives — and the 'general equilibrium' framing for this gap is a useful, less commonly articulated concept in this specific debate.
To check: Whether any study has since attempted to measure creator-side content adaptation to ranking incentives at the ecosystem level.
- 27¶
Hedged conclusion: stated-preference ranking may increase pro-attitudinal exposure compared with engagement ranking.
“Overall, these results suggests that it is possible that ranking by stated preferences could heighten users’ exposure to content that reinforces their pre-existing beliefs, relative to the engagement-based algorithm.”
assertion · unclear
consistent · medium confidence · novel — An appropriately hedged conclusion drawn directly from the Figure 2 decomposition (claim 20) and consistent with prior research on pro-attitudinal preference (Bakshy, Messing & Adamic 2015 on Facebook exposure), cited by the authors themselves.
To check: Figure 2 in-group/out-group tweet-share breakdown for the SP timeline versus the other two timelines.
- 28¶
Contrary to Agan et al.'s Facebook result, political in-group preference appears partly deliberate rather than purely automatic.
“However, our results reveal an opposite trend when defining in-groups and out-groups by political affiliation. This might indicate that political in-group preferences are not solely “system 1” biases but also involve more deliberate “system 2” preferences.”
contrarian · unclear
plausible · low confidence · novel — A genuinely novel synthesis contrasting with Agan et al.'s NBER finding (race/religion in-groups on Facebook), but it rests on a single small-sample, small-tweet-pool exploratory analysis (only ~20 candidate tweets per user) and a domain difference (political vs. racial/religious in-groups) that could equally reflect measurement or context differences rather than a genuine System-1/System-2 distinction; the authors themselves hedge with 'might indicate.'
To check: Replication of the Agan et al. Facebook in-group/stated-preference design specifically on political in-groups, with a larger candidate-tweet pool.
- 29¶
The stated-preference timeline was built from only ~20 candidate tweets, far fewer than a platform's real candidate pool.
“An important caveat in interpreting our results is that we only had a pool of approximately twenty tweets—the top ten in the users’ chronological and engagement timelines—to choose from when ranking by stated preferences.”
assertion · unclear
consistent · high confidence — An honest, directly-quoted limitation acknowledged by the authors themselves; it's a real methodological constraint that appropriately tempers claims 27 and 28.
To check: Compare the SP timeline construction (top 10 of ~20 candidates) against a real deployment with a much larger retrieval pool.
- 30¶
Breaking ties by out-group animosity rather than randomly yields a stated-preference feed that cuts hostility without increasing in-group exposure.
“Compared to the engagement-based and reverse-chronological timeline, we find that this modified SP timeline greatly reduces the amplification of out-group hostile content while avoiding heightened exposure to in-group content.”
assertion · unclear
plausible · low confidence · novel — This references Supplementary Section S4.9, which is outside the excerpted text range provided, so I cannot check the underlying numbers; the qualitative claim (tie-breaking on out-group-animosity to mitigate the in-group-reinforcement problem found in claim 20) is a sensible design tweak but I have no independent way to verify the reported magnitude.
To check: The full S4.9 supplementary table (not included in the provided excerpt) comparing the modified SP timeline's animosity and in-group-exposure metrics.
- 31¶
Stated-preference signals are sparse in production, so engagement remains the dominant ranking objective.
“On real-world platforms, users’ stated preferences (in the form of user surveys or user controls like a “Do not recommend” button) are rarely observed, making engagement-based signals still the predominant target for ranking”
assertion · unclear
consistent · medium confidence — Matches general industry knowledge: engagement/behavioral signals (clicks, dwell time, likes) are abundant and cheap to observe at scale, while explicit stated-preference signals are sparse and costly to collect, which is why engagement-based ranking has dominated production recommender systems historically.
To check: Cunningham et al.'s cited review [19], 'What We Know About Using Non-Engagement Signals in Content Ranking' (2024), for a survey of industry practice.
- 32¶
The literature on stated-preference ranking rests on an empirically untested assumption that such alignment improves outcomes.
“A fundamental, yet largely untested, assumption in this line of work is that aligning algorithms with users’ stated preferences can lead to improved outcomes for both individuals and society.”
assertion · unclear
plausible · medium confidence — A fair characterization of a nascent CS subfield (stated-preference/'System-2' recommenders), though notably the corresponding author (Milli) is a co-author of one of the cited prior works [36] proposing exactly this kind of method — worth flagging that the framing of an 'untested assumption' also motivates and validates the authors' own prior research agenda.
To check: Whether any deployed platform has run a large-scale field experiment directly comparing stated-preference-aligned ranking against engagement ranking on downstream individual/societal wellbeing outcomes.
- 33¶
Conditional prediction: if creators adapt to the algorithm, the measured amplification understates long-run effects.
“If users have learned to produce more of the content that the engagement-based algorithm incentivizes [13, 16] , then its long-term effects on the content-based outcomes we measure (emotions, partisanship, and out-group animosity in tweets) may be even greater.”
prediction · unclear
unverifiable · medium confidence — An explicitly conditional, forward-looking claim about feedback loops between ranking incentives and content production that the study itself did not test; plausible in principle given documented creator adaptation to platform incentives generally, but not something the paper's cross-sectional audit design can confirm or refute.
To check: A longitudinal study tracking whether creator posting behavior shifts toward higher-engagement (angrier/more partisan) content over time on a given platform.
- 34¶
Users view about 30 tweets per Twitter session, so a top-ten window is a defensible slice.
“Looking at the first ten tweets might be a reasonable approximation though—Bouchaud et al. (2023) found that users viewed an average of 30 tweets in one session of using Twitter [9] .”
quantity · unclear
plausible · medium confidence — Cites a real, concurrent crowdsourced Twitter audit (Bouchaud, Chavalarias & Panahi, Scientific Reports 2023) listed in the bibliography; the specific 30-tweets-per-session statistic is plausible for that study's design but I cannot independently confirm the exact figure from memory.
To check: Bouchaud et al. (2023), 'Crowdsourced audit of Twitter's recommender systems,' Scientific Reports 13(1).
- 35¶
The recruited sample skews younger and more Democratic than ANES's Twitter-using population.
“our population tended to be younger (53% of our study were aged 18-34 years old, compared to 33% in the ANES study) and more likely to affiliate with the Democratic Party (56% Democrat in our study versus 43% in the ANES study)”
quantity · unclear
consistent · high confidence — Matches the paper's own Table 12/14 demographic comparison tables shown later in the source; an honestly reported sampling limitation (crowdworker recruitment skews younger and more Democratic than the ANES benchmark of Twitter users).
To check: Table 12 (Political Party) and Table 14 (Age Group) comparisons to ANES 2020.
- 36¶
The analysed dataset is 1,730 timeline-pair responses from 806 participants after attention-check exclusions.
“Our final data set consisted of 1730 responses from 806 unique participants.”
quantity · unclear
consistent · high confidence — Matches Table 2's attrition table exactly (total column: 1730 passed attention checks, from Table 1's 806 unique participants).
To check: Table 2's attrition totals and Table 1's unique-participant count.
- 37¶
Binarised, anger appears in 62% of engagement-timeline political tweets versus 52% chronologically.
“In particular, 62 percent of political tweets in the engagement timeline expressed anger, compared to 52 percent in the chronological timeline.”
quantity · unclear
consistent · high confidence — Matches Table 24's binarized political-tweet anger figures (62.312% engagement vs. 52.151% chronological, for author-expressed anger).
To check: Table 24, Binarized Emotions (Political), Author Angry row.
- 38¶
Unwanted political tweets are markedly more common in the engagement feed, 22% versus 16%.
“However, when we restrict to only political tweets, 22 percent of tweets in the engagement timeline are unwanted while only 16 percent of tweets in the chronological timeline are unwanted.”
quantity · unclear
consistent · high confidence — Matches Table 35 (No/'unwanted' responses: 21.96% engagement vs. 16.21% chronological for political tweets).
To check: Table 35, Reader stated preference restricted to political tweets.
- 39¶
For left-leaning users, engagement ranking doubles exposure to right-leaning political tweets.
“The engagement-based algorithm doubles the amount of cross-cutting content that left-leaning users see (16 percent compared to 8 percent in the chronological timeline).”
quantity · unclear
consistent · high confidence · novel — Consistent with Table 28's leaning distribution for left-leaning users (Right+Far right: ~8.4% chronological vs. ~15.7% engagement) — a specific and somewhat counterintuitive finding (engagement ranking increasing cross-cutting exposure for left-leaning users specifically) that adds useful nuance beyond the paper's headline polarization story.
To check: Table 28's Right/Far-right percentages for left-leaning users across timelines.
- 40¶
Musk's account was amplified far more than any other in the engagement timeline during the study window.
“Elon Musk also stands out as a significant outlier, receiving higher amplification than any other account by a wide margin.”
assertion · unclear
consistent · high confidence — Matches Table 18 directly (Musk's engagement-minus-chronological difference of 1236.0, vastly larger than the next account, @RonFilipkowski at 144.0).
To check: Table 18's Most Amplified column.
- 41¶
The audit data are consistent with reports that Musk's reach was artificially boosted, despite his denial.
“Though Musk denied such claims [23] , our data seems to suggest that he did receive an abnormally high amount of amplification, relative to other accounts.”
contrarian · unclear
consistent · medium confidence — This aligns with real, contemporaneous reporting I recall: Platformer (Zoë Schiffer and Casey Newton) reported in February 2023 that Twitter engineers deployed a system to artificially boost Musk's tweet visibility after his Super Bowl tweet underperformed, which Musk publicly denied; the paper's own amplification data (claim 41) is at least directionally consistent with that reporting, though it can't distinguish an explicit code-level boost from organic algorithmic amplification of a very high-engagement account.
To check: The February 2023 Platformer report (Schiffer & Newton) versus Musk's public denial tweet, cross-referenced with the study's Table 18 amplification figures.
- 42¶
News organisations are the most de-amplified accounts, probably because posting volume advantages them only in a chronological feed.
“Consistent with prior findings [6] , the accounts that were most de-amplified were news outlets. This is likely because news outlets post more frequently than ordinary accounts.”
assertion · unclear
consistent · medium confidence — Cites Bandy & Diakopoulos (2021, CSCW), a real prior audit finding that algorithmic curation reduces media/link exposure relative to chronological feeds; the explanatory mechanism ('post more frequently') is offered as a likely but not directly tested causal account, appropriately hedged with 'likely.'
To check: Table 17/18's news-outlet posting frequency and de-amplification figures, cross-referenced with Bandy & Diakopoulos (2021).
- 43¶
The engagement algorithm does not simply privilege the largest or verified accounts.
“necessarily favor the most “popular” accounts: authors shown by the algorithm have a lower median number of followers and are less likely to be verified.”
contrarian · unclear
consistent · high confidence — Matches Table 16 exactly (median followers: 139,874 chronological vs. 92,959 engagement; verified share: 50.4% vs. 39.1%). A useful counter to the naive assumption that engagement ranking simply amplifies the biggest accounts.
To check: Table 16's Author's # followers (p50) and Is the author verified? rows.
- 44¶
Re-labelling tweet-level outcomes with GPT-4 reproduces the human-labelled effects qualitatively.
“All the results based on GPT-4 judgments are notably similar to the results based on reader judgments.”
assertion · unclear
consistent · medium confidence — Matches the direction and rough magnitude of Table 36's GPT-4-labeled ATEs compared to Table 15's human-labeled ATEs (e.g., Author Angry 0.549 vs 0.473, Partisanship 0.191 vs 0.244), though 'notably similar' somewhat understates the meaningfully larger political-anger divergence flagged in claim 46.
To check: Side-by-side comparison of Table 15 (human) and Table 36 (GPT-4) ATEs.
- 45¶
GPT-4 scores the algorithm's political selections as angrier than readers do (1.47 SD versus 0.75 SD).
“The main difference in results is that GPT-4 judges the political tweets chosen by the engagement-based algorithm to be even angrier than how humans judged the tweets to be.”
quantity · unclear
consistent · high confidence — Matches Table 36 exactly (Author Angry, political tweets: 1.4653 SD under GPT-4 vs. 0.754 SD under human readers in Table 15) — a large and correctly characterized divergence, interesting as a caution about LLM-based content labeling potentially amplifying certain effect sizes.
To check: Table 36's Author Angry (political tweets) row versus Table 15's corresponding human-labeled row.
- 46¶
Survey work finds Americans both perceive divisive amplification and normatively reject it, which fits the paper's stated-preference result.
“Rathje et al. [42] found that U.S. adults believe social media platforms amplify negative, emotionally-charged, and out-group hostile content but should not.”
assertion · unclear
consistent · medium confidence — Matches the cited paper's title and thrust as listed in the bibliography ('People Think That Social Media Platforms Do (but Should Not) Amplify Divisive Content,' Perspectives on Psychological Science, 2024) by the same research group behind the well-known 2021 PNAS finding that out-group animosity drives engagement; consistent with my knowledge of that research line.
To check: Rathje, Robertson, Brady & Van Bavel (2024), Perspectives on Psychological Science 19(5).
- 47¶
The YouTube radicalisation-pipeline hypothesis was substantially undercut by evidence that users arrive via their own links and subscriptions.
“However, subsequent studies revealed that users often reach extreme channels through external links and subscriptions, suggesting that intentional human choice and demand might be an alternative explanation for the consumption of extreme content [17, 43, 31, 32] .”
assertion · unclear
consistent · high confidence — This matches a well-established correction in the YouTube radicalization literature: Hosseinmardi et al. (2021, PNAS), Chen, Nyhan, Reifler, Robertson & Wilson (2023, Science Advances), Ribeiro, Veselovsky & West (2023), and Hosseinmardi et al. (2024, PNAS) collectively pushed back on the earlier 'rabbit hole' narrative (Tufekci 2018) by showing much extreme-content consumption arrives via external links/subscriptions rather than in-platform recommendations — this is now a fairly well-known corrective in the field.
To check: The four cited papers' findings on external referral versus in-platform recommendation as pathways to extreme YouTube content.
- 48¶
Agan et al.'s thesis: ranking trained on impulsive behaviour reproduces System 1 cognitive biases.
“Similarly, Agan et al. [1] suggest that since cognitive biases are more likely to be triggered during fast, intuitive thinking (“system 1”), algorithms trained on more passive or impulsive behaviors—like time spent watching a video or lingering on a tweet—are likely to replicate these biases.”
assertion · unclear
plausible · low confidence — Accurately attributed to Agan, Davenport, Ludwig & Mullainathan's NBER working paper 'Automating Automaticity'; I recall this paper exists and engages with dual-process framing of algorithmic bias, but I have lower confidence in the fine details of its Facebook race/religion in-group findings, which this paper later contrasts its own political in-group results against (claim 28).
To check: Agan, Davenport, Ludwig & Mullainathan, NBER working paper 'Automating Automaticity: How the Context of Human Choice Affects the Extent of Algorithmic Bias' (2023).
- 49¶
The authors report that GPT-4 produced validly formatted labels for roughly 88% of the tweets submitted to it (24,998 of 28,301).
“Out of the 28,301 unique tweets that we submitted to GPT-4 to label, it returned validly formatted responses for 24,998 tweets.”
quantity · unclear
plausible · medium confidence — An ~88% valid-JSON-format rate for GPT-4 (circa 2023, without a strict JSON mode) on a bespoke multi-question prompt is consistent with known GPT-4 behavior of the era, where malformed or partial JSON responses were common before OpenAI's function-calling/JSON-mode features matured. I cannot verify the exact count since it comes from the authors' own unpublished labelling run, but the failure rate is unsurprising.
To check: The authors' released code/data (if public) showing the raw GPT-4 responses and parse failures for the 28,301-tweet batch.
- 50¶
The authors claim that restricting the comparison to tweets with valid GPT-4 responses makes the human-vs-GPT-4 ATE comparison apples-to-apples.
“This was done to ensure that any observed difference in ATEs between the human and GPT-4 labels was not influenced by using different sets of tweets for each.”
assertion · unclear
consistent · high confidence — This is standard methodological hygiene: holding the sample fixed across two labeling conditions is the correct way to isolate the effect of the labeling method itself rather than sample composition. It's a sound experimental-design principle, not an empirical claim to verify.
To check: N/A — this is a logical/methodological claim, verified by inspecting whether the same tweet IDs were in fact used for both ATE computations in the released code.
- 51¶
A stated-preference ranking lowers partisan animosity compared with Twitter's engagement-based timeline.
“Ranking by stated preferences (SP) reduced the amount of partisan animosity, relative to the engagement-based timeline.”
assertion · unclear
unverifiable · low confidence — This is a proprietary randomized-feed experiment I have no independent access to. It is directionally consistent with the recommender-systems literature showing engagement-optimized ranking tends to surface more emotionally activating/divisive content than preference- or relevance-based ranking (e.g., work on 'meaningful social interactions' metrics and outrage amplification), but I cannot confirm the specific magnitude or significance reported here.
To check: Table 37's reported p-values and standardized effects for partisanship/out-group animosity, cross-checked against the paper's public replication data if released.
- 52¶
The animosity reduction under the SP timeline came mostly from filtering out attacks on the reader's own side, not attacks on the other side.
“However, this primarily occurred through a reduction in animosity towards the reader’s in-group, rather than animosity towards the reader’s out-group.”
contrarian · unclear
unverifiable · low confidence · novel — This in-group/out-group asymmetry is a specific, somewhat counterintuitive empirical decomposition from the authors' own data that I cannot independently confirm. It is an interesting framing — that people tolerate hostility toward the outgroup but object to hostility toward their own side — which is a plausible extension of known findings (e.g., Rathje, Van Bavel & van der Linden 2021, PNAS, showing out-group animosity strongly predicts engagement) but I have not seen this exact in-group/out-group decomposition reported elsewhere.
To check: Figure 51's breakdown of animosity by target (in-group vs out-group) across the three timelines.
- 53¶
Optimising for stated preferences maximised in-group content and minimised out-group content relative to the engagement and chronological baselines.
“The SP timeline had the highest proportion of in-group content and the lowest proportion of out-group content (relative to both the engagement and chronological timeline).”
assertion · unclear
unverifiable · low confidence — A specific empirical distributional claim from the authors' proprietary dataset; I cannot verify the percentages, though it is a coherent consequence of ranking by stated preference (users tend to prefer in-group-aligned content), consistent with known selective-exposure / echo-chamber literature.
To check: Figure 51's left panel, showing the distribution of political tweets by in-group/out-group/moderate across timelines.
- 54¶
The SP-OA ranking applies a fixed 0.5-point penalty to tweets labelled as containing out-group animosity.
“In the SP-OA timeline, if a tweet was labeled as having out-group animosity, then we bumped its score down by $0.5$ points.”
quantity · unclear
consistent · high confidence — This is simply the authors describing their own ranking construction, directly stated in the source text with no reason to doubt it as a factual description of the method (as opposed to its effects, which are empirical).
To check: The paper's methods code/replication package implementing the scoring formula.
- 55¶
Because the animosity penalty is smaller than the gap between stated-preference scores, SP-OA cannot lose preference satisfaction relative to SP.
“Since the presence of out-group animosity is effectively only used to break ties among tweets with the same stated preference, the SP-OA timeline will satisfy users’ preferences to the same extent that the regular SP timeline does.”
assertion · unclear
consistent · high confidence — Given the stated scoring scheme (preference scores of 1/0/-1, i.e., integer gaps of at least 1, and an animosity penalty of only 0.5), the penalty can never flip the ranking between two tweets with different stated-preference scores — this is a straightforward arithmetic consequence, not something requiring external verification.
To check: Re-deriving the ranking algebraically from the stated scoring rule (1/0/-1 preference scale, 0.5 penalty) confirms the penalty cannot cross a full preference-score gap.
- 56¶
SP-OA produced less angry, partisan and out-group-hostile content than all three comparison timelines.
“The SP-OA yielded the lowest instances of angry, partisan, and out-group hostile content when stacked against the chronological, engagement, and SP timelines.”
assertion · unclear
unverifiable · low confidence — A comparative empirical claim resting on the authors' own bootstrap ATE estimates (Table 38, Figure 52) that I cannot independently confirm, though it is internally consistent with the negative standardized effects for Author Angry (-0.058, political: -0.275) and Out-group Animosity (-0.252) reported in the same table excerpt.
To check: Table 38 and Figure 52's full set of standardized effects compared across all four timelines simultaneously.
- 57¶
Out-group-directed animosity in political tweets fell from 34%/33% under engagement/SP ranking to 17% under SP-OA.
“In the engagement timeline and SP timeline, respectively 34 percent and 33 percent of political tweets contained animosity towards the users’ out-group. In contrast, the SP-OA timeline halved this proportion: only 17 percent contained animosity towards the users’ out-group.”
quantity · unclear
unverifiable · low confidence — A specific percentage from a figure not included in the excerpted tables (Figure 51), taken from proprietary experimental data I cannot check directly, though the 'roughly halved' framing is arithmetically accurate given the stated numbers (33 to 17 is a reduction of ~48%, close to half).
To check: Figure 51's right panel showing the percentage of political tweets with out-group-directed animosity by timeline.
- 58¶
Table 38 reports a standardized treatment effect of -0.252 on out-group animosity for the SP-OA timeline, with p = 0.0002.
“Out-group Animosity: -0.252 · -0.033 · 0.085 · 0.054 · 0.0002”
quantity · unclear
consistent · high confidence — This is a direct, verbatim transcription of the row from Table 38 as it appears in the source text — it matches exactly, so as a transcription/reporting claim it checks out; I cannot independently verify the underlying statistical computation itself.
To check: Cross-reference against the original Table 38 in the source PDF/arxiv paper.
- 59¶
The authors conclude that a tie-breaking penalty on out-group animosity achieves preference satisfaction, less divisive amplification, and no in-group reinforcement simultaneously.
“In conclusion, the SP-OA timeline has high satisfaction of users’ stated preferences, mitigates the amplification of divisive content, and does so without reinforcing in-group bias.”
assertion · unclear
plausible · low confidence · novel — This is the authors' own summary conclusion about their proposed intervention, and it is internally consistent with the reported statistics (e.g., Table 38's non-significant In-group Perception effect of -0.004, p=0.83, supports 'without reinforcing in-group bias'; the large Reader Pref effects of ~1.1 standardized support 'high satisfaction'). However, it is the researchers evaluating their own newly-invented ranking scheme, so the favorable framing deserves scrutiny — 'high satisfaction' and 'mitigates amplification' are qualitative judgments layered on quantitative effects whose real-world magnitude/durability I cannot assess.
To check: Whether the SP-OA timeline's preference-satisfaction, animosity, and in-group-bias effects (Table 38) all point the claimed direction and remain robust to multiple-testing correction and out-of-sample replication.
- 60¶
The 0.5-point animosity penalty is by construction never large enough to override a user's stated preference ordering.
“Note that even if a tweet has out-group animosity, if the user stated that they valued the tweet, it will always be ranked higher than a tweet that they are indifferent to or do not value.”
assertion · unclear
consistent · high confidence — Same arithmetic as claim 6: a 0.5-point penalty applied on a 1/0/-1 preference scale cannot overturn the ordering between tweets with different stated-preference scores, since the minimum gap between preference tiers (1 point) exceeds the penalty (0.5 points). This is a valid deduction from the stated rule.
To check: Algebraic verification: for any preference scores p1 > p2 (integers from {1,0,-1}), p1 - 0.5 > p2 is guaranteed only if p1 - p2 ≥ 1, which holds for all distinct integer preference tiers.
The engagement algorithm amplifies emotional, divisive content beyond what users' own follow choices or tweet-level stated preferences would explain.stands on 3 consistent steps · weakest link: 1 plausible inference
premise · consistent — Methodological premise: using each user's own reverse-chronological feed as control holds follow choices fixed. · claim 7
evidence · consistent — Engagement ranking raised tweet partisanship and out-group animosity each by about a quarter of a standard deviation. · claim 14
evidence · consistent — Across all tweets, the algorithm amplified expressed anger most strongly, with smaller effects for sadness and anxiety. · claim 16
inference · plausible — Misalignment with stated preference rules out deliberate user choice as a full explanation of the algorithm's content effects. · claim 22
conclusion · consistent — The audit establishes amplification of emotional content over and above both follow choices and tweet-level stated preferences. · claim 23
The engagement algorithm optimises revealed preference while failing stated preference, especially on political content.stands on 2 consistent evidence · weakest link: 1 plausible premise
premise · plausible — Internal randomised Twitter experiments show the engagement timeline raises time spent versus reverse-chronological. · claim 9
evidence · consistent — For political tweets specifically, users valued the algorithm's picks significantly less than the chronological ones. · claim 19
evidence · consistent — Across all tweets, stated preference for engagement-timeline tweets is only marginally higher. · claim 18
conclusion · consistent — The algorithm satisfies revealed preference (time spent) while failing stated preference, most sharply on political content. · claim 21
Ranking by stated preferences depolarises content mainly by removing the out-group, so it may increase pro-attitudinal exposure.stands on 2 consistent steps
premise · consistent — The constructed stated-preference timeline lowered negative emotion and raised happiness relative to the engagement timeline, for both authors and readers. · claim 20
evidence · consistent — The stated-preference timeline's apparent depolarisation comes from filtering out-group voices and hostility aimed at the reader's own side, not hostility generally.
conclusion · consistent — Hedged conclusion: stated-preference ranking may increase pro-attitudinal exposure compared with engagement ranking. · claim 27
Engagement ranking selects more polarising content than the user's follow graph alone would yield.stands on 2 consistent evidence
evidence · consistent — Engagement-selected tweets moved readers' feelings against their out-group and mildly toward their in-group. · claim 15
evidence · consistent — Engagement ranking raised tweet partisanship and out-group animosity each by about a quarter of a standard deviation. · claim 14
conclusion · consistent — Engagement ranking is more polarising than the user's own follow graph would produce. · claim 24
The assumption that retention-optimised ranking should make readers feel better was the one prediction the data refuted.stands on 1 consistent premise · weakest link: 1 unverifiable inference
premise · consistent — The authors assumed retention-optimised ranking would favour positive emotion in readers, since happy users return. · claim 10
inference · unverifiable — Pre-registered Hypothesis 5 predicted readers would feel happier and less negative on the engagement-based timeline. · claim 11
conclusion · consistent — Hypothesis 5 was the single pre-registered prediction the data contradicted; readers reported more of all four emotions, happiness included. · claim 12
Because correlational evidence cannot separate user demand from algorithmic amplification, a paired audit with stated-preference measurement is needed.stands on 1 consistent evidence · weakest link: 2 plausible steps
premise · plausible — The observed emotion–engagement correlation in prior work is confounded by users' own choice of accounts to follow. · claim 6
evidence · consistent — The YouTube radicalisation-pipeline hypothesis was substantially undercut by evidence that users arrive via their own links and subscriptions. · claim 47
inference · plausible — Even a reverse-chronological control is insufficient, because the algorithm could be serving fine-grained preferences users cannot express through follows. · claim 8
- ¶
conclusion · ungraded — “To provide a stronger grounding and justification to these methods, it is crucial to empirically test this hypothesis, which is the aim of our work.”
Political in-group preference is not purely an automatic System 1 bias; it survives reflective elicitation.stands on 1 consistent evidence · weakest link: 1 plausible premise
premise · plausible — Agan et al.'s thesis: ranking trained on impulsive behaviour reproduces System 1 cognitive biases. · claim 48
evidence · consistent — The stated-preference timeline's apparent depolarisation comes from filtering out-group voices and hostility aimed at the reader's own side, not hostility generally.
conclusion · plausible — Contrary to Agan et al.'s Facebook result, political in-group preference appears partly deliberate rather than purely automatic. · claim 28
Adding an out-group-animosity tie-break to a stated-preference ranking removes the in-group bias that pure preference ranking introduces, without costing preference satisfaction.stands on 2 consistent steps · weakest link: 4 unverifiable steps
premise · unverifiable — A stated-preference ranking lowers partisan animosity compared with Twitter's engagement-based timeline. · claim 51
premise · unverifiable — The animosity reduction under the SP timeline came mostly from filtering out attacks on the reader's own side, not attacks on the other side. · claim 52
evidence · unverifiable — Optimising for stated preferences maximised in-group content and minimised out-group content relative to the engagement and chronological baselines. · claim 53
- ¶
inference · ungraded — “Is it possible to satisfy users’ stated preferences without inducing in-group bias? To investigate this, we considered a variant of the SP timeline that used the presence of out-group animosity to break ties between tweets.”
premise · consistent — The SP-OA ranking applies a fixed 0.5-point penalty to tweets labelled as containing out-group animosity. · claim 54
inference · consistent — Because the animosity penalty is smaller than the gap between stated-preference scores, SP-OA cannot lose preference satisfaction relative to SP. · claim 55
evidence · unverifiable — Out-group-directed animosity in political tweets fell from 34%/33% under engagement/SP ranking to 17% under SP-OA. · claim 57
conclusion · plausible — The authors conclude that a tie-breaking penalty on out-group animosity achieves preference satisfaction, less divisive amplification, and no in-group reinforcement simultaneously. · claim 59
Restricting to tweets with valid GPT-4 output makes the human-label and GPT-4-label treatment effects directly comparable.stands on 1 plausible premise
premise · plausible — The authors report that GPT-4 produced validly formatted labels for roughly 88% of the tweets submitted to it (24,998 of 28,301). · claim 49
- ¶
inference · ungraded — “To estimate the average treatment effects displayed in Figure 18, we used only the tweets with valid GPT-4 responses (for both the ATEs based on human labels and those based on GPT-4 labels).”
conclusion · consistent — The authors claim that restricting the comparison to tweets with valid GPT-4 responses makes the human-vs-GPT-4 ATE comparison apples-to-apples. · claim 50