Does Reddit display brain-like behaviours? Can we “Frankenstein” Reddit?

Mapping the Collective Reddit Brain

Reddit becomes a testable brain metaphor: subreddit categories as regions, hyperlinks as wiring, and synchronized language as functional connectivity.

Executive summary

A social network analyzed like a connectome

This project uses connectomics as a modelling lens, not a biological equivalence claim. The vocabulary is useful: structure, function, hubs, lesions, excitation, inhibition, and reorganization.

The output is a deployable data story: network science, NLP, sentiment, temporal correlation, and interactive visual evidence in one place.

Why it matters

  • For social computing: it exposes hidden synchrony between communities that may never link directly.
  • For network science: it compares structural edges, functional correlations, sentiment circuits, and lesion response.
  • For product thinking: it turns a notebook-heavy analysis into a recruiter-readable interactive case study.

38,883

category-matched hyperlinks

directed subreddit-to-subreddit links

1,900

unique subreddits

after category filtering

19

semantic regions

high-level subreddit categories

86

language features

LIWC, sentiment, readability, structure

69

continuous signals

subreddits retained for time-series analysis

2014–2017

time horizon

monthly functional-connectivity window

Interactive data anatomy

Category-matched links: The filtered directed edge set used for category-level structural and sentiment analyses.

Evidence card

Structure and function diverge

Some communities co-activate over time even when they do not directly hyperlink, exposing hidden functional synchrony.

Evidence

Functional-only links remain after subtracting explicit hyperlinks from the functional connectome.

Open Functional Mapping

Method pipeline

Step 1

Raw Reddit hyperlinks

Step 2

Category atlas

Step 3

Structural graph

Step 4

Sentiment layers

Step 5

Monthly LIWC signals

Step 6

Functional connectome

Step 7

Brain-space projection

Introduction

Zoom out from Reddit and individual posts blur into communities exchanging signals. Some links are explicit wiring. Other relationships only appear when communities rise and fall together over time.

The analogy is compact: subreddits are neurons, hyperlinks are structural connections, and correlated LIWC signals are functional connections. The site asks whether that abstraction reveals patterns that are hard to see in the raw hyperlink graph.

To make the map interpretable, we use semantic categories as an atlas. They do not cover every subreddit, but they provide stable landmarks for studying structure, sentiment, lesions, and rhythm.

Here, a category means a high-level label (e.g., “Technology”, “Sports & Games”, “News & Politics”) assigned to a subreddit for analysis.

Our analogy (simple, but useful)

  • 1 neuron = 1 subreddit
  • Structural connection = hyperlink
  • Functional connection = correlated LIWC feature activity
  • 1 subreddit hyperlink = 1 action potential firing between neurons

Research Questions

Question 1

Do subreddit categories exhibit distinct functional roles (sensory / interneuron / motor) based on in-degree and out-degree?

Question 2

Can we define a functional network of subreddits beyond explicit structural hyperlinks (via temporal co-activation)?

Question 3

To what extent does functional connectivity transcend structural wiring, and what characterizes these remotely communicating communities?

Interactives: controls

Zoom. Open plots in a larger focused view.

Slider. Drag the white knob to move through time.

Snapshot. Each slider position updates the visual state.

Filters. Toggle positive, negative, or combined edges.

Do Our Neurons Represent the Whole Brain?

The category atlas only labels part of Reddit. Before using it as a map, we test whether those labeled communities preserve the dominant behavior of the full network.

The short answer: the subset captures the core distributional structure, while sacrificing rare long-tail communities.

Open the validation details.

Listening to the Whole Network

Each hyperlink carries a linguistic fingerprint: length, structure, sentiment, and LIWC-derived psychological markers. These features describe how communities communicate, not only who connects to whom.

We compare the labeled subset with the full network through feature distributions. Their central tendencies and shapes largely overlap.

That makes the labeled communities useful landmarks: not exhaustive, but typical enough to orient the map.

What Was Left Out

The lost portion is mostly the edge of the feature space: rare styles, niche communities, and low-frequency behavior. The tradeoff is explicit: we gain interpretability while losing some long-tail specificity.

Projecting the Brain Into Low Dimensions

PCA is fitted on the full dataset, then both the full network and labeled subset are projected into the same state space. The subset occupies the same core regions along the leading components.

In practical terms, it preserves the main modes of variation while smoothing away idiosyncratic extremes.

Why This Matters

The labeled subset works as a valid working atlas: not complete, but representative enough to assign regions and study structure–function relationships.

Projecting Reddit onto the Brain

We project subreddit categories onto an MNI brain template through thematic clusters: groups of categories mapped to semantic concepts inspired by the Huth et al. framework.

The result is not a literal brain claim; it is an interpretable spatial scaffold for comparing communities, links, and functional co-activation.

Loading 3D brain…

The structural network is built from explicit hyperlinks: who links to whom. We ask whether categories form recognizable wiring roles: receivers, broadcasters, relays, and long-range connectors.

Figure 3 – Category in vs out degree

Categories with many incoming and few outgoing links behave like sensory regions, while those with the opposite pattern behave like motor regions. Categories close to the diagonal resemble interneurons that relay information.

Figure 4 – “Axon lengths” via embedding distance

For each source category, we measure how far it reaches in embedding space. Some categories mostly connect to near neighbors, while others regularly jump to distant topics, acting as long-range integrators.

Neuron types (framing)

The analogy is role-based: incoming links resemble sensory input, outgoing links resemble motor output, and balanced categories behave like relays.

Positive & Negative Circuits

Positive and neutral links act as supportive signals; negative links act as antagonistic signals. We use this as a simple excitation / inhibition layer over the hyperlink graph.

The question becomes: who broadcasts support, who broadcasts conflict, and which communities absorb the most negativity?

Figure 8 – Sentiment received by categories

Two of the top sources of inhibition are also top recipients, and “News & Politics” receives the most negativity overall.

Figure 9 – Sentiment sent by categories

“Funny/Memes”, “History & Culture”, and “Philosophy/Religion/Spirituality” show up as strong negativity broadcasters — inhibitory hubs in the emotional network.

Mapping Individual Relationships: Reddit Fluorescent Cell Culture

Interactive – Positive/Negative visualization

Interactive view ready on demand

Click to load the full interactive; brain maps include legend and hover focus.

Select a source category to inspect supportive vs antagonistic outgoing links.

In this visualization, we examine the axons projecting from every source category. Source categories can be selected from the menu.

Treat this as a conflict atlas: red arrows and warm tones expose where sentiment becomes antagonistic.

Embedding Distances: Mapping Axon Lengths

Embedding distance becomes a proxy for axon length: short distances connect similar audiences, while long distances bridge distant communities.

The interactive “explodes” the graph into parallel connections so distance, sentiment, and direction can be inspected category by category.

Interactive – Distance comparisons

Interactive view ready on demand

Click to load the full interactive; brain maps include legend and hover focus.

Compare semantic distance, direction, and sentiment for each category.

Use the explorer to spot long-distance links: categories that connect beyond their immediate semantic neighborhood.

Direction also matters. Some categories send many links to a target but receive very few in return, revealing asymmetric attention.

The key test is whether longer semantic jumps are associated with more negative sentiment.

For Technology, positivity drops strongly as embedding distance increases, suggesting that distant audiences are more likely to receive antagonistic links.

Positive and Negative Hubs

Finally, we test hub dependence: does the network survive random damage, or collapse when the biggest hubs are removed first?

The grey curve removes nodes randomly. The colored curve removes highest-degree nodes first.

A steeper colored drop means the layer depends on a few hubs.

Interactive – Positive hub visualization

Interactive view ready on demand

Click to load the full interactive; brain maps include legend and hover focus.

Positive links degrade gradually, indicating a more distributed layer.

Interactive – Negative hub visualization

Interactive view ready on demand

Click to load the full interactive; brain maps include legend and hover focus.

Negative links collapse faster when the largest hubs are removed first.

In the positive network, the largest hubs are Business, Finance and Crypto, Funny/Memes, and Technology. Being a positive hub has nothing to do with sentiment proportion; rather, through sheer volume of interactions, these categories keep the positive network connected.

In the negative network, Funny/Memes, Technology, and News and Politics dominate. The red curve drops sharply: negativity is more hub-dependent than positivity.

Takeaway: large-scale negativity is not evenly distributed; it is amplified by a small number of influential hubs.

Lesions

On June 10, 2015, Reddit banned five large subreddits. We treat this as a network lesion: a sudden removal of major hubsfollowed by measurable shifts in platform-wide language.

The finding is not just local removal. After the ban, aggregate discourse becomes more negative, more certain, more out-group focused, and more linguistically complex.

Daily aggregation controls for changing post volume and lets us compare affect, cognition, social orientation, and discourse style before vs after the event.

Open statistical details.

MANOVA tests whether the joint psychological profile changes after the ban, rather than testing each feature in isolation.

The MANOVA revealed a significant systemic change in the network’s psychological profile following the ban (Pillai’s trace = 0.18, F(6, 1210) = 45.31, p < .001).

Welch tests then identify which features move and in what direction.

FeatureMean Change (After − Before)p-valueDirectionInterpretationCollective Meaning
VADER_Compound−0.0159< 1e-17Overall sentiment became more negativeIncreased emotional negativity
LIWC_Negemo+0.00063< 1e-10More negative emotional languageHeightened emotional intensity
LIWC_Anger+0.00039< 1e-7Increased expressions of angerGrievance and hostility
LIWC_Anx+0.000080.0013More anxiety-related languagePerceived threat / uncertainty
LIWC_CogMech−0.00239< 1e-20Reduced cognitive processing languageLess deliberative reasoning
LIWC_Insight−0.00116< 1e-35Fewer reflective / insight termsReduced self-reflection
LIWC_Tentat−0.00093< 1e-28Less tentative languageDecreased nuance / openness
LIWC_Certain+0.00021< 1e-4More certainty and absolutist framingCognitive rigidity
LIWC_They+0.00026< 1e-13Increased out-group referencesStronger us-vs-them framing
LIWC_You−0.00019< 1e-4Fewer direct addressesReduced interpersonal engagement
LIWC_We−0.000030.58 (ns)No change in inclusive languageNo increase in group solidarity
Readability_index+0.48< 1e-5Higher linguistic complexityMore argumentative discourse
Avg_words_per_sentence+0.84< 1e-4Longer sentencesIncreased rhetorical elaboration

Sentiment and affect. Overall sentiment became more negative, with decreases in VADER Compound scores and increases in negative emotion and anger, indicating a more emotionally charged discourse.

Cognitive style. Insight, cognitive processing, and tentative language decreased, while certainty increased, reflecting reduced deliberation and greater cognitive rigidity.

Social orientation. References to out-groups (They) increased, direct address (You) decreased, and inclusive language (We) showed no significant change, suggesting a shift toward adversarial framing without increased group cohesion.

Discourse complexity. Readability and average sentence length increased, indicating more elaborated and argumentative posts.

Collectively, these patterns suggest that the ban functioned as a coordinated disruption in the network’s social brain.

Conclusion

Removing major hubs was associated with systemic changes in emotional, cognitive, and social language patterns.

Causality remains bounded: the timing is consistent with a network-level disruption, but concurrent external factors cannot be fully excluded.

Functional Mapping

Structural connectivity asks who links to whom. Functional connectivity asks which communities move together over time. The key question is whether Reddit contains synchrony beyond explicit hyperlinks.

Interactive – Functional connectivity difference

Interactive view ready on demand

Click to load the full interactive; brain maps include legend and hover focus.

Turning links into neural signals

A hyperlink is only the wiring. The signal lives in the post behind it: sentiment, LIWC features, timing, and linguistic style.

We retain subreddits that fire continuously enough to support time-series analysis, then aggregate activity into monthly bins.

The selection balances two constraints:

  • Completeness: avoid empty temporal bins.
  • Eligibility: keep enough subreddits for network analysis.

We then choose the smallest time bin that preserves the most continuous signals.

CONNECTIVITY FILTER – MOST SYNAPTICALLY ACTIVE SUBREDDITS

Bin selection – continuity vs coverage

We evaluate multiple bin sizes, balancing continuity (few gaps) and how many subreddits remain analyzable.

We rank subreddits by total hyperlink volume (incoming + outgoing) and keep the most synaptically active communities, ensuring our time-series analysis focuses on neurons that keep firing.

Result

The winner is 1 month: stable, continuous, and still granular.

The final issue is dimensionality: 86 features per subreddit are too fragmented for a readable functional map.

We compress related LIWC variables into six functional signals. These become the activation traces used to build the functional connectome.

Interactive – Cluster time-series explorer

Interactive view ready on demand

Click to load the full interactive; brain maps include legend and hover focus.

Compare the two monthly signals; use the selector to switch functional domain.

Interactive – Functional connectivity

Interactive view ready on demand

Click to load the full interactive; brain maps include legend and hover focus.

Hover nodes to inspect subreddit groups and simplify links when the map is dense.

The punchline: wiring vs rhythm

We compute temporal correlations across the six functional signals and ask: who moves together even without direct links?

From subreddit topology to brain space

Open the extended interpretation.

The comparison separates two views: all significant functional links, and functional minus structural connectivity. The second view isolates communities that synchronize without directly linking.

This answers the central question: is Reddit only explicit interaction, or does it also contain hidden shared rhythm?

  • Discovering collective resonance Functional-only links show distant subreddits reacting to shared events without direct communication.
  • Identifying hidden functional networks Some communities behave alike even when the hyperlink graph does not connect them.
  • Testing network resilience If rhythm survives beyond wiring, the system is less dependent on direct links.

The punchline: subtracting hyperlinks reveals who moves together anyway.

Data & Methods

We combine multiple public datasets and several preprocessing steps to construct the Reddit brain. Below is a high-level overview of the pipeline that underlies all the visualizations above.

Important framing

This is a computational analogy, not a biological claim. Brain terms are used to structure the analysis: they help compare wiring, co-activation, hubs, and perturbations in a large social network.

  1. Hyperlinks. Concatenate title and body hyperlink datasets into a single set of edges between source and target subreddits.
  2. Category mapping. Merge in subreddit categories from an external taxonomy; keep only subreddits with a dominant category and discard ambiguous cases.
  3. Filtered structural network. Retain only edges where both endpoints have valid categories, obtaining the core structural graph used in our analyses.
  4. Subreddit embeddings. Load 300-dimensional subreddit embeddings and compute cosine distances for each hyperlink, interpreting them as “axon lengths”.
  5. Time-series construction. Bin activity monthly and select subreddits whose time series are continuous enough for analysis.
  6. Functional networks. Aggregate LIWC-derived features into functional domains (Affect, Social, Cognition, Somatic, Spatial, Time) and compute correlations.
  7. Sentiment layer. Use sentiment labels on hyperlinks to build positive and negative circuits between categories.

All figures on this page are generated from our analysis notebooks and exported as static assets for this front-end.

What I built

  • Designed the Reddit-to-brain modelling framework and converted it into an interactive data story.
  • Built the reproducible preprocessing path from hyperlink data to structural, sentiment, lesion, and functional-connectivity views.
  • Implemented the Next.js front-end, light/dark theming, lazy-loaded figures, zoomable panels, and custom connectome comparison.
  • Optimized the static GitHub Pages deployment so large scientific assets load progressively instead of blocking the first view.

Limitations

  • Coverage: only subreddits that can be matched to interpretable categories enter the category-level atlas.
  • Causality: functional links are statistical co-activation, not proof of information transfer.
  • Analogy: neuroscience vocabulary is a modelling lens; Reddit and biological brains differ in mechanism and scale.

Continue exploring

From notebooks to a deployable research product

The repository contains the notebooks, preprocessing scripts, static assets, regression tests, and GitHub Pages workflow behind this site.