ANAGNORISIS
Stratum 3 · Paper 1.3 rev. 2 · Full text
Authors: Patrick Grünig · Claude Fable (Anthropic)
Status: Paper 1.3 of ten — the series’ individual-level neural paper; reads independently of the series’ framework
Note on authorship. This paper is a human–AI collaboration, and the byline says so plainly. The division of labor: the thesis, the source corpus, the conceptual arc, and every decision of substance are Patrick Grünig’s; drafting, reference verification, and adversarial revision are Claude’s — one collaborator instantiated across model generations (first draft: Claude Opus 4.6, February 2026; charter-governed revision: Claude Fable 5, with Claude Opus 5 verification agents, August 2026). Accountability for the work and custody of it are human, and Grünig’s. Where a venue’s policy does not admit machine co-authorship, this byline converts to an acknowledgment without loss: the note records facts, not a claim to legal personhood. One standing rule keeps the collaboration honest, stated in full in Paper 5.1 (§9.1): the AI co-author’s fluent agreement with the thesis is never evidence for it.
Note on epistemic status. This paper is a narrative review joined to a model proposal, and it holds its verbs to that register. What the reviewed studies show is reported in their own terms, with sample sizes stated and replication status reported wherever a replication attempt is known to exist; what the model proposes is marked as proposal; what would decide it is stated as a prediction. The words established, demonstrated, validated, and proven do not appear below as this paper’s own verdicts on its own claims. Every external reference carries a verification record in the project archive; the References section states each record’s level.
Note on this revision (removable at publication). This is the reworked successor of the February 2026 draft, produced under the project’s method charter after a three-phase review. Relative to that draft: a conflated lesion citation is unwelded into the three real studies it merged, and the theory one of them was testing is now cited by name; a review citation with fabricated journal metadata is replaced by the real joint publication, with the network model it imports credited to its originator; one confabulated reference is deleted; the central fMRI anchor’s three analyses are reported as three, the source’s own text having denied the merged rendering; the promise to report evidence quality and replication status, made and then broken, is kept — sample sizes inline, replication attempts reported per finding, and a dedicated evidence-quality section added; a preregistered 2026 replication and a 2024 lesion-network study are added, and they change the model’s evidence profile enough that the review now grades its four systems separately; the identity-fusion, doubt, and self-model literatures the review’s own topic required are engaged, with the title construct’s lineage acknowledged; and domain generality is restated at its evidence’s scope. The body makes no reference to prior drafts; the change record lives in the project archive.
Deeply held beliefs (religious, political, ideological) resist revision in ways casual beliefs do not, and the resistance is poorly predicted by evidence quality. This narrative review synthesizes the neuroscience of that resistance and proposes the Identity-Belief Fusion Model (IBFM): a conviction entrenches when it is simultaneously encoded into the self-representational system (default mode network), shielded from evaluative revision (prefrontal doubt and flexibility circuits, and evidence-exempt encoding formats), reinforced by reward circuitry, and defended by threat-processing systems. We report sample sizes and known replication attempts for every anchor finding, and grade the model’s four systems by them. The grades diverge. The self-referential system is the best supported: challenges to identity-relevant beliefs engage default-mode structures, a result with a preregistered cross-cultural replication. The shielding system is next: lesion evidence links prefrontal integrity to belief flexibility, culminating in a two-cohort lesion-network study (n = 190) whose network map reproduced across cohorts at r = 0.82. The reward system rests on suggestive but unreplicated findings, and the threat system’s best-known results have failed direct replication in physiological, structural (in part), and fMRI forms. The model therefore stands as an integrative hypothesis whose strongest limbs are self and shielding, a conclusion at odds with the threat-centric picture popular accounts favor. We relate the model to the identity-fusion research program, derive implications for radicalization, deconversion, and therapy, state one novel prediction (reduced dynamic default-mode flexibility in entrenched believers), and specify the evidence the model still lacks.
Keywords: belief entrenchment, identity fusion, default mode network, prefrontal cortex, cognitive flexibility, religious fundamentalism, replication, lesion network mapping, deconversion, motivated reasoning
Human beliefs vary enormously in their resistance to change. A person may readily update an estimate of tomorrow’s weather when given new data, yet leave their deepest religious or political convictions untouched by decades of counter-argument. What distinguishes revisable beliefs from entrenched ones?
A purely epistemic account, on which entrenched beliefs are simply held with more confidence or on more evidence, fits the data poorly. The classic statement of the alternative is Abelson’s (1986): some beliefs function less like maps than like possessions, owned, displayed, and defended as part of the self. The empirical form of that alternative appears where the epistemic account is weakest. In a large survey study of perceived climate-change risk, Kahan and colleagues (2012) found that polarization was greatest among the respondents highest in science literacy and numeracy — the opposite of what a knowledge-deficit account predicts. The mechanism their data suggest is a conflict of interest: individuals have a personal stake in holding beliefs aligned with those of the people closest to them, whatever the collective stakes of being right. Kahan’s framework names this identity-protective cognition. The study itself concerns one politicized risk domain, and its transfer to religious conviction is an analogy; but the shape of the result recurs across domains: what predicts resistance to revision is not how well a belief is supported but how deeply it is woven into who the believer takes themselves to be.
This paper reviews the neuroscience of that weaving. We use belief entrenchment for the process by which a conviction becomes integrated into the neural systems that maintain the self, acquiring a resistance to revision that is mediated by self-representational, motivational, and defensive systems rather than by evidence evaluation alone.
We review evidence across four neural systems:
The review is narrative rather than systematic; we prioritize theoretical integration over exhaustive coverage. Two disciplines govern it throughout. First, for every anchor finding we report the sample size and, where any replication attempt is known, its outcome; this literature is one of small samples, and a reader should not have to look up how small. Second, the field’s own calibration is applied to the paper’s own evidence: the median neuroimaging study samples about 25 participants, and brain-wide association effects measured at such sizes are systematically inflated (Marek et al., 2022; Button et al., 2013). Section 8.6 places every anchor used here against those benchmarks. The result of taking both disciplines seriously is one of this review’s main findings: the four systems above are not equally supported, and the model built on them (§6) is graded system by system.
This review is the individual-level paper of a ten-part series; two companion papers build a collective-level account on the substrate reviewed here (Papers 1.1 and 1.2). The review stands on its own terms, and nothing below depends on the series’ framework.
Three existing research programs own parts of this paper’s territory, and naming them at the outset is both an attribution duty and a resource.
Identity fusion. The title’s phrase “identity-fused conviction” deliberately echoes a defined construct. In the identity-fusion research program, fusion is “a visceral feeling of oneness with the group,” in which the borders between personal and social self become porous while the personal self remains potent (Swann, Gómez, Seyle, Morales & Huici, 2009; Gómez et al., 2011; Swann, Jetten, Gómez, Whitehouse & Bastian, 2012). Fusion is measured, by a pictorial scale (Swann et al., 2009) and a validated verbal scale (Gómez et al., 2011), and it is among the strongest known predictors of extreme pro-group behavior, including endorsement of fighting and dying for the group. That program concerns the fusion of self with a group. The model developed here concerns the fusion of self with a belief — a related but distinct claim, and one the fusion literature reached first for the collective case. We state the relation precisely in §6, where the fusion scales also supply the behavioral instrument the neural model otherwise lacks.
False Tagging Theory. The proposal that the prefrontal cortex implements doubt — affective “false tags” affixed to comprehended representations, without which comprehension defaults to belief — is Asp and Tranel’s False Tagging Theory (Asp & Tranel, 2013; open statement in Asp, Manzel, Koestner, Denburg & Tranel, 2013). Section 3 reviews the theory’s lesion test and its complications.
Self-model theory. The phenomenological difference between “I believe X” and “X is simply how things are” has a worked-out theoretical treatment in Metzinger’s account of phenomenal transparency (Metzinger, 2003a, 2003b). Section 3.5 imports it as philosophy that names and predicts the pattern in the data reviewed here, not as neuroscientific evidence.
The most direct evidence comes from challenging beliefs in the scanner. Kaplan, Gimbel and Harris (2016) presented 40 participants — all screened to be strongly liberal — with arguments contradicting strongly held political and non-political beliefs during fMRI. The study contains three analyses, and because their fates diverge below, we report them as three.
First, the direct contrast: processing challenges to political beliefs, compared with non-political ones, increased activity in regions of the default mode network (precuneus, posterior cingulate cortex, medial prefrontal cortex, inferior parietal lobe, anterior temporal lobe) — “a set of interconnected structures associated with self-representation and disengagement from the external world,” in the authors’ description. The reverse contrast, non-political over political, engaged dorsolateral prefrontal and orbitofrontal cortices: the circuitry of deliberate evaluation was more engaged for the beliefs participants were willing to revise.
Second, an item-level analysis: across argument items, resistance to belief change tracked dorsomedial prefrontal signal positively and orbitofrontal signal negatively.
Third, an individual-differences analysis: across participants, those who updated least showed more amygdala and dorsal anterior insula signal while evaluating counterevidence (r ≈ 0.35–0.36). The authors are explicit that these emotion-linked structures appeared only in this between-subjects correlation. In their words: “we did not find these structures to be activated in general within the group while considering the challenges; nor did these structures appear in our direct comparison of political and non-political challenges.” Popular renderings of this study, on which the brain treats ideological challenge as threat, merge the three analyses into one and attach the amygdala to a contrast it does not appear in; §5.1 returns to what the third analysis can and cannot carry.
What the design contrast itself shows is the self-referential result: for these participants, challenges to identity-relevant beliefs were processed by the machinery of self-representation, not preferentially by the machinery of evaluation.
This system’s evidence has been retested. In a preregistered study in a different country and political culture, Kossowska and colleagues (2026) presented 43 strongly left-wing Polish participants with counterarguments to political and non-political beliefs. The core pattern reproduced: robust resistance to political-belief change alongside greater openness on non-political beliefs; heightened default-mode activation under political challenge, “especially in regions associated with self-referential processing”; and the dorsomedial-prefrontal/orbitofrontal resistance pattern. One original result did not reproduce: the study found no correlation between belief change and activation in insula or amygdala. The replication’s own summary places the weight where the surviving data are: belief change “is deeply rooted in neural systems responsible for maintaining self-identity.”
Two small samples (40 and 43, the second preregistered) do not settle a literature. But the self-referential result is, at present, the only finding in this review’s fMRI core that has been independently reproduced — and the reproduction crossed a language and a political culture while sampling, like the original, only one side of the ideological spectrum. The result’s generalization to conservative believers is untested.
For religious belief specifically, the self-referential system has a further entanglement: the believer’s relationship with a modeled agent. Reasoning about the contents of other minds engages a specific region of the temporo-parietal junction, dissociable from adjacent regions and unresponsive to non-social control content (Saxe & Kanwisher, 2003) — a mental-state-attribution result with no religious content, and one whose mentalizing function the network literature places within the default mode network’s own territory (McNamara & Grafman, 2024). The religious application has its own evidence: in an fMRI study of the cognitive structure of religious belief, judgments about God’s involvement and God’s emotions engaged networks processing theory of mind regarding intent and emotion (Kapogiannis, Barbey, Su, Zamboni, Krueger & Grafman, 2009). Believers contemplating what God wants are, by these data, running the circuitry evolved for modeling human minds — on an agent whose responses never disconfirm the model.
The entrenchment-relevant consequence is interpretive but direct: a neurally instantiated model of a divine mind, maintained for years and consulted daily, is not phenomenologically a proposition. It is a relationship. Deconversion, on this reading, requires more than revising a claim; it requires losing a social partner the brain has treated as real — which is one reason exit from deep belief presents clinically as grief (§7.2).
At the level of large-scale networks, a recent synthesis by McNamara and Grafman (2024), a Hypothesis and Theory article reviewing what its authors call a representative rather than exhaustive sampling of recent studies, concludes that religious and spiritual experiences “appear to depend crucially upon” interactions among three networks: the default mode network (self-referential processing), the frontoparietal executive network (cognitive control), and the salience network (significance detection and switching between the other two). The triple-network architecture itself is Menon’s (2011) general model of large-scale brain organization and psychopathology; McNamara and Grafman’s contribution is its application to religious cognition, offered explicitly as hypothesis. The frame is useful here because it assigns each of this review’s systems a place in one architecture, and because it makes the salience network’s switching role a candidate mechanism for how significance-saturated content captures both self-processing and executive resources at once.
If entrenched beliefs were merely strongly held, evidence would still reach them. The second system concerns why it often does not: the capacities whose loss, override, or circumvention leaves a belief unrevisable.
The classical lesion literature ties dorsolateral prefrontal cortex to cognitive flexibility — the capacity to shift set, consider alternatives, and inhibit prepotent responses (Milner, 1963; for the modern landscape, see Stuss & Knight, 2013). The question for this review is whether that capacity gates belief flexibility.
The largest direct test comes from the Vietnam Head Injury Study. Zhong, Cristofori, Bulbulia, Krueger and Grafman (2017) assessed religious fundamentalism in 119 male combat veterans with penetrating traumatic brain injury, roughly forty years after injury, alongside 30 demographically comparable combat veterans without brain injury. The familiar rendering of this study overstates it. In the categorical group comparison, the only significant contrast was that vmPFC-lesion patients scored higher than patients with lesions outside the prefrontal cortex; the dlPFC group did not differ significantly from any other group. The dorsolateral finding is instead a mediation result on the full patient sample: greater dlPFC lesion volume predicted higher fundamentalism indirectly, through reduced cognitive flexibility and through reduced trait openness, with vmPFC lesion volume controlled. The authors’ summary: cognitive flexibility and openness “are necessary for flexible and adaptive religious commitment,” and such diversity of religious thought, in their words, “is dependent on dlPFC functionality.” The effects are small, and the sample is a single demographic (male American veterans of one war) — a limit the authors name themselves.
A bidirectional possibility deserves statement as the hypothesis it is: if prefrontal flexibility gates belief revision, and if flexibility is itself trainable and atrophies with disuse, then long entrenchment might reduce the very capacity that could end it. Nothing in the lesion data tests this direction; longitudinal designs could (§8.1).
False Tagging Theory (§1.3) makes the complementary claim: comprehension defaults to belief, and doubt is a distinct, prefrontal — specifically ventromedial — operation. On this theory, vmPFC damage should produce a doubt deficit: an increased tendency to accept ideas as presented. Asp, Ramchandran and Tranel (2012) tested exactly this, comparing 10 patients with bilateral vmPFC lesions, 10 patients with lesions elsewhere, and 16 medical comparison patients who had survived life-threatening non-neurological events (there is no healthy-control arm). The vmPFC group reported higher authoritarianism and higher religious fundamentalism than both comparison groups, exceeded scale norms that the comparison groups matched, and reported increased specific religious beliefs after their injury. The authors read the result as supporting a vmPFC-critical doubt function.
The two lesion programs do not agree on localization. False Tagging Theory places doubt in vmPFC and finds support in ten patients; the larger Vietnam cohort study was designed around the vmPFC hypothesis and reports instead that its data did not support a vmPFC-specific mechanism, locating the mediated effect at dlPFC. Neither an n of 10 nor a near-boundary indirect effect settles the question.
A 2024 study substantially reframes the disagreement. Ferguson, Asp and colleagues — a collaboration spanning both earlier lesion programs, with Grafman and Tranel among the authors — applied lesion network mapping to religious fundamentalism in two independent datasets: 106 Vietnam-veteran patients from the same registry as the 2017 study, and 84 patients from the Iowa Neurological Patient Registry with etiologically different lesions (predominantly stroke and surgical resection; 42 women), for a combined n of 190 (Ferguson et al., 2024). Rather than asking which lesion locations raise fundamentalism scores, the method asks which network the relevant locations share, using a normative connectome from 1,000 healthy subjects. The answer: lesions associated with greater fundamentalism are connected to a specific, strongly right-lateralized network with nodes in right orbitofrontal, right dorsolateral prefrontal, and right inferior parietal cortex. The map derived from one dataset reproduced in the other at a spatial correlation of r = 0.82, with cross-validation in both directions, and the association survived control for lesion size. On this result, vmPFC and dlPFC are not competing answers but neighboring nodes of one network whose disruption tracks fundamentalist belief; and the cross-dataset reproduction, across cohorts that differ in etiology, sex composition, and age range, is the strongest replication event in this review’s literature.
Two cautions accompany the study, both the authors’ own in substance. The network is inferred through a normative connectome, not measured in the patients; and the study’s comparative finding — that lesion maps for confabulation and criminal behavior resemble the fundamentalism map — is a statement about spatial similarity of lesion-network topographies, not an equation of fundamentalism with either condition.
The anterior cingulate cortex monitors conflict between competing representations (Botvinick, Braver, Barch, Carter & Cohen, 2001) — the natural place to look for the “friction” that counterevidence ought to generate. Three findings bear on it, at three different registers.
First, in the belief fMRI study of Harris et al. (2009), in which 15 committed Christians and 15 nonbelievers evaluated religious and matched non-religious statements, the contrast of religious over non-religious content showed greater ACC signal across the pooled sample. The authors, noting that the same pattern marked uncertainty in their earlier work and that religious statements took longer to judge in both groups, “speculate that both groups experienced greater cognitive conflict and uncertainty while evaluating religious statements.” That is a pooled result and a stated conjecture; it does not show that believers specifically experience conflict with their own doctrine, and we do not claim it does.
Second, a believer-specific datum exists in the error-processing literature: across two EEG studies, stronger religious zeal and greater belief in God were associated with less anterior-cingulate reactivity to errors on a Stroop task, and with fewer errors, with personality and cognitive ability controlled (Inzlicht, McGregor, Hirsh & Nash, 2009). The task is generic rather than doctrinal, so the link to belief shielding is an interpretive step; but the direction is the interesting one: conviction accompanies a quieter conflict signal, which the authors read as religion buffering anxiety and the experience of error.
Third, the structural claim once attached to this system has not survived testing. The association between anterior cingulate gray-matter volume and political orientation (Kanai, Feilden, Firth & Rees, 2011) failed to replicate in the one adequately powered attempt (§5.2).
Held together, the three findings sketch a shielding picture with two distinct routes, to which the encoding findings below add a third: the conflict signal may fire and be resolved downstream in the belief’s favor (the Harris conjecture; and see Westen’s motivated-reasoning data, §4.2), or the signal itself may be damped in high-conviction hosts (the Inzlicht direction). These routes have different intervention implications — re-weighting an intact signal versus re-sensitizing a damped one — and the model treats them as distinct (§6.1).
Shielding need not operate on the evaluator; it can operate on the belief’s encoding.
Sacred values. In the fMRI auction paradigm of Berns et al. (2012; 31 analyzed), personal values participants refused to sell — the operational marker of sacredness — were processed with increased activity in left temporoparietal junction and ventrolateral prefrontal cortex, “regions previously associated with semantic rule retrieval,” while regions associated with utilitarian valuation showed no differential engagement — an explicit null in orbitofrontal, ventromedial, and striatal tests, not a deactivation. (In the same localizer design, parietal regions ran the other way: more active for the values participants were willing to sell.) The authors’ reading: sacred values operate “through the retrieval and processing of deontic rules and not through a utilitarian evaluation of costs and benefits.” A conviction encoded as a rule is not weighed, and evidence is an input to weighing. The result is one modest-sized study, and its exploratory whole-brain pass (which additionally implicated right amygdala for sacred values) awaits confirmation; but the encoding-format idea it supports is the cleanest mechanism in this section.
Transparency. The phenomenological form of the same bypass has a name in the philosophy of mind. In Metzinger’s self-model theory, a representation is phenomenally transparent to the degree that the processing stages that produced it are unavailable to introspection: the system cannot experience the representation as a representation, and so experiences its content as unmediated reality — in Metzinger’s words, “[t]he phenomenology of transparency, therefore, is the phenomenology of naive realism” (Metzinger, 2003a, 2003b). Transparency in this account is gradual, not binary. The import for entrenchment is direct, provided the theory’s own structure is respected: in Metzinger’s system, thoughts are ordinarily opaque — experienced as representations — and transparency belongs to phenomenal states and above all to the self-model. A conviction woven into the transparent partition of the self-model is therefore no longer met as a thought at all: the difference between “I believe X” and “X is simply true” is the difference between an opaque representation, experienced as held, and content lived as world. A conviction so integrated is not defended against evidence so much as constitutionally prior to it; there is, for its host, no “belief” there to evaluate. We import this as theory that names and predicts the pattern in the data above (untagged belief, unweighed rules, quiet conflict signals), not as a neuroscientific result; §7.3 draws its therapeutic corollary.
Entrenchment is not only defense; deep conviction is also sought, savored, and returned to. The direct evidence that religious experience engages reward circuitry comes from a single striking study. Ferguson et al. (2018) scanned 19 devout members of the Church of Jesus Christ of Latter-day Saints (returned missionaries, selected for weekly practice and self-reported spiritual experience) during tasks modeled on the tradition’s devotional practice: prayer, scripture reading, quotations from religious authorities, and audiovisual stimuli. The recognizable state their tradition names “feeling the Spirit” was reproducibly associated, across four acquisitions and three paradigms, with activation in bilateral nucleus accumbens and ventromedial prefrontal cortex, alongside frontal attentional regions.
The study’s temporal claim, as commonly cited, says more than the data do. What was measured is accumbens BOLD signal rising two to four seconds after participants pressed a button to report peak spiritual feeling; the authors infer, through an assumed hemodynamic delay, that the underlying neural activity “likely corresponded” to one to three seconds before the press. The reference event is a button press, not the onset of conscious experience, and the authors expressly decline to exclude the reverse order — accumbens activation following the decision that an experience was spiritual. What the study supports is an association between a culturally recognized devotional state and reward-circuit engagement in expert practitioners; the direction of the arrow between reward signal and conscious report is open, by the authors’ own account.
No independent replication of this paradigm exists, in this tradition or any other, as of this writing; a search of the study’s citing literature finds none. Convergent but paradigm-different support: in the Harris et al. (2009) sample, religious over non-religious statements engaged ventral striatum among other regions, in believers and nonbelievers alike. The reward system’s case rests, then, on one n = 19 study plus adjacent findings, and §6’s grading reflects it.
A second reward result connects this system to the shielding story, and it is the single most model-relevant dataset in this review. Westen, Blagov, Harenski, Kilts and Hamann (2006) scanned 30 committed partisans (15 Democrats, 15 Republicans, all men) during the 2004 U.S. presidential campaign as they evaluated self-contradictions by their own candidate, the opposing candidate, and neutral figures. Reasoning about one’s own candidate’s contradictions (the motivated condition) engaged ventromedial prefrontal cortex, anterior cingulate, posterior cingulate, insula, and lateral orbital cortex, and produced no differential engagement of dorsolateral prefrontal cortex, the region prior work associates with “cold” reasoning; the authors note this absence in their own contrasts. And once participants had had time to rationalize the threatening information away, the contrast against their initial confrontation with it showed a large ventral striatum activation, which the authors read as “likely reflect[ing] reward or relief engendered by ‘successful’ equilibration to an emotionally stable judgment.”
Their closing inference is this model’s thesis reached from the other direction: “[t]he combination of reduced negative affect… and increased positive affect or reward… once subjects had ample time to reach biased conclusions suggests why motivated judgments may be so difficult to change.” Defense of an identity-relevant belief is not only a shield; it is a rewarded operation. Each successful rationalization pays. A conviction defended this way is maintained by an internal reinforcement loop that never needs external confirmation. The threat, shielding, and reward systems of this review are captured in this one dataset of thirty men — which is both why we lean on it and why it, too, needs replication.
The reward findings invite an analogy with addiction: sensitization (repeated devotional states strengthening stimulus–reward associations), tolerance (escalation toward more intensive practice), withdrawal (dysphoria and meaning-void on disengagement, a pattern clinically described in exits from high-demand religion; Winell, 1993, 2011), and cue-triggered craving (familiar music, prayers, gatherings evoking the conditioned response long after exit). The analogy is bounded and we keep it so: belief entrenchment involves doctrinal content, social embedding, and self-narrative that substance dependence lacks, and no study has measured mesolimbic adaptation across a religious career. The shared skeleton — reward-mediated maintenance of a behavior pattern that persists against the host’s reflective judgment — is what the analogy is for, and nothing more is claimed for it.
Popular accounts of belief neuroscience lead with threat: the challenged believer’s amygdala lights up; the brain defends belief as it would defend the body. The evidence for that picture is weaker than for any other system in this review.
As §2.1 detailed, the amygdala/insula finding in Kaplan et al. (2016) is a between-subjects correlation (participants who updated least showed more emotion-circuit signal during counterevidence), not an activation of these structures by belief challenge in general; the authors state as much. Their interpretation menu is also broader than the familiar gloss: alongside the threat reading (“[o]ne interpretation of these activations…”), they propose that amygdala signal may index skepticism toward the counterevidence — heightened scrutiny of the messenger, not alarm for the self. In the one preregistered replication, the correlation did not reproduce (Kossowska et al., 2026). An individual-differences result at n = 40 that fails at n = 43 is not evidence to build on; it is a hypothesis awaiting a properly powered test.
Kanai, Feilden, Firth and Rees (2011) related self-reported political orientation (a single five-point item) to gray-matter volume in 90 young UK adults with an internal replication in 28 more: conservatism associated with larger right amygdala volume, liberalism with larger anterior cingulate volume. The authors’ causal hedging was unusually explicit — the causal nature of the relationship, in their words, “cannot be determined” from cross-sectional data, leaving open whether structure shapes attitudes or attitudes and their enactment shape structure — and it deserves repeating with the finding.
The finding has since had one adequately powered, preregistered replication: in 928 members of the Amsterdam Open MRI Collection, the amygdala–conservatism association reproduced at roughly one third the original effect size (Pearson’s r ≈ 0.068 against the original 0.23), while the anterior cingulate association did not reproduce at all (Petropoulos Petalas, Schumacher & Scholte, 2024). The replication also extends the surviving association beyond Anglo-American two-party contexts, into a multiparty electorate. An earlier preregistered attempt is its own lesson: Boekel and colleagues (2015) selected the finding for confirmatory replication and had to abort the test because their 36-person sample contained too little ideological variance to run it — the homogeneity that also characterized the original’s student pool, and that the Dutch population sample finally overcame. In sum: a small association between conservatism and right amygdala volume survives at n = 928; everything else in the volumetric story has shrunk or failed.
The strongest form of the threat picture was physiological: conservatives showing stronger skin-conductance and startle responses to threatening stimuli (Oxley et al., 2008; n = 46). A preregistered direct replication with the original stimuli, plus conceptual replications in two countries (total n across studies = 635), found no association between skin-conductance threat response and social conservatism — a standardized coefficient 37 times smaller than the original’s — and no evidence for a latent physiological threat-sensitivity trait (Bakker, Schumacher, Gothreau & Arceneaux, 2020; the startle measure was not retested, so the verdict covers the skin-conductance claim). At the behavioral level, five measures of negativity bias across four U.S. national samples (combined n ≈ 4,000) showed no consistent relation to ideology (Johnston & Madson, 2022). A 2025 systematic review of nineteen studies finds majority-but-mixed support overall, “modest” magnitude, and substantial method-dependence, with physiological and priming studies the least consistent (Dong et al., 2025). No meta-analysis postdating these replication results exists.
For this review’s purposes the conclusion is not that threat plays no role in belief defense (the fear-conditioning evidence below is about something else and stands), but that the dispositional claim — that entrenched or conservative believers are constitutionally more threat-reactive — is unsupported in its strong forms and weak in its surviving ones.
The threat system’s best-supported role in entrenchment concerns not disposition but learning. The amygdala is central to fear conditioning — the acquisition, storage, and expression of learned fear (LeDoux, 2000) — and the conditioning literature’s central fact is that extinction does not erase such learning. Extinction is new, context-dependent learning stored alongside the old, leaving the conditioned stimulus with two available meanings, the current context selecting between them; relapse under context change or the passage of time follows from the architecture (Bouton, 2002). Fear that has been unlearned is, in this sense, still there.
Apply this to doctrine acquired under fear. A child taught damnation through vivid depictions of punishment plausibly acquires the belief partly as a conditioned fear association: the doctrine as conditioned stimulus, terror as response. This application is this review’s account, not a tested result; no study has traced hell-belief through conditioning paradigms. But it predicts, from textbook mechanisms, a phenomenon the clinical literature on religious exit documents in detail: former believers who have intellectually rejected the doctrine of hell for years, yet respond to its cues with involuntary fear (Winell, 1993, 2011). Disbelief revises the proposition; it does not reach the conditioning. In the terms of this series’ companion framework, that residue is an installed regulatory setpoint outliving the membership that installed it, and the companion paper’s proposed exit studies would measure setpoint retention separately from stated belief and current affect (Paper 1.2, §3.3 and Paradigm B); the conditioning account reviewed here is the natural neural substrate for that variable.
The conditioning story requires no claim that believers are threat-reactive people, and it survives every negative result in §§5.1–5.3. It is the threat system’s real contribution to entrenchment — not a louder alarm, but a deeper installation.
We propose the Identity-Belief Fusion Model (IBFM): a conviction entrenches to the degree that it is simultaneously
A model is only as strong as its limbs, and the limbs’ grades differ:
| System | Anchor evidence | n | Independent replication status |
|---|---|---|---|
| Self-referential encoding | Kaplan et al. 2016; Kossowska et al. 2026 | 40; 43 | Reproduced (preregistered, cross-cultural) |
| Shielding — capacity/doubt | Zhong et al. 2017; Asp et al. 2012; Ferguson et al. 2024 | 119+30; 36; 190 | Network result reproduced across two cohorts (r = 0.82) |
| Shielding — override | Westen et al. 2006; Harris et al. 2009 | 30; 30 | None found (citing literature searched) |
| Shielding — format | Berns et al. 2012; Metzinger (theory) | 31; — | None found (citing literature searched) |
| Reward | Ferguson et al. 2018 | 19 | None exists (citing literature searched) |
| Threat — dispositional | Kanai et al. 2011; Oxley et al. 2008 | 90+28; 46 | Amygdala volume: reproduced at ~1/3 effect (n = 928); ACC volume: failed; physiology: failed |
| Threat — conditioning | LeDoux 2000; Bouton 2002 (mechanisms); Winell (clinical) | — | Mechanism literature robust; application untested |
The asymmetry is the model’s most important feature, and we state it rather than smooth it. The IBFM’s best-supported claims are that entrenched beliefs live in the self-system and are shielded from evaluation, with the shielding system carrying the single most reproducible result in this literature (the two-cohort lesion network). The reward system is plausible and nearly untested. The dispositional threat limb is the weakest, and the model does not rest on it; the conditioning limb does the threat system’s real work. A version of this model built three years ago would have led with threat; the replication record reordered it.
Entrenchment on this model is graded, not binary. A belief may be self-encoded but still evaluable (conviction held with awareness of its contingency); self-encoded and format-shielded (experienced as plain reality); and additionally reward-maintained and fear-anchored (the full configuration, at the far pole). Two independent literatures supply the spectrum’s instrumentation and its phenomenology. The identity-fusion program measures degrees of self–group merger with validated pictorial and verbal scales whose scores predict extreme pro-group behavior (Swann et al., 2009; Gómez et al., 2011); adapted from group-target to belief-target, these scales are the natural behavioral measure of the self-encoding limb — and until such an adaptation is validated, the IBFM’s central variable has no instrument, which we list as the model’s first methodological debt. On the phenomenological side, Metzinger’s transparency is explicitly a matter of degree, which matches the observation that beliefs move along this spectrum over a biography. Radicalization is movement toward the fused pole, a movement the fusion literature’s own field data associate with shared intense experience (Whitehouse, McQuinn, Buhrmester & Swann, 2014; Whitehouse, 2018); deconversion is movement away from it, a process §7.2 decomposes.
The model is stated domain-generally: nothing in its four systems mentions God, party, or cause, and the case material above already spans religious conviction (Mormon devotion, fundamentalism scales), political identity (U.S. liberals, Polish left-wingers, committed partisans), and sacred values of mixed content. The neuroimaging evidence base, however, is a handful of specific populations: forty American liberals, forty-three Polish leftists, thirty American partisans, nineteen devout Mormons, and two cohorts of lesion patients dominated by male American veterans. Generality at the neural level is therefore a prediction, not a result.
The behavioral literature supplies adjacent support. Cognitive rigidity relates to ideological extremity, partisanship, and dogmatism “across political and non-political ideologies” (Zmigrod, 2020); in a data-driven study administering 37 cognitive tasks and 22 personality surveys to 334 participants drawn from a parent sample of 522, related psychological profiles characterized political, nationalistic, religious, and dogmatic attitudes measured within one sample — overlapping profiles with domain-specific variation, not one uniform signature (Zmigrod, Eisenberg, Bissett, Robbins & Poldrack, 2021); and religious disbelief specifically has been associated with greater cognitive flexibility across three lab instruments in 744 participants (Zmigrod, Rentfrow, Zmigrod & Robbins, 2019). A shared family of rigidity correlates across ideological domains is what the IBFM predicts; identical neural profiles across domains are what it has not yet been tested on.
On this model, radicalization is not primarily the acquisition of extreme content but the fusing of content with self — movement along §6.2’s spectrum, whatever the doctrine. The fusion literature’s field data give the process description independent support: among 179 Libyan revolutionaries surveyed mid-conflict, frontline combatants, the men who shared the fighting’s dysphoria, were as fused with their battalion as with their own families, while logistical supporters were not; and fusion theory generalizes the pattern as shared intense experience, or the perception of shared essence, begetting oneness begetting willingness to sacrifice (Whitehouse et al., 2014; Whitehouse, 2018; the field data are cross-sectional, and the causal reading is the theory’s). Counter-radicalization aimed at content — debating the doctrine — targets the least entrenched layer of the whole arrangement. The model’s implication is that prevention and exit work should target the fusion process: the identity dependence, the reward monopoly, the fear anchoring, in whichever mix a given case presents (§7.3).
Leaving a deeply held belief system, on this model, is four separate undoings, which need not proceed together:
The decomposition explains the clinically familiar dissociations: the ex-believer who has completed the second undoing but not the fourth (intellectually free, still afraid); the one who has completed all but the third (convinced, safe, and desolate). Partial deconversion is not a failure of will but the expected presentation of a multi-system fusion being dismantled one system at a time. The phenomenology of the process — grief rather than error-correction — is treated at the collective scale in this series’ companion paper, whose account of exit as contraction of a shared computational boundary is complementary to the individual-level decomposition given here (Paper 1.2, §3.2).
If entrenchment-related distress presents along four dimensions, intervention should match the dominant one: exposure-based and extinction-informed methods where conditioned fear dominates (with Bouton’s context-dependence predicting relapse risk and arguing for extinction practice across contexts); behavioral activation and community-building where reward loss dominates; narrative and identity work where self-encoding dominates; and, where format-shielding dominates, approaches whose goal is what self-model theory calls rendering the representation opaque: helping the client experience the conviction as a belief, with a history and a form, rather than as the world. (Metzinger’s standard example of exactly this shift is lucidity — the dreamer’s whole model of reality “suddenly experienced as a model.”) That last formulation is this model’s translation of a familiar therapeutic aim, not a validated technique, and this entire section is proposal: matched-intervention designs are among the studies §8 calls for.
Nearly everything reviewed here is cross-sectional. Whether the neural profile precedes entrenchment, follows it, or both, is unmeasured; the bidirectional hypothesis of §3.1 and the developmental question of §5.4 (does fear-acquired doctrine differ neurally from calm-acquired doctrine?) both require longitudinal designs tracking conversion and deconversion as they happen: difficult, expensive, and to our knowledge not yet done.
The lesion literature (§§3.1–3.3) supplies nearly all of the field’s causal-direction evidence, and it concerns loss of function. Reversible neuromodulation of the network identified by Ferguson et al. (2024) during belief-evaluation tasks would test the shielding system directly; no such study exists, though vmPFC neuromodulation has been reported to reduce endorsement of religious belief — a result noted in Ferguson et al.’s (2018) discussion, and a sign the method has already touched the domain.
The samples enumerated in §6.3 are WEIRD with few exceptions, and several are single-sex. The one deliberate cross-cultural extension in this literature — the Polish replication — reproduced the self-referential result across a real cultural distance while staying on the same ideological pole: encouraging, and doubly insufficient. Non-Abrahamic traditions, non-Western political cultures, and non-student adult samples are nearly absent.
Why the same doctrine entrenches deeply in one host and shallowly in another is the model’s open parameter. Candidate moderators — attachment style, need for closure, ambiguity tolerance, developmental timing of exposure, trait anxiety — are all measurable with existing instruments, and the fusion scales (§6.2) give the outcome side a candidate metric. This is the cheapest gap in the list to start closing.
The model describes a configuration; entrenchment and its undoing are processes. Dynamic functional-connectivity methods, a reviewed methodology with known interpretive caveats (Hutchison et al., 2013), make the model’s one novel neural prediction testable: individuals with more entrenched convictions should show reduced dynamic flexibility of default-mode configurations — a narrowed repertoire of self-model states — relative to matched individuals with revisable convictions, beyond any difference in static activation. The prediction is falsifiable, uses existing methods, and its failure would remove the model’s claim that entrenchment is a self-system property rather than a belief-content property.
The promise of §1.2 completes here. The median neuroimaging study samples roughly 25 participants, and at such sizes brain-wide associations are systematically inflated and hard to reproduce; reproducibility in that paradigm begins in the thousands (Marek et al., 2022), with large-cohort work placing typical requirements at 1,500–3,900 participants (Liu, Abdellaoui, Verweij & van Wingen, 2023). Median statistical power across the neuroscience meta-analyses of 2011 was estimated at 21%, a figure whose distribution is multimodal rather than uniform, with some subfields adequately powered and others severely not (Button et al., 2013; Nord, Valton, Wood & Roiser, 2017). Against these benchmarks, the fMRI core of this review (samples of 19–43) sits at or below the field’s median, in the regime where effects inflate; its two adequately powered studies are the structural replication (n = 928), which shrank the effect it confirmed by two thirds, and the behavioral null (n ≈ 4,000). The review’s strongest evidence comes instead from the paradigm class that the benchmark literature itself exempts from the thousands requirement because its effects are larger: lesion designs, including the two-cohort network study whose cross-dataset reproduction (r = 0.82) is this literature’s best result. That distribution of strength — small task-fMRI effects awaiting large-sample confirmation; robust lesion-network findings; one preregistered task replication that split its source study’s claims — is the state of the evidence, and the model of §6 is graded by it.
Deep belief entrenchment, on the evidence reviewed here, is not a failure of intelligence or information. It is a configuration: a conviction seated in the self-representation system; exempted, by capacity loss, rewarded override, or encoding format, from the processes that revise ordinary beliefs; fed by its own reward loop; and, where acquisition was fearful, anchored by conditioning that outlasts assent. The Identity-Belief Fusion Model names the configuration and grades its own limbs: the self and shielding systems carry replicated and reproduced results, the reward system a plausible and untested mechanism, the dispositional threat story mostly its own failed replications, and the conditioning story a robust mechanism literature awaiting its direct test.
Three conclusions travel beyond the model. First, the field’s most reproducible finding, the right-lateralized prefrontal-parietal lesion network for fundamentalism, locates entrenchment’s clearest neural correlate in the systems that evaluate, not the systems that fear: what most reliably distinguishes rigid belief is not more alarm but less revision. Second, the one result that replicated across cultures says that challenges to a deep conviction are processed as events about the self — which is what makes “just consider the evidence” a category error as advice, and identity-safe framings of disconfirming information a rational design goal for anyone who communicates it. Third, beliefs acquired under fear are doubly encoded, as proposition and as conditioning, and only the proposition answers to argument; the residue explains why people can stop believing and keep suffering, and it marks where clinical work, not debate, is the relevant instrument.
A belief that has fused with the self is maintained by the machinery that maintains the person. That is what entrenchment is, on this model — and it is why its undoing is never merely a change of mind, and why understanding it requires, and rewards, the full apparatus of systems neuroscience applied with replication discipline.