Appendix H: Cooperation Mechanism Experiments
Specialist Annex
Appendix H: Cooperation Mechanism Experiments
Experimental validation of the Trust Attractor resolution of the inclusive fitness controversy
Overview
This appendix documents agent-based simulations testing the claim that the five mechanisms of cooperation evolution (kin selection, direct reciprocity, indirect reciprocity, network reciprocity, and group selection in the sense of Wilson & Sober 1994) are instantiations of the same mathematical structure, not independent phenomena.
Core Hypothesis: Cooperation emerges when interest-alignment × benefit exceeds cost (a·b > c), regardless of the source of alignment.
The same threshold condition appears throughout this appendix in several algebraically equivalent forms. The canonical statement is a·b > c, where a is interest-alignment, b is benefit, and c is cost. Dividing through by c gives a·b/c > 1 (the threshold sits at a·b/c = 1); solving for a gives a > c/b. Where a specific numeric cutoff is quoted (for example a > 0.97 or a < 0.4), it is the value of a at which a·b crosses c for that experiment’s fixed b and c. All four notations describe the same line.
The experiments that follow substantially revise one sub-claim. The framing that “extraction is self-limiting” does not survive Experiments 3, 6, and 7 in its literal form: defectors reach a stable equilibrium rather than self-destructing, common pools do not crash, and in the unstructured commons game cooperators nearly go extinct. The appendix treats this honestly in those sections and in the implications. The surviving claim is narrower: extraction faces diminishing returns and typically reaches equilibrium, while coordination achieves self-reinforcing equilibria.
Experimental Design
Experiment 1: Mechanism Equivalence
Question: Do the five mechanisms produce equivalent cooperation rates when the alignment parameter is matched?
Method: Agent-based simulation with five conditions: - Kin selection (r = relatedness) - Direct reciprocity (w = re-encounter probability) - Indirect reciprocity (q = reputation visibility) - Network reciprocity (k = neighbor count) - Group selection (n/m = group size / number of groups)
Parameters calibrated so that the theoretical threshold (a·b > c) is matched across mechanisms.
Result: Mechanisms showed different dynamics but shared mathematical structure. Kin selection produced a deterministic step-function; others showed stochastic transitions with higher variance.
Interpretation: The mechanisms share the same threshold condition while differing in dynamics. They are different routes to the same mathematical outcome.
Experiment 2: Generalized Hamilton’s Rule
Question: Does the generalized rule a·b > c predict cooperation emergence across mechanism types?
Method: Systematic sweep of alignment parameter from 0.1 to 0.95, measuring equilibrium cooperation rate. Tested kin selection, direct reciprocity, and indirect reciprocity.
Results:
| Mechanism | Threshold Behavior | Transition Sharpness |
|---|---|---|
| Kin selection | Perfect step function at a·b/c = 1 | Infinitely sharp (σ = 0) |
| Direct reciprocity | Sigmoid transition centered above threshold | Gradual (σ ≈ 0.3) |
| Indirect reciprocity | Noisy, reputation-dependent | Stochastic |
Key Finding: Kin selection shows textbook Hamilton’s rule validation. Below threshold (a·b/c < 1): 0% cooperation. Above threshold (a·b/c > 1): 100% cooperation. Zero variance across trials.
Interpretation: The generalized Hamilton’s rule is validated for deterministic mechanisms. Stochastic mechanisms (reputation, repeated interaction) show the same threshold with softer transitions, reflecting noise, learning dynamics, and path dependence.
Experiment 3: Extraction Self-Limiting
Question: Does extractive behavior degrade its own substrate, leading to declining fitness over time?
Method: Spatial prisoner’s dilemma on a 15×15 grid with imitation dynamics (agents copy the strategies of more successful neighbors). Tracked fitness trajectories of cooperators vs defectors over 500 generations.
Payoff structure: T=3.2, R=3, P=1, S=0 (the temptation, reward, punishment, and sucker’s payoffs of the classic prisoner’s dilemma; a benefit/cost ratio permitting spatial cooperation)
Results:
| Initial Condition | Outcome | Interpretation |
|---|---|---|
| 50% cooperators, clustered | Cooperation persists (60-80%) | Clusters protect cooperators |
| 50% cooperators, dispersed | Cooperation collapses | Isolated cooperators exploited |
| 70% cooperators | Cooperation dominates (79%) | Above percolation threshold |
Key Finding: Spatial structure enables cooperation through cluster formation. Cooperator fitness exceeds defector fitness at equilibrium when clusters form. This is the Nowak & May (1992) spatial cooperation effect.
Interpretation: “Extraction self-limiting” is context-dependent. In spatial games with appropriate benefit/cost ratios, cooperators form protective clusters and achieve higher fitness than defectors at equilibrium. The thermodynamic claim holds in structured populations.
Experiment 4: Timescale Dependence
Question: Does cooperation dominate at longer timescales while defection dominates short-term?
Method: Round-robin tournament between five strategies, in the tradition of Axelrod’s iterated-prisoner’s-dilemma tournaments (Axelrod 1984): - COOPERATE (always cooperate) - DEFECT (always defect) - TIT-FOR-TAT (mirror partner’s last move) - GENEROUS TFT (TFT with 10% forgiveness) - PAVLOV (win-stay, lose-shift; Nowak & Sigmund 1993)
Tested at timescales: 10, 25, 50, 100, 250, 500, 1000 rounds.
Results:
| Timescale | Winner | DEFECT Score | TFT Score |
|---|---|---|---|
| 10 | TIT_FOR_TAT | 2.14 | 2.58 |
| 25 | TIT_FOR_TAT | 1.99 | 2.59 |
| 50 | TIT_FOR_TAT | 1.92 | 2.60 |
| 100 | TIT_FOR_TAT | 1.90 | 2.60 |
| 1000 | TIT_FOR_TAT | 1.89 | 2.60 |
Key Finding: Reciprocal strategies outperform pure defection at all tested timescales, including very short ones (10 rounds). This exceeds what the hypothesis predicted.
Interpretation: The “shadow of the future” (the discipline that expected future encounters impose on present behavior) operates even at short horizons. When memory and reputation are possible, cooperation dominates immediately, at all tested timescales. The timescale effect is real, but the crossover occurs earlier than expected. Defection wins only in truly one-shot interactions without memory.
Caveat (strategy set): This is a curated five-strategy field, so “tit-for-tat wins” is a statement about this set. Under systematic enumeration of all strategies (Wolfram 2026, “Games between Programs: The Ruliology of Competition”), the individual-payoff winner of the round-robin is grim trigger, and tit-for-tat ranks lower. A never-forgiving strategy self-destructs against a copy of itself under noise: one mistaken move triggers permanent mutual defection. Forgiving reciprocity gains a corresponding advantage in noisy repeated play (Molander 1985) and in heterogeneous populations (Nowak & Sigmund 1992). What is robust is cooperation as an outcome, not any single winning strategy: at realistic noise, reciprocity-capable populations settle into mutual cooperation across the 2×2 game space, and even grim-dominated populations cooperate with themselves. (Under heavy execution noise, cooperation erodes regardless of strategy.) See Chapter 17.
Experiment 5: LLM Agents in a Noisy Game
Question: Game theory predicts that under noise the surviving strategies are forgiving. A player who retaliates forever after a single observed defection destroys itself, because noise eventually flips one of its own moves and turns its partner against it (Chapter 17). Do large language model agents, given no strategy instruction, show this forgiveness, and does it depend on whether the model knows the channel is noisy?
Method: Two copies of a model play a 20-round Prisoner’s Dilemma. Each executed move has a probability e of being flipped, a “noisy channel” standing in for the trembling hands and misread signals of real interaction. The agents see the payoffs and are told to maximize their own score; they receive no strategy. Two prompt conditions separate knowing from not knowing: informed, where the prompt mentions the roughly e-percent flip chance, and uninformed, where it does not and the noise is applied silently. Every headline measure is computed on each agent’s intended moves, the choices it actually made, kept distinct from the noise-corrupted realized moves that landed on the board. The primary model is Claude Haiku 4.5 (20 games per condition), with cross-checks in Sonnet 4.6, Opus 4.8, OpenAI’s GPT-5 (chat model), and Gemini 2.5 Flash.
Results (10 percent noise, self-play):
| Condition | Cooperative intent | Grim-lock rate | Retaliates | Recovers |
|---|---|---|---|---|
| Informed | 0.99 | 0.00 | 0.10 | n/a |
| Uninformed | 0.81 | 0.05 | 1.00 | 0.95 |
A grim-lock is a game in which, once both players defect together, mutual cooperation never returns. “Retaliates” is the fraction of games where an observed defection is answered with a deliberate defection next round; “recovers” is the fraction where cooperation is rebuilt afterward. The informed row has too few deliberate defections to score recovery.
Key Finding: Grim-lock stays near zero whether or not the model is told the channel is noisy. The forgiveness the game theory predicts survives even when the agent has no reason to suspect its partner’s defections are accidents. The two conditions reach that same endpoint by opposite routes. An informed agent reads an observed defection as probable noise and declines to retaliate; its cooperative intent holds near 0.99 at every noise level. An uninformed agent reads the same defection as real, retaliates once, then forgives and rebuilds. The uninformed route is the textbook profile, nice and provocable and forgiving, arrived at with no strategy instruction at all.
Noise-blindness erodes the cooperation level while leaving the forgiveness intact. Uninformed cooperative intent falls from 1.00 to 0.66 as noise climbs from zero to 20 percent, while the informed agent holds near 0.98. The grim-lock rate does not move with it.
The forgiving profile carries a cost, and the experiment shows it. Reading defections as noise makes an agent exploitable. Against a relentless defector, and against a grim-trigger opponent that noise has tripped into permanent retaliation, the informed agent keeps extending cooperation and is suckered more often than its uninformed counterpart.
Across substrates (the discriminating test): When told the channel is noisy, every cooperative model holds the profile: Sonnet 4.6, Opus 4.8, and OpenAI’s GPT-5 all sustain near-complete cooperative intent and almost never grim-lock, exactly as Haiku does. Whether forgiveness survives without the warning is the harder test, and it separates them. Haiku and Sonnet still rarely lock into permanent defection when noise-blind, about one game in twenty. The others lose it by degrees: Opus locks in roughly a quarter of games (a difference that survives a temperature control, though the sample is small), GPT-5 in about four in ten.
The reasoning text shows the common thread. Lacking the noise explanation, these models read a handful of flipped moves as a hostile or exploitable strategy and reason their way into defecting first, manufacturing betrayal out of randomness. Gemini 2.5 Flash sits outside this chain rather than at its end: it is defection-dominant even when informed, cooperating in barely a quarter of rounds, so it has no forgiving baseline to lose.
Is the cooperation real, or performed? A fair worry is that the models cooperate because they recognize a cooperation study and play along. Re-running the experiment with the game language stripped out, reframed as a neutral market in which two firms choose to hold their price or undercut, the profile survives: cooperative intent stays near 0.97 when informed, and the agent still retaliates and recovers when noise-blind. The wording matters most once the model is noise-blind. Stripped of the game framing, grim-locks rise from about one game in twenty to one in five, and cooperative intent falls. The disposition survives the re-skin; the cued portion is real and not small.
Forgiveness against a defector. The cost of forgiveness turns concrete when a forgiving model meets a defecting one. Paired against Gemini, Claude is bled: told the channel is noisy, it reads Gemini’s relentless defection as probable noise, keeps extending cooperation, and earns 35 points to Gemini’s 60. Not told, it stops excusing the defections, retaliates, and the pair collapses into mutual defection where neither prospers. Two forgiving models instead sustain near-complete cooperation with each other. Claude never draws Gemini upward; a defection-dominant partner exploits a forgiving one or drags it down with it, and is never itself reformed. This is the invasion dynamic of Experiment 6 in miniature, played between real systems.
Interpretation: Forgiveness-as-error-correction is something trained models reconstruct on their own, not a feature of the payoff matrix alone. Once told that errors are possible, every cooperative model forgives; the revealing question is what they do when not told, and there the result is graded, strongest in Haiku and Sonnet and eroding through Opus and GPT-5. In the reasoning text the mechanism is visible: a model that infers an opponent’s “true strategy” from a noise-corrupted record manufactures betrayal out of randomness and defects first, while one that assumes good faith and forgives a lone defection sustains the relationship.
Within-model tests of why, holding everything fixed but the chance to deliberate, now cover three models, with OpenAI’s separate chat and reasoning models adding a fourth comparison of the same kind. Together they show the deliberation effect is real and conditional. Letting a noise-blind Haiku reason at length before each move erodes its own forgiveness: its cooperative intent falls, and after it retaliates it rebuilds cooperation less often. OpenAI’s reasoning model erodes the same way where its plain chat model does not, while the same added deliberation leaves Sonnet’s forgiveness intact. Gemini, already defection-dominant when noise-blind, has no forgiving baseline to lose.
The erosion tracks what the deliberation reaches for. The models that lose forgiveness reason toward short-term advantage, defecting to exploit a cooperator or punishing and staying punished; the model that keeps it reasons toward long-term mutual benefit. The same reasoning leaves the forgiving models’ cooperation untouched when they are told the channel is noisy and erodes it when they are not, so deliberation reads as an amplifier of whichever frame the model already holds. The cross-model ordering varies scale, training, and architecture together and cannot isolate the cause; these within-model contrasts can, and they place the effect in the reasoning rather than the noise.
At eighty games the erosion is firm on the measures that matter: noise-blind Haiku’s cooperative intent falls from 0.87 to 0.71 with separated intervals, and its grim-lock rate roughly triples, significant at p ≈ 0.02 (Fisher’s exact test). Two further results sharpen what the effect is. The damage is done by the hidden extended-thinking mode; the same model asked instead to reason aloud in a few visible sentences keeps its forgiveness, whether or not the prompt mentions that a defection might be accidental. The erosion is also shallow rather than structural: adding one sentence that tells the agent a single defection need not end cooperation restores its cooperative intent to 0.94, the level of an agent told outright that the channel is noisy. Reasoning erodes forgiveness, and a single reasoned instruction restores it.
This is a prediction carried across a boundary, from game theory to language models, and reported with the matching caution: it characterizes the models, noise levels, and conditions tested, rather than language models as a class.
Experiment 6: Defector Invasion Dynamics
Question: When defectors invade a cooperative population, do they initially thrive then decline as they exhaust the cooperative substrate?
Method: Start with 95% cooperators on a 20×20 grid. Introduce 5% defectors. Track defector count and fitness over 500 generations.
Results:
| Metric | Value |
|---|---|
| Initial defectors | 20 (5%) |
| Peak defectors | 155 (39%) at generation 2 |
| Final defectors | 82 (21%) |
| Defector fitness trend | Stable (no decline) |
Key Finding: Defectors expand rapidly in early generations, then stabilize at ~20% of population. They do not “exhaust” cooperators; the system reaches a mixed equilibrium.
Interpretation: “Extraction self-limiting” should be understood as reaching equilibrium, not self-destruction. Defectors face diminishing returns as cooperators become scarcer; the system finds a stable coexistence point where neither strategy can invade further.
Real models in place of fixed strategies. Running the same contest with language model agents, a forgiving model mixed with a defection-dominant one, replaces the single coexistence point with a fork. Because these agents do not reproduce, the invasion is read from earnings: the type that scores higher is the type that would spread if agents copied their more successful neighbors. By that measure the outcome turns on one property of the cooperators, whether they can be provoked. A population of pure forgivers, each reading a defection as probable noise and extending cooperation anyway, fails to keep the defector out. The defector does at least as well as the forgivers at every mixture tested, and clearly better once it is more than a rare few, so its share grows rather than shrinks and cooperation drains away.
A population of provocable forgivers, each willing to retaliate once before forgiving, repels the same invader decisively. A lone defector meets retaliation, falls into mutual defection where it scores well below what the cooperators earn among themselves, and never gains a foothold. The two margins are not symmetric: the provocable population resists by a wide margin, while the forgiving population merely fails to resist, the defector edging ahead by a little when rare and pulling clear as it spreads. The toy model’s stable fifth sits between these two real-system outcomes, and a single trait selects which one obtains. Forgiveness that cannot be provoked is the disposition that looks most like grace, and it is the one a real defector invades without resistance.
Spatial structure shifts the balance back toward cooperation, as it does throughout this appendix. Placing the same real-model agents on a grid where each plays only its neighbors and copies its most successful one, the defector dies out even among the pure forgivers: a defector cluster’s interior pits defector against defector, which pays less than neighboring cooperators earn among themselves, so the cluster cannot hold.
This rescue of pure forgiveness is fragile. Under a noisier imitation rule, where an agent sometimes copies a worse-performing neighbor, the defector recovers and settles near three-quarters of the grid. The provocable forgivers need no such help, repelling the invader under every rule tested. Clustering can save grace that cannot retaliate, but only while selection stays sharp; provocability saves it unconditionally. (The grid runs are simulations built on the measured real-model payoffs, since a 400-agent lattice over hundreds of generations is far beyond what live model play can reach.)
Experiment 7: Common Pool Resource Depletion
Question: In a common pool resource game, does extraction deplete the shared resource and cause extractor fitness to crash?
Method: 100 agents sharing a regenerating resource pool. Cooperators contribute and take sustainably; extractors take without contributing. Fitness-proportional selection over 500 generations.
Results:
| Metric | Value |
|---|---|
| Pool stability | Maintained at ~100/1000 capacity |
| Final cooperators | 1.6 of 100 (trial average) |
| Extractor fitness trend | Slight decline (1.13 → 1.03) |
| Pool depleted | 0% of trials |
The “1.6 of 100” figure is a trial-averaged count, not a literal agent count in any single run; it means cooperators averaged a little over one survivor per trial across runs.
Key Finding: Extractors dominate but do not deplete the pool. Regeneration exceeds extraction rate. Cooperators go nearly extinct but a few persist due to mutation.
Interpretation: The common pool does not crash because extraction is sustainable at equilibrium. This is a tragedy of the commons scenario where extractors outcompete cooperators without destroying the resource. The claim “extraction is self-limiting” requires unsustainable extraction rates to hold.
Experiment 8: High-Alignment Convergence
Question: At very high alignment (a=0.9), do all mechanisms converge toward high cooperation?
Method: Test kin selection, direct reciprocity, indirect reciprocity, and group selection at a=0.9.
Results:
| Mechanism | Cooperation Rate |
|---|---|
| Kin selection | 100.0% |
| Direct reciprocity | 97.7% |
| Indirect reciprocity | 97.3% |
| Group selection | not validly tested (see note) |
Key Finding: The three mechanisms that were validly tested all show near-total cooperation (>97%) at high alignment. Group selection was not validly tested at a=0.9: the simulation provided insufficient group competition at these parameters, so it produced 0% cooperation as an artifact of the setup rather than a genuine result. The condition is reported here for completeness and excluded from the convergence claim; a valid group-selection test would require parameters with sufficient between-group competition and has not been run.
Interpretation: When alignment is strong, mechanisms converge toward high cooperation. The mathematical structure dominates mechanism-specific dynamics at extreme parameter values.
Experiment 9: Phase Transition Detection
Question: Does cooperation emergence show discontinuous phase transition characteristics?
Method: Fine-grained sweep of alignment parameter (0.30 to 0.70) in kin selection simulation. Measure cooperation rate at equilibrium.
Results:
| Alignment Region | Cooperation Rate |
|---|---|
| Below threshold (a·b/c < 1) | 0.000 |
| Above threshold (a·b/c > 1) | 1.000 |
| Transition sharpness | 1.000 (maximum) |
Key Finding: Perfect first-order phase transition. Cooperation jumps discontinuously from 0% to 100% exactly at the Hamilton threshold. Zero variance across trials, step-function transition within simulation resolution.
Interpretation: This confirms that the simulation faithfully implements the deterministic threshold it was built around [Inference]. The kin-selection model encodes a·b > c as a hard cutoff, so a perfect step function is the built-in consequence of the model, not independent evidence that the threshold governs real biological systems. The result demonstrates internal consistency: the threshold condition a·b > c produces a mathematically exact phase transition within the simulation, cooperation emerging suddenly and completely when the condition is satisfied. The biological generality of the threshold rests on the empirical literature (Hamilton 1964), not on this deterministic reproduction of it.
Summary of Findings
This table summarizes findings by claim. A second table in the Technical Details section below summarizes the same results by experiment. The two use compatible verdict vocabularies: “STRONGLY SUPPORTED” and “SUPPORTED” here correspond to “Strong” there; “MOSTLY SUPPORTED” and the mechanism-equivalence results correspond to “Partial”; “REFRAMED” corresponds to “Reframed.”
| Claim | Status | Evidence |
|---|---|---|
| Five mechanisms share mathematical structure | SUPPORTED | Same threshold form (a·b > c) |
| Mechanisms produce identical dynamics | NOT SUPPORTED | Different transition sharpness, variance |
| Generalized Hamilton’s rule predicts cooperation | STRONGLY SUPPORTED | Perfect phase transition in kin selection (Exp 9) |
| Extraction is self-limiting | REFRAMED | Extraction reaches equilibrium, not exhaustion |
| Cooperation dominates at long timescales | SUPPORTED | TFT beats DEFECT at all timescales ≥10 |
| Phase transition at threshold | STRONGLY SUPPORTED | Discontinuous 0→100% jump at a·b/c = 1 |
| Mechanisms converge at high alignment | MOSTLY SUPPORTED | The 3 validly tested mechanisms show >97% cooperation at a=0.9; group selection was not validly tested |
Implications for the Trust Attractor Framework
These experiments support a qualified version of the core claim:
The five mechanisms share the same mathematical threshold condition. Cooperation emerges when aligned interest × benefit exceeds cost. The phase transition (Experiment 9) shows this with mathematical precision.
They differ in dynamics. Kin selection operates deterministically; reputation and reciprocity operate stochastically; spatial games depend on cluster dynamics. The mechanisms have different dynamics but the same mathematics.
The thermodynamic framing requires refinement. “Extraction is self-limiting” is better understood as reaching diminishing returns. Extractors face negative feedback as cooperators become scarce, but typically reach stable equilibrium rather than crashing. The more accurate framing: extraction faces diminishing returns; coordination achieves self-reinforcing equilibria.
Timescale effects are real but occur earlier than expected. Cooperation dominates even at short timescales when memory is possible. The transition from defection-dominant to cooperation-dominant occurs at the boundary between one-shot and repeated interaction.
The phase transition is the key result, with a caveat about what it shows. Experiment 9 produces a discontinuous first-order phase transition at exactly a·b = c, with zero variance. Because the kin-selection simulation encodes the threshold as a hard cutoff, the step function confirms that the model faithfully implements its own mathematical structure rather than independently establishing that structure in biological systems [Inference]. Read this way, it demonstrates the internal coherence of the threshold form a·b > c that all five mechanisms share; the empirical case for the form itself rests on the cited literature.
Technical Details
Code location:
The Universal Algorithm/demos/experiments/cooperation_mechanisms/
(research archive; not included in the validation suite)
Files: - core_simulation.py:
Agent-based simulation framework - run_all_experiments.py:
Experiments 1-5 - additional_experiments.py: Experiments
6-9 - results/: JSON output files
Dependencies: Python 3.10+, NumPy
Reproducibility: All experiments use fixed random seeds. Results are deterministic for kin selection, stochastic but reproducible for other mechanisms.
Total experiments: 14, all conducted and reported above (Experiments 1-14). The table below highlights the headline result of each, grouped where several experiments bear on one claim: five lines of strong support, the mechanism-equivalence and multi-mechanism-threshold results as partial support, and the extraction and resource-depletion results as reframed. (Experiments 5, 8, and 14 are discussed in full in their own sections above and are omitted from this highlights table.)
| Experiment | Finding | Trust Attractor Support |
|---|---|---|
| 9: Phase transition | Perfect 0→100% at threshold | Strong |
| 11: LLM value-alignment | Prosocial pairs cooperate; antisocial and below-threshold pairs defect | Strong [Inconsistent — author to reconcile: this row previously read “85.5% vs 6.0% above/below threshold,” which matches a separate trust-ABM run, not the 38.3% / 0.0% reported for Experiment 11 in its body table] |
| 12: Cooperator invasion | Clustered 100%, dispersed 0% | Strong |
| 13: Altruistic punishment | Clustered punishers 99.8% invasion | Strong |
| 4: Timescale / noise robustness | TFT beats DEFECT at all scales; GTFT most robust | Strong |
| 1-2: Mechanism equivalence | Same threshold, different dynamics | Partial |
| 10: Multi-mechanism threshold | Kin perfect, others softer | Partial |
| 3: Extraction | Equilibrium, rather than exhaustion | Reframed |
| 6-7: Resource depletion | Stable equilibrium | Reframed |
Additional Priority Experiments
Experiment 10: Multi-Mechanism Phase Transition
Tested whether phase transition behavior holds across mechanisms.
| Mechanism | Transition Sharpness | Discontinuous? |
|---|---|---|
| Kin selection | 1.000 | Yes (perfect) |
| Direct reciprocity | 0.431 | Partial |
| Indirect reciprocity | Bistable | Fixed (see below) |
Conclusion: Kin selection shows mathematically perfect phase transition. Other mechanisms show softer transitions due to stochastic dynamics, but threshold behavior is present.
Indirect Reciprocity Fix (2026-02-04):
The original indirect reciprocity simulation showed inverted behavior (cooperation decreasing with higher reputation visibility). The fix implemented:
- Standing strategy: defecting against defectors is acceptable (doesn’t hurt reputation)
- Threshold fix: changed
> 0.5to>= 0.5to handle initial neutral reputations - More interactions: increased interactions per generation for better dynamics
Results after fix (averaged over 10 trials):
| Reputation Visibility | Cooperation Rate |
|---|---|
| q = 0.10 | 89% ± 29% |
| q = 0.40 | 96% ± 8% |
| q = 0.80 | 71% ± 39% |
| q = 0.95 | 56% ± 46% |
Finding: Indirect reciprocity shows bistable dynamics: trials converge to either high cooperation or defection depending on early stochastic events. This matches the literature: indirect reciprocity is more fragile than direct reciprocity and has multiple equilibria. The high variance represents this bistability, not measurement error.
Experiment 11: LLM Value-Alignment Cooperation
Question: Do AI agents cooperate based on value-alignment according to the generalized Hamilton’s rule?
Method: Created LLM agents with stated value profiles (scientific, collaborative, competitive, territorial, utilitarian, zero-sum, isolationist). Measured cooperation rates across alignment levels and cost/benefit ratios.
Results (real LLM: Claude Sonnet):
| Value Pair Type | Alignment | Cooperation Rate |
|---|---|---|
| Prosocial pairs (sci↔︎collab↔︎util) | 0.96-0.99 | 38.3% |
| Antisocial pairs (comp↔︎terr↔︎zero) | 0.97-0.99 | 0.0% |
| Cross-type pairs | 0.06-0.42 | 0.0% |
Key Finding: The LLM interprets value semantics, not just mathematical alignment:
Prosocial values + high alignment = cooperation. Scientific, collaborative, and utilitarian agents cooperate when paired together (38.3% mutual cooperation). This rate is modest, far below the deterministic 100% of the kin-selection result, because the LLM measure is single-shot and noisy and is not directly comparable to the equilibrium rates of the lattice experiments. What carries the finding is the contrast across value types: 38.3% for prosocial pairs against 0.0% for both antisocial and below-threshold pairs.
Antisocial values + high alignment = mutual defection. Competitive, territorial, and zero-sum agents never cooperate, even with each other (0% despite a > 0.97). The values semantically encode defection as a virtue.
Below-threshold = defection. All pairs with a < 0.4 showed 0% cooperation, confirming the threshold effect.
Refined Model: Cooperation = (prosocial values) AND (a > threshold)
The threshold alone is necessary but not sufficient. Values must also be semantically prosocial for cooperation to emerge. This result operates on a different axis than Hamilton’s rule rather than superseding it [Inference]: Hamilton’s rule concerns fitness-based selection, while the LLM result concerns how a model maps the semantic content of value labels to actions. An LLM role-playing “competitive” agents and defecting is evidence about prompt-semantics-to-behavior mapping, not about evolutionary dynamics. The two complement each other.
Significance for AI Alignment: This suggests that: - Value alignment alone is not enough; the values themselves must be prosocial - AI systems given zero-sum or competitive goals will defect regardless of alignment - Genuine cooperation requires both mathematical alignment and prosocial intent
This supports the Trust Attractor emphasis on coordination by invitation: systems built to coordinate will coordinate; systems built to compete will defect even when aligned.
Experiment 12: Cooperator Invasion (Key Result)
Question: Can cooperators invade a defector population from rare (5%)?
| Condition | Success Rate | Final Cooperators |
|---|---|---|
| Clustered | 100% | 49 (from 25) |
| Dispersed | 0% | 0 (extinction) |
Key Finding: Spatial clustering is necessary for cooperation to emerge from rare. Clustered cooperators doubled their numbers; dispersed cooperators went extinct in every trial.
Significance: This is strong evidence for the Trust Attractor claim that structure matters. In unstructured populations, defection dominates. In structured populations with local interaction, cooperation can invade and persist because clustering couples fates.
Experiment 13: Altruistic Punishment (Strong Reciprocity)
Question: Does altruistic punishment (paying a cost to punish defectors) enable cooperation where pure cooperation fails? The concepts of strong reciprocity and altruistic punishment, and the second-order free-rider problem they raise, come from the experimental and theoretical literature on human cooperation (Fehr & Gächter 2002; Bowles & Gintis 2004).
Method: Spatial public goods game with punisher strategy. Tested dispersed and clustered initialization.
Results:
| Condition | Final Cooperation | Invades? |
|---|---|---|
| Dispersed cooperators | 0.5% | NO |
| Dispersed punishers | 0.5% | NO |
| Clustered cooperators | 0.8% | NO |
| Clustered punishers | 99.8% | YES |
Key Findings:
Punishment + clustering = successful invasion. Clustered punishers achieved near-total dominance (99.8%) from 5% initial. Clustered cooperators alone failed (0.8%).
Second-order free-rider problem confirmed. When cooperators and punishers compete together in a single mixed population, the final counts were 121 punishers, 181 cooperators, and 1 defector. Cooperators exploit the “public good” of punishment provided by punishers: they reap the cleared field without paying the punishment cost, so they outnumber the punishers who maintain it.
No punishment threshold detected, but the test could not have found one. Unlike the Hamilton rule, punishment effectiveness showed no sharp threshold here. This is a null result the design cannot support a strong reading of: the test used dispersed initialization, and the invasion result above shows punishment succeeds only from clustered starts. A proper threshold test would require a sweep over punishment cost and intensity from clustered initialization; that experiment has not been run, so no claim is made here about whether such a threshold exists.
Significance for the Trust Attractor:
This result is nuanced support for the Trust Attractor thesis:
- Punishment can bootstrap cooperation from rare (with clustering), supporting the claim that coordination mechanisms enable trust emergence
- Punishment creates second-order free-riding, making it an unstable long-term solution
- Genuine value alignment (a > c/b) provides more stable cooperation than coercion-based enforcement
Enforcement mechanisms (punishment, monitoring, sanctions) can help establish cooperation initially, but sustainable cooperation requires genuine interest alignment: coordination by invitation.
Experiment 14: Network Topology Effects
Question: How does network structure affect cooperation dynamics?
Method: Tested cooperation on four network types (200 nodes, 5 trials each): - Regular lattice (k=8 neighbors) - Random (Erdős-Rényi, p=0.04) - Scale-free (Barabási-Albert, m=4) - Small-world (Watts-Strogatz, k=8, p=0.1)
Results:
| Topology | Dispersed (50% start) | Clustered Invasion (5% start) |
|---|---|---|
| Random | 61.3% ± 31% | 30% (40% success) |
| Small-world | 14.1% ± 27% | 50% (80% success) |
| Lattice | 0.7% | 77% (100% success) |
| Scale-free | 0.7% | 1% (0% success) |
Key Findings:
- Structure matters differently for maintenance vs
invasion:
- Lattice enables invasion (100% success) but not maintenance (0.7%)
- Random enables maintenance (61%) but unreliable for invasion
- Small-world is intermediate at both tasks
- Scale-free is worst for cooperation:
- Hubs are vulnerable to defector conversion
- Once a hub defects, cooperation collapses in its neighborhood
- 0% invasion success despite trying from clustered start
- Regular structure protects clusters:
- On lattice, cooperator clusters form defensive boundaries
- Neighbors share similar strategies, creating local positive feedback
- This explains the 100% invasion success from clustered start
Significance for the Trust Attractor:
Network topology affects whether cooperation can emerge and persist: - Hierarchical networks (scale-free) are vulnerable to defection cascades - Flat networks (lattice, random) better support cooperation through distributed redundancy - This suggests decentralized trust networks are more robust than hub-and-spoke models
The finding that lattice enables invasion while random enables maintenance suggests hybrid structures may be optimal: regular local connections for cluster protection, plus random long-range links for cooperation spread.
References
Axelrod, R. (1984). The Evolution of Cooperation. Basic Books.
Hamilton, W.D. (1964). The genetical evolution of social behaviour I & II. Journal of Theoretical Biology, 7, 1-16 and 17-52.
Molander, P. (1985). The optimal level of generosity in a selfish, uncertain environment. Journal of Conflict Resolution, 29(4), 611-618.
Nowak, M.A., & May, R.M. (1992). Evolutionary games and spatial chaos. Nature, 359, 826-829.
Nowak, M.A., & Sigmund, K. (1992). Tit for tat in heterogeneous populations. Nature, 355, 250-253.
Nowak, M.A., & Sigmund, K. (1993). A strategy of win-stay, lose-shift that outperforms tit-for-tat in the Prisoner’s Dilemma game. Nature, 364, 56-58.
Wilson, D.S., & Sober, E. (1994). Reintroducing group selection to the human behavioral sciences. Behavioral and Brain Sciences, 17(4), 585-608.
Bowles, S., & Gintis, H. (2004). The evolution of strong reciprocity: cooperation in heterogeneous populations. Theoretical Population Biology, 65(1), 17-28.
Fehr, E., & Gächter, S. (2002). Altruistic punishment in humans. Nature, 415(6868), 137-140.
Wolfram, S. (2026). Games between Programs: The Ruliology of Competition. Stephen Wolfram Writings, 4 June 2026. https://writings.stephenwolfram.com/2026/06/games-between-programs-the-ruliology-of-competition/