Bold Conjectures
Source: Mark Spitznagel, *Safe Haven: Investing for Financial Storms*, Chapter 6, “Bold Conjectures”; original teaching treatment with recalculated explanations and interactive reconstructions of all eight chapter figures.
The enterprise problem and today’s slice
Canonical safe havens are often accepted by reputation—cash is stable, Treasuries rally in panics, trend followers find crisis alpha, and gold protects against disorder. The consequence is a portfolio built from labels whose historical protection may be too costly, too weak, or too dependent on a vanished regime.
Enterprise problem: an allocator or operator needs to compare real protective candidates on one falsifiable portfolio scoreboard while preserving uncertainty about whether their past payoff will recur.
Whole-course context: earlier days derived geometric compounding and the one-path problem, then built a taxonomy and cost-effectiveness plane; today subjects real-world candidates to that framework.
Today’s slice: Chapter 6 evaluates Treasury bills, longer Treasuries, commodity-trading-advisor strategies, and gold, examines payoff warping, and ends with a safe-haven frontier.
End-of-day evidence: you will operate all eight source-figure labs, reproduce the reported cost and net coordinates, and write a provisional accept-or-reject record that separates book benchmarks from illustrative scenarios.
Still unsolved: historical corroboration cannot guarantee a future payoff; implementation costs, taxes, liquidity, mandate constraints, and individual suitability remain outside these teaching models.
Key terms for the empirical contest
Real-world tests become misleading when “safe,” “worked,” and “significant” are left undefined. These terms make the chapter's claim, data transformation, and rejection rule inspectable by a beginner.
| Term | Plain meaning |
|---|---|
| SPX | The S&P 500 index return series used as the risky portfolio proxy in the chapter |
| Payoff profile | Candidate return summarized conditionally on the stock-market state |
| Return bin | One of five SPX annual-return ranges used to group candidate returns from the same years |
| Bootstrap | Resampling observed annual outcomes to construct many synthetic multi-year paths |
| CAGR | Compound annual growth rate, the constant annual rate connecting starting and ending wealth |
| 5th-percentile CAGR | A lower-tail growth result that only 5% of modeled paths fall below |
| Arithmetic cost | Change in the combined portfolio's arithmetic average return versus SPX alone |
| Geometric effect | Compound-growth benefit obtained by reducing damaging paths |
| Net portfolio effect | Geometric effect after arithmetic cost; equivalently, the combined portfolio's geometric CAGR difference versus SPX |
| Mechanical payoff | Protection produced by a contractual or structural mechanism, conditional on its terms being honored |
| Statistical payoff | Protection observed as a historical tendency that is not required to recur |
| Basis risk | Risk that the protection fails to track the actual exposure being protected |
| Warping | Change in a conditional payoff profile across eras or environments |
| Rejection region | Results sufficiently below the baseline that the chapter rejects the stated cost-effectiveness hypothesis |
| Frontier | A line of candidates providing a matched amount of lower-tail mitigation at different arithmetic costs |
“Fails to reject” is intentionally weaker than “proves.” A candidate that passes one historical test has survived that test only; a new return, a better model, or a changed payoff can move it.
The chapter’s argument and experimental design
A real asset can outperform for reasons unrelated to protection, so observing a high return cannot establish a safe-haven mechanism. The chapter starts from a prior hypothesis and asks whether the predicted portfolio consequence appears in historical experiments.
The null hypothesis is that adding a candidate lowers lower-tail risk and, through that reduction, raises the portfolio's geometric or median compound growth. The test uses the same target for each candidate: increase the modeled 25-year 5th-percentile CAGR from 2.7% for SPX alone to 4.8%. The candidate allocation is varied until that target is reached, with a maximum allocation of 50%. If the candidate cannot reach the target by the cap, the test uses the capped weight and records the shortfall.
The source method groups each candidate's annual returns by the simultaneous SPX annual return:
| Bin | SPX annual-return state | Why it matters |
|---|---|---|
| 1 | Below -15% | Severe stock loss; the primary protection state |
| 2 | -15% to 0% | Ordinary down market |
| 3 | 0% to 15% | Moderate positive market |
| 4 | 15% to 30% | Strong market |
| 5 | Above 30% | Very strong market |
The candidate's return range and mean within each simultaneous bin create an empirical conditional payoff profile. The source then resamples annual stock outcomes and draws a candidate outcome from the corresponding bin, building 10,000 synthetic paths of 25 years. SPX uses the same long history across candidate comparisons; candidate histories differ according to data availability.
This is a disciplined comparison, not a time machine. Resampling assumes the empirical bins are informative, and annual observations conceal intrayear liquidity and timing. The labs below recreate the teaching relationships and exact reported summary coordinates, not the book's proprietary raw bootstrap paths. Every moved control is therefore labeled as an illustrative calculation.
Why Neptune and Vulcan matter
A payoff can change because a real hidden variable moved it, or because the original relationship was never reliable. The consequence is that a failed test may reject a useful mechanism, while a passed test may preserve an impostor.
The chapter borrows an astronomy analogy. An unexplained deviation in Uranus's orbit led to the discovery of Neptune: the underlying gravitational theory survived after adding a real missing body. A similar conjecture about a planet Vulcan did not explain Mercury; the theory itself needed revision. In safe-haven analysis, inflation, real interest rates, policy, positioning, market structure, and valuation can be Neptune-like variables. A post hoc story with no stable causal role can be Vulcan.
This creates two costly error classes. A type I error rejects a candidate whose mechanism remains sound under a missing condition. A type II error retains a candidate because a story excuses evidence that should have killed the hypothesis. The latter can be especially dangerous: protection is relied on in a bad state and fails there.
Mechanical payoffs reduce this ambiguity but do not eliminate it. A contractual option has a state-contingent formula, yet counterparty default, trading halts, basis mismatch, expiration timing, and implementation price still matter. Treasuries, CTA strategies, and gold rely substantially on statistical relationships. Their profile can look protective without being compelled to remain so.
Figure 1 lab — three-month Treasury-bill profile
Cash-like Treasury bills are expected to hold value rather than explode in a crash, so their main risk is the large allocation required to improve the lower tail. The first figure shows that a stable component can reduce loss and still lower the portfolio's compound growth.
Purpose and axes. The horizontal axis contains the five annual SPX return bins, from below -15% to above 30%. The vertical axis is conditional annual return in percentage points. A dashed square line shows SPX bin averages, a dotted circle line shows the three-month Treasury-bill profile, and a solid diamond line shows their blended portfolio.
Controls. Allocation scale multiplies the source 50% bill weight; 100% of book size means the source comparison weight, not a portfolio recommendation. Regime stability controls how much of the displayed conditional shape is retained. At 100%, metrics are source benchmarks. Any lower setting compresses the traced payoff around its mean and becomes an illustrative stability challenge.
Protocol. Reset first. Record the source allocation, arithmetic cost, geometric effect, and net effect. Reduce allocation scale while leaving regime stability fixed and observe the crash-bin loss and ordinary-state return together. Restore allocation, reduce regime stability, and record whether a flatter payoff changes the net effect. Finally, increase allocation above the source size and ask whether more apparent safety raises or lowers the cost paid in all states.
Worked interpretation. The book reports a 4.8% standalone bill arithmetic average. A 50% SPX + 50% bill portfolio has an 8.1% arithmetic average and 7.7% geometric average, versus 11.4% and 9.5% for SPX. Thus the exact arithmetic difference is -3.3 percentage points, while the exact net geometric difference is -1.8. The implied geometric effect is +1.5, insufficient to offset the arithmetic cost. The bill also failed to reach the 4.8% lower-tail target before the 50% cap.
Assumptions and non-inferences. The plotted bin means are source-shaped teaching coordinates; they are not reconstructed Treasury observations. The comparison ignores current yields, inflation, tax, reinvestment terms, default conventions, currency, and liability matching. It does not say Treasury bills are unsafe in nominal terms. It says the source strategic SPX comparison rejected their cost-effectiveness under its stated objective and period.
Figure 2 lab — Treasuries across the yield curve
Longer maturity makes Treasury prices more sensitive to rate changes and historically more responsive to flight-to-quality demand. The consequence is a more protective crash profile coupled to greater regime risk from growth and inflation expectations.
Purpose and axes. The horizontal axis again uses the five SPX annual-return bins. The vertical axis is the conditional annual return of the blended SPX/Treasury portfolio in percent. Circle, square, and diamond markers distinguish three-month bills, ten-year notes, and twenty-year bonds; dash patterns preserve the distinction without color.
Controls. Allocation scale multiplies each source comparison allocation: 50% for bills, 37% for ten-year notes, and 34% for twenty-year bonds. Regime stability compresses each source-shaped Treasury profile toward its own mean. Only the reset state displays book-benchmark provenance.
Protocol. Reset and compare crash-bin height, allocation, and net effect across maturity. Hold stability at 100% and scale all allocations down to test whether the maturity ordering remains. Restore allocation and reduce stability to challenge the flight-to-quality shape. Name the candidate that loses least geometric CAGR, then state why “least negative” is a relative result rather than proof of an absolute haven.
Worked interpretation. The source bill case has a -3.3 arithmetic cost and -1.8 net effect. The ten-year case uses 37%, with a 7.4% standalone arithmetic average, 9.9% combined arithmetic average, 9.1% combined geometric average, -1.5 cost, and -0.4 net. The twenty-year case uses 34%, with 8.4% standalone arithmetic, 10.4% combined arithmetic, 9.4% combined geometric, -1.0 cost, and -0.1 net. Moving outward improves the historical scoreboard, but the full-period twenty-year result sits approximately on the source rejection boundary rather than clearly above it.
Assumptions and non-inferences. Maturity is not the only changing variable; coupon, duration, starting yield, inflation, policy, and sample composition differ. The teaching profiles approximate the visual relationships, not tradable bond indexes. A liability-matching investor may rationally hold a Treasury even when this SPX-relative CAGR test is negative, because the objective and protected exposure differ.
Figure 3 lab — the warped twenty-year Treasury
A full-period average can hide eras in which the same candidate changes from helpful to harmful. The consequence is a strategic allocation whose apparent protection depends on selecting the right regime in advance.
Purpose and axes. The horizontal axis contains the five SPX bins and the vertical axis shows blended annual return in percent. The three lines correspond to 1973–1988, 1989–2004, and 2005–2020. The conceptual SPX resampling baseline is held fixed while only the twenty-year Treasury conditional profile changes.
Controls. Allocation scale multiplies the source 34% twenty-year bond weight in each era. Regime stability compresses each era's conditional shape toward its mean; it is a sensitivity device, not a probability that an era repeats. Reset restores the three source net values.
Protocol. Reset and record all three net effects before moving a control. Compare only the crash-bin position first, then inspect the other four bins where cost accumulates. Lower regime stability and note whether apparent era differences shrink. Restore stability, vary allocation, and determine whether any single weight makes all three source-shaped eras positive. End by writing two competing explanations: one named Neptune-like condition and one possibility that the relationship is mostly noise.
Worked interpretation. The book reports period net portfolio effects of -0.8, +0.6, and -0.9 percentage points. Their arithmetic costs are -1.6, -0.4, and -2.1; combined geometric CAGRs are 8.7%, 10.1%, and 8.6%. The sign reversal occurs even though the SPX bootstrap distribution is held constant in the chapter's experiment. A single full-period -0.1 net therefore conceals instability large enough to change the decision.
Assumptions and non-inferences. Three 16-year splits produce small samples inside five bins. The boundary dates are a teaching partition, not discovered regimes, and the displayed conditional points are approximate visual reconstructions. The lab cannot identify whether inflation, starting yields, policy, or noise caused the warping. Selecting the positive middle window after seeing it would be hindsight, not a deployable rule.
Figure 4 lab — CTA crisis-alpha profile
Commodity trading advisors, or CTAs, often use trend-following rules that can become long or short after price movement persists. The attraction is return-seeking protection, but the consequence is lag: an abrupt reversal or short crisis may arrive before a trend system builds the needed position.
Purpose and axes. The x-axis contains SPX annual-return bins and the y-axis is conditional annual return in percent. The CTA line slopes upward from a relatively strong severe-down-market bin, while the blended portfolio shows the result of combining the source CTA index with SPX.
Controls. Allocation scale multiplies the source capped 50% CTA weight. Regime stability weakens the conditional crisis-alpha shape toward a flatter payoff. Both controls generate illustrative scenarios; the exact source metrics appear only after Reset.
Protocol. Reset and compare the CTA's severe-down-market return with its results in the other bins. Record the 50% allocation, arithmetic cost, geometric effect, and net. Reduce allocation to test whether lower carry cost also removes too much downside response. Restore allocation and reduce regime stability to model weaker alignment between trends and stock crises. State whether any attractive standalone CTA return is relevant without the combined portfolio result.
Worked interpretation. The source CTA index has a 4.0% standalone arithmetic average. At 50%, the combined portfolio arithmetic average is 7.7% and geometric average is 7.3%. Against SPX, the exact arithmetic cost is -3.7 and net geometric effect is -2.2, implying a +1.5 geometric benefit that does not cover the cost. The source allocation also fails to reach the required 4.8% 5th-percentile CAGR before the cap.
Assumptions and non-inferences. CTA programs differ by market universe, horizon, volatility scaling, fees, leverage, execution, and model changes. The lab is not a backtest of a current fund and does not assert all trend following is identical. It demonstrates that a negative crash correlation is insufficient when sizing and standalone return make the whole portfolio compound worse.
Figure 5 lab — gold's full-period profile
Gold is treated as monetary insurance because it can respond to fear, real rates, inflation expectations, and distrust of financial institutions. The consequence is a noisy conditional payoff with no contractual requirement to rise when stocks fall.
Purpose and axes. The x-axis is the five SPX annual-return bins and the y-axis is conditional gold or combined-portfolio annual return in percent. The gold severe-loss bin is visibly higher than the corresponding Treasury and CTA profiles. The source figure also shows wide within-bin ranges; the lab focuses on the mean relationship and explains that omitted dispersion explicitly.
Controls. Allocation scale multiplies the source 20% gold weight. Regime stability compresses the source-shaped conditional profile toward its mean, challenging how much crash alignment survives. Reset is the only state whose summary metrics are marked as book benchmarks.
Protocol. Reset and record the four summary metrics. Compare the severe-loss bin with the remaining bins rather than looking only at gold's total return. Reduce allocation and observe whether the positive net persists when less crash payoff reaches the portfolio. Restore allocation, weaken regime stability, and find the point at which the illustrative net crosses zero. Record that crossing as sensitivity evidence, not as an estimated forecast.
Worked interpretation. For 1973–2020, the source reports gold's standalone arithmetic average as 9.1%. An 80% SPX + 20% gold portfolio has 10.9% arithmetic and 9.8% geometric averages. The exact arithmetic cost versus SPX is -0.5, while the exact net geometric effect is +0.3; the implied geometric benefit is +0.8. Under the full-period test, the hypothesis is not rejected. The source also reports severe-down-market gold outcomes spanning roughly +5% to +70%, a warning that the mean is not a guaranteed crash payoff.
Assumptions and non-inferences. The chart starts after gold became freely floating and ends in 2020. It ignores storage, custody, tax, currency, spreads, and the distinction between bullion, futures, miners, and funds. It does not establish gold as an inflation hedge, a permanent strategic haven, or a suitable allocation. The later era split is essential evidence, not an optional footnote.
Figure 6 lab — the warped gold profile
Gold's positive full-period result can be dominated by one early era, so the aggregate can preserve a reputation that later observations do not support. The consequence is a strategic claim that quietly requires a tactical inflation or real-rate forecast.
Purpose and axes. The horizontal axis contains SPX annual-return bins and the vertical axis shows the blended SPX/gold annual return in percent. Three lines represent 1973–1988, 1989–2004, and 2005–2020, all using the source 20% gold allocation at Reset.
Controls. Allocation scale changes that common comparison weight. Regime stability compresses each period profile toward its period mean, testing dependence on the conditional shape. It does not interpolate the macroeconomic world or predict which window comes next.
Protocol. Reset and record the three exact net values. Identify which period creates the full-sample positive result. Compare the crash bin and ordinary bins for that era, then repeat for the later eras. Lower regime stability and observe whether the first era's advantage survives. Vary allocation and ask whether one fixed strategic weight produces a positive result in all three displayed eras.
Worked interpretation. The source net portfolio effects are +1.5, -1.1, and -0.1 percentage points. Arithmetic differences versus SPX are +0.7, -2.0, and -0.8; combined geometric CAGRs are 11.0%, 8.4%, and 9.4%. Gold's strong 1973–1988 record offsets two later windows that fail the strategic cost-effectiveness test. Calling the entire period positive is mathematically correct and operationally incomplete.
Assumptions and non-inferences. The split is small, annual, and retrospective. No causal variable is fitted, and the lab does not turn inflation, real rates, or central-bank policy into a timing rule. A tactical hypothesis would need a prior signal, trading costs, and out-of-sample validation. Without that, “gold works in the right regime” merely restates the observed result.
Figure 7 lab — all candidates on the cost-effectiveness plane
Separate payoff charts are difficult to compare because allocation, carry, and crash response differ. The cost-effectiveness plane reduces each experiment to the two forces that determine the net result.
Purpose and axes. The horizontal axis is arithmetic cost in percentage points, shown as a positive amount paid relative to SPX. The vertical axis is geometric effect in percentage points. The diagonal effect = cost is the SPX break-even baseline. A point above it has positive net effect; a point below it lies in the rejection region.
Controls. Allocation scale moves each point outward or inward by changing both cost and effect. Regime stability retains less of the candidate's geometric effect while leaving its arithmetic cost explicit. The control-generated positions are illustrative; Reset restores the source coordinates for bills, notes, bonds, CTA, gold, and the three Chapter 5 cartoon prototypes.
Protocol. Reset and locate gold, the three Treasuries, and CTA relative to the diagonal. Calculate net = effect - cost for one point directly from its coordinates. Compare the real candidates with the cartoon store-of-value, alpha, and insurance points. Lower regime stability and observe which points cross the diagonal first. Then reduce allocation scale and explain why moving a negative-net point toward the origin makes the error smaller without making it cost-effective.
Worked interpretation. Three exact source calculations anchor the plane. Treasury bills have cost 3.3, effect 1.5, and net -1.8. CTA has cost 3.7, effect 1.5, and net -2.2. Gold has cost 0.5, effect 0.8, and net +0.3. The longer Treasuries sit closer to break-even: the ten-year note at approximately cost 1.5, effect 1.1, net -0.4; the twenty-year bond at cost 1.0, effect 0.9, net -0.1.
Assumptions and non-inferences. Collapsing a path distribution into two coordinates loses information about tail depth, estimation error, liquidity, and regime stability. Distance from the diagonal is not a confidence interval. Points from different histories are placed together for a common decision question, not because their data quality and implementation are identical.
Figure 8 lab — the safe-haven frontier
A binary pass/fail decision hides useful relative progress among imperfect candidates. The frontier asks which payoff produces a matched lower-tail improvement with the least arithmetic cost.
Purpose and axes. The axes remain arithmetic cost and geometric effect in percentage points. The solid diagonal is SPX break-even. The dotted frontier passes through the three idealized cartoon payoffs that reach the source 4.8% 5th-percentile CAGR target from a 2.7% unmitigated baseline. Labeled points are source case studies; the additional diamonds are deterministic illustrative candidates that recreate the chapter's “graveyard” relationship without claiming to transcribe its roughly forty assets.
Controls. Allocation scale changes the amount of each labeled source-shaped candidate. Regime stability weakens geometric effects and pushes illustrative candidates downward. The frontier itself remains a visible benchmark; a moved point is a scenario, not new book evidence.
Protocol. Reset and trace the frontier from the store-of-value cartoon toward insurance. Observe that moving left lowers arithmetic cost while retaining the matched downside target. Identify real candidates below the line and explain that their allocations delivered less lower-tail improvement for the cost. Reduce regime stability and watch the candidate set fall away from the frontier. Finally, compare a point's distance from the diagonal with its distance from the frontier: the first answers absolute net effect, the second answers matched-downside efficiency.
Worked interpretation. The source frontier is not a generic efficient frontier of mean and volatility. It holds the 5th-percentile CAGR improvement fixed. More explosive, conditionally timed payoffs require smaller allocations, reducing the arithmetic drag paid during normal markets. The insurance cartoon can therefore occupy the left end despite a low standalone average. It protects the part of the path where logarithmic damage is greatest with a small weight.
Assumptions and non-inferences. The dotted line is a pedagogical relationship anchored to the chapter's prototypes and bootstrap target. The unlabeled points in this lab are synthetic, not asset claims. The frontier does not provide an allocation, forecast returns, or guarantee that a mechanically attractive hedge can be bought at the assumed price. A real frontier must be rebuilt for the actual exposure, mandate, data, and costs.
What the eight figures establish together
One figure can be dismissed as a special case, but the sequence reveals a stable decision logic. The consequence is a more useful conclusion than a ranking of popular assets.
The Treasury-bill test shows that stability can be too expensive. Longer Treasuries move toward a crash-sensitive profile, but their result is marginal and warped across time. CTA trend following produces a recognizable crisis shape, yet its low standalone average and large required allocation overwhelm the geometric effect in the source test. Gold passes the full-period test narrowly, then fails as a stable strategic claim when its history is split.
The combined plane shows why. Cost-effective protection depends not on a prestigious name or even on a negative crash correlation, but on the ratio of conditional effect to the allocation and cost required. The frontier then turns that ratio into a direction for improvement: mitigate the worst paths with less capital and less ordinary-state drag.
Three chapter conclusions follow:
- Cost-effective strategic protection is difficult; most canonical candidates in the source comparison fail the absolute test.
- Historical successes can be temporary hybrids of strategic and tactical behavior because their payoff profiles warp.
- A meaningful frontier still exists: payoff shape can lower the worst paths with sufficiently little cost to raise median compound wealth.
The graveyard of rejected hypotheses is progress. A rejected Treasury or CTA test does not make the asset useless. It sets a better baseline, clarifies the objective it did not meet, and redirects the search toward a more conditional or less costly payoff.
Economics application: policy regimes warp defensive assets
Macroeconomic labels fail when inflation, growth, and policy move the same asset in opposing directions. The consequence is a “safe” public or pension portfolio whose bond ballast works in one regime and compounds losses in another.
Consider a pension with equity exposure and nominal liabilities. Long Treasuries may rally during a deflationary recession as growth and rates fall. In an inflationary sell-off, stock prices and long-bond prices can fall together. A full-period correlation averages those mechanisms and can conceal the very state in which liabilities and collateral become difficult to fund.
Apply the chapter's method before choosing a label. Define the liability-relevant bad state, bucket candidate outcomes by that state, and preserve the period splits. Match lower-tail funding improvement before comparing cost. Then write the Neptune variables—starting yield, inflation, real rates, policy reaction—and the Vulcan falsifier that would show the story is merely retrospective. This is a framework for institutional analysis, not a current bond recommendation.
Startup application: build a runway frontier
Founders often choose between idle cash, diversified revenue, and insurance-like contractual protection without a common scoreboard. The consequence is either excessive reserve that starves growth or concentrated exposure that ends the company after one shock.
Define the risky “SPX” analogue as the startup's core growth plan and the lower-tail target as months of runway after a severe sales, platform, or financing shock. Candidate havens might include cash, a committed credit line, annual prepaid contracts, a second distribution channel, staged hiring, cloud-spend caps, or key-person coverage. Each has an arithmetic cost in normal months and a conditional effect in the bad state.
Construct a startup frontier by sizing each candidate to the same minimum-runway target. A large cash reserve may reach it but impose a high growth cost. A smaller contract that releases cash only after a named disruption may be more efficient, provided counterparty and basis risk are controlled. Test across financing and demand regimes; a facility that disappears when markets close is a statistical hopeful haven, not mechanical protection.
Business application: compare continuity controls as whole-system payoffs
Operational resilience budgets are frequently allocated by compliance category rather than by the loss they prevent. The consequence is multiple expensive controls that all fail against one shared bottleneck.
For a manufacturer, compare inventory reserve, dual sourcing, business-interruption insurance, and flexible production capacity. For a software service, compare spare capacity, multi-region failover, rollback, incident staffing, and contractual liability coverage. Translate every control into the same units: ordinary operating cost, conditional loss avoided, recovery time, and residual failure state.
Match a downside outcome first—for example, no more than two days of unavailable delivery—then compare cost. A dual supplier may sit below the frontier if it shares the same port. A tested offline restore may move left because a small recurring drill prevents a large irreversible data loss. Periodically split results by architecture or vendor era to detect warping after systems change.
Daily-life application: protection should preserve choices
Personal decisions have one lived path and incomplete probabilities, so importing asset rankings is especially hazardous. The consequence is paying for a reputation instead of protecting the household obligation that matters.
Start with the bad state: income interruption, urgent repair, caregiving need, disability, or a time-critical move. Cash is a store of value with an opportunity cost. Insurance is conditional but contains exclusions and counterparty terms. A flexible contract, transferable skill, or reliable support network can also preserve a future choice, though its payoff is difficult to price.
Size candidates to the same practical floor—months of essential expenses, time available, or ability to remain housed—then compare recurring cost. Stress inflation and access, not just nominal account value. The framework does not prescribe cash, bonds, gold, or any percentage. It asks whether the chosen protection works for the named household state and remains affordable through ordinary years.
A reviewable safe-haven experiment
An analysis cannot be challenged if its data, target, and rejection rule are chosen after the result. The consequence is a polished backtest that only records the researcher's preferences.
Write a one-page experiment record before running a candidate:
| Field | Required entry |
|---|---|
| Protected whole | Portfolio, liability, runway, or operating capability |
| Bad state | Observable event and loss threshold |
| Candidate mechanism | Mechanical terms and statistical assumptions |
| Data window | Start, end, frequency, missing observations, and reason |
| Conditional bins | Predeclared states used to pair exposure and candidate |
| Downside target | Same lower-tail outcome used for all candidates |
| Allocation cap | Maximum acceptable exposure and why |
| Cost measure | Carry, opportunity cost, fees, liquidity, and implementation |
| Rejection rule | Portfolio consequence that causes the claim to be rejected |
| Regime challenge | At least two prior period splits and a causal hypothesis |
| Non-inference | What the experiment cannot establish |
Operate the eight labs only after completing the first four rows. Copy source metrics from Reset, then put every moved-control result in a separate “illustrative scenario” column. A reviewer should be able to reproduce the direction of the conclusion and identify where an omitted mechanism could reverse it.
Limitations and model risk
The chapter's scoreboard is more economically meaningful than a standalone-return ranking, but it is not complete. The consequence of treating it as final truth would violate the falsification discipline it teaches.
Annual binning loses sequence inside the year, drawdown timing, path-dependent rebalancing, and crisis liquidity. A bootstrap treats the historical support as available for reuse and may understate outcomes never observed. Five bins trade resolution for sample size; some contain few candidate observations. Different candidates have different histories, and period splits make those samples smaller still.
The 5th percentile is a proxy, not a worst case. It says nothing about the mean loss beyond the cutoff, and its estimate is fragile in fat tails. Median CAGR does not capture consumption needs, collateral calls, taxes, fees, leverage limits, or legal mandates. SPX is one risky exposure; another portfolio has different bad states and basis risk.
The lab controls deliberately avoid pretending to solve those problems. Allocation scale and regime stability generate deterministic teaching relationships. They do not rerun the book's historical bootstrap, and their smooth response is not evidence that real portfolios change smoothly. Exact book values are shown only at Reset and explicitly marked; every alternative is illustrative.
Finally, a source-era failure does not mean an asset has no legitimate role. Treasury bills can fund near-term liabilities; long bonds can match duration; CTAs can diversify other exposures; gold can serve preferences unrelated to SPX CAGR. The test rejects a specific strategic cost-effectiveness claim under a specific objective and dataset.
Sources and further study
The chapter's historical claims require the original book and the scientific method that motivates its provisional language. These sources support the ideas; they do not validate a current allocation.
- Mark Spitznagel, Safe Haven: Investing for Financial Storms, Wiley, 2021, Chapter 6.
- Karl Popper, The Logic of Scientific Discovery, for falsifiability, provisional corroboration, and the limits of naïve falsification.
- Richard Feynman's account of guessing, deriving consequences, and comparing them with experiment, used in the chapter as the research sequence.
- Daniel Bernoulli's geometric wealth framework and the Petersburg merchant example, developed earlier in the book and reused here as the mechanism behind cost-effective protection.
The eight graph relationships were visually inspected in the user-provided PDF at physical pages 191, 194, 196, 198, 200, 202, 204, and 206. Reported cost and net values are preserved as source benchmarks. Conditional line coordinates and the unlabeled frontier candidate set are deterministic teaching reconstructions, not extracted raw data.
Key takeaways
Chapter 6 replaces the folklore of safe assets with a provisional empirical contest. The durable result is not that one asset always wins, but that the whole portfolio must earn more geometric effect than the arithmetic cost required to protect its lower tail.
- A real safe-haven claim starts from a prior mechanism and a falsifiable portfolio consequence.
- Treasury bills are stable but costly in the source SPX comparison; longer maturity reduces the drag while adding regime sensitivity.
- CTA crisis alpha has a protective shape but fails the source cost-effectiveness test at the capped allocation.
- Gold narrowly passes the full-period test and then reveals substantial era dependence.
- Mechanical payoffs are generally more dependable than statistical associations, but implementation and basis risks remain.
- Warping can create both mistaken rejection and mistaken acceptance; period splits expose instability without identifying its cause.
- The cost-effectiveness plane separates arithmetic cost, geometric effect, and net result.
- The safe-haven frontier holds downside improvement constant and rewards protection achieved with less capital and less carry.
- Most rejected candidates can still have other valid roles; the rejection belongs to the stated objective and experiment.
Checklist
Mastery means being able to reconstruct the chapter's reasoning without promoting a historical winner into a permanent haven. Complete every item with explicit units and provenance.
- [ ] I can define SPX bins, bootstrap paths, median CAGR, and 5th-percentile CAGR in plain language.
- [ ] I can state the null hypothesis and the observation that rejects it.
- [ ] I can distinguish arithmetic cost, geometric effect, and net portfolio effect.
- [ ] I can reproduce the T-bill
-3.3 / -1.8, CTA-3.7 / -2.2, and gold-0.5 / +0.3source coordinates. - [ ] I can explain the longer-Treasury maturity comparison without calling marginal improvement proof.
- [ ] I can report the warped Treasury net sequence
-0.8 / +0.6 / -0.9. - [ ] I can report the warped gold net sequence
+1.5 / -1.1 / -0.1. - [ ] I can distinguish a Neptune-like missing condition from a Vulcan-like rescue story.
- [ ] I can operate all eight labs and keep Reset benchmarks separate from illustrative scenarios.
- [ ] I can read absolute break-even and matched-downside frontier distance as different questions.
- [ ] I can transfer the method to economics, startups, business continuity, and daily life with a named protected whole.
- [ ] I can list the sampling, liquidity, basis, counterparty, and regime limits before making a decision.