The Case Against the Modern Theory of Finance
Source: Benoit Mandelbrot and Richard L. Hudson, *The (Mis)Behaviour of Markets*, Chapter 5, “The Case Against the Modern Theory of Finance”; original teaching treatment with further sources below.
The enterprise problem and today’s slice
Organizations often validate calculations while leaving the model’s representation of reality untested. Correct arithmetic cannot rescue a decision built on thin tails, stable parameters, or frictionless actions that the operating environment does not provide.
Enterprise problem: leaders need a systematic way to detect model misspecification before exposure concentrates around a false assumption.
Whole-course context: after constructing survival, dependence, diffusion, and portfolio baselines, the course now audits the foundations those models share.
Today’s slice: this chapter tests Gaussian tails, independence, stationarity, continuity, liquidity, and parameter certainty against empirical behavior.
End-of-day evidence: you will quantify a tail mismatch, operate comparison and error-surface labs, and write a model card tied to a decision.
Still unsolved: rejecting one model does not identify a unique replacement; later chapters develop turbulence, roughness, memory, and multifractal time.
Key terms for model criticism
Criticism becomes useful only when it identifies a failed claim and consequence. These terms separate model error from ordinary sampling noise.
| Term | Plain meaning |
|---|---|
| Model misspecification | The model omits or misstates a material mechanism |
| Residual | Observed value minus the value explained by a fitted model |
| Stylized fact | Empirical pattern recurring across many instruments or periods |
| Excess kurtosis | More mass near the center and tails than a matching Gaussian |
| Volatility clustering | Large magnitudes tending to occur near other large magnitudes |
| Regime change | Shift in the process generating observations |
| Jump | Discontinuous move not captured by a continuous diffusion |
| Liquidity | Ability to transact size without unacceptable delay or price impact |
| Backtest | Historical simulation of a rule or model |
A failed prediction is not automatically misspecification; finite samples vary. Evidence strengthens when several diagnostics fail in related ways and the mismatch persists out of sample.
The chapter’s empirical case
Modern theory gained power by simplifying, but Mandelbrot argues that several simplifications remove market turbulence rather than approximate it. Large moves occur too often, volatility varies through time, and discontinuities defeat smooth rebalancing.
The case should not be reduced to “bell curves are bad.” Gaussian models can describe measurement error and ordinary fluctuations well. The problem is using a thin-tailed, independent, stationary model to set exposure where extremes, clusters, and changing liquidity determine survival.
The disciplined response is empirical: list assumptions, derive observable consequences, compare them with data, and measure decision sensitivity. Replacing one doctrine with “fractals explain everything” would repeat the same error.
Mechanics: compare predicted and observed exceedances
Tail mismatch becomes concrete when a model predicts an event count. Under a standard normal distribution, an absolute observation beyond three standard deviations has probability about 0.27%.
In 10,000 independent observations, the Gaussian expectation is about 10,000 × 0.0027 = 27 three-sigma exceedances. Suppose a cleaned return series contains 160. The ratio 160 / 27 ≈ 5.9 indicates the fitted Gaussian understates that threshold’s frequency.
This calculation is only a diagnostic. Volatility estimation, time variation, dependence, and multiple testing affect uncertainty. Standardizing every observation by one global variance can itself create tail excess when volatility changes. A conditional volatility model may explain part of the mismatch without proving Gaussian innovations.
The assumption audit
Every modeling assumption should map to a diagnostic and a decision consequence. An audit prevents vague dissatisfaction from becoming unfalsifiable rhetoric.
Diagnostics should be chosen before inspecting the most dramatic observations. Otherwise analysts can always find a threshold or window that supports the preferred story.
Worked numerical example: a false diversification promise
Misspecification becomes costly when several assumptions fail together. Consider a two-asset portfolio sized using 10% volatility for each asset and correlation 0.2.
With equal weights, portfolio variance is 0.25(0.10²) + 0.25(0.10²) + 0.5(0.2)(0.10)(0.10) = 0.006, giving volatility about 7.75%.
During stress, suppose each asset’s volatility doubles to 20% and correlation rises to 0.9. Variance becomes 0.25(0.20²) + 0.25(0.20²) + 0.5(0.9)(0.20)(0.20) = 0.038, giving about 19.49%.
The stressed volatility is 2.5 times the ordinary estimate. If positions were leveraged to the first number and liquidity simultaneously worsened, realized loss can exceed even the stressed variance calculation. The issue is not one bad parameter; it is a state-dependent system.
The 2D model-error lab protocol
Side-by-side curves reveal where a fitted model and observed pattern disagree. The 2D lab varies one assumption-error control against a fixed comparison.
Use the lab as an assumption test, not a market forecast:
- Press Reset and note the Assumption error percentage.
- Identify the Model and Observed curves from their labels and distinct encodings.
- Increase Assumption error one step and inspect where separation becomes largest across the input axis.
- Record the Maximum model miss metric at low, default, and high settings.
- Reset and verify that the default curves and metric are reproduced.
The simulations are deterministic teaching models. Their purpose is to expose sensitivity, not to declare the rougher line true by construction.
The 3D model-error surface protocol
Model error grows through interactions that a single backtest hides. The 3D surface maps error across tail severity and dependence under a controllable exposure.
Read it carefully:
- Read the legend: x is tail severity, y is dependence, and z is model error.
- Move the Exposure slider and compare the surface range at low and high exposure.
- Select tail severity with the x focus slider and dependence with the y focus slider; record selected x, y, and z.
- Hold one focus value fixed while moving the other to locate where exposure amplifies disagreement.
- Orbit by pointer or keyboard, then Reset the exposure, focus, and camera before another comparison.
Surface height is a calculated error metric on a finite grid. It does not measure every omitted mechanism, and visual smoothness does not imply stable real-world parameters.
Fractal lens: expose assumption error across scales
A model can match coarse observations while omitted detail expands at finer scales, causing controls to fail exactly where confidence is highest. This fractal lens draws one recursively refined coastline and measures its root-mean-square deviation from the straight line joining its endpoints; it does not display a fitted baseline model.
Parameters. Resolution depth d ranges from 2 to 8 and exposes finer coastline detail. Assumption gap g ranges from 0.10 to 1.00 and increases deterministic midpoint displacement from the endpoint line, revealing more constructed roughness.
Protocol. Reset the Model-error coastline lens and record Resolution depth, Assumption gap, Hidden error, and mean absolute change across displayed lags. Raise only Assumption gap, restore it, then add resolution levels; distinguish greater displacement at a fixed construction from detail revealed by refinement. Finish with both high and state which real omitted mechanism needs an independent test.
Assumptions. The path is a finite deterministic teaching construction with fixed endpoints and one refinement rule. It omits a fitted comparator, sampling error, structural breaks, data quality, alternative rough mechanisms, implementation failure, and feedback; the rough path is not declared true by construction.
Interpretation. Hidden error is 100 times the path's root-mean-square deviation from its endpoint line. A larger value demonstrates sensitivity inside this construction only. It does not validate a replacement model, estimate a real loss, or offer financial advice; use it to prioritize held-out tests and reversible decisions.
Source figure lab: The Old Master
A long level chart can hide most of its own history because later values occupy nearly all the vertical scale. That compression matters whenever growth makes early failures look insignificant even though people living through them experienced large relative changes.
Original argument and source trace. The book's “The Old Master” figure (printed p. 89; supplied PDF p. 211) plots daily Dow index levels from 1916 through the early-2000s bear market. Its linear nominal scale makes the 1990s rise dominate, while the 1987 fall is visible and much earlier turbulence appears compressed.
Interactive adaptation. This is a conceptual reconstruction, not historical Dow data. It starts one shared, labeled synthetic Dow proxy used by the next five figure labs. Proxy volatility changes ordinary movement, Tail intensity changes scheduled extremes, Crisis intensity changes the principal crash cluster, and Scenario seed changes bounded innovations.
Protocol and readout. Reset, record level range and largest change, then raise only Crisis intensity. Next raise Tail intensity while holding the seed fixed, and finally change the seed to distinguish a parameter effect from one path realization. The live readout reports all four inputs plus the level range, largest point change, and maximum standardized magnitudes for the proxy and its matched Brownian comparator.
Assumptions and falsifier. The upward drift, event positions, and volatility feedback are teaching choices. This lab would fail as an explanation of scale distortion if linear and logarithmic views preserved the same visual ranking of early and late proportional events, or if the derived views later in this day could not be regenerated exactly from this proxy.
| Domain | Application | Decision evidence |
|---|---|---|
| Software engineering | Plot cumulative traffic or storage beside relative change before sizing capacity. | Early incidents remain visible after log or rate normalization. |
| LLMs | Compare cumulative token volume with per-request token changes. | A model-growth story separates scale from workload volatility. |
| AI agents | Track cumulative tool calls and step-to-step bursts. | Run limits respond to burst size, not only lifetime totals. |
| Startups | Separate cumulative users from weekly growth changes. | A growth chart no longer hides early churn shocks. |
| Business | Pair nominal revenue levels with percentage change. | Planning sees operational shocks behind a larger base. |
| Daily life | Compare cumulative commitments with weekly additions and removals. | Calendar load reflects new pressure rather than only the total. |
Source figure lab: Looking Closer
Level charts answer “where did the total end up?” but can conceal “what changed today?” The consequence is a decision system calibrated to cumulative progress while short-interval instability remains unmeasured.
Original argument and source trace. “Looking closer” (printed p. 90; supplied PDF p. 212) transforms the Dow level series into one-day index-point changes. The source labels the span 1916–2003 and uses the same observations as the preceding chart, so differences—not a replacement dataset—create the new view.
Interactive adaptation. The lab derives signed first differences directly from the shared synthetic Dow proxy. Its four controls are identical to The Old Master so a selected parameter state has one reproducible meaning across level and change views; it never silently normalizes the differences into returns.
Protocol and readout. Fix the seed, note the largest signed excursion, then vary Proxy volatility, Tail intensity, and Crisis intensity one at a time. Compare this lab's displayed-largest-change readout with The Old Master’s displayed-level-range readout at the same settings; each reports only evidence visible in its own graph while both remain transformations of one proxy. Reset must restore the complete path and metric.
Assumptions and falsifier. Synthetic index-point differences are scale-dependent and are not evidence about the historical Dow. The adaptation is falsified if any plotted difference is not exactly level[t] - level[t-1], if changing a control leaves the readout unchanged, or if an apparently extreme point exists only because a second hidden series was used.
| Domain | Application | Decision evidence |
|---|---|---|
| Software engineering | Difference cumulative request and error counters. | Incident detection uses per-window changes derived from the same counters. |
| LLMs | Difference cumulative cost and token telemetry. | Budget alarms identify sudden consumption rather than a large lifetime total. |
| AI agents | Difference cumulative tool and retry counts by step. | A runaway branch is visible at the step where amplification begins. |
| Startups | Difference cumulative signups into daily net additions. | Acquisition spikes and reversals become reviewable. |
| Business | Difference inventory or cash balances. | Operational movement reconciles exactly to the balance series. |
| Daily life | Difference sleep debt, tasks, or spending totals. | Intervention targets the day creating the change. |
Source figure lab: Looking Under the Varnish
Absolute units can make equal percentage changes look unequal across decades. Without a consistent relative scale, teams may mistake a larger denominator for a more turbulent process.
Original argument and source trace. “Looking under the varnish. Two charts here:” (printed p. 91; supplied PDF pp. 214–215) couples the Dow in logarithmic level scale with daily log changes. The source uses the pair to restore the importance of 1929, the Depression, World War II, 1987, and alternating narrow and wide volatility bands.
Interactive adaptation. Both panels come from the shared synthetic Dow proxy: the top applies a natural logarithm to every positive level, and the bottom differences those logarithms. Proxy volatility, Tail intensity, Crisis intensity, and Scenario seed rebuild the one underlying path before either transform is applied.
Protocol and readout. Hold the seed fixed and locate the largest cluster in the log-change panel, then find its corresponding movement in the log-level panel. Reduce Crisis intensity to zero and test whether other tail events remain. The readout supplies the common input state and extrema; use it to verify that panel differences are transforms, not separate scenarios.
Assumptions and falsifier. Log changes approximate percentage changes only when levels stay positive, which the bounded proxy enforces. The figure's mechanism is falsified if the panels lose timestamp alignment, if log(level[t]) - log(level[t-1]) does not reproduce the lower panel, or if scale normalization does not alter the historical emphasis of the level view.
| Domain | Application | Decision evidence |
|---|---|---|
| Software engineering | Plot log request volume above log changes. | A ten-percent surge has comparable visual weight at small and large scale. |
| LLMs | Compare log token throughput with changes in throughput. | Capacity planning separates growth from volatility clusters. |
| AI agents | Normalize tool traffic before comparing early and late runs. | Equal proportional retry storms appear comparable. |
| Startups | Use log revenue plus growth-rate changes. | Early and mature-stage shocks share one relative scale. |
| Business | Compare proportional demand shifts across product sizes. | Small products are not dismissed only because their units are smaller. |
| Daily life | Compare percentage budget or workload changes over time. | A proportional overload is visible despite a changing baseline. |
Source figure lab: The Reproduction
A model can produce a plausible cumulative path while generating implausible local motion. Looking only at the top line therefore permits a wrong mechanism to pass an “eyeball” test.
Original argument and source trace. “The reproduction” (printed p. 92; supplied PDF pp. 217–218) shows a computer-simulated Bachelier Brownian price series above its increments. The book contrasts its independent, bell-shaped, evenly dispersed changes with the clustered Dow changes on the preceding pages.
Interactive adaptation. A seeded Brownian comparator is matched to the shared proxy in number of increments, sample mean, and sample standard deviation. The same four controls rebuild the proxy first and then rematch the comparator; Brownian observations are always labeled synthetic and are never presented as historical prices.
Protocol and readout. Reset and compare the smooth plausibility of the level panel with the even spread of the increment panel. Raise Tail intensity and Crisis intensity: the comparator's scale changes because matching is fair, but its independent temporal order remains. The readout reports only the displayed Brownian level range and largest increment. Change Scenario seed to test whether the qualitative absence of clustering survives another synthetic sample.
Assumptions and falsifier. Matching the first two moments does not make two processes equivalent, and a single simulation cannot estimate a universal tail. The comparison is falsified if length, mean, or standard deviation do not match the proxy differences, or if repeated seeds systematically produce the same clustered timing without a dependence mechanism.
| Domain | Application | Decision evidence |
|---|---|---|
| Software engineering | Compare real latency with an independent baseline matched on mean and variance. | Run-length and burst diagnostics, not average fit, decide adequacy. |
| LLMs | Match synthetic request sizes to real mean and variance. | Queue simulations are rejected if temporal concentration is missing. |
| AI agents | Shuffle or independently resample tool costs. | Branch-risk controls depend on observed sequencing beyond totals. |
| Startups | Compare sales paths with an independently ordered baseline. | Campaign clustering is not dismissed as average variation. |
| Business | Stress an inventory model with matched but independent demand. | Failure under real ordering identifies chronology risk. |
| Daily life | Compare actual workload order with a shuffled schedule. | Recovery needs reflect clustering rather than total hours. |
Source figure lab: Original vs. Reproduction—Through the Analyzer
Raw units make rare events difficult to compare across models and periods. Standardizing magnitude exposes how far an observation lies from the variability its own model calls ordinary.
Original argument and source trace. “Original vs. reproduction—through the analyzer” (printed p. 93; supplied PDF pp. 219–220) places Dow change magnitudes above Brownian magnitudes in units of standard deviation, sigma. The source highlights multiple roughly 10-sigma Dow moves and the 1987 move near 22 sigma versus a Brownian series concentrated near one to three sigma.
Interactive adaptation. The top panel standardizes absolute deviations of shared proxy differences; the bottom applies the same operation to the matched Brownian comparator. Both use the identical fixed 0–8σ vertical domain and tick marks, so a taller point means a larger standardized magnitude rather than a separately rescaled panel. Controls change the proxy and its matched baseline, while labels state that height represents unusualness and does not preserve return sign.
Protocol and readout. With one seed, lower and raise Tail intensity and record both maximum-sigma values. Then alter Crisis intensity alone and locate the change in the proxy panel. Use several seeds before making a qualitative claim; the readout makes both maxima and all inputs explicit.
Assumptions and falsifier. One global standard deviation can itself exaggerate tails when volatility changes, so the lab diagnoses mismatch rather than proving a unique cause. The claim is weakened if conditioning on local volatility removes the proxy–Brownian gap, and the implementation fails if sigma uses inconsistent centering or scaling between plotted observations and readout.
| Domain | Application | Decision evidence |
|---|---|---|
| Software engineering | Express latency and error spikes in baseline deviations. | Alerts distinguish a large unit value from a statistically unusual one. |
| LLMs | Standardize cost or latency by model and workload class. | Router thresholds expose rare misses without mixing scales. |
| AI agents | Score step cost relative to the plan's normal variation. | Outlier actions trigger review before budgets are exhausted. |
| Startups | Standardize daily acquisition or churn shocks. | A campaign anomaly is compared with the right operating baseline. |
| Business | Normalize demand surprises across products. | Risk review ranks unusualness rather than raw volume. |
| Daily life | Normalize sleep, spending, or workload by personal variability. | Intervention responds to unusually large changes. |
Source figure lab: Two Into One Will Not Go
Chronological charts show clustering but make the tail shape hard to compare. Ranking magnitudes discards time order deliberately so the empirical and modeled extremes can be overlaid on one frame.
Original argument and source trace. “Two into one will not go” (printed p. 94; supplied PDF pp. 221–222) sorts and counts standardized Dow and Brownian changes by size. The book's Brownian gray bars taper near five sigma while Dow observations continue into a fat tail, including the roughly 22-sigma 1987 move.
Interactive adaptation. The lab sorts the exact standardized magnitudes from the preceding two-panel analyzer on the same fixed 0–8σ scale. The matched Brownian uses a long dash; the shared proxy uses a short dash plus point markers, so identity never depends on color. Proxy volatility, Tail intensity, Crisis intensity, and Scenario seed rebuild the observations; sorting changes only order.
Protocol and readout. Start at the left and follow both ranked curves toward the largest observations, using dash and marker patterns rather than color alone. Raise Tail intensity, then Crisis intensity, and record the displayed top-rank gap and both maxima. Repeat with a second seed to avoid treating one Brownian maximum as a theoretical limit.
Assumptions and falsifier. Ranking removes temporal information, so this lab cannot diagnose clustering or direction. A held-out thin-tailed model with conditional volatility could reduce the displayed gap; the implementation is falsified if ranked values are not exactly the analyzer magnitudes sorted ascending or if one seed is presented as a population probability.
| Domain | Application | Decision evidence |
|---|---|---|
| Software engineering | Rank request latency or incident loss against a matched baseline. | Tail capacity uses observed exceedances, not average latency alone. |
| LLMs | Rank per-request cost and hallucination severity. | Guardrails target the tail of failures rather than mean quality. |
| AI agents | Rank run cost, steps, and irreversible-action severity. | Budgets cover rare runaway trajectories. |
| Startups | Rank customer-acquisition cost and churn shocks. | Unit-economics plans include extreme cohorts. |
| Business | Rank supplier delays or daily P&L impacts. | Contingency stock follows tail separation. |
| Daily life | Rank unusually costly or overloaded days. | Buffers are set from extremes without claiming their timing. |
Source figure lab: No Bell Curve
A distribution can contain too many tiny observations and too many extremes at the same time. Fitting only mean and variance then misses both inactivity and crisis, with too little probability assigned between them.
Original argument and source trace. “No bell curve” (printed p. 98; supplied PDF pp. 228–229) compares the frequency of sterling–guilder exchange-rate changes from 1609–2000, attributed to DeVries 2002, with a standard bell curve. The empirical record has excess mass near zero and in the tails.
Interactive adaptation. The historical observations were not transcribed, so the lab clearly displays a synthetic frequency curve beside a theoretical Gaussian. Center mass raises near-zero concentration, Tail mass raises extremes, and Histogram bins changes inspection resolution without changing the stated provenance.
Protocol and readout. Set Tail mass low while varying Center mass, then reverse the experiment. Increase the number of bins and check whether the center and tail gaps retain their signs. The live readout reports both input masses, resolution, and differences from the bell curve at the center and six-sigma edge.
Assumptions and falsifier. The synthetic mixture demonstrates shape, not DeVries's measured frequencies, and the connected points approximate histogram bins. The argument would be weakened if held-out historical data with consistent cleaning fit the Gaussian in both center and tails; the lab fails if copy or labels imply the displayed values are the 1609–2000 record.
| Domain | Application | Decision evidence |
|---|---|---|
| Software engineering | Model many tiny latency changes plus rare stalls. | SLO budgets include both idle periods and severe tail events. |
| LLMs | Inspect token-length and response-latency distributions. | Batching policy handles many short and a few enormous requests. |
| AI agents | Measure mostly small tool costs with rare long branches. | Hard run caps are based on the observed tail. |
| Startups | Model many low-spend users plus rare high-support accounts. | Pricing and support capacity avoid average-customer fiction. |
| Business | Compare order-size or delay histograms with a bell curve. | Inventory buffers reflect tail excess. |
| Daily life | Inspect many trivial expenses plus rare large bills. | Emergency reserves target the actual distribution shape. |
Source figure lab: A Wild Market
Aggregate profit and loss can swing violently when institutions respond to one crisis at the same time. A smooth risk estimate can therefore fail exactly when correlation, liquidity, and leverage tighten together.
Original argument and source trace. “A wild market” (printed p. 107; supplied PDF pp. 247–248), attributed to Medova 2000, shows aggregate daily profits and losses for four anonymized large banks during Russia's 1998 debt-default episode. The source uses the signed swings to make market turbulence operational rather than theoretical.
Interactive adaptation. The anonymous observations and units were not transcribed. This lab uses an explicitly synthetic signed aggregate-bank P&L proxy: Baseline variation controls ordinary noise, Crisis amplification controls stress magnitude, Recovery width controls persistence, and Scenario seed changes bounded innovations. A dashed crisis-magnitude guide is displayed on the same fixed axis; it is part of the evidence rather than a hidden model curve.
Protocol and readout. Hold the seed fixed and compare low and high Crisis amplification. Increase Recovery width and observe how long the dashed guide stays elevated, then change the seed to separate that stable envelope effect from one noisy sequence. The reported recovery half-life is derived from the guide’s first post-peak crossing halfway back to its displayed baseline; the peak magnitude and mean are derived from the displayed signed series.
Assumptions and falsifier. This is neither the four-bank historical series nor a loss forecast, and its arbitrary units cannot support capital advice. The teaching mechanism is weakened if aggregate institutional data show independence and stable liquidity through comparable crises; the implementation fails if it implies bank identity, original units, or Medova observations.
| Domain | Application | Decision evidence |
|---|---|---|
| Software engineering | Aggregate service gains and losses during a regional incident. | Recovery planning measures signed oscillation and persistence. |
| LLMs | Aggregate latency and cost deviations across providers. | Routing detects correlated provider stress. |
| AI agents | Sum tool successes, retries, and rollback costs during one run. | A circuit breaker activates on trajectory volatility. |
| Startups | Combine revenue, refunds, support, and cloud cost during a launch. | Runway alerts include synchronized operational effects. |
| Business | Aggregate desk, supplier, or regional exposure during stress. | Crisis limits respond to common movement and liquidity. |
| Daily life | Track signed cash-flow swings during disruption. | A reserve is tested against sequences, not only totals. |
Assumptions, limits, and replacement risk
Critiquing a baseline does not license an unconstrained alternative. Rich models have more parameters, greater estimation uncertainty, and more ways to overfit.
A power-law tail may fit only a limited range. Long-memory estimates can be distorted by structural breaks. A jump model may confuse data errors with events. A flexible simulator can reproduce selected facts while failing an unused statistic. Therefore compare alternatives on withheld periods and decision-relevant metrics.
Robust decisions often need less model confidence, not more complexity: cap leverage, hold liquidity, diversify mechanisms, define stop conditions, and monitor assumption drift. This chapter teaches model-risk practice, not personalized financial advice.
Engineering applications: audit the model before trusting the benchmark
Engineering evidence becomes unsafe when a capacity test, LLM benchmark, or agent simulator omits the mechanism that governs deployment failure. A shared assumption audit makes each model state its sampling process, dependence, state transitions, and irreversible consequences.
For every engineering model, record one prediction, one held-out test, one plausible alternative, and the action taken if the prediction fails. Run a coupled stress in which traffic, evaluation error, and tool recovery worsen together, because shared infrastructure can invalidate all three baselines at once. The sections below retain the detailed fixes; the roll-up ensures a green metric cannot hide an omitted boundary.
LLM systems: benchmark misspecification
LLM evaluation can be mathematically clean yet operationally wrong when the task distribution, scoring rule, or judge omits costly failures. Average accuracy compresses severity and dependence.
A model may score 90% on independently sampled questions but repeatedly fail one protected-language group or long-context pattern. An LLM judge can share the evaluated model’s bias, creating correlated evaluation error. Prompt leakage can make a backtest look predictive.
Publish category-level outcomes, severe-failure counts, confidence intervals, and temporal holdouts. Test with human or rule-based checks that fail differently. The goal is not one universal benchmark; it is evidence aligned with the deployment decision.
AI agents: a simulator can omit irreversible state
Agent evaluations often model tool calls as independent successes and score only final task completion. That specification ignores partial writes, permission escalation, and errors propagated through memory.
An agent may complete nine of ten steps but corrupt state at step three. A retry then repeats the mutation. Final text quality cannot represent the damage. The model needs state transitions, idempotency, rollback, and policy boundaries.
Build adversarial traces with shared-tool outages and misleading observations. Verify external state after action, not only the agent’s report. A high simulated completion rate is useful only if the simulator contains the failure modes governing deployment.
Startup applications: spreadsheet precision can hide regime risk
Startup plans become misspecified when smooth conversion, churn, and financing assumptions ignore concentration and feedback. More decimal places do not create more information.
A model may assume 3% monthly churn independently across customers. If one platform policy affects half the customers, churn is correlated and runway collapses faster. Fundraising availability may worsen precisely when revenue disappoints.
Create named downside regimes, vary several linked inputs together, and define action triggers before observing results. Maintain a simple base case, but make commitments survive plausible model error.
Business applications: challenge the operating model at its decision boundary
Business dashboards can be internally consistent while omitting returns, service delays, contract concentration, or correlated supplier failure. The error matters when the dashboard authorizes a commitment that cannot be reversed.
Choose one operating decision and trace every input to evidence, update frequency, and accountable owner. Perturb linked assumptions together, compare the result with an independently constructed scenario, and define the observation that suspends the decision. More detailed forecasts earn trust only when their additional assumptions are also testable.
Daily life applications: plans omit recovery and rare interruptions
Personal schedules often model every day as stationary and every task duration as mild. Illness, caregiving, travel, and administrative work create bursts and recovery costs.
A plan using average task time can remain overloaded because the longest tasks determine missed deadlines. Tracking only completed items hides the queue of deferred work. Adding more optimization cannot fix missing slack.
Compare planned and actual duration distributions, preserve context about interruptions, and limit concurrent commitments. The model should support a humane buffer rather than convert life into a brittle productivity score.
Decision exercise: create an assumption ledger
A model audit fails when assumptions are remembered only after an adverse result. An assumption ledger records each claim, diagnostic, consequence, and owner before the model governs exposure.
Create rows for distributional shape, independence, parameter stability, continuity, liquidity, execution cost, and data quality. For each row, write one observable prediction and one test. “Returns are Gaussian” becomes a predicted exceedance count; “liquidity is adequate” becomes executable size and price-impact bounds; “parameters are stable” becomes a rolling-window tolerance.
Use the 2D lab to compare Model and Observed curves under low and high Assumption error, recording Maximum model miss. Use the 3D surface to identify combinations of tail severity and dependence that cross the decision tolerance as Exposure changes. Do not choose an alternative merely because it is harsher. Compare it on an unused statistic and record estimation uncertainty.
Use this ledger structure:
| Assumption | Observable implication | Test | Failure cost | Response |
|---|---|---|---|---|
| Thin tail | Few high-threshold exceedances | Out-of-sample count | Undersized buffer | Stress empirical tail |
| Independence | Short runs and weak lag dependence | Shuffle control | Clustered depletion | Add state model |
| Stability | Rolling parameters agree within uncertainty | Window comparison | Stale limits | Shorten validity |
| Continuity | No material gaps | Cleaned gap audit | Failed hedge | Jump scenario |
| Liquidity | Planned size executes near quoted price | Size-impact test | Forced loss | Cap size and hold cash |
Rank rows by decision sensitivity, not statistical novelty. A tiny but significant deviation may be immaterial, while a poorly estimated liquidity cliff can dominate survival. For each red row, state whether the safest response is better data, a richer model, lower exposure, or operational protection. Some uncertainty cannot be modeled away.
Add a response column with three options: monitor, reduce exposure, or revise mechanism. Assign a trigger and owner. For an LLM release, a severe-failure excess may reduce autonomy; for a service, correlated tail latency may cap retries; for a startup, a runway stress may delay a fixed commitment. End with a review date because an assumption can become false without a code change. The ledger turns criticism into governance instead of rhetoric.
Keep two records that teams often collapse: evidence quality and consequence severity. Sparse evidence may justify another measurement when failure is reversible, but the same uncertainty may justify a hard limit when failure is ruinous. Record reversibility, detection delay, and recovery cost beside each response. This prevents a familiar error: treating “not statistically settled” as “safe enough to expose.” The ledger should become stricter as losses become harder to reverse.
Synthesis: criticize by consequence
Model criticism earns value when it changes an exposure, guardrail, or experiment. Naming a failed assumption is the bridge between data and action.
assumption
-> observable prediction
-> diagnostic mismatch
-> decision sensitivity
-> revise model or reduce exposure
-> monitor out of sample
The best model is not the most realistic in every detail. It is adequate for the decision, explicit about omissions, and paired with safeguards for uncertainty. A model card should state data, horizon, assumptions, rejected alternatives, sensitivity, and owner.
Sources and further study
The empirical case should be read through primary studies that document tails, dependence, and changing volatility. These sources also show that alternative models solve different parts of the problem.
- Benoit Mandelbrot and Richard L. Hudson, The Misbehavior of Markets, Basic Books.
- Benoit Mandelbrot, “The Variation of Certain Speculative Prices”, The Journal of Business 36(4), 1963.
- Robert F. Engle, “Autoregressive Conditional Heteroscedasticity with Estimates of the Variance of United Kingdom Inflation”, Econometrica 50(4), 1982.
- Rama Cont, “Empirical Properties of Asset Returns: Stylized Facts and Statistical Issues”, Quantitative Finance 1(2), 2001.
Key takeaways
Correct calculations do not guarantee an adequate model. Assumptions must produce observable claims and survive decision-relevant tests.
- Gaussian center fit does not validate tail risk.
- Independence can fail in magnitude even when signs look unpredictable.
- Stable historical parameters can break across regimes.
- Continuous models omit gaps and execution discontinuities.
- Diversification estimates fail when volatility and correlation change together.
- Robustness combines alternative models, exposure limits, and monitoring.
Checklist
The case against a model is complete only when the failed assumption and decision impact are explicit. Use these checks on one live model.
- [ ] I can list every material distributional and operational assumption.
- [ ] I can compare predicted and observed exceedance counts.
- [ ] I can test dependence, stability, jumps, and liquidity separately.
- [ ] I can operate both labs and interpret the error readouts.
- [ ] I can identify overfitting risk in a richer replacement.
- [ ] I can name a decision that changes under a plausible alternative.
- [ ] I can define monitoring for assumption drift.
- [ ] I can explain why teaching simulations are not forecasts or advice.