Skip to main content

Taming the "Fat Tail": Decrypting Climate Disaster Costs

In the world of risk modeling, natural disasters are notoriously difficult to quantify. While frequency is relatively predictable, economic impact is chaotic. A single "Black Swan" event—like the 2011 Tohoku Earthquake or the 2004 Indian Ocean Tsunami—can cause more economic damage in an afternoon than thousands of smaller events combined over a decade.

I analyzed global disaster data from EM-DAT (2000-2025) to understand these patterns. Below, I look at the geography of these events and then at how I estimate losses for the many events that report none. Update: an earlier version of this post described a hand-calibrated formula. I have since replaced it with the standard damage regression from the disaster-economics literature and, more importantly, tested it properly: every score below comes from years the model never saw.

The Geography of Risk

To understand the scope, I first look at where these events occur. As the data shows, the distribution is far from uniform.

Figure 1: Natural Disasters by Region. Asia is the undisputed global epicenter of natural disaster frequency, accounting for nearly double the event count of the Americas.

However, frequency tells only half the story. The type of disaster varies radically by region, dictating the kind of economic models we need to build.

Figure 2: Disaster Type vs. Region Heatmap. The "Risk Fingerprint": Note the dark red clusters. Asia’s primary challenge is Riverine Floods (1,713 events), while the Americas face a massive concentration of Storms (881 events).

The "Missing Data" Problem

While we have solid data on event counts (like those above), reliable economic loss data is scarce. Of the 10,902 natural-disaster records in my EM-DAT export, only 3,312 (30%) report a total damage figure. Together those 3,312 events report $5.2 trillion in damage (2025 US$). Every other record is a gap.

To fill it, I built a Parametric Loss Estimator. The method itself is not new; the contribution is the evaluation. On 687 events from 2020–2025, held out of training, the estimator lands within 2× of reported damage 34.5% of the time and within 10× 87.0% of the time. The best naive baseline (the median loss for the event's type and income group) manages 23.9% and 73.5%.

Under the Hood: A Standard Damage Regression, Tested Out of Time

The core is a log-linear damage regression, the workhorse specification behind the income–damage and hazard–damage findings of Kahn (2005), Toya and Skidmore (2007), Nordhaus (2010) and Bakkensen and Mendelsohn (2016). It produces a central estimate for each event. A Log-Normal/Pareto error model then turns that estimate into a range, because disaster losses follow two distinct sets of rules.

1. The Central Estimate (Log-Linear Regression)

For each event, the model predicts the logarithm of damage from what is known after the event, fitted on the 3,312 events that do report damage:

  • Inputs: the Disaster Type, People Affected, Deaths, Hazard Intensity (earthquake magnitude, storm wind speed, or affected area), and the country's GDP per Capita in the event year (World Bank, constant 2015 US$).
  • The Formula: slopes vary by disaster type t, with a time trend and an "unknown" indicator for each missing input:
ln(Loss) = at + bt ln(1 + Affected) + dt ln(1 + Deaths) + c ln(GDP per capita) + ht Hazard + trend
  • Elasticities: the fitted effects are modest. Damage rises about 0.2% for each 1% more people affected (0.15–0.38 by type) and about 1.2% for each 1% higher GDP per capita; deaths add 0.14–0.71 depending on type. Type-specific slopes are shrunk toward pooled ones, so thin types such as volcanic activity (20 events) borrow strength from the rest.
  • ML correction: a Random Forest (300 trees, depth 6, settings fixed before any test results were seen) learns the regression's systematic errors. It adds 3–4 points within 2× in the two later test periods and nothing in the earliest one.

2. The Range (Log-Normal Body, Pareto Tail)

A point estimate alone is misleading. Standard models treat a $100 billion hurricane as statistically impossible, even though history proves they happen.

So each estimate comes with a 95% interval. The lower side is Log-Normal; above the 90th percentile of errors, a Pareto Tail takes over. Both are fitted to out-of-fold errors, not in-sample ones.

  • The Alpha: a Hill estimate on the out-of-fold errors gives a Pareto tail index of about 1.2, a very heavy tail.
  • The Result: on 2020–2025 events, 96.7% of actual losses fall inside the 95% interval. The price is width: intervals typically span two orders of magnitude, so each estimate is an order of magnitude with a range, not a point figure.

Visualizing the Volatility

Why go to all this trouble to model the "Fat Tail"? Because the historical data proves that economic damage is defined by spikes, not averages.

Figure 3: Economic Damage Over Time (Adjusted). The massive spikes you see—2011 (Tohoku Earthquake/Thai Floods) and 2017 (Hurricanes Harvey/Irma/Maria)—are exactly why a simple average fails. A standard average would predict a smooth line; the Pareto tail in the error model anticipates these mountains.

Summary

A standard damage regression, a small ML correction and a Log-Normal/Pareto error model together beat naive baselines in every test period (+6 to +9 points within 2×, +10 to +19 within 10×) and produce 95% intervals that hold up. But absolute accuracy is modest: even the best model misses by more than 2× on about two events in three. The bigger caveat is who the gaps are. Events with no reported damage are poorer and smaller than the ones I could test on: 18% of gaps are in low-income countries, against 3% of test events, so accuracy there is essentially untested. The estimates are useful for filling historical gaps and comparing many events at order-of-magnitude precision, always with the interval; they are not suitable for single-event decisions, insurance or official loss reporting. Loss data: EM-DAT, CRED / UCLouvain (public.emdat.be).

Comments

Popular posts from this blog

Mapping the Blast Radius: What Happens When the "World's Factory" Stops?

In global trade, "efficiency" often masks "fragility." We know that China is central to the global electronics supply chain, but how central? And if that node were to go dark, who would feel the shockwaves first? To answer this, I moved beyond standard trade statistics and built a network simulation using the OECD Inter-Country Input-Output (ICIO) Tables (2023 Edition) . This dataset maps the DNA of the global economy, tracking every dollar of input across 66 countries and 45 industries. 1. The Blast Radius: Tracing the Contagion I treated the global economy as a directed graph and simulated a total supply shock to Chinese Electronics (CHN_C26) . By tracing the flow of inputs across three tiers of buyers, I visualized the "Blast Radius" of this disruption. Fig 1: The Supply Chain Cascade. The shock originates in China (Red) and immediately hits "Tier 1" assembly hubs (Dark Blue) before cascading to global consumers...

Turkey's Informality Tax

A general-equilibrium model of Turkey's dual labour market says the cost of taxing formal work doesn't show up where we usually look for it. The bottomline: Making Turkey's transfer to the unemployed a third more generous raises unemployment from 8.6% to 9.2% . That is the honest cost, and it is not large. How you pay for it matters more than whether you pay for it. Funded by payroll taxes, the poor end up 0.6% worse off than before the transfer was raised — the policy defeats itself. Funded by VAT, they are 0.5% better off . The reason is not unemployment, which is nearly identical under both. It is informality . A higher payroll tax pushes formal jobs into the unregistered sector, where the wage is 43% lower and nothing is taxed. Turkey is not on the wrong side of the payroll-tax Laffer curve — revenue peaks around 53%, well above today's ~37.5% wedge. But the marginal cost...

Mapping the Matrix: Malaysia's Corporate Network

The complex interconnections that define modern society are readily visible in our Facebook or LinkedIn networks. But what about the corporate world? I recently began mapping the web of cross-shareholding among Malaysia's listed corporations. Using Social Network Analysis (SNA), I visualized the ecosystem to identify key "communities" and power brokers. The goal? To see if a company's position in this web predicts its financial destiny. The State of Play (2015) Figure 1: Malaysia's Corporate Network in 2015 (Created via Gephi) The map above visualizes the network as of 2015. While I have excluded specific company names, the "super-nodes" (the largest circles) represent the most connected shareholders in the economy—primarily government-linked investment companies (GLICs), fund managers, and international asset managers. A Decade of Growing Complexity ...