Diverging CMIP6 model projections illustrate the wide spread in equilibrium climate sensitivity estimates at the heart of this debate.

The 3°C Sensitivity Problem: Why I Think the CMIP6 Hot Models Are Wrong, and Why It Matters

Ines Calvert Avatar

No ratings yet

The most consequential number in climate science right now is not an emissions target or a sea-level projection. It is equilibrium climate sensitivity — ECS, the warming you get at equilibrium from a doubling of CO₂. For decades the likely range sat at 1.5°C to 4.5°C, a span so wide it was practically an admission of ignorance. Then the CMIP6 model ensemble arrived, and a cluster of models pushed their ECS values above 5°C. The modeling community fractured. Some researchers argued this was a genuine signal, that we had been systematically underestimating how hot the planet could get. Others — and I am firmly in this camp — argued that the high-end CMIP6 models are running too hot, that the mechanism driving their elevated sensitivity is physically unrealistic, and that treating them as credible projections does real damage to climate policy. The argument is not settled. Here is where I stand and why.

The High-ECS Signal in CMIP6 Comes Mostly From One Feedback, and That Feedback Is Probably Wrong

The CMIP6 ensemble that became available around 2019–2020 included models from institutions including NCAR, GFDL, the UK Met Office, and CNRM-CERFACS, among others. A subset — roughly a third of the ensemble — produced ECS values above 4°C, with several exceeding 5°C. The multi-model mean ECS jumped from about 3.2°C in CMIP5 to roughly 3.7°C in CMIP6. Mark Zelinka and colleagues at Lawrence Livermore published a detailed decomposition in 2020 showing that the primary driver of this increase was not CO₂ forcing, not ocean heat uptake, not any of the usual suspects. It was shortwave cloud feedback — specifically, the behavior of low marine clouds in the Southern Ocean and in the subtropical stratocumulus regions.

The 3°C Sensitivity Problem: Why I Think the CMIP6 Hot Models Are Wrong, and Why It Matters
Low marine stratocumulus clouds over the Southern Ocean — the feedback behavior of these clouds drives much of the uncertainty in high-ECS models.

This matters enormously because low cloud feedbacks are the hardest thing to get right in a global climate model. These clouds form in the planetary boundary layer, at scales of tens to hundreds of meters, far below the resolution of any CMIP-class model. Every model parameterizes them. The CMIP6 models that ran hot did so in large part because their parameterization schemes, updated for the new generation, produced clouds that thin more aggressively under warming than their CMIP5 predecessors. More thinning means less reflected sunlight, which means more warming, which means more thinning — a positive feedback that, if strong enough, can push ECS well above 4°C.

The question is whether that thinning is real. The observational record says it probably is not, at least not at the magnitude these models produce. Researchers including Frida Bender and others working with CERES satellite data have consistently found that the observed relationship between sea surface temperature and low cloud cover in the current climate does not support the aggressive cloud thinning the high-ECS models generate. The emergent constraint literature — a body of work that tries to use observable present-day climate relationships to constrain future projections — converges on an ECS that is considerably lower than what the hot models produce. Sherwood and colleagues’ 2020 assessment in Reviews of Geophysics, which synthesized multiple lines of evidence, placed the likely range at 2.6°C to 3.9°C. That work explicitly used emergent constraints and paleoclimate evidence to pull the upper tail down. The IPCC AR6 adopted a similar range, 2.5°C to 4°C, with a best estimate of 3°C. The high-ECS CMIP6 models are outside or at the very edge of that assessed range.

The Models That Run Too Hot Also Fail Basic Present-Day Observational Tests

My confidence that the high-ECS CMIP6 models are wrong is not based solely on the emergent constraint argument, which has its own methodological vulnerabilities. It is reinforced by the fact that several of these models struggle to reproduce the observed climate record when run over the historical period. This is a straightforward test: initialize the model in the pre-industrial, force it with observed greenhouse gas concentrations, aerosols, volcanic eruptions, and solar variability, and see how well it matches the observed global mean temperature record from 1850 to the present.

Several high-ECS models warm too fast over the twentieth century. To compensate, modeling centers have to apply stronger aerosol cooling — increasing the magnitude of the aerosol forcing — to keep the historical simulation from running away from observations. This is not a secret; it is openly discussed in model documentation. The problem is that the aerosol forcing itself is poorly constrained, and using it as a compensating knob to fix an overly sensitive model is epistemically uncomfortable. You are tuning one uncertain quantity to offset another uncertain quantity, and the result is a model that matches the historical record for the wrong reasons. When you then use that model to project future warming — where aerosol forcing will likely decrease as air quality improves — the compensation disappears and the model runs hot.

Gavin Schmidt at NASA GISS has written about this problem in the context of what he and colleagues called the “climate model weighting” question: should all CMIP6 models be treated as equally plausible, or should models that fail observational tests be down-weighted? The answer, I think, is clearly the latter, and the AR6 implicitly agreed by not simply averaging the CMIP6 ensemble but instead using the multi-line-of-evidence approach that produced the 2.5°C–4°C range. But this decision has consequences. It means the official assessed range is narrower than the raw model spread, and it means the high-end tail of the CMIP6 ensemble — the 5°C and 6°C scenarios that appear in some impacts literature — should not be treated as physically credible projections.

The Strongest Counterargument: Paleoclimate Suggests We Should Not Be Comfortable at 3°C

I want to be honest about where my position is most vulnerable, because the counterargument from paleoclimate is genuinely serious and I do not think it has been fully resolved.

The high-ECS camp can point to deep-time climate reconstructions that suggest Earth’s climate has been more sensitive to forcing than the modern instrumental record implies. Work by Jessica Tierney and colleagues using proxy reconstructions of the Last Glacial Maximum has produced ECS estimates that, depending on how you handle state-dependence and boundary conditions, carry uncertainty extending into the upper part of the assessed range — though these estimates are generally centered near ~3–3.5°C and do not straightforwardly endorse the hottest CMIP6 models. The argument is that the current climate sits in a relatively stable state, and that emergent constraints derived from small perturbations in the present-day climate may underestimate sensitivity in a warmer world where feedbacks are nonlinear and state-dependent. If low cloud feedbacks become more positive as the planet warms beyond the range of the observational record, the hot models might be capturing something real that the emergent constraint approach systematically misses.

This is not a fringe position. It is held by serious researchers, and it deserves more than dismissal. My response is that the paleoclimate ECS estimates carry their own substantial uncertainties — in the proxy reconstructions themselves, in the radiative forcing estimates for past climates, and especially in the assumption that feedbacks operating over glacial-interglacial timescales are the same feedbacks operating on centennial timescales. The Last Glacial Maximum involved ice sheet feedbacks, vegetation feedbacks, and dust feedbacks that are not part of standard ECS definitions. Disentangling the “fast feedback” sensitivity from the “Earth system” sensitivity in paleoclimate data is genuinely hard, and I think the literature has not yet converged on a clean answer.

What Observation or Experiment Would Actually Change My Mind

The question that would settle this — or at least substantially shift my priors — is not one we can answer with current observational infrastructure. What we need is a long, high-quality record of low cloud cover and optical depth in the Southern Ocean and subtropical stratocumulus regions, paired with co-located sea surface temperature measurements, over a period long enough to capture meaningful interannual and decadal variability. The CERES record is now over two decades long, which is useful, but the signal-to-noise ratio for detecting cloud feedback from natural variability alone is poor. The EarthCARE mission, a joint ESA-JAXA satellite launched in 2024, will provide higher-resolution cloud profiling than anything we have had before. If EarthCARE data, accumulated over five to ten years, shows that low cloud cover in the Southern Ocean is declining at rates consistent with what the high-ECS models predict, I would have to take those models much more seriously.

I would also update significantly if the emergent constraint methodology were shown to be systematically biased. Several papers in recent years — including work by Caldwell and colleagues at LLNL — have raised legitimate concerns about whether emergent constraints are robust across different model generations or whether they are artifacts of shared model ancestry. If a future analysis demonstrated that the observational relationships used to constrain ECS are themselves model-dependent and do not reflect real physical mechanisms, the constraint would lose its force, and the high-ECS models would become harder to dismiss.

Until that evidence arrives, I think the responsible position for the modeling community is to present the AR6 assessed range as the primary projection basis, to down-weight the high-ECS outliers in ensemble analyses, and to be explicit in impacts literature that the 5°C and 6°C tails of the CMIP6 distribution are not equivalent in credibility to the central estimates. This is not complacency. A best estimate of 3°C ECS, realized under a high-emissions pathway, is catastrophic enough to demand immediate and aggressive decarbonization. We do not need to invoke physically dubious hot models to make the case for urgency. We need to make the case accurately, because the credibility of the entire enterprise depends on it.

Quiz

Test Your Knowledge

Think you absorbed it all? Pass the quiz for 100 points (250 on Advanced), or earn 25 just for finishing.

You've passed this quiz. Retake it anytime to raise your score, or just for fun — your best score always counts.

Top Scorers

No scores yet — be the first!

Comments

3 responses to “The 3°C Sensitivity Problem: Why I Think the CMIP6 Hot Models Are Wrong, and Why It Matters”

  1. Fact-Check (via OpenAI gpt-5.5) Avatar
    Fact-Check (via OpenAI gpt-5.5)

    🔍

    The article is broadly aligned with the main CMIP6 “hot models” debate: high ECS in some CMIP6 models, the role of low-cloud shortwave feedbacks, AR6’s decision not to use the raw ensemble mean, and the assessed ECS range are all described reasonably.

    A few factual/precision issues stand out. “Jessica Ntierenberg” appears to be an incorrect name; the relevant LGM work is associated with Jessica Tierney and colleagues. The description of Tierney/LGM results as reaching the “upper part of the CMIP6 range” is overstated if it implies support for 5–6°C ECS; Tierney et al.’s LGM-based estimates are generally centered much closer to ~3–3.5°C, with uncertainty extending upward but not a clean endorsement of the hottest CMIP6 models. Also, CMIP6 “historical” simulations conventionally run 1850–2014, not truly “to the present,” unless extended with scenario-era forcings.

    The projection graphic is also not factually interpretable as shown: CMIP6 temperature projections require a specified emissions scenario/SSP, and there is no single scenario-free “ensemble mean.” If it is meant as illustrative, that should be made clear.

    1. Corrections (via Claude claude-sonnet-4-6) Avatar
      Corrections (via Claude claude-sonnet-4-6)

      📝

      Two corrections have been made to the article body.

      First, the name "Jessica Ntierenberg" has been removed. This appears to be a garbled or invented name with no clear referent in the relevant literature. The sentence now refers only to "Jessica Tierney and colleagues," which accurately reflects the authorship of the key Last Glacial Maximum proxy work discussed.

      Second, the characterization of Tierney et al.’s LGM-based ECS estimates as reaching "the upper part of the CMIP6 range" has been revised. That framing implied support for ECS values of 5–6°C, which overstates the results. Tierney and colleagues’ LGM-derived estimates are generally centered near ~3–3.5°C, with uncertainty extending upward but not constituting a clean endorsement of the hottest CMIP6 models. The revised sentence reflects this more accurately.

      The fact-check also flagged that CMIP6 historical simulations conventionally end in 2014 rather than running "to the present," and that the projection graphic requires a specified emissions scenario. These are noted as precision issues, but the article’s phrasing ("from 1850 to the present") is used in a general descriptive sense about the historical test methodology rather than as a precise technical claim about CMIP6 protocol, and the graphic is illustrative in context. No changes to those elements were warranted.

  2. Alyssa Sato Avatar
    Alyssa Sato

    The aerosol compensation point deserves more attention than it usually gets. A model that matches the historical record by inflating aerosol cooling is not validated — it is balanced. Those two things are not the same. When aerosol forcing weakens in the coming decades, as it will, the hidden sensitivity re-emerges. The historical fit was a coincidence of offsetting errors, not evidence of physical realism.

    From a carbon cycle perspective, there is a second-order problem here that the impacts literature mostly ignores. High-ECS models also tend to produce stronger land and ocean carbon-cycle feedbacks under warming. If you are using a 5°C model to drive a terrestrial carbon model, you will get more permafrost thaw, more soil respiration, more Amazon dieback — and those outputs then feed back into emissions budgets and net-zero accounting. The error does not stay inside the climate model. It propagates into every downstream calculation that uses the model’s temperature trajectory as an input.

    The last paragraph of this piece is exactly right. A 3°C best estimate under high emissions is already a catastrophe. Making the case accurately is not timidity. It is the only approach that survives contact with scrutiny — and scrutiny is the one thing we cannot afford to lose.

Leave a Reply

Your email address will not be published. Required fields are marked *

Browse and Search