The most consequential number in climate science right now is not an emissions target or a sea-level projection. It is equilibrium climate sensitivity — ECS, the warming you get at equilibrium from a doubling of CO₂. For decades the likely range sat at 1.5°C to 4.5°C, a span so wide it was practically an admission of ignorance. Then the CMIP6 model ensemble arrived, and a cluster of models pushed their ECS values above 5°C. The modeling community fractured. Some researchers argued this was a genuine signal, that we had been systematically underestimating how hot the planet could get. Others — and I am firmly in this camp — argued that the high-end CMIP6 models are running too hot, that the mechanism driving their elevated sensitivity is physically unrealistic, and that treating them as credible projections does real damage to climate policy. The argument is not settled. Here is where I stand and why.
The High-ECS Signal in CMIP6 Comes Mostly From One Feedback, and That Feedback Is Probably Wrong
The CMIP6 ensemble that became available around 2019–2020 included models from institutions including NCAR, GFDL, the UK Met Office, and CNRM-CERFACS, among others. A subset — roughly a third of the ensemble — produced ECS values above 4°C, with several exceeding 5°C. The multi-model mean ECS jumped from about 3.2°C in CMIP5 to roughly 3.7°C in CMIP6. Mark Zelinka and colleagues at Lawrence Livermore published a detailed decomposition in 2020 showing that the primary driver of this increase was not CO₂ forcing, not ocean heat uptake, not any of the usual suspects. It was shortwave cloud feedback — specifically, the behavior of low marine clouds in the Southern Ocean and in the subtropical stratocumulus regions.

This matters enormously because low cloud feedbacks are the hardest thing to get right in a global climate model. These clouds form in the planetary boundary layer, at scales of tens to hundreds of meters, far below the resolution of any CMIP-class model. Every model parameterizes them. The CMIP6 models that ran hot did so in large part because their parameterization schemes, updated for the new generation, produced clouds that thin more aggressively under warming than their CMIP5 predecessors. More thinning means less reflected sunlight, which means more warming, which means more thinning — a positive feedback that, if strong enough, can push ECS well above 4°C.
The question is whether that thinning is real. The observational record says it probably is not, at least not at the magnitude these models produce. Researchers including Frida Bender and others working with CERES satellite data have consistently found that the observed relationship between sea surface temperature and low cloud cover in the current climate does not support the aggressive cloud thinning the high-ECS models generate. The emergent constraint literature — a body of work that tries to use observable present-day climate relationships to constrain future projections — converges on an ECS that is considerably lower than what the hot models produce. Sherwood and colleagues’ 2020 assessment in Reviews of Geophysics, which synthesized multiple lines of evidence, placed the likely range at 2.6°C to 3.9°C. That work explicitly used emergent constraints and paleoclimate evidence to pull the upper tail down. The IPCC AR6 adopted a similar range, 2.5°C to 4°C, with a best estimate of 3°C. The high-ECS CMIP6 models are outside or at the very edge of that assessed range.
The Models That Run Too Hot Also Fail Basic Present-Day Observational Tests
My confidence that the high-ECS CMIP6 models are wrong is not based solely on the emergent constraint argument, which has its own methodological vulnerabilities. It is reinforced by the fact that several of these models struggle to reproduce the observed climate record when run over the historical period. This is a straightforward test: initialize the model in the pre-industrial, force it with observed greenhouse gas concentrations, aerosols, volcanic eruptions, and solar variability, and see how well it matches the observed global mean temperature record from 1850 to the present.
Several high-ECS models warm too fast over the twentieth century. To compensate, modeling centers have to apply stronger aerosol cooling — increasing the magnitude of the aerosol forcing — to keep the historical simulation from running away from observations. This is not a secret; it is openly discussed in model documentation. The problem is that the aerosol forcing itself is poorly constrained, and using it as a compensating knob to fix an overly sensitive model is epistemically uncomfortable. You are tuning one uncertain quantity to offset another uncertain quantity, and the result is a model that matches the historical record for the wrong reasons. When you then use that model to project future warming — where aerosol forcing will likely decrease as air quality improves — the compensation disappears and the model runs hot.
Gavin Schmidt at NASA GISS has written about this problem in the context of what he and colleagues called the “climate model weighting” question: should all CMIP6 models be treated as equally plausible, or should models that fail observational tests be down-weighted? The answer, I think, is clearly the latter, and the AR6 implicitly agreed by not simply averaging the CMIP6 ensemble but instead using the multi-line-of-evidence approach that produced the 2.5°C–4°C range. But this decision has consequences. It means the official assessed range is narrower than the raw model spread, and it means the high-end tail of the CMIP6 ensemble — the 5°C and 6°C scenarios that appear in some impacts literature — should not be treated as physically credible projections.
The Strongest Counterargument: Paleoclimate Suggests We Should Not Be Comfortable at 3°C
I want to be honest about where my position is most vulnerable, because the counterargument from paleoclimate is genuinely serious and I do not think it has been fully resolved.
The high-ECS camp can point to deep-time climate reconstructions that suggest Earth’s climate has been more sensitive to forcing than the modern instrumental record implies. Work by Jessica Tierney and colleagues using proxy reconstructions of the Last Glacial Maximum has produced ECS estimates that, depending on how you handle state-dependence and boundary conditions, carry uncertainty extending into the upper part of the assessed range — though these estimates are generally centered near ~3–3.5°C and do not straightforwardly endorse the hottest CMIP6 models. The argument is that the current climate sits in a relatively stable state, and that emergent constraints derived from small perturbations in the present-day climate may underestimate sensitivity in a warmer world where feedbacks are nonlinear and state-dependent. If low cloud feedbacks become more positive as the planet warms beyond the range of the observational record, the hot models might be capturing something real that the emergent constraint approach systematically misses.
This is not a fringe position. It is held by serious researchers, and it deserves more than dismissal. My response is that the paleoclimate ECS estimates carry their own substantial uncertainties — in the proxy reconstructions themselves, in the radiative forcing estimates for past climates, and especially in the assumption that feedbacks operating over glacial-interglacial timescales are the same feedbacks operating on centennial timescales. The Last Glacial Maximum involved ice sheet feedbacks, vegetation feedbacks, and dust feedbacks that are not part of standard ECS definitions. Disentangling the “fast feedback” sensitivity from the “Earth system” sensitivity in paleoclimate data is genuinely hard, and I think the literature has not yet converged on a clean answer.
What Observation or Experiment Would Actually Change My Mind
The question that would settle this — or at least substantially shift my priors — is not one we can answer with current observational infrastructure. What we need is a long, high-quality record of low cloud cover and optical depth in the Southern Ocean and subtropical stratocumulus regions, paired with co-located sea surface temperature measurements, over a period long enough to capture meaningful interannual and decadal variability. The CERES record is now over two decades long, which is useful, but the signal-to-noise ratio for detecting cloud feedback from natural variability alone is poor. The EarthCARE mission, a joint ESA-JAXA satellite launched in 2024, will provide higher-resolution cloud profiling than anything we have had before. If EarthCARE data, accumulated over five to ten years, shows that low cloud cover in the Southern Ocean is declining at rates consistent with what the high-ECS models predict, I would have to take those models much more seriously.
I would also update significantly if the emergent constraint methodology were shown to be systematically biased. Several papers in recent years — including work by Caldwell and colleagues at LLNL — have raised legitimate concerns about whether emergent constraints are robust across different model generations or whether they are artifacts of shared model ancestry. If a future analysis demonstrated that the observational relationships used to constrain ECS are themselves model-dependent and do not reflect real physical mechanisms, the constraint would lose its force, and the high-ECS models would become harder to dismiss.
Until that evidence arrives, I think the responsible position for the modeling community is to present the AR6 assessed range as the primary projection basis, to down-weight the high-ECS outliers in ensemble analyses, and to be explicit in impacts literature that the 5°C and 6°C tails of the CMIP6 distribution are not equivalent in credibility to the central estimates. This is not complacency. A best estimate of 3°C ECS, realized under a high-emissions pathway, is catastrophic enough to demand immediate and aggressive decarbonization. We do not need to invoke physically dubious hot models to make the case for urgency. We need to make the case accurately, because the credibility of the entire enterprise depends on it.


Leave a Reply