ESG scoring: build a defensible methodology with your criteria and apply it at scale

ESG scoring: how to build a defensible methodology (and apply it at scale)

July 27, 2026

Every serious sustainability decision eventually becomes a scoring problem. Whether you are comparing portfolio companies, screening a target, or assessing members of a coalition, at some point the judgment has to compress into something comparable: a score. And that is where the trouble starts, because the third-party scores most teams inherit disagree with each other, hide their methods, and rarely measure the things your organization actually cares about.

The disagreement is not an impression; it is one of the best documented facts in sustainable finance. The foundational MIT research on the topic, the “Aggregate Confusion” study, found an average correlation of just 0.54 between the ESG ratings of six major providers, meaning two respected raters routinely reach different conclusions about the same company. A 2026 systematic review in the Review of Managerial Science confirms the divergence remains unresolved. For credit ratings, by comparison, agencies agree almost perfectly. This guide covers how ESG scoring actually works, why ratings diverge, and how to build a scoring methodology of your own that you can defend and apply at scale.

Key takeaways

  • ESG scoring converts sustainability evidence into comparable numbers; its value depends entirely on the transparency and consistency of the methodology behind it.
  • Third-party ESG ratings disagree with each other (average correlation 0.54 in the foundational MIT research), which is why organizations increasingly build their own scoring.
  • A defensible methodology has four properties: criteria that reflect your priorities, weights you can explain, evidence you can trace, and consistent application across every company and cycle.
  • AI removes the historical constraint on in-house scoring, which was never methodology design but the reading capacity to apply it at scale.

What is ESG scoring?

ESG scoring is the process of evaluating a company’s environmental, social, and governance performance against defined criteria and expressing the result as a comparable measure: a number, grade, or tier. A score can come from a third-party ratings provider applying its own methodology, or from an organization applying its own criteria to the evidence in company disclosures.

The mechanics are the same either way. A methodology defines what is measured (criteria), how much each criterion matters (weights), what evidence counts (sources), and how raw findings convert into the final score. Every debate about ESG scores, including the famous disagreements between raters, is really a debate about one of those four choices.

ESG scores vs. ESG ratings: what’s the difference?

In practice the terms blur, but the useful distinction is who owns the methodology. ESG ratings usually refer to third-party assessments from providers like MSCI or Sustainalytics, where the methodology is proprietary to the rater and applied identically to thousands of companies. ESG scoring, especially in-house scoring, means your organization defines the criteria and owns the logic.

Third-party ratings offer coverage and convenience, and for many workflows that is enough. The trade-offs are the documented divergence between providers, methodologies you cannot fully inspect, and criteria set for a general market rather than your mandate. That is why stewardship teams, coalitions, and advisors increasingly treat external ratings as one input and build their own assessment on top.

Why organizations build their own ESG scoring

Three situations push teams from consuming scores to building them. The first is mandate mismatch: your investment thesis, membership framework, or client methodology weighs things the big raters do not, so an off-the-shelf score answers a question you did not ask. The second is defensibility: when a score drives an engagement conversation, a committee decision, or a published report, you need to show exactly why a company scored what it did, which is impossible with a black-box rating. The third is coverage: private companies, members, and suppliers often are not rated at all, so if you want a score, you have to produce it.

One global sustainability coalition we work with faced all three at once, scoring more than 200 member companies against its own proprietary framework, and found the real constraint was never the methodology. It was applying that methodology consistently across hundreds of companies and multiple reviewers, a problem it solved with AI-assisted assessment that cut the cycle from months to weeks while making scoring more consistent.

How to build an ESG scorecard: five steps

A scorecard is your methodology made concrete: the criteria, weights, and scoring logic on one page. Building one is less mysterious than it sounds.

  1. Start from decisions, not data. List the decisions the score will inform (invest, engage, escalate, admit, renew) and work backward to the criteria that actually change those decisions. A scorecard built from available data measures what is convenient; one built from decisions measures what matters.
  2. Define criteria you can evidence. Every criterion needs an answerable question and an evidence source: disclosures, policies, targets, performance data. If no evidence could ever settle it, it is an opinion, not a criterion.
  3. Set weights you can explain. Weights encode your priorities, and someone will eventually ask why governance is 40% and biodiversity is 5%. Write the rationale down when you set them, because that document is what makes the scorecard defensible later.
  4. Score against evidence, with sources attached. Every scored point should trace to the passage that justifies it. This is the difference between a score you can publish and one you have to caveat.
  5. Apply it consistently, then version it. The same criteria for every company, every reviewer, every cycle. When the methodology evolves (and it should), version it explicitly so year-over-year comparisons stay honest.

💡 Manifest Climate applies your scorecard (criteria, weights, and all) across any number of companies, with source-linked evidence behind every scored point. Explore our Benchmarking solution.

Applying a methodology at scale: the real bottleneck

Designing a methodology takes a workshop. Applying it takes a workforce, or it used to. Scoring one company against 50 criteria means finding evidence across hundreds of pages of disclosure; scoring a portfolio, a membership, or a sector multiplies that by hundreds. This reading capacity problem, not methodology design, is why most organizations historically settled for third-party ratings despite their limitations.

Purpose-built AI changes the economics. It reads the disclosure record, extracts the evidence for each criterion, and applies the scoring logic identically every time, while analysts review the evidence and own the judgment. The output is the thing manual processes struggle to produce at scale: a consistent, source-traceable score for every company, comparable across the whole set. Teams then spend their time where scores disagree with expectations, which is where the insight usually lives.

Score with confidence using Manifest Climate

Manifest Climate is the AI-powered assessment engine for sustainability, built for organizations that put their name behind their scores. It applies your scoring methodology (your criteria, your weights, your thresholds) across portfolios, memberships, and sectors, and returns source-linked, audit-ready results your team can defend in front of committees, members, and the press.

Investors use it to score holdings against their house view, coalitions use it to score members against proprietary frameworks, and advisors use it to run client methodologies without drowning in documents. The methodology stays yours; the reading stops being the bottleneck.

If your scoring ambitions have been limited by what your team can read, book a demo and bring your scorecard.

Frequently asked questions

What is ESG scoring?
ESG scoring is the process of evaluating a company’s environmental, social, and governance performance against defined criteria and expressing the result as a comparable measure. Scores can come from third-party ratings providers or from an organization’s own methodology applied to company disclosures.

Why do ESG ratings disagree?
Foundational MIT research found an average correlation of just 0.54 between six major ESG raters. They disagree because each provider chooses different criteria, weights, and measurement approaches, so the same company can legitimately score well with one rater and poorly with another.

What makes an ESG scoring methodology defensible?
Four properties: criteria that reflect the decisions the score informs, weights with a written rationale, evidence traced to specific sources for every scored point, and consistent application across every company, reviewer, and cycle.

What is an ESG scorecard?
An ESG scorecard is a scoring methodology made concrete: the defined criteria, their weights, the evidence sources, and the scoring logic, usually summarized in a single framework a team can apply repeatedly to different companies.

Can AI do ESG scoring?
AI applies a scoring methodology at scale: reading disclosures, extracting evidence for each criterion, and applying the logic consistently across hundreds of companies. The methodology and final judgment stay with the organization; AI removes the reading bottleneck that historically made in-house scoring impractical.