How do you assess a grant whose intended result is a change in how an industry behaves, and whose evidence will not arrive for another four years?

Most funders answer with indicators. Count the people trained, the policies referenced, the reports published, the convenings held. The count is defensible, comparable across a portfolio, and easy to audit. It is also, quite often, a record of activity rather than of change.

Laudes Foundation has published a different answer, and published it in enough detail that anyone can pick it up and use it.

What the system is

Twenty-one rubrics, grouped into four categories. Category A covers process: design, implementation, monitoring and adaptation, communication and learning, and organisational and network capacity. Category B covers early and later changes, from stakeholder-informed policies through to redefined value and compelling narratives. Categories C and D cover 2025 outcomes and 2030 impact, with deliberate overlap between the three.

Each rubric is rated on the same five-point scale: harmful, unconducive, partly conducive, conducive, thrivable.

Rubrics apply to every Laudes grant above EUR 100,000. During application, a partner selects one to four rubrics from categories B and C that match what the initiative is actually trying to do, and self-assesses rubric A5 on the organisational and network capacity needed to get there. The partner rates the initiative at baseline and again in learning reports.

That is the whole architecture. It fits on a page, which is part of the point.

Why a scale, and not a count

Laudes states the reasoning plainly: change “cannot be captured by numbers alone because metrics put the focus on what can be counted, not always what’s most important.”

A rubric is not a softer version of an indicator. It is a different instrument. An indicator asks how much of something happened. A rubric asks what standard of practice the work has reached, and requires you to describe that standard in advance, in words, before anyone is being assessed against it. The discipline sits in the definition, not in the data collection.

The rating scale carries a design decision worth noticing. Laudes says its top two levels were calibrated to be “a serious stretch target, but achievable with a sustained effort”, and its lowest level is named “harmful” rather than “insufficient” or “developing”. Most performance scales top out at compliance and bottom out at politeness. This one makes it possible to record that an initiative made things worse, and makes it unlikely that anyone reaches the ceiling in year one.

That changes the meaning of a middling score. “Partly conducive” is described as minimally acceptable practice with room to scale, which is a normal and reportable position to be in. On a five-point scale where the top is easy, the same score reads as failure.

The part that is harder than it looks

The partner does the rating. That single design choice is where most of the practical difficulty sits, and Laudes does not hide it.

The framework’s answer is triangulation. Any rating must rest on more than one source of evidence, and Laudes specifies “preferably contrasting evidence and sources of data that establish independent confirmation”. The guidance then goes further than most measurement documentation is willing to:

Robust conclusions and ratings are not purely a matter of validity; credibility is also important. Consider how others would read your assessment and what additional supporting evidence is needed to convince them the rating is justified and not just an opinion.

That is an honest statement of the problem. A self-assessed qualitative rating is only worth what an outside reader will accept, and the burden of making it acceptable falls on the person doing the rating. It shifts effort from counting to evidencing, which is not obviously less work. It is different work, requiring judgement about what a sceptical reader needs rather than accuracy in a spreadsheet.

For a foundation with programme staff who know each partner, that trade is manageable. For a fund with forty portfolio companies and one impact lead, it is not, at least not without deciding in advance which rubrics matter and who verifies them.

What this would take for a fund

Three things, none of them technical.

Choose the rubrics before the pipeline, not after. Laudes has partners select from a published set at application stage. A fund that writes its rubrics after the first reporting cycle has written a justification, not a standard.

Decide what independent confirmation means for you, in advance. Not whether you will triangulate, but what a second source looks like for each rubric, and who is allowed to supply it.

Accept that a mid-scale rating is information. If a system cannot record that something is only partly working, it will not record it, and everything will report as conducive by the second year.

The question that started this was how you assess something whose result arrives in 2030. Laudes’ rubrics answer it well enough to borrow. The question they leave open, and the one worth working on, is not how to describe good practice. It is who the description has to convince, and what evidence that person will accept from an organisation reporting on itself.