WeMatter
Scoring Analysis

How Calibration Works

A step-by-step explanation of the statistical calibration used by Scoring Analysis, from full reviewer means to leave-one-out percentiles.

The Calibration Process: Step by Step

Scoring Analysis applies eight carefully designed calibration steps to produce explainable score adjustments.


Step 1: Calculate Organisation-Wide Patterns

Before we can adjust individual scores, we need to understand what's "normal" for the organisation as a whole.

What We Calculate

  • Organisation Average: The mean across all valid scores of the same response type (uMatter and WeMatter separately)
  • Organisation Spread: How much variation exists in scores across the organisation

Why this matters: The organisation average provides the common benchmark for reviewer offsets. It uses the full same-response distribution, including the target score. Each reviewer's mean also uses all of their valid scores for that response type.

Technical note on spread calculation: We use a trimmed standard deviation rather than a regular standard deviation. For n sorted scores, we remove floor(n × 0.1) values from each tail, then calculate population standard deviation over the retained scores. Cohorts below 10 therefore remove no values. Spread uses the complete reviewer and organisation cohorts before trimming.


Step 2: Analyse Each Reviewer's Patterns

For every reviewer in the system, we examine their full rating behaviour for the response type across two dimensions:

Central Tendency

What it is: The full reviewer mean compared with the full same-response organisation mean. For example, a reviewer mean of 75 is five points above an organisation mean of 70.

Why we measure it: The difference supplies the reviewer offset input. It can reflect rating behaviour, participant case mix, or both; Calibration does not identify the cause.

Spread (Variability)

What it is: How much a reviewer's scores vary from their own average. A reviewer with a large spread gives scores across a wide range (e.g., from 40 to 90). A reviewer with a small spread clusters their scores tightly (e.g., from 65 to 75).

Why we measure it: The spread ratio supplies the optional scale input. Differences can reflect rating behaviour or participant case mix, so inspect the underlying cohort before drawing conclusions about a reviewer.

How we measure it: Like the organisation spread, we remove floor(n × 0.1) values from each sorted tail before calculating population standard deviation. This reduces, but does not eliminate, sensitivity to unusual values; cohorts below 10 are untrimmed.


Step 3: Calculate Consistent Reviewer Offsets

Reviewer offsets use full means rather than recalculating a different mean for each target.

Full-Mean Contract

For an eligible reviewer, Calibration compares their full mean with the full organisation mean for the same response type. The target is included in both means where it belongs.

Because the two means and the reviewer's score count do not change from target to target, one reviewer has the same offset for every one of their targets of that response type. A target's effective adjustment can still differ because it has a different combination of reviewers, spread correction changes its distance from the organisation mean, or the final score is clamped to 0-100.

Practical Example

Imagine a reviewer has given 10 scores: [65, 70, 72, 68, 75, 71, 69, 73, 67, 70]. The average is 70.

When we adjust the score of 75:

  • We use the full reviewer mean of 70
  • The score of 75 remains included in that mean
  • We compare 70 with the full same-response organisation mean

When we adjust the score of 68:

  • We still use the full reviewer mean of 70
  • The score of 68 remains included in that mean
  • The reviewer therefore contributes the same offset to both targets

Why this matters: The offset represents one observed reviewer-cohort mean difference rather than a target-specific recalculation. Leave-one-out is used later for raw, adjusted, and reviewer percentiles only.


Step 4: Apply Shrinkage

Raw statistics can be misleading when we do not have much data. A reviewer with only two valid scores receives no offset adjustment because Calibration requires at least three total valid reviewer scores. After that threshold, presets with a positive offsetK shrink early adjustments towards zero. Very Aggressive uses offsetK = 0, so it deliberately applies the full eligible offset.

The Problem with Small Samples

With very few reviews, a reviewer mean can be highly sensitive to the participants in that reporting-period cohort. The formula does not estimate whether the difference comes from reviewer behaviour or case mix.

The solution: Preset-controlled shrinkage

We use a mathematical technique called shrinkage that balances between:

  • The reviewer's personal rating pattern
  • The organisation's overall pattern

How it works:

The adjustment we apply is a weighted average:

  • With a positive offsetK, more data gives the reviewer pattern more weight
  • With a positive offsetK, less data keeps the offset closer to zero
  • With offsetK = 0, every eligible reviewer gets full offset weight

The formula involves a parameter called offsetK (the shrinkage factor). The weight given to the reviewer's individual pattern is:

For a target score, nEff is the reviewer's total valid score count minus one:

nEff = Reviewer Score Count - 1

Weight = nEff / (nEff + offsetK)

A reviewer must have at least three total valid scores before any offset is applied. Subtracting one is a deliberately conservative choice for shrinkage and multi-reviewer aggregation; it does not mean the target is removed from either mean.

Very Aggressive is the explicit exception to small-sample shrinkage: its offsetK is zero, so the weight is 1 as soon as the three-score eligibility threshold is met.

Example: Balanced Preset (offsetK = 12)

  • Reviewer with 3 total valid scores: nEff = 2; weight = 2/(2+12) = 0.143
  • Reviewer with 13 total valid scores: nEff = 12; weight = 12/(12+12) = 0.5
  • Reviewer with 49 total valid scores: nEff = 48; weight = 48/(48+12) = 0.8

Why we use this heuristic:

The fixed weight is inspired by formal shrinkage methods, but it is not itself a fitted empirical Bayes or James-Stein estimator. It provides a transparent way to reduce offsets from small eligible cohorts in presets with a positive offsetK. It does not remove participant case-mix confounding.


Step 5: Calculate Adjustments

Now we can calculate how much to adjust each score. This happens in two stages:

Offset Adjustment (Aligning Mean Differences)

What we calculate: The difference between the full same-response organisation mean and full reviewer mean, multiplied by the nEff shrinkage weight.

offset = (organisationMean - reviewerMean) * nEff / (nEff + offsetK)

Example:

  • Full same-response organisation average: 70
  • Full reviewer average: 78
  • Shrinkage weight: 0.5 (based on number of reviews)
  • Offset adjustment: (70 - 78) × 0.5 = -4

For a raw score of 75:

  • Adjusted score: 75 + (-4) = 71

The reviewer mean is 8 points above the organisation mean. This example's formula weight is 0.5, so the reviewer contributes a -4 offset.

That reviewer contributes the same -4 offset to each of their targets for this response type. Their targets' effective adjustments may nevertheless differ through reviewer combinations, spread correction, or clamping.

Spread Adjustment (Aligning Spread Differences)

This adjustment is optional and only applied in certain presets.

What we calculate: The ratio between the organisation's spread and the reviewer's spread.

Example:

  • Organisation spread: 15 points
  • Reviewer spread: 10 points
  • Spread ratio: 15/10 = 1.5
  • Shrinkage weight: 0.6 (based on number of reviews)
  • Scale factor: (1 - 0.6) × 1.0 + 0.6 × 1.5 = 0.4 + 0.9 = 1.3

For a score of 75 with a full organisation mean of 70:

  • Distance from the organisation mean: 75 - 70 = 5
  • After spread adjustment: 5 × 1.3 = 6.5
  • Adjusted score: 70 + 6.5 = 76.5

The reviewer spread is lower than the organisation spread, so this scale factor expands the score's distance from the organisation mean. The calculation does not identify whether the spread difference comes from reviewer behaviour or participant case mix.

Combined Adjustment

When both offset and spread adjustments are enabled, they work together:

  1. Calculate each eligible reviewer's offset from the full reviewer and same-response organisation means.
  2. Combine reviewer offsets using nEff as the weight.
  3. Calculate each qualifying reviewer's scale after removing floor(n × 0.1) scores from each distribution tail, then combine scales using nEff. Reviewers below the preset's spread threshold do not enter this average; their neutral default cannot dilute qualifying spread evidence.
  4. Scale the target's distance from the full organisation mean, including the combined offset.

Mathematical form: adjusted = fullOrganisationMean + combinedScale x (raw - fullOrganisationMean + combinedOffset)

Without spread adjustment, adjusted = raw + combinedOffset. The final result is clamped to 0-100.

Limits on spread adjustment:

We never allow spread adjustments to become extreme. Scale factors are constrained to reasonable ranges (typically between 0.7 and 1.35, depending on preset). This prevents over-correction and ensures adjustments remain proportionate and interpretable.


Step 6: Handle Multiple Reviewers

Many reviews involve multiple reviewers (for example, a person might have 3 managers reviewing them). When this happens, we need to combine the adjustments from each reviewer.

How we combine adjustments:

We calculate a weighted average of all reviewers' adjustments. Each reviewer's weight is nEff, their total valid score count minus one. The subtraction is a conservative weighting choice; all valid scores still contribute to the full reviewer mean.

Example

A person is reviewed by 3 managers:

  • Manager A: Has 21 total valid scores (nEff = 20), suggests offset of -3, scale factor of 1.1
  • Manager B: Has 6 total valid scores (nEff = 5), suggests offset of +2, scale factor of 1.0
  • Manager C: Has 16 total valid scores (nEff = 15), suggests offset of -1, scale factor of 1.05

Weights: 20, 5, 15 (total: 40)

The displayed combined evidence weight is therefore 40. It is a sum of reviewer-specific nEff weights, not 40 unique or independent people; the same score can contribute through more than one reviewer cohort.

Combined offset:

  • (-3 × 20 + 2 × 5 - 1 × 15) / 40 = (-60 + 10 - 15) / 40 = -1.625

Combined scale factor:

  • Manager B has too few scores for spread correction under any enabled preset, so only Managers A and C enter the scale average
  • (1.1 × 20 + 1.05 × 15) / 35 = 1.079

All three eligible reviewer offsets contribute to the combined offset. Scale uses only reviewers who separately meet the preset's spread threshold.


Step 7: Ensure Scores Stay Valid

After all adjustments, we ensure scores remain within the valid range.

Clamping:

All final scores are constrained to fall between 0 and 100. If an adjustment would push a score below 0, we set it to 0. If it would push above 100, we set it to 100.

Why this is necessary:

Mathematical adjustments can sometimes produce values outside the valid range, especially when correcting for reviewers with extreme biases or spreads. Clamping ensures all final scores are interpretable and comparable.

Monitor Clamping

Review clamping rates for each reporting period. Frequent clamping can signal an aggressive preset, unusual reviewer patterns, or a cohort that needs closer inspection.


Step 8: Calculate Percentile Rankings

After calibration, we calculate where each score sits relative to others in the organisation.

Percentile ranking: A percentile counts lower comparator scores plus half of comparator scores tied with the target, divided by the number of comparators. With no ties, a 75th percentile means 75% of comparator scores are lower. With ties, it represents that midpoint-tie position rather than a strictly-lower percentage.

We Calculate Three Percentiles

  1. Raw Score Percentile: Where the original, unadjusted score ranks
  2. Adjusted Score Percentile: Where the calibrated score ranks
  3. Reviewer Percentile: The average of the score's leave-one-out percentile within each eligible reviewer's own scoring history

Why all three?

Comparing these percentiles shows the impact of calibration and provides different contexts for understanding scoring outcomes.

Leave-one-out percentiles:

Leave-one-out applies to percentiles, not to the reviewer or organisation means used for offsets. A score's raw, adjusted, and reviewer percentiles are calculated against the relevant reference set with that target excluded.

Ranking settings

The Analysis response setting selects one reviewer-backed set of ratings per participant and controls Rankings, Distribution, and Adjustment. The preferred response is considered first. The fallback orders are:

  • Feedback Workshop: Feedback Workshop, then Manager Review
  • Manager Review: Manager Review, then Feedback Workshop

Self-Reflection is neither a choice nor a fallback. Participants without a scorable reviewer-backed response are excluded, and the reporting-period preview identifies iMatter-only, unscorable, and superseded scorecards. The resolved source is shown for every included participant. Distribution includes one resolved score per participant, so its organisation view can mix Feedback Workshop and Manager Review. Adjustment shows one combined row per linked reviewer. Its Scores value counts linked resolved ratings, while its Adjustment and Spread scale are usage-weighted descriptive summaries of the internal response-specific factors. Scoring tendency and Scale use compare combined selected raw distributions with the unique organisation selected-raw benchmark. None of these combined presentation values replaces the separate internal calculation cohorts or participant mathematics.

Score weighting blends raw and adjusted score points:

WeightingRawAdjusted
Raw only100%0%
Favour raw75%25%
Equally weighted50%50%
Favour adjusted25%75%
Adjusted only0%100%

Ties use the underlying score, percentile, and evidence values for deterministic ordering. Equivalent results share the best tier reached by their tied group. Tier percentages divide the final ordered list; changing ranking settings does not recalculate raw or adjusted scores.

Last updated on