WeMatter
Scoring Analysis

Interpreting Results

Understand Scoring Analysis adjusted scores, leave-one-out percentiles, and reviewer statistics.

Overview

Scoring Analysis shows several pieces of information for each score. This guide explains what each metric means and how to interpret it.


Understanding Adjustment Factors

For each adjusted score, we report several key metrics:

Reviewer Adjustment

What It Means

The Rankings table shows the combined reviewer offset before spread correction and the final 0-100 score limit. For multiple eligible reviewers, their offsets are combined using their nEff evidence weights.

Example interpretations:

  • Reviewer adjustment of -5: Reviewer offsets contribute 5 points downwards
  • Reviewer adjustment of +3: Reviewer offsets contribute 3 points upwards
  • Reviewer adjustment of 0: Eligible reviewer offsets make no contribution

One reviewer contributes a uniform offset across their targets because it uses their full mean and the full same-response organisation mean. Effective adjustments can still differ through reviewer combinations, explicit spread correction, or clamping to 0-100.

Spread correction is shown in its own participant column. The CSV also includes Total effective adjustment, which is Adjusted minus Raw after spread correction and clamping. These values can therefore differ and should not be used interchangeably.

Combined Evidence Weight

For one reviewer, this is nEff = reviewer valid score count - 1. A reviewer needs at least three total valid scores before an offset is applied. All valid scores, including the target, enter the full reviewer mean; the minus one is retained as a conservative shrinkage and aggregation weight. For multiple reviewers, Calibration sums their eligible nEff values. The result is an evidence weight, not a count of distinct or independent comparison people, because reviewer cohorts can overlap. Higher values generally mean the calculation applies less shrinkage, but should not be read as a unique sample count.

Evidence weight is not a confidence score. With a positive offsetK, a larger single-reviewer nEff produces less offset shrinkage. Very Aggressive uses offsetK = 0, so eligible offsets are unshrunk regardless of evidence weight.


Understanding Percentiles

Percentiles provide context about where a score sits relative to others in the organisation. Each raw, adjusted, and reviewer percentile excludes the target score from its own comparison set.

Raw Score Percentile

Shows where the original score ranked before calibration.

Example: A raw score percentile of 55 represents the target's position after counting lower comparator scores plus half of tied comparator scores.

Adjusted Score Percentile

Shows where the adjusted score ranks after applying the selected calibration policy.

Example: An adjusted score percentile of 68 represents the adjusted target's midpoint-tie position within the other adjusted scores.

Comparing Percentiles

The difference between raw and adjusted percentiles reveals the impact of calibration: - Higher adjusted percentile: The policy moved the score up relative to the cohort - Lower adjusted percentile: The policy moved the score down relative to the cohort - Similar percentiles: The relative position changed little

Reviewer Percentile

Shows the average of the score's leave-one-out percentile within each eligible reviewer's own scoring history.

This provides additional context about where the score sits in each reviewer's distribution. It is not a percentile against one combined cohort of people with the same reviewer set.


Complete Example Interpretation

Let's interpret a complete set of results:

Raw Score: 72
Adjusted Score: 79
Reviewer Adjustment: +5
Spread Correction: +2
Total Effective Adjustment: +7
Combined Evidence Weight: 18

Raw Score Percentile: 55th
Adjusted Score Percentile: 68th
Reviewer Percentile: 51st

What This Tells Us

The Score:

  • Original score of 72 was adjusted up to 79 (a +7 adjustment)
  • The eligible reviewers' full means were below the organisation benchmark

The Evidence Weight:

  • The combined nEff evidence weight is 18
  • This is a sum of reviewer-specific weights, not a confidence interval or a count of independent people

The Percentiles:

  • Raw (55th): The original score was slightly above average
  • Adjusted (68th): After calibration, the score moved to a higher relative position
  • Reviewer (51st): On average, the score was typical within the eligible reviewers' individual scoring histories

Conclusion: The calibration policy moved this person to a higher relative position. Their reviewer percentile shows that the score was around the middle of the eligible reviewers' individual scoring histories on average. Reviewer case mix must be considered before interpreting the cause of the movement.


When Calibration Makes the Biggest Difference

Adjustments will be largest when:

  1. Large observed mean differences: Eligible reviewer means differ materially from the organisation mean
  2. Larger score counts: Positive-offsetK presets apply less offset shrinkage as nEff grows
  3. Aggressive presets: More aggressive calibration settings are in use
  4. Spread differences: Eligible reviewer spreads differ from organisation spread when scale adjustment is enabled

Example: Large Adjustment

Raw Score: 68
Adjusted Score: 81
Adjustment: +13

Full reviewer average: 63 (full same-response organisation average: 75)
Valid reviewer scores: 25
Preset: Aggressive

This large adjustment arises because:

  • The full reviewer mean is 12 points below the organisation mean
  • The reviewer has 25 valid scores in the reporting period
  • Aggressive preset applies fuller corrections

Adjustments will be smallest when:

  1. Aligned means: Reviewer means are close to the organisation mean
  2. Limited data: Positive-offsetK presets shrink offsets from small eligible cohorts; Very Aggressive does not
  3. Conservative presets: Conservative settings limit adjustment magnitude
  4. Opposing offsets: Multiple reviewer offsets can partially cancel each other out

Example: Minimal Adjustment

Raw Score: 75
Adjusted Score: 76
Adjustment: +1

Full reviewer average: 74 (full same-response organisation average: 75)
Valid reviewer scores: 8
Preset: Balanced

This small adjustment makes sense because:

  • The full reviewer mean is close to the organisation mean
  • Moderate data availability
  • Balanced preset applies shrinkage appropriately

Reviewer Adjustment Dashboard

To promote transparency and help reviewers improve, we provide a statistics dashboard showing each reviewer's patterns.

The Adjustment table shows one combined row for each reviewer linked to at least one score in the resolved Analysis dataset. Scores counts those links, not the size of the reviewer's underlying calibration cohorts. A co-reviewed participant counts once in each linked reviewer's row and once in the organisation benchmark.

These combined values are presentation summaries. Participant calibration still uses separate internal cohorts, eligibility rules, offsets and scales, and the combined summaries are never fed back into participant calculations.

Metrics Displayed

Score Count

Number of resolved selected scores currently linked to this reviewer.

Scoring Tendency

The reviewer's combined selected raw mean compared with the organisation's unique selected raw benchmark.

Scale Use

The reviewer's combined selected raw middle-half spread compared with the organisation benchmark.

Adjustment

Usage-weighted eligible internal offsets. Missing or offset-ineligible contributions are neutral.

Spread Scale

Usage-weighted display scale when at least one contribution is eligible; otherwise Not eligible.

Scoring tendency and Scale use use the same raw-distribution profiles and five-score minimum as Distribution. They describe selected raw scores, not pooled calibration estimates or reviewer bias. The Spread scale's neutral 1.00 for an ineligible internal contribution is a display convention only; participant calibration excludes ineligible co-reviewer scales from its denominator rather than diluting eligible scales.


Using Reviewer Statistics

For Reviewers

Inspect your attributed cohort:

  • Compare your average to the organisational average
  • Compare your observed spread with the organisational spread
  • Review participant case mix before interpreting either difference

Example: If your reviewer mean is 65 and the organisation mean is 72, the selected-period cohort differs by 7 points. The statistic alone does not show whether reviewer behaviour or participant assignment caused that difference.

For Review Administrators

Identify areas for review:

  • Investigate large or changing cohort differences
  • Review scorecard quality, reviewer assignments, and case mix
  • Use underlying reviews, not aggregate statistics alone, before offering guidance

Example: A reviewer cohort with scores between 65 and 70 has low observed spread. Inspect the underlying participants and reviews before concluding that the reviewer is underusing the scale.

For Transparency

Build trust:

  • Demonstrate adjustment behaviour with examples
  • Show the formula, assumptions, and limitations
  • Allow stakeholders to verify calculations

Reading the Dashboard

Example Dashboard Entry

Reviewer: Sarah Johnson
Resolved scores: 24
Scoring tendency: Slightly lenient
Reviewer selected raw average: 77.3
Organisation selected raw average: 72.0
Scale use: Narrow
Reviewer selected raw middle-half spread: 8.2
Organisation selected raw middle-half spread: 12.5
Adjustment: +1.8
Spread scale: ×1.08

Interpretation:

  • Sarah is linked to 24 resolved selected scores; this is not the sum of her internal calibration-statistic counts
  • Her combined selected raw average is 5.3 points above the unique organisation selected-raw benchmark, producing the Slightly lenient profile
  • Her selected raw middle-half spread is narrower than the organisation benchmark, producing the Narrow profile
  • The +1.8 Adjustment is a resolved-usage-weighted summary of eligible internal offset contributions
  • The ×1.08 Spread scale is a resolved-usage-weighted display summary with at least one eligible spread contribution; when none is eligible, the dashboard shows Not eligible instead

The profile directions and Adjustment can disagree, as in this example. Scoring tendency and Scale use describe the combined selected raw distribution and can be influenced by which selected scores make up the view and by reviewer assignment case mix. Adjustment instead combines response-specific internal calibration offsets by resolved usage. These values are descriptive summaries, not pooled calibration estimates, evidence of reviewer bias, or mean participant outcomes.

The final change to any one of Sarah's targets can differ from the displayed summaries because of co-reviewers, response-specific calibration cohorts, spread application, and the 0-100 clamp.

Feedback for Sarah: Review the underlying scorecards and assignment case mix before drawing a conclusion about Sarah's rating behaviour. The statistics describe the selected cohort; they do not identify its cause.


Best Practices for Interpretation

1. Look at Both Raw and Adjusted Scores

Always review both values to understand the magnitude of calibration's impact.

Important

Do not use adjusted scores as the sole basis for high-stakes decisions. Review raw scores, adjusted scores, evidence, reviewer assignments, and case mix together.

2. Consider Evidence Weight and Score Counts

The first eligible adjustment has three total valid reviewer scores and nEff = 2. With a positive offsetK, it receives a smaller formula weight than a much larger nEff; Very Aggressive applies full offset weight at either count. All valid reviewer scores still enter the full reviewer mean.

3. Use Percentiles

Percentiles are often more intuitive than point scores for comparing relative score positions.

Why percentiles help:

  • Express relative cohort position on a common 0-100 scale
  • Provide relative context
  • Work well across different score ranges

4. Check Reviewer Percentiles

They provide context about where a score sits within each eligible reviewer's individual scoring history on average.

5. Don't Over-Interpret Small Differences

Small percentile differences may not be practically meaningful. Calibration does not calculate confidence intervals or significance tests, so interpret them alongside sample size and broader patterns.


Common Scenarios

Scenario 1: Large Percentile Shift

Raw Percentile: 45th
Adjusted Percentile: 72nd

Meaning: The selected calibration policy moved this score substantially upwards based on reviewer and organisation patterns. Check reviewer evidence and case mix before attributing the shift to reviewer strictness.

Scenario 2: Minimal Change

Raw Percentile: 63rd
Adjusted Percentile: 61st

Meaning: The adjusted relative position stayed close to the raw position.

Scenario 3: Downward Adjustment

Raw Percentile: 78th
Adjusted Percentile: 65th

Meaning: The selected calibration policy moved this score down relative to the cohort. Check reviewer evidence and case mix before attributing the shift to reviewer leniency.

Last updated on