Interpreting Results
Understand Scoring Analysis adjusted scores, leave-one-out percentiles, and reviewer statistics.
Overview
Scoring Analysis shows several pieces of information for each score. This guide explains what each metric means and how to interpret it.
Understanding Adjustment Factors
For each adjusted score, we report several key metrics:
Reviewer Adjustment
What It Means
The Rankings table shows the combined reviewer offset before spread correction
and the final 0-100 score limit. For multiple eligible reviewers, their
offsets are combined using their nEff evidence weights.
Example interpretations:
- Reviewer adjustment of -5: Reviewer offsets contribute 5 points downwards
- Reviewer adjustment of +3: Reviewer offsets contribute 3 points upwards
- Reviewer adjustment of 0: Eligible reviewer offsets make no contribution
One reviewer contributes a uniform offset across their targets because it uses their full mean and the full same-response organisation mean. Effective adjustments can still differ through reviewer combinations, explicit spread correction, or clamping to 0-100.
Spread correction is shown in its own participant column. The CSV also includes Total effective adjustment, which is Adjusted minus Raw after spread correction and clamping. These values can therefore differ and should not be used interchangeably.
Combined Evidence Weight
For one reviewer, this is nEff = reviewer valid score count - 1. A reviewer
needs at least three total valid scores before an offset is applied. All valid
scores, including the target, enter the full reviewer mean; the minus one is
retained as a conservative shrinkage and aggregation weight. For multiple
reviewers, Calibration sums their eligible nEff values. The result is an
evidence weight, not a count of distinct or independent comparison people,
because reviewer cohorts can overlap. Higher values generally mean the
calculation applies less shrinkage, but should not be read as a unique sample
count.
Evidence weight is not a confidence score. With a positive offsetK, a larger
single-reviewer nEff produces less offset shrinkage. Very Aggressive uses
offsetK = 0, so eligible offsets are unshrunk regardless of evidence weight.
Understanding Percentiles
Percentiles provide context about where a score sits relative to others in the organisation. Each raw, adjusted, and reviewer percentile excludes the target score from its own comparison set.
Raw Score Percentile
Shows where the original score ranked before calibration.
Example: A raw score percentile of 55 represents the target's position after counting lower comparator scores plus half of tied comparator scores.
Adjusted Score Percentile
Shows where the adjusted score ranks after applying the selected calibration policy.
Example: An adjusted score percentile of 68 represents the adjusted target's midpoint-tie position within the other adjusted scores.
Comparing Percentiles
The difference between raw and adjusted percentiles reveals the impact of calibration: - Higher adjusted percentile: The policy moved the score up relative to the cohort - Lower adjusted percentile: The policy moved the score down relative to the cohort - Similar percentiles: The relative position changed little
Reviewer Percentile
Shows the average of the score's leave-one-out percentile within each eligible reviewer's own scoring history.
This provides additional context about where the score sits in each reviewer's distribution. It is not a percentile against one combined cohort of people with the same reviewer set.
Complete Example Interpretation
Let's interpret a complete set of results:
Raw Score: 72
Adjusted Score: 79
Reviewer Adjustment: +5
Spread Correction: +2
Total Effective Adjustment: +7
Combined Evidence Weight: 18
Raw Score Percentile: 55th
Adjusted Score Percentile: 68th
Reviewer Percentile: 51stWhat This Tells Us
The Score:
- Original score of 72 was adjusted up to 79 (a +7 adjustment)
- The eligible reviewers' full means were below the organisation benchmark
The Evidence Weight:
- The combined
nEffevidence weight is 18 - This is a sum of reviewer-specific weights, not a confidence interval or a count of independent people
The Percentiles:
- Raw (55th): The original score was slightly above average
- Adjusted (68th): After calibration, the score moved to a higher relative position
- Reviewer (51st): On average, the score was typical within the eligible reviewers' individual scoring histories
Conclusion: The calibration policy moved this person to a higher relative position. Their reviewer percentile shows that the score was around the middle of the eligible reviewers' individual scoring histories on average. Reviewer case mix must be considered before interpreting the cause of the movement.
When Calibration Makes the Biggest Difference
Adjustments will be largest when:
- Large observed mean differences: Eligible reviewer means differ materially from the organisation mean
- Larger score counts: Positive-
offsetKpresets apply less offset shrinkage asnEffgrows - Aggressive presets: More aggressive calibration settings are in use
- Spread differences: Eligible reviewer spreads differ from organisation spread when scale adjustment is enabled
Example: Large Adjustment
Raw Score: 68
Adjusted Score: 81
Adjustment: +13
Full reviewer average: 63 (full same-response organisation average: 75)
Valid reviewer scores: 25
Preset: AggressiveThis large adjustment arises because:
- The full reviewer mean is 12 points below the organisation mean
- The reviewer has 25 valid scores in the reporting period
- Aggressive preset applies fuller corrections
Adjustments will be smallest when:
- Aligned means: Reviewer means are close to the organisation mean
- Limited data: Positive-
offsetKpresets shrink offsets from small eligible cohorts; Very Aggressive does not - Conservative presets: Conservative settings limit adjustment magnitude
- Opposing offsets: Multiple reviewer offsets can partially cancel each other out
Example: Minimal Adjustment
Raw Score: 75
Adjusted Score: 76
Adjustment: +1
Full reviewer average: 74 (full same-response organisation average: 75)
Valid reviewer scores: 8
Preset: BalancedThis small adjustment makes sense because:
- The full reviewer mean is close to the organisation mean
- Moderate data availability
- Balanced preset applies shrinkage appropriately
Reviewer Adjustment Dashboard
To promote transparency and help reviewers improve, we provide a statistics dashboard showing each reviewer's patterns.
The Adjustment table shows one combined row for each reviewer linked to at least one score in the resolved Analysis dataset. Scores counts those links, not the size of the reviewer's underlying calibration cohorts. A co-reviewed participant counts once in each linked reviewer's row and once in the organisation benchmark.
These combined values are presentation summaries. Participant calibration still uses separate internal cohorts, eligibility rules, offsets and scales, and the combined summaries are never fed back into participant calculations.
Metrics Displayed
Score Count
Number of resolved selected scores currently linked to this reviewer.
Scoring Tendency
The reviewer's combined selected raw mean compared with the organisation's unique selected raw benchmark.
Scale Use
The reviewer's combined selected raw middle-half spread compared with the organisation benchmark.
Adjustment
Usage-weighted eligible internal offsets. Missing or offset-ineligible contributions are neutral.
Spread Scale
Usage-weighted display scale when at least one contribution is eligible; otherwise Not eligible.
Scoring tendency and Scale use use the same raw-distribution profiles and five-score minimum as Distribution. They describe selected raw scores, not pooled calibration estimates or reviewer bias. The Spread scale's neutral 1.00 for an ineligible internal contribution is a display convention only; participant calibration excludes ineligible co-reviewer scales from its denominator rather than diluting eligible scales.
Using Reviewer Statistics
For Reviewers
Inspect your attributed cohort:
- Compare your average to the organisational average
- Compare your observed spread with the organisational spread
- Review participant case mix before interpreting either difference
Example: If your reviewer mean is 65 and the organisation mean is 72, the selected-period cohort differs by 7 points. The statistic alone does not show whether reviewer behaviour or participant assignment caused that difference.
For Review Administrators
Identify areas for review:
- Investigate large or changing cohort differences
- Review scorecard quality, reviewer assignments, and case mix
- Use underlying reviews, not aggregate statistics alone, before offering guidance
Example: A reviewer cohort with scores between 65 and 70 has low observed spread. Inspect the underlying participants and reviews before concluding that the reviewer is underusing the scale.
For Transparency
Build trust:
- Demonstrate adjustment behaviour with examples
- Show the formula, assumptions, and limitations
- Allow stakeholders to verify calculations
Reading the Dashboard
Example Dashboard Entry
Reviewer: Sarah Johnson
Resolved scores: 24
Scoring tendency: Slightly lenient
Reviewer selected raw average: 77.3
Organisation selected raw average: 72.0
Scale use: Narrow
Reviewer selected raw middle-half spread: 8.2
Organisation selected raw middle-half spread: 12.5
Adjustment: +1.8
Spread scale: ×1.08Interpretation:
- Sarah is linked to 24 resolved selected scores; this is not the sum of her internal calibration-statistic counts
- Her combined selected raw average is 5.3 points above the unique organisation selected-raw benchmark, producing the Slightly lenient profile
- Her selected raw middle-half spread is narrower than the organisation benchmark, producing the Narrow profile
- The +1.8 Adjustment is a resolved-usage-weighted summary of eligible internal offset contributions
- The ×1.08 Spread scale is a resolved-usage-weighted display summary with at least one eligible spread contribution; when none is eligible, the dashboard shows Not eligible instead
The profile directions and Adjustment can disagree, as in this example. Scoring tendency and Scale use describe the combined selected raw distribution and can be influenced by which selected scores make up the view and by reviewer assignment case mix. Adjustment instead combines response-specific internal calibration offsets by resolved usage. These values are descriptive summaries, not pooled calibration estimates, evidence of reviewer bias, or mean participant outcomes.
The final change to any one of Sarah's targets can differ from the displayed summaries because of co-reviewers, response-specific calibration cohorts, spread application, and the 0-100 clamp.
Feedback for Sarah: Review the underlying scorecards and assignment case mix before drawing a conclusion about Sarah's rating behaviour. The statistics describe the selected cohort; they do not identify its cause.
Best Practices for Interpretation
1. Look at Both Raw and Adjusted Scores
Always review both values to understand the magnitude of calibration's impact.
Important
Do not use adjusted scores as the sole basis for high-stakes decisions. Review raw scores, adjusted scores, evidence, reviewer assignments, and case mix together.
2. Consider Evidence Weight and Score Counts
The first eligible adjustment has three total valid reviewer scores and nEff = 2. With a positive offsetK, it receives a smaller formula weight than a much
larger nEff; Very Aggressive applies full offset weight at either count. All
valid reviewer scores still enter the full reviewer mean.
3. Use Percentiles
Percentiles are often more intuitive than point scores for comparing relative score positions.
Why percentiles help:
- Express relative cohort position on a common 0-100 scale
- Provide relative context
- Work well across different score ranges
4. Check Reviewer Percentiles
They provide context about where a score sits within each eligible reviewer's individual scoring history on average.
5. Don't Over-Interpret Small Differences
Small percentile differences may not be practically meaningful. Calibration does not calculate confidence intervals or significance tests, so interpret them alongside sample size and broader patterns.
Common Scenarios
Scenario 1: Large Percentile Shift
Raw Percentile: 45th
Adjusted Percentile: 72ndMeaning: The selected calibration policy moved this score substantially upwards based on reviewer and organisation patterns. Check reviewer evidence and case mix before attributing the shift to reviewer strictness.
Scenario 2: Minimal Change
Raw Percentile: 63rd
Adjusted Percentile: 61stMeaning: The adjusted relative position stayed close to the raw position.
Scenario 3: Downward Adjustment
Raw Percentile: 78th
Adjusted Percentile: 65thMeaning: The selected calibration policy moved this score down relative to the cohort. Check reviewer evidence and case mix before attributing the shift to reviewer leniency.
Last updated on