Technical Reference
Statistical foundations, exact full-mean offset calculations, leave-one-out percentiles, assumptions, and limitations for Scoring Analysis.
Statistical Foundations
Scoring Analysis applies a calibration method built from shrinkage, full-mean reviewer offsets, leave-one-out percentiles, and robust spread estimates.
Why This Approach?
Strengths of Our Method
Explicit Eligibility
Requires three valid reviewer scores before applying an offset
Bounded Corrections
Preset scale limits and final clamping bound the possible adjusted score
Preset Evidence Weighting
Uses an explicit score-count weight for offset shrinkage and reviewer aggregation
Consistent Reviewer Offsets
Full reviewer and organisation means give one reviewer a uniform offset across their targets
Transparent & Explainable
Each adjustment can be traced back to specific patterns and benchmarks
Handles Multiple Reviewers
Weighted averaging appropriately combines information from multiple sources
Calculation Design
1. Fixed Shrinkage Heuristic
The offset uses a fixed, preset-controlled shrinkage heuristic:
- Compare the full reviewer mean with the full organisation mean
- Reduce that difference according to
nEffand the preset'soffsetK - Apply no offset shrinkage when
offsetK = 0
Mathematical form:
weight = nEff / (nEff + offsetK)
offset = (organisationMean - reviewerMean) * weightWhere:
nEff= reviewer valid score count minus oneoffsetK= shrinkage factor- Higher nEff → trust individual more
- Higher offsetK → trust organisational more
Both means use all valid scores for the response type, including the target
where it belongs. The minus one in nEff is retained as a conservative
shrinkage weight, not as an instruction to exclude the target from either mean.
2. Relation to Formal Shrinkage Methods
Shrinking reviewer-mean differences towards zero is conceptually related to empirical Bayes and James-Stein methods. The implemented fixed formula is not a fitted empirical Bayes model or the James-Stein estimator, and it has no general proof of lower error for these reviewer cohorts.
Key insight:
Formal shrinkage methods motivate caution with noisy small cohorts. Here,
offsetK is a transparent product parameter rather than an estimated model
parameter, and case-mix confounding remains.
3. Full-Mean Reviewer Offsets
For each response type, Calibration calculates the full organisation mean and each reviewer's full mean:
reviewerMean = reviewerSum / reviewerCount
organisationMean = organisationSum / organisationCount
nEff = reviewerCount - 1
offset = (organisationMean - reviewerMean) * nEff / (nEff + offsetK)A reviewer contributes an offset only with at least three total valid scores.
The reviewer and organisation means include the target score. nEff remains one
less than the reviewer's count to keep shrinkage conservative and to weight
multi-reviewer aggregation.
One reviewer therefore has a uniform offset across all of their targets for the same response type. Effective target adjustments can still differ because targets can have different reviewer combinations, explicit spread correction acts on their distance from the organisation mean, and final scores are clamped to 0-100.
4. Leave-One-Out Percentiles
Leave-one-out is reserved for raw organisation, adjusted organisation, and reviewer-history percentiles. Each target is removed from its own relevant reference distribution before ranking.
5. Robust Statistics
Using trimmed standard deviations is a form of robust statistics that:
- Reduces sensitivity to outliers
- Provides more stable estimates
- Better represents typical variation
Implementation:
For n sorted values, we remove floor(n × 0.1) observations from each tail,
then calculate population standard deviation over the retained values. Counts
below 10 therefore remove no observations. Spread uses the complete reviewer and
organisation cohorts before trimming, not target-specific LOO cohorts.
Exact per-target calculation
For each eligible reviewer attached to a target score:
reviewerMean = reviewerSum / reviewerCount
organisationMean = organisationSum / organisationCount
nEff = reviewerCount - 1
offset = (organisationMean - reviewerMean) * nEff / (nEff + offsetK)Reviewers with fewer than three total valid scores are skipped. Both means use
all valid scores of the relevant response type. Where a target has multiple
eligible reviewers, Calibration combines their offsets using nEff as the
weight.
When the preset enables spread correction and the reviewer meets that preset's spread sample threshold:
spreadWeight = reviewerCount / (reviewerCount + scaleK)
scale = clamp(1 + spreadWeight × (organisationSpread / reviewerSpread - 1), minScale, maxScale)Both spreads use the same floor-based trimming rule. Multiple reviewer scales
are combined using nEff, but only reviewers meeting the preset's scaleMinN
threshold enter that scale average. A non-qualifying reviewer's neutral default
does not dilute a qualifying reviewer's spread evidence. The final score is:
adjusted = fullOrganisationMean + combinedScale × (raw - fullOrganisationMean + combinedOffset)If spread correction is disabled, it is raw + combinedOffset. Scores are then
clamped to 0-100. Percentiles exclude the target from the relevant raw,
adjusted, or reviewer reference set and split ties at their midpoint.
6. Hierarchical Modelling
Our approach of modelling individual reviewer patterns within an organisational context is conceptually similar to hierarchical or multilevel statistical models.
Structure:
- Level 1: Individual scores
- Level 2: Reviewer patterns
- Level 3: Organisational patterns
This hierarchy naturally incorporates both individual and group information.
Methods We Explicitly Avoid
Understanding what we don't do is as important as understanding what we do.
Z-Score Normalisation
Not Used
Converting scores to standard deviations from the mean.
Why we don't use it:
| Problem | Impact |
|---|---|
| Unstable with small samples | Reviewer cohorts can be small |
| Assumes normal distributions | Organisational scores often skewed or clustered |
| Can produce extreme adjustments | Small samples lead to unstable z-scores |
| Doesn't account for uncertainty | Treats all estimates as equally reliable |
Our approach instead:
A fixed shrinkage heuristic that reduces small-cohort offsets in presets with a
positive offsetK. It does not estimate uncertainty or remove case-mix
confounding.
Simple Centring
Not Used
Just subtracting the reviewer's average and adding the organisation average.
Why we don't use it:
- Treats all reviewer averages as equally reliable, regardless of sample size
- Can over-correct when sample sizes are small
- No mechanism to handle edge cases
Our approach instead:
Preset-controlled shrinkage that reduces corrections for small eligible cohorts
when offsetK is positive. Very Aggressive deliberately does not shrink offsets.
Linear Regression Adjustment
Not Used
Using regression models to predict expected scores and adjust based on residuals.
Why we don't use it:
- Requires enough representative data for the chosen model and covariates
- Assumes linear relationships between predictors and scores
- selection and overfitting are constant concerns
- Difficult to explain to non-technical stakeholders
- Requires choosing which variables to include
Our approach instead:
Simpler, more transparent calculations that are easier to validate, explain, and implement reliably.
Ranking-Based Methods
Not Used
Converting all scores to ranks and working with ordinal positions.
Why we don't use it:
- Loses information about score magnitudes
- Can't distinguish between small and large performance gaps
- Makes it difficult to interpret individual scores
- Less intuitive for stakeholders
Our approach instead:
We preserve magnitude information while still calculating percentiles (which are rank-based) as supplementary information.
Comparison Matrix
| Method | Pros | Cons | Our decision |
|---|---|---|---|
| No calibration | Simple, preserves raw data | Does not address reviewer-mean differences | Keep raw scores visible |
| Z-scores | Well-known, mathematically elegant | Unstable with small samples, assumes shape | Do not use |
| Simple centring | Easy to understand | Gives every eligible cohort full weight | Add preset shrinkage |
| Regression models | Can model covariates and complex patterns | Requires modelling choices and more evidence | Use simpler calculations |
| Ranking-based | Less sensitive to score scale | Loses magnitude information | Preserve score magnitudes |
| Our method | Transparent, bounded, configurable | Confounds reviewer pattern with case mix | Use with stated limitations |
Assumptions and Limitations
Assumptions
Our method assumes:
Key Assumptions
1. Consistent reviewer behaviour
Reviewers maintain relatively stable rating patterns over time.
2. Representative organisational average
The organisation-wide average represents a meaningful benchmark.
3. Sufficient variety
Reviewers review a somewhat diverse group. If a reviewer only reviews exceptional performers, adjustments may be inappropriate.
4. Comparable scales
All reviewers interpret the rating scale in roughly comparable ways. Major semantic differences would require additional calibration.
5. Independence of reviews
Reviews are independent of each other. No collusion or copying between reviewers.
Limitations
Important Limitations
1. Requires sufficient data
An offset requires at least three total valid scores for a reviewer. Fewer valid scores produce no reviewer adjustment.
2. Reporting-period sensitivity
Reviewer patterns use valid scores inside the selected reporting period. Recent changes can be blended with earlier scores in that period, while older scores outside it contribute nothing. Changing the period can change reviewer eligibility, means, and adjustments.
3. Cannot correct systematic organisational biases
If everyone in an organisation rates high, we have no external benchmark to adjust against.
4. Assumes linearity
Adjustments assume bias is constant across the score range. A reviewer who is +5 points lenient at 60 is assumed to be +5 points lenient at 80.
5. Requires stable population
Works best when reviewer pools and organisational patterns are relatively stable. Major organisational changes may require recalibration.
Best Practices and Recommendations
For Organisations Implementing Calibration
1. Start with Balanced Preset
Begin with the Balanced preset and gather evidence before adjusting. It is the documented default, not a guarantee of statistical validity for every case mix.
2. Communicate Clearly
Explain to all stakeholders what calibration is, why it is beneficial, and how it works. Transparency builds trust.
3. Show Examples
Use hypothetical data to demonstrate behaviour, assumptions, and limitations before people see their actual scores.
4. Provide Manager Statistics
Make reviewer statistics available to promote transparency and help reviewers calibrate their ratings.
5. Review Periodically
Assess the system quarterly or biannually to ensure it continues working well as the organisation evolves.
6. Consider Context
If decisions are high-stakes (promotions, redundancy), lean towards conservative presets to minimise risk.
7. Train Reviewers
Even with calibration, better-trained reviewers produce more reliable data. Invest in reviewer training.
For Interpreting Results
-
Look at both raw and adjusted scores to understand the magnitude of calibration's impact
-
Consider evidence weight: the first eligible single-reviewer adjustment has three total valid reviewer scores and
nEff = 2; larger values carry more weight. For multiple reviewers, the displayed combined evidence weight is the sum of theirnEffvalues, not a count of unique independent comparisons.nEffdoes not describe which scores enter the full means. -
Use percentiles: They express relative score positions more intuitively than point scores
-
Check reviewer percentiles: They show where a score sits within each eligible reviewer's individual scoring history on average
-
Don't over-interpret small differences: Calibration does not calculate uncertainty or practical significance, so assess nearby percentiles in context
For Reviewers
-
Use the full rating scale: Compressing ratings into a narrow range reduces their informational value and leads to larger spread adjustments
-
Be consistent: The system assumes your rating pattern is stable over time
-
Review the feedback: Check your manager statistics to understand your patterns relative to peers
-
Calibrate with peers: Discuss scoring approaches with other reviewers to build shared understanding
-
Focus on quality: Calibration adjusts observed group patterns, but it cannot fix low-quality, rushed, or uninformed reviews
System Health Indicators
Monitor these metrics to ensure calibration is working properly:
1. Reviewer Variance
Compare variance between reviewers' average raw and adjusted scores. A large change confirms that calibration is materially reshaping reviewer groups, but there is no universal healthy reduction target because case mix differs.
Calculation:
Variance Reduction = 1 - (Variance of Adjusted Scores / Variance of Raw Scores)2. Reasonable Adjustment Distribution
Monitor the distribution of effective adjustments, clamping, and largest movers. Compare periods with similar cohorts rather than applying fixed healthy percentages. Material or sudden shifts should trigger review of presets, case mix, assignments, and source data.
3. Stable Reviewer Statistics
Reviewer patterns should be relatively stable over time.
Monitor:
- Reviewer means and spreads across comparable reporting periods
- Eligibility changes caused by different reporting-period cohorts
Large changes suggest:
- Data quality issues
- Inappropriate preset
- Genuine shifts in reviewer behaviour requiring investigation
4. Stable Reviewer Percentile Coverage
Reviewer percentiles average one leave-one-out percentile from each eligible reviewer's scoring history. Averaging can naturally concentrate values towards the centre, so the resulting distribution is not expected to be uniform.
Monitor: Track missing reviewer percentiles, their distribution, and material shifts between reporting periods. Sudden changes can indicate reviewer reassignment, small eligible cohorts, or changes in scoring behaviour.
5. Fairness Perception Surveys
Regularly survey stakeholders about perceived comparability and trust:
- Do people believe scores are comparable across reviewers?
- Do reviewers understand their patterns?
- Do people trust the adjustment process?
Treat survey results as product evidence to review over time, not as proof that the statistical assumptions hold.
Implementation Checklist
Before implementing calibration, ensure:
- Reviewer and organisation score counts inspected for the selected period
- Stable review process (consistent scoring rubrics)
- Stakeholder buy-in (leadership understands and supports)
- Communication plan prepared (documentation, FAQs, training)
- Manager statistics dashboard ready
- Preset selected based on organisational context
- Monitoring plan established (quarterly reviews)
- Feedback mechanisms in place (how stakeholders can raise concerns)
- Technical validation completed (spot-check calculations)
- Rollout strategy defined (pilot vs. full deployment)
Glossary of Terms
Adjustment: The amount added to or subtracted from a raw score to produce an adjusted score.
Bias: A systematic tendency to rate higher or lower than average. Also called "offset" or "central tendency difference."
Clamping: Constraining scores to remain within the valid range (0-100).
Effective Count (nEff): The reviewer's total valid score count minus one. It is a conservative shrinkage and multi-reviewer aggregation weight, not the count used to calculate the full reviewer mean.
Leave-One-Out (LOO): A percentile technique where the target is excluded from its raw, adjusted, or individual reviewer-history reference set before ranking. LOO is not used for offset means or spread distributions.
Calibration: The process of adjusting scores under the selected reviewer and organisation comparison policy.
Offset Adjustment: A uniform reviewer-level shift calculated from the full same-response organisation mean and full reviewer mean, then shrunk using nEff.
Organisation Average: The full mean across all valid organisation scores of the same response type, including the target where it belongs.
Organisation Spread: The amount of variation in scores across the organisation, measured using trimmed standard deviation.
Percentile: The target's position after counting lower comparator scores plus half of tied comparator scores, divided by the comparator count.
Preset: A predefined configuration of calibration parameters (very-conservative, conservative, balanced, aggressive, very-aggressive).
Raw Score: The original, unadjusted score as submitted by the reviewer.
Reviewer Average: The full mean across all of a reviewer's valid scores of the same response type, including the target where it belongs.
Scale Factor: A multiplier applied to adjust for differences in rating spread. Values above 1.0 expand compressed ratings; values below 1.0 compress expanded ratings.
Shrinkage: A statistical technique that balances between individual patterns and group patterns, with the balance determined by the amount of available data. Also called "regression towards the mean."
Spread: The amount of variation in a set of scores, measured using trimmed standard deviation. Also called "variability" or "dispersion."
Spread Adjustment: An adjustment that expands or compresses the distance of a score from a reference point to match organisational spread patterns.
Trimmed Standard Deviation: Population standard deviation after removing
floor(n × 0.1) values from each sorted tail. Counts below 10 are untrimmed.
Weight: In multi-reviewer situations, the relative influence of each reviewer's adjustment. It uses nEff = reviewer valid score count - 1 without excluding the target from the offset means.
Further Reading
For background on related statistical ideas, not descriptions of this exact heuristic:
- Empirical Bayes Methods: Efron & Morris (1977), "Stein's Paradox in Statistics"
- Shrinkage Estimators: James & Stein (1961), "Estimation with Quadratic Loss"
- Cross-Validation: Stone (1974), "Cross-Validatory Choice and Assessment of Statistical Predictions"
- Robust Statistics: Huber (1981), "Robust Statistics"
- Hierarchical Models: Gelman & Hill (2006), "Data Analysis Using Regression and Multilevel/Hierarchical Models"
Conclusion
Scoring Analysis uses a transparent, bounded policy for adjusting observed reviewer-mean and spread differences. It preserves raw scores, exposes evidence, and offers configurable correction strength.
It does not estimate causal reviewer bias, participant ability, confidence intervals, or statistical significance. Use adjusted results alongside raw scores, reviewer assignments, case mix, and human review.
Last updated on