WeMatter
Scoring Analysis

Technical Reference

Statistical foundations, exact full-mean offset calculations, leave-one-out percentiles, assumptions, and limitations for Scoring Analysis.

Statistical Foundations

Scoring Analysis applies a calibration method built from shrinkage, full-mean reviewer offsets, leave-one-out percentiles, and robust spread estimates.


Why This Approach?

Strengths of Our Method

Explicit Eligibility

Requires three valid reviewer scores before applying an offset

Bounded Corrections

Preset scale limits and final clamping bound the possible adjusted score

Preset Evidence Weighting

Uses an explicit score-count weight for offset shrinkage and reviewer aggregation

Consistent Reviewer Offsets

Full reviewer and organisation means give one reviewer a uniform offset across their targets

Transparent & Explainable

Each adjustment can be traced back to specific patterns and benchmarks

Handles Multiple Reviewers

Weighted averaging appropriately combines information from multiple sources


Calculation Design

1. Fixed Shrinkage Heuristic

The offset uses a fixed, preset-controlled shrinkage heuristic:

  • Compare the full reviewer mean with the full organisation mean
  • Reduce that difference according to nEff and the preset's offsetK
  • Apply no offset shrinkage when offsetK = 0

Mathematical form:

weight = nEff / (nEff + offsetK)
offset = (organisationMean - reviewerMean) * weight

Where:

  • nEff = reviewer valid score count minus one
  • offsetK = shrinkage factor
  • Higher nEff → trust individual more
  • Higher offsetK → trust organisational more

Both means use all valid scores for the response type, including the target where it belongs. The minus one in nEff is retained as a conservative shrinkage weight, not as an instruction to exclude the target from either mean.


2. Relation to Formal Shrinkage Methods

Shrinking reviewer-mean differences towards zero is conceptually related to empirical Bayes and James-Stein methods. The implemented fixed formula is not a fitted empirical Bayes model or the James-Stein estimator, and it has no general proof of lower error for these reviewer cohorts.

Key insight:

Formal shrinkage methods motivate caution with noisy small cohorts. Here, offsetK is a transparent product parameter rather than an estimated model parameter, and case-mix confounding remains.


3. Full-Mean Reviewer Offsets

For each response type, Calibration calculates the full organisation mean and each reviewer's full mean:

reviewerMean = reviewerSum / reviewerCount
organisationMean = organisationSum / organisationCount
nEff = reviewerCount - 1
offset = (organisationMean - reviewerMean) * nEff / (nEff + offsetK)

A reviewer contributes an offset only with at least three total valid scores. The reviewer and organisation means include the target score. nEff remains one less than the reviewer's count to keep shrinkage conservative and to weight multi-reviewer aggregation.

One reviewer therefore has a uniform offset across all of their targets for the same response type. Effective target adjustments can still differ because targets can have different reviewer combinations, explicit spread correction acts on their distance from the organisation mean, and final scores are clamped to 0-100.


4. Leave-One-Out Percentiles

Leave-one-out is reserved for raw organisation, adjusted organisation, and reviewer-history percentiles. Each target is removed from its own relevant reference distribution before ranking.


5. Robust Statistics

Using trimmed standard deviations is a form of robust statistics that:

  • Reduces sensitivity to outliers
  • Provides more stable estimates
  • Better represents typical variation

Implementation:

For n sorted values, we remove floor(n × 0.1) observations from each tail, then calculate population standard deviation over the retained values. Counts below 10 therefore remove no observations. Spread uses the complete reviewer and organisation cohorts before trimming, not target-specific LOO cohorts.

Exact per-target calculation

For each eligible reviewer attached to a target score:

reviewerMean = reviewerSum / reviewerCount
organisationMean = organisationSum / organisationCount
nEff = reviewerCount - 1
offset = (organisationMean - reviewerMean) * nEff / (nEff + offsetK)

Reviewers with fewer than three total valid scores are skipped. Both means use all valid scores of the relevant response type. Where a target has multiple eligible reviewers, Calibration combines their offsets using nEff as the weight.

When the preset enables spread correction and the reviewer meets that preset's spread sample threshold:

spreadWeight = reviewerCount / (reviewerCount + scaleK)
scale = clamp(1 + spreadWeight × (organisationSpread / reviewerSpread - 1), minScale, maxScale)

Both spreads use the same floor-based trimming rule. Multiple reviewer scales are combined using nEff, but only reviewers meeting the preset's scaleMinN threshold enter that scale average. A non-qualifying reviewer's neutral default does not dilute a qualifying reviewer's spread evidence. The final score is:

adjusted = fullOrganisationMean + combinedScale × (raw - fullOrganisationMean + combinedOffset)

If spread correction is disabled, it is raw + combinedOffset. Scores are then clamped to 0-100. Percentiles exclude the target from the relevant raw, adjusted, or reviewer reference set and split ties at their midpoint.


6. Hierarchical Modelling

Our approach of modelling individual reviewer patterns within an organisational context is conceptually similar to hierarchical or multilevel statistical models.

Structure:

  • Level 1: Individual scores
  • Level 2: Reviewer patterns
  • Level 3: Organisational patterns

This hierarchy naturally incorporates both individual and group information.


Methods We Explicitly Avoid

Understanding what we don't do is as important as understanding what we do.

Z-Score Normalisation

Not Used

Converting scores to standard deviations from the mean.

Why we don't use it:

ProblemImpact
Unstable with small samplesReviewer cohorts can be small
Assumes normal distributionsOrganisational scores often skewed or clustered
Can produce extreme adjustmentsSmall samples lead to unstable z-scores
Doesn't account for uncertaintyTreats all estimates as equally reliable

Our approach instead:

A fixed shrinkage heuristic that reduces small-cohort offsets in presets with a positive offsetK. It does not estimate uncertainty or remove case-mix confounding.


Simple Centring

Not Used

Just subtracting the reviewer's average and adding the organisation average.

Why we don't use it:

  • Treats all reviewer averages as equally reliable, regardless of sample size
  • Can over-correct when sample sizes are small
  • No mechanism to handle edge cases

Our approach instead:

Preset-controlled shrinkage that reduces corrections for small eligible cohorts when offsetK is positive. Very Aggressive deliberately does not shrink offsets.


Linear Regression Adjustment

Not Used

Using regression models to predict expected scores and adjust based on residuals.

Why we don't use it:

  • Requires enough representative data for the chosen model and covariates
  • Assumes linear relationships between predictors and scores
  • selection and overfitting are constant concerns
  • Difficult to explain to non-technical stakeholders
  • Requires choosing which variables to include

Our approach instead:

Simpler, more transparent calculations that are easier to validate, explain, and implement reliably.


Ranking-Based Methods

Not Used

Converting all scores to ranks and working with ordinal positions.

Why we don't use it:

  • Loses information about score magnitudes
  • Can't distinguish between small and large performance gaps
  • Makes it difficult to interpret individual scores
  • Less intuitive for stakeholders

Our approach instead:

We preserve magnitude information while still calculating percentiles (which are rank-based) as supplementary information.


Comparison Matrix

MethodProsConsOur decision
No calibrationSimple, preserves raw dataDoes not address reviewer-mean differencesKeep raw scores visible
Z-scoresWell-known, mathematically elegantUnstable with small samples, assumes shapeDo not use
Simple centringEasy to understandGives every eligible cohort full weightAdd preset shrinkage
Regression modelsCan model covariates and complex patternsRequires modelling choices and more evidenceUse simpler calculations
Ranking-basedLess sensitive to score scaleLoses magnitude informationPreserve score magnitudes
Our methodTransparent, bounded, configurableConfounds reviewer pattern with case mixUse with stated limitations

Assumptions and Limitations

Assumptions

Our method assumes:

Key Assumptions

1. Consistent reviewer behaviour

Reviewers maintain relatively stable rating patterns over time.

2. Representative organisational average

The organisation-wide average represents a meaningful benchmark.

3. Sufficient variety

Reviewers review a somewhat diverse group. If a reviewer only reviews exceptional performers, adjustments may be inappropriate.

4. Comparable scales

All reviewers interpret the rating scale in roughly comparable ways. Major semantic differences would require additional calibration.

5. Independence of reviews

Reviews are independent of each other. No collusion or copying between reviewers.


Limitations

Important Limitations

1. Requires sufficient data

An offset requires at least three total valid scores for a reviewer. Fewer valid scores produce no reviewer adjustment.

2. Reporting-period sensitivity

Reviewer patterns use valid scores inside the selected reporting period. Recent changes can be blended with earlier scores in that period, while older scores outside it contribute nothing. Changing the period can change reviewer eligibility, means, and adjustments.

3. Cannot correct systematic organisational biases

If everyone in an organisation rates high, we have no external benchmark to adjust against.

4. Assumes linearity

Adjustments assume bias is constant across the score range. A reviewer who is +5 points lenient at 60 is assumed to be +5 points lenient at 80.

5. Requires stable population

Works best when reviewer pools and organisational patterns are relatively stable. Major organisational changes may require recalibration.


Best Practices and Recommendations

For Organisations Implementing Calibration

1. Start with Balanced Preset

Begin with the Balanced preset and gather evidence before adjusting. It is the documented default, not a guarantee of statistical validity for every case mix.

2. Communicate Clearly

Explain to all stakeholders what calibration is, why it is beneficial, and how it works. Transparency builds trust.

3. Show Examples

Use hypothetical data to demonstrate behaviour, assumptions, and limitations before people see their actual scores.

4. Provide Manager Statistics

Make reviewer statistics available to promote transparency and help reviewers calibrate their ratings.

5. Review Periodically

Assess the system quarterly or biannually to ensure it continues working well as the organisation evolves.

6. Consider Context

If decisions are high-stakes (promotions, redundancy), lean towards conservative presets to minimise risk.

7. Train Reviewers

Even with calibration, better-trained reviewers produce more reliable data. Invest in reviewer training.


For Interpreting Results

  1. Look at both raw and adjusted scores to understand the magnitude of calibration's impact

  2. Consider evidence weight: the first eligible single-reviewer adjustment has three total valid reviewer scores and nEff = 2; larger values carry more weight. For multiple reviewers, the displayed combined evidence weight is the sum of their nEff values, not a count of unique independent comparisons. nEff does not describe which scores enter the full means.

  3. Use percentiles: They express relative score positions more intuitively than point scores

  4. Check reviewer percentiles: They show where a score sits within each eligible reviewer's individual scoring history on average

  5. Don't over-interpret small differences: Calibration does not calculate uncertainty or practical significance, so assess nearby percentiles in context


For Reviewers

  1. Use the full rating scale: Compressing ratings into a narrow range reduces their informational value and leads to larger spread adjustments

  2. Be consistent: The system assumes your rating pattern is stable over time

  3. Review the feedback: Check your manager statistics to understand your patterns relative to peers

  4. Calibrate with peers: Discuss scoring approaches with other reviewers to build shared understanding

  5. Focus on quality: Calibration adjusts observed group patterns, but it cannot fix low-quality, rushed, or uninformed reviews


System Health Indicators

Monitor these metrics to ensure calibration is working properly:

1. Reviewer Variance

Compare variance between reviewers' average raw and adjusted scores. A large change confirms that calibration is materially reshaping reviewer groups, but there is no universal healthy reduction target because case mix differs.

Calculation:

Variance Reduction = 1 - (Variance of Adjusted Scores / Variance of Raw Scores)

2. Reasonable Adjustment Distribution

Monitor the distribution of effective adjustments, clamping, and largest movers. Compare periods with similar cohorts rather than applying fixed healthy percentages. Material or sudden shifts should trigger review of presets, case mix, assignments, and source data.


3. Stable Reviewer Statistics

Reviewer patterns should be relatively stable over time.

Monitor:

  • Reviewer means and spreads across comparable reporting periods
  • Eligibility changes caused by different reporting-period cohorts

Large changes suggest:

  • Data quality issues
  • Inappropriate preset
  • Genuine shifts in reviewer behaviour requiring investigation

4. Stable Reviewer Percentile Coverage

Reviewer percentiles average one leave-one-out percentile from each eligible reviewer's scoring history. Averaging can naturally concentrate values towards the centre, so the resulting distribution is not expected to be uniform.

Monitor: Track missing reviewer percentiles, their distribution, and material shifts between reporting periods. Sudden changes can indicate reviewer reassignment, small eligible cohorts, or changes in scoring behaviour.


5. Fairness Perception Surveys

Regularly survey stakeholders about perceived comparability and trust:

  • Do people believe scores are comparable across reviewers?
  • Do reviewers understand their patterns?
  • Do people trust the adjustment process?

Treat survey results as product evidence to review over time, not as proof that the statistical assumptions hold.


Implementation Checklist

Before implementing calibration, ensure:

  • Reviewer and organisation score counts inspected for the selected period
  • Stable review process (consistent scoring rubrics)
  • Stakeholder buy-in (leadership understands and supports)
  • Communication plan prepared (documentation, FAQs, training)
  • Manager statistics dashboard ready
  • Preset selected based on organisational context
  • Monitoring plan established (quarterly reviews)
  • Feedback mechanisms in place (how stakeholders can raise concerns)
  • Technical validation completed (spot-check calculations)
  • Rollout strategy defined (pilot vs. full deployment)

Glossary of Terms

Adjustment: The amount added to or subtracted from a raw score to produce an adjusted score.

Bias: A systematic tendency to rate higher or lower than average. Also called "offset" or "central tendency difference."

Clamping: Constraining scores to remain within the valid range (0-100).

Effective Count (nEff): The reviewer's total valid score count minus one. It is a conservative shrinkage and multi-reviewer aggregation weight, not the count used to calculate the full reviewer mean.

Leave-One-Out (LOO): A percentile technique where the target is excluded from its raw, adjusted, or individual reviewer-history reference set before ranking. LOO is not used for offset means or spread distributions.

Calibration: The process of adjusting scores under the selected reviewer and organisation comparison policy.

Offset Adjustment: A uniform reviewer-level shift calculated from the full same-response organisation mean and full reviewer mean, then shrunk using nEff.

Organisation Average: The full mean across all valid organisation scores of the same response type, including the target where it belongs.

Organisation Spread: The amount of variation in scores across the organisation, measured using trimmed standard deviation.

Percentile: The target's position after counting lower comparator scores plus half of tied comparator scores, divided by the comparator count.

Preset: A predefined configuration of calibration parameters (very-conservative, conservative, balanced, aggressive, very-aggressive).

Raw Score: The original, unadjusted score as submitted by the reviewer.

Reviewer Average: The full mean across all of a reviewer's valid scores of the same response type, including the target where it belongs.

Scale Factor: A multiplier applied to adjust for differences in rating spread. Values above 1.0 expand compressed ratings; values below 1.0 compress expanded ratings.

Shrinkage: A statistical technique that balances between individual patterns and group patterns, with the balance determined by the amount of available data. Also called "regression towards the mean."

Spread: The amount of variation in a set of scores, measured using trimmed standard deviation. Also called "variability" or "dispersion."

Spread Adjustment: An adjustment that expands or compresses the distance of a score from a reference point to match organisational spread patterns.

Trimmed Standard Deviation: Population standard deviation after removing floor(n × 0.1) values from each sorted tail. Counts below 10 are untrimmed.

Weight: In multi-reviewer situations, the relative influence of each reviewer's adjustment. It uses nEff = reviewer valid score count - 1 without excluding the target from the offset means.


Further Reading

For background on related statistical ideas, not descriptions of this exact heuristic:

  • Empirical Bayes Methods: Efron & Morris (1977), "Stein's Paradox in Statistics"
  • Shrinkage Estimators: James & Stein (1961), "Estimation with Quadratic Loss"
  • Cross-Validation: Stone (1974), "Cross-Validatory Choice and Assessment of Statistical Predictions"
  • Robust Statistics: Huber (1981), "Robust Statistics"
  • Hierarchical Models: Gelman & Hill (2006), "Data Analysis Using Regression and Multilevel/Hierarchical Models"

Conclusion

Scoring Analysis uses a transparent, bounded policy for adjusting observed reviewer-mean and spread differences. It preserves raw scores, exposes evidence, and offers configurable correction strength.

It does not estimate causal reviewer bias, participant ability, confidence intervals, or statistical significance. Use adjusted results alongside raw scores, reviewer assignments, case mix, and human review.

Last updated on