Frequently Asked Questions
Common questions about Scoring Analysis and its statistical calibration method.
Common Questions
This page addresses frequently asked questions about Scoring Analysis and the calibration method it applies. If you cannot find your question here, contact your system administrator.
About Fairness
"Isn't this unfair to people with strict reviewers?"
Compare Both Scores
Calibration is intended to reduce the effect of established reviewer-mean differences. Always compare adjusted results with raw scores, reviewer evidence, and assignment context.
When reviewer case mix is reasonably representative, calibration can improve comparability by adjusting established reviewer-mean differences towards the organisation benchmark. It does not prove that a difference was caused by the reviewer.
Example:
- Person A scores 65 in a reviewer cohort with a mean of 60
- Person B scores 75 in a reviewer cohort with a mean of 80
- Calibration can move Person A up and Person B down relative to the organisation benchmark
- The adjusted values support a policy-based comparison; they do not prove equal underlying performance
"What if a reviewer is strict because they only review high performers?"
This is an important limitation. Calibration observes score patterns; it cannot separate reviewer strictness from the performance mix assigned to that reviewer.
What shrinkage does:
- Keeps corrections small when a reviewer has little evidence
- Gives persistent reviewer-mean differences more weight as evidence grows
- Does not determine whether that difference comes from reviewer behaviour or participant case mix
If a reviewer systematically receives unusually strong or weak participants, their offset can be inappropriate even with many scores. Review assignments, reviewer statistics, and raw scores should be reviewed alongside adjusted results, especially for high-stakes decisions.
Calibration improves comparability under a reasonably representative case-mix assumption; it is not a causal estimate of reviewer bias.
About Gaming
"Can people game the system by choosing lenient reviewers?"
Calibration Is Not a Manipulation-Control System
Calibration can reduce the advantage of an established lenient rating pattern, but it cannot guarantee that strategic reviewer selection has no effect.
What Calibration does:
- Compares each eligible reviewer's full mean with the full same-response organisation mean
- Shrinks offsets from smaller eligible cohorts when
offsetKis positive - Combines offsets when a target has multiple reviewers
- Reports reviewer percentiles for additional context
The target remains part of the full reviewer and organisation means. Calibration does not make that target independent of its own offset calculation and must not be treated as proof that strategic reviewer selection cannot matter.
What organisations still need:
- Clear reviewer-assignment rules
- Monitoring for unusual selection or scoring patterns
- Multiple reviewers where appropriate
- Review and audit processes for high-stakes decisions
About New Reviewers
"What happens when a new reviewer joins with no history?"
New reviewers start with zero adjustment because we have no data to base an adjustment on. Their first reviews receive no correction.
Gradual evidence weighting:
- After 3 total valid scores: The first possible adjustment uses the full reviewer mean with
nEff = 2; presets with a positiveoffsetKshrink it towards zero, while Very Aggressive does not - After 5-10 reviews: Positive-
offsetKpresets give the observed reviewer-mean difference more weight - After 20+ reviews: Positive-
offsetKpresets move closer to their full eligible offset
Why this approach:
This progression makes positive-offsetK presets more conservative with small
cohorts. Very Aggressive is intentionally different and should only be used with
deliberate evidence and oversight.
New Reviewer Protection
The three-score eligibility threshold applies to every preset. Additional
small-sample shrinkage applies only when the selected preset has a positive
offsetK.
About Self-Assessments
"Why are iMatter scores never adjusted?"
iMatter scores represent self-reflection—how the person sees their own performance.
They are excluded from Scoring Analysis because the dashboard compares and calibrates reviewer-backed responses. iMatter-only scorecards remain visible as excluded records in the reporting-period preview so their omission is explicit.
Key reasons:
-
No external reviewer bias: There's no external reviewer introducing rating tendencies that need correction
-
Different purpose: Self-assessments serve a different function (self-awareness and development) than external reviews (evaluation and comparison)
-
Preserves genuine perception: The score represents the individual's genuine self-assessment, and adjusting it would distort their self-perception rather than correct a measurement bias
Example:
If someone rates themselves 65 (they're self-critical) and the organisation average is 72:
- Adjusting this to 72 would misrepresent their self-perception
- The value is in understanding their self-view, not in having a "calibrated" self-view
About Score Changes
"Can calibration ever make a score worse?"
Yes, and this is intentional.
It Changes the Comparison Basis
A reviewer mean above the organisation benchmark can produce a downward offset. This aligns an observed reviewer-mean difference with the organisation benchmark; it does not prove the reviewer caused that difference.
How to interpret it:
The adjusted score is a policy-based comparison using reviewer and organisation patterns from the selected reporting period. Review raw scores, reviewer case mix, and evidence alongside it.
Bidirectional adjustment:
Observed low reviewer means can produce upward offsets, while high reviewer means can produce downward offsets. That symmetry does not establish equivalent participant performance.
Example:
- High reviewer-mean example: 82 → 75
- Low reviewer-mean example: 68 → 75
- Result: Both align to the same policy-adjusted score, without proving equal underlying performance
About System Health
"How do I know if calibration is working properly?"
Inspect several indicators together; none proves that calibration is correct for a particular case mix:
Reviewer Variance
Compare raw and adjusted reviewer-group variance to see how materially the policy reshapes groups.
Stakeholder Understanding
Check whether stakeholders understand the formula, assumptions, and limitations.
Adjustment Distribution
Monitor effective adjustments, largest movers, and clamping between comparable periods.
Stable Manager Statistics
Reviewer patterns should be relatively stable over time. Wild swings suggest data quality issues.
Stable Reviewer Percentile Coverage
Reviewer percentile availability and distributions should remain stable between comparable reporting periods.
About Presets
"What's the difference between 'balanced' and 'aggressive' presets in practice?"
For a reviewer mean of 62, organisation mean of 70, and 15 valid reviewer scores:
Balanced Preset:
- Reviewer offset: +4.31 points (
8 × 14 / (14 + 12)) - Spread correction requires 12 scores and is clamped to 0.88-1.15
- More offset shrinkage and narrower scale limits
Aggressive Preset:
- Reviewer offset: +5.89 points (
8 × 14 / (14 + 5)) - Spread correction requires 10 scores and is clamped to 0.80-1.25
- Less offset shrinkage and wider scale limits
Effective target adjustments can differ from these reviewer offsets because of co-reviewers, spread correction, or clamping.
Choosing Between Them
Use Balanced as default. Only move to Aggressive when you have clear evidence of substantial reviewer inconsistency that Balanced doesn't adequately address.
About Transparency
"How can I verify the calculations are correct?"
Reviewer Adjustment table:
Every linked reviewer has one combined summary row showing:
- Resolved score count
- Selected-raw scoring tendency and scale use
- Usage-weighted Adjustment
- Eligible Spread scale or Not eligible
These are descriptive presentation summaries. Participant details retain the response-specific adjustment, spread correction, evidence and final adjusted score used for that person.
Reproducible calculations:
All adjustments follow documented formulas that can be verified:
- Calculate organisation average and spread
- Calculate the full same-response organisation mean and each full reviewer mean; calculate spread after removing
floor(n × 0.1)values from each sorted tail - Apply shrinkage based on sample size
- Calculate offset and scale adjustments
- Apply to raw scores
Audit trail:
Each adjusted score shows:
- Raw score
- Adjustment factor
- Sample size used
- Final adjusted score
Anyone can work backwards from these values to verify the calculation.
About Edge Cases
"What happens if everyone in the organisation rates high?"
Limitation
Calibration cannot correct for systematic organisational biases. If everyone rates high, we have no external benchmark to adjust against.
What this means:
- If the entire organisation averages 85 when the intended scale midpoint is 70, calibration won't fix this
- Calibration adjusts for differences between reviewers, not for organisation-wide patterns
- Addressing systematic organisational bias requires recalibration or training
Example:
- Organisation A averages 85 (everyone rates high)
- Organisation B averages 60 (everyone rates low)
- Within each organisation, calibration applies that organisation's selected comparison policy
- Between organisations, scores aren't directly comparable
"What about reviewers with very few reviews?"
Reviewers with fewer than three total valid scores receive no adjustment. From
three valid scores onwards, positive-offsetK presets shrink early adjustments.
Very Aggressive applies full offset weight immediately after eligibility:
Protection mechanisms:
- Eligibility threshold: Fewer than three valid scores produce no offset
- Preset shrinkage: Positive
offsetKvalues reduce early eligible offsets - Explicit exception: Very Aggressive uses
offsetK = 0and no offset shrinkage
Example:
Reviewer with 3 total valid scores, whose full mean is 85 while the full same-response organisation mean is 72:
- Apparent bias: +13 points
- nEff: 2
- Shrinkage weight (
offsetK = 12): 2/(2+12) = 0.14 - Offset: (72 - 85) × 0.14 ≈ -1.8 points
Under Balanced, the three-score cohort receives only 14% of the full observed mean difference. This is formula shrinkage, not a confidence estimate.
The target remains included in both full means. nEff = 2 is the conservative
shrinkage weight; it does not mean that only two scores were used to calculate
the reviewer mean.
About Implementation
"Should we tell people their scores have been adjusted?"
Transparency Recommendation
Yes, transparency is recommended. Hiding calibration creates suspicion and reduces trust.
Best practice:
- Show both raw and adjusted scores
- Explain why calibration is used (comparability)
- Provide access to this documentation
- Make manager statistics available
- Encourage questions and discussions
Communication approach:
"Your raw score was 72. After applying the selected calibration policy to your reviewers' scoring patterns relative to the organisation benchmark, your adjusted score is 79. Review both values alongside the evidence and reviewer context."
Last updated on