HR & Onboarding · Performance Management

Aggregating Performance Review Calibration Data

Calibration meetings are supposed to run on real data, but most teams walk in with a spreadsheet someone stitched together manually from exported ratings, showing raw scores with no context on which managers tend to rate high or low, which scores are true outliers, or what the written feedback behind a number actually says. Without that context, calibration conversations end up litigating individual scores in the room instead of working from a shared, adjusted view, and small teams get held to the same forced-distribution expectations as large ones even though a bell curve is meaningless at a sample size of three. The whole exercise takes longer and produces worse outcomes than the data itself should allow.

STARTING PRICE

From €799

Complex tier · Multi-system orchestration, custom logic, and higher-volume or higher-risk processing.

Get a quote →

Saves roughly 8-12 hrs per review cycle for HR/People Ops building calibration materials.

How the automation works

We aggregate ratings across managers and teams from Lattice and 15Five into a calibration-ready view that surfaces manager rating tendency — which managers run systematically high or low — so the conversation can adjust for that pattern instead of arguing about individual numbers in isolation. Distribution expectations scale to actual team size, so a team of three isn't measured against a bell curve that only makes statistical sense at larger scale, and every score in the view carries its underlying written feedback alongside it, so an outlier rating comes with the context that explains it rather than a bare number. Visibility into the aggregated data is scoped so a manager in a shared calibration session sees their own team's detail and cross-team summary context, without full access to another manager's individual employee ratings.

Process flow

Aggregating Performance Review Calibration Data — process diagram Flow diagram: Cycle ratings finalized → Compute manager rating tendency → Scale distribution expectations to team size → Attach written feedback to outlier scores → Scope visibility by session role → Generate calibration-ready view. Cycle ratingsfinalizedTRIGGERCompute managerrating tendencyAIScaledistributionAIAttach writtenfeedback toAIScopevisibility byAIGeneratecalibration-readyOUTPUT
  1. 01

    Cycle ratings finalized trigger

    Self-review, manager-review and peer-feedback scores are pulled from Lattice or 15Five once the review cycle closes, replacing a manual export-and-merge step.

  2. 02

    Compute manager rating tendency ai

    Each manager's historical rating pattern is analyzed against peer managers to surface systematic leniency or severity bias, so calibration can adjust for the pattern instead of treating every raw score as equally comparable.

  3. 03

    Scale distribution expectations to team size ai

    Rating distribution guidance adjusts for actual team size, so a three-person team isn't evaluated against a forced-distribution curve that only makes statistical sense with a larger sample.

  4. 04

    Attach written feedback to outlier scores ai

    Outlier and boundary-case scores are surfaced with the written feedback that accompanies them, so calibration discussion works from the qualitative context behind a number, not the number alone.

  5. 05

    Scope visibility by session role ai

    Access to individual employee ratings is limited to a manager's own team plus cross-team summary-level context, so a shared calibration session doesn't expose another manager's team detail beyond what's needed.

  6. 06

    Generate calibration-ready view output

    A single view combining adjusted scores, tendency flags, distribution context and written feedback replaces the manually built spreadsheet going into the calibration meeting.

Get a quote for this automation →

Inputs

  • Finalized cycle ratings by employee and manager (Lattice/15Five)
  • Manager rating history across prior cycles
  • Team size and roster data (Workday)
  • Written review feedback text
  • Calibration session participant/role list

Outputs

  • Calibration-ready aggregated rating view
  • Manager rating tendency (leniency/severity) flags
  • Team-size-adjusted distribution guidance
  • Outlier scores with linked written feedback
  • Role-scoped calibration data access

Works with

Prefer a fully custom build instead of an off-the-shelf integration? We scope both options during your free consultation — most jobs like this one work fine on standard connectors, but higher-volume or non-standard systems sometimes need bespoke API work, reflected in the complex tier.

Where this goes wrong if you get it wrong

  • Some managers rate systematically higher or lower than others, and raw aggregated scores without adjusting for that leniency or severity bias push calibration conversations to argue about the wrong thing — the tooling needs to surface the manager-tendency pattern directly, not just hand the room a table of unadjusted numbers.
  • Small team sizes make rating distributions statistically meaningless — a team of three 'should' show a bell curve under forced-distribution logic, but that expectation is nonsensical at that sample size, and calibration tooling that applies the same distribution guidance regardless of team size produces bad, misleading direction for those managers.
  • Aggregating ratings into a single number before factoring in the qualitative written feedback strips out the context that explains an outlier score, so the tool needs to surface the written feedback alongside the rating, not replace the narrative with a number that looks precise but isn't self-explanatory.
  • Calibration data is HR-sensitive, and cross-team visibility needs to be scoped carefully — a manager in a shared calibration session shouldn't be able to see another team's individual employee ratings in full detail just because they're in the room for the broader distribution discussion.

Frequently asked questions

How does this account for managers who consistently rate higher or lower than their peers?

Manager rating tendency is computed from historical patterns and surfaced as a flag in the calibration view, so the conversation can adjust for known leniency or severity bias instead of treating every raw score as directly comparable.

Does this force every team into the same rating distribution regardless of size?

No — distribution guidance scales to actual team size, so a small team isn't held to a forced bell-curve expectation that only makes statistical sense with a larger sample.

Can calibration participants see the written feedback behind a score, not just the number?

Yes, outlier and boundary scores carry their underlying written feedback in the same view, so the qualitative context behind a rating isn't lost in aggregation.

Will a manager in a calibration session see another manager's full team ratings?

No — visibility is scoped so a manager sees their own team's full detail plus summary-level cross-team context, not another manager's individual employee ratings.

Does this replace the calibration meeting itself?

No, it replaces the manual data-prep work that goes into the meeting — the actual calibration conversation and final rating decisions still happen with the managers in the room.