Skip to main content

How individual Impact is measured, evidenced, and overridden

For privacy impact assessments and AI governance reviews: what individual-level outputs Pensero produces, the evidence behind each number, where a human can confirm or correct it, and why this is decision support.

Written by Wayne

This article answers a question that comes up in privacy impact assessments, DPIAs and AI governance reviews: does Pensero use AI or automated decision-making to evaluate employees?

The short answer: Pensero measures the impact of delivered work, at the level of the individual contributor. It does not build profiles of people: it does not assess personal characteristics, traits or conduct, and it makes no prediction about a person. It also does not perform automated decision-making in the sense of Article 22 GDPR. Every number Pensero shows is backed by the specific work items that produced it, and an authorised human can confirm or correct each underlying assessment, with the original and the correction both kept on the record.

Does Pensero generate scores or rankings for individuals?

Yes. Pensero is an engineering measurement platform, so it necessarily works at the level of the individual contributor. The individual-level outputs are:

  • Value Delivered points per work item and per person over a period.

  • Percentile score and ranking (Pensero Score / CodeRank / CollabRank), including a position on an anonymous cross-customer leaderboard.

  • Quality metrics: defect cost, rework cost, number of defective pull requests.

  • Capability evidence: what a person has demonstrated, per technology, problem, product area and way of working, and in which role (author, reviewer, communicator).

  • Comparative benchmarks: how a person compares to their team, their level, a saved cohort, the whole organisation, or industry aggregates.

  • Forward-looking estimates and suggestions: how hard a proposed piece of work looks, and who could be a good fit to do or review it.

Some of these are produced by deterministic arithmetic (points, penalties, percentiles, trends) and some by an AI model assessing the work artifact (how large a change is, how complex it was, what category of work it is, what a day of work evidences about capability). Both kinds are described below, together with the evidence they rest on.

Does Pensero make decisions about employees?

No. Nothing in Pensero produces a decision, a classification or a consequence for a person:

  • No output is written back to an HR system, payroll system, identity provider, Git provider or ticketing system as an assessment of a person. Pensero has no integration with HR, performance management, compensation or disciplinary systems.

  • There is no rule or threshold anywhere that labels a person (for example "under-performing", "high potential", "at risk") or triggers an action about them.

  • Compensation figures are entered by your own finance leads and used only for cost-per-delivery and R&D capitalisation reporting. Pensero never generates, recommends or adjusts a compensation figure.

  • Where suggestions exist (for example who could review a proposed change), they are suggestions only. Pensero does not assign work or reviewers in your provider.

Pensero is a decision-support tool. If a manager uses Pensero output in a performance conversation or a staffing choice, that assessment is made by a human, on the evidence Pensero surfaces, under your organisation's own process.

The evidence behind every number

A Pensero Score is not a model output you have to take on faith. It is the top of a chain you can open up at every level, in the product and through the API:

  • the work artifact itself: pull request, ticket, document, review, review comment or agent session;

  • the feature or epic it belonged to;

  • its size (T-shirt magnitude), with the written explanation of why that size was chosen;

  • its complexity per competency on your organisation's own rubric scale, each with its own written explanation;

  • the resulting Value Delivered points before penalties;

  • each penalty applied, pointing at the record that caused it: the later pull request that reworked or reverted the work, the duplicate, the release merge, or the bug fix, together with the share of the code involved;

  • the final points, their split per technology, and only then the period total, percentile and ranking position.

Capability evidence works the same way. Each capability is written up from a single day of one person's delivered work, keeps the model's own summary and justification, records which model wrote it and when, and links to the exact work items behind it — so a capability entry reads as "fourteen authored items in this area, six reviews", never as an opinion with nothing under it. The model is explicitly instructed to judge only the work in front of it and not to carry over reputation or seniority, and not to infer capability from file paths or folder names.

Where a human can confirm or correct an assessment

Every AI-produced input to a score is reviewable, and the review is part of the record:

  • Work size: confirm it, or correct it with a comment. The original value is retained, so the difference between what Pensero said and what a human decided stays visible.

  • Complexity per competency: confirm or correct each competency's level individually.

  • Defect attribution: challenge whether a defect belongs to this work at all.

  • Category of work (new capability, improvement, keeping the lights on, productivity): override it; the original category and its explanation are kept.

  • Whole repositories: exclude a repository from scoring, with a mandatory reason, recorded against the person who excluded it.

  • Rubric level: a person's claimed level comes from the level their manager assigned, and Pensero keeps it separate from what the work has proven rather than merging the two.

When a score is challenged, a review agent re-examines the original assessment and records whether it was justified, unfair because data was missing, or unfair for another reason — including what data was missing and whether it could have been fetched from a connected source. Challenges therefore improve the system rather than only the single score.

Any correction recalculates everything above it: points, penalties, percentiles, rankings and capability evidence all follow from the corrected evidence.

Safeguards you can configure

  • Role-based access: individual contributors see their own report; managers see their reporting line; administrators configure broader access.

  • Visibility settings per organisation for everything comparative: relative delivery, relative collaboration, the collaboration network, the Pensero Score and ranking, comparisons against level, team, organisation, saved cohorts and custom fields, the quality and collaboration quadrants, and the AI Impact tab.

  • Anonymisation: names, emails and usernames can be replaced with anonymised identifiers across every surface, including the API, with only authorised administrators able to map them back.

  • Off by default: cross-customer scoring, forward-looking planning estimates and score visibility to individual contributors are all opt-in.

  • Audit trail: every change a user makes is recorded with who, what and when; AI assessments keep the model used and its justification, so a change of model or prompt is visible per day instead of silently rewriting history.

  • No third-party identities: the cross-customer leaderboard returns only a position, a percentile and a population size — never anybody from another organisation.

Two deliberate design choices are worth knowing about for governance reviews: a capability counts as proven only after several scored items, because one item is an anecdote; and demonstrated strength and later faults are kept as separate figures rather than netted against each other, so a prolific person with two bugs never looks the same as someone with no track record.

Suggested wording for a PIA or DPIA response

You are welcome to reuse the following:

Pensero evaluates work artifacts produced by employees in the organisation's engineering tools and derives individual delivery scores, percentile rankings, quality and rework metrics, capability evidence and comparative benchmarks, using a combination of deterministic calculations and AI assessment of the artifacts themselves. The measurement is of the delivered work and its impact: Pensero does not assess personal characteristics, traits or conduct, and builds no behavioural profile of an individual.

It does not carry out automated decision-making within the meaning of Article 22. No output produces a decision, classification or consequence for an individual; nothing is written back to HR or payroll systems; and every figure is presented together with the specific artifacts, penalties and written justifications that produced it. Authorised users can confirm or correct each underlying assessment — work size, complexity per competency, defect attribution, category of work and repository inclusion — with both the original and the corrected value retained and attributed. Where managers use Pensero output in an assessment or decision about an employee, that assessment is made by a human on the evidence Pensero surfaces, and the organisation remains controller of that process.

Related articles

Did this answer your question?