Contents
What Pensero Is
Pensero is the engineering intelligence platform for understanding the real impact of work across humans, AI-assisted engineers, and AI agents.
Pensero connects to the systems where engineering work happens, including code, tickets, documents, communication, AI coding tools, quality data, and more. It brings together what belongs together to understand the actual outcomes being delivered.
Instead of simply counting activity like commits, tickets, or lines of code, Pensero understands the work and translates it into a consistent view of delivery, quality, AI impact, and cost.
As engineering moves from human-written code to AI-assisted development and increasingly agentic workflows, Pensero helps leaders answer an important question:
We’re investing in AI. Is it actually making us more effective, and is it under control?
The objective is not another dashboard. It is to give teams the context to understand what changed, why it changed, and where to act.
The questions it answers
Every metric links to a leadership question. The ones you will use most:
Conversation | Metrics that drive it |
Are we delivering enough? | Delivery per HC, Active headcount |
Is code quality declining? | Defect rate, Rework rate, Code coverage |
Why are projects taking so long? | Cycle time, Time to merge, Waste rate |
Is our AI investment working? | AI-assisted %, Delivery lift, Tokens per delivery |
Are we too siloed? | Collaboration ratio, Knowledge gaps, PR pairing rate |
Are we strategic or reactive? | Roadmap alignment, New stuff %, KTLO % |
Do we have succession risks? | Knowledge gaps, Code Rank density |
How do we compare to others? | Benchmark percentiles across all metrics |
Reading metrics in context
Numbers without context can mislead. The same number can be fine for one team and a problem for another. For example, 30% of work spent on maintenance may be normal for a platform team, a concern for a product team, and a serious problem for an early startup. Pensero gives you that context with benchmarks by company stage and team type, trends over time, and known patterns. The What Pensero Measures section below lists the healthy ranges and the warning signs to watch.
Principle: data helps you decide. It does not decide for you. Use it for supportive conversations and to fix systems, never to police people.
Navigating Pensero
Three controls do most of the work: the date range (which period), the scope filter (which people), and the navigation menu (which metric). Set the first two and every page follows them. This section shows how to get to the right view fast.
Date range
The date selector at the top of each metrics page controls the period you are viewing. Choose a Week, Month, Quarter, Cycle, rolling period, or custom date range.
Your selection stays active as you navigate Pensero. All dates follow your organization’s timezone, and most metrics compare your selected range with the previous equivalent period.
“Week” vs. “Days to date: 7”: Week is a fixed Monday to Sunday that you can step through. Days to date is the last 7 days, always ending today. Use Week to compare week to week. Use Days to date for an always current view.
Scope filter: managers and executives
The filter limits all content to the people you choose.
Managers see their reporting line.
Executives see the whole organization.
Individual contributors do not have the filter and see only their own data.
Open it and pick from four tabs: Contributors, Teams, Cohorts, or the All search. Click Apply and every chart and table updates. The button says “All N people” when no filter is set and “N people matched” when one is active.
The filter stays on as you move between pages and sessions. It is saved for about 30 days, per user and per device. It does not sync between computers, and it resets if you clear your browser data or use incognito mode. “Clear selection” inside the filter window removes it. One thing to know when sharing: the date range is in the URL, but the filter is not. To give a colleague the same view, save the group as an Organization cohort and share its name.
Cohorts: reusable groups
A cohort is a saved group defined by rules, for example “Senior ICs”, “Backend engineers on product teams”, or “New hires under six months”. Build it once and it updates itself as people's details change. In the Cohorts tab, “Add cohort” opens a builder. Conditions in one rule group must all be true (role is IC and level is senior or staff). Separate rule groups work as OR (one group or another). You can also always include or always exclude specific people. You can filter by contributors, teams, role, level, technology, location, tenure, employment type, and work mode. A live preview shows who matches as you build.
Save a cohort as Personal (only you) or Organization (everyone can use it, but only the owner can edit it). Cohorts save time for anything you check often. Build it in two minutes and reuse it in every review.
Finding your way around
The left sidebar groups pages like this:
Signals; AI adoption (AI intelligence, Agents)
Engineering intelligence (Work, Quality, Efficiency, Reviews)
Company (Contributors with a Teams tab, Benchmark, Impact, Calibrate, CapEx)
Settings (Org settings, Integrations with a Repositories tab for managers, Data health, Help center).
Some pages depend on your role. See the role notes below. Most metric pages look the same: a Summary tab with cards (a main number, a trend arrow, and “vs. previous period”) and charts, a Contributors tab with a table per person, and detail tabs for specific work types (PRs, tickets, documents). Breadcrumbs at the top (for example Work › Contributors › Alice) take you back to any level.
How to read the visuals: a green arrow means the metric is moving in the good direction, red means the bad direction, and gray means no change. So “down” is green for defect rate. Quadrant charts put one metric on each axis with one dot per person. Top right is strong on both. Bottom left needs a look. The other two corners are trade-offs. Hover over any chart to see exact values. Click a dot to go to that person.
The Contributors table
The Contributors table shows metrics by person, while the Teams tab shows team-level results for managers and directors.
Use the active filters and date range, sort or search the table, open a person for more detail, or export the data to CSV. Annotations are also visible and included in exports.
Which columns show, the default sort, and the available groupings are set for the whole organization in Settings › Org settings, not per person. Managers see their reporting line. Executives see everyone. An IC sees only themselves.
Individual contributors (ICs) see their own metrics and anonymous comparisons against their level, team, or company. They cannot see other people’s data.
They can also use Annotations and My data in Data Health to review their context and connected accounts.
Annotations
Annotations add context to a person’s timeline, such as leave, onboarding, on-call, or incidents. Depending on your workspace settings, they can be added by managers or individual contributors and appear on charts and exports.
See Using Annotations.
Data health
Data health (Settings group, heart icon) checks that activity from your tools is captured and linked to the right person.
Everyone has the My data tab to check and link their own accounts. Managers and directors also have My reports data for their reporting line. Executive roles have Org data for the whole organization.
See User & Team Setup.
Agents
AI coding agents (Devin, Cursor, Codex, and similar) can be added as contributors and marked as AI agents. Their work adds up to an Agentic delivery total on the Agents page (AI adoption group).
Impact, CapEx and Repositories
Impact (Company group) shows each person's impact, the evidence behind it, and any overrides. See How individual Impact is measured.
CapEx (Company group) reports engineering work that can be capitalized, for finance. See Engineering CapEx in Pensero.
Managers and directors also have a Repositories tab at the top of the Integrations page. It lists the repositories their team works on. It is useful when delivery data looks incomplete.
Practical habits
A few habits help. Start wide (all contributors, full team), find the outlier, then filter in to look closer. Combine a date range with a filter for focused questions. “Backend team, last quarter” is one period and one cohort away. Bookmark the views you use often (the URL keeps the page, tab, and date range). To compare two groups side by side, use the Calibrate page instead of switching filters. If a number looks old, a hard refresh (Ctrl/Cmd-Shift-R) loads fresh data.
What Pensero Measures
31 core metrics in 8 categories. Each table gives the definition and which direction is better. Direction is a guide, not a goal. Read every metric with the context above and the patterns at the end of this section.
Delivery: output and capacity
Metric | What it measures | Direction |
Total delivery | Story points completed in a period. Absolute output volume; use for capacity and workload. | ↑ higher |
Delivery per headcount | Average weekly delivery per active engineer. Normalizes output across team sizes. | ↑ higher |
Active headcount | Engineers with completed work in the period. Team size, not performance. | neutral |
Quality: code quality and technical excellence
Metric | What it measures | Direction |
Defect rate | Share of delivery spent fixing bugs the team introduced. Direct quality signal. | ↓ lower |
Rework rate | Share of delivery later rewritten. Churn, shifting requirements, or debt. | ↓ lower |
Revert rate | Share of delivery rolled back. Production stability and risk. | ↓ lower |
Duplicate code rate | Share of delivery with duplicated patterns. Debt and maintainability. | ↓ lower |
Code coverage | Share of code with automated tests. Testing discipline. Needs coverage integration. | ↑ higher |
Efficiency: speed and bottlenecks
All time-based efficiency metrics use P90 (90th percentile). This shows the typical delays, not rare extremes.
Metric | What it measures | Direction |
Cycle time | P90 ticket-assignment to PR-merge. End-to-end delivery speed. Needs ticketing integration. | ↓ lower |
Time to merge | P90 PR-creation to merge. Review-process speed after code is written. | ↓ lower |
Time to approve | P90 PR-creation to first approval. Review responsiveness. | ↓ lower |
Time to comment | P90 PR-creation to first comment. Earliest responsiveness signal. | ↓ lower |
Waste rate | Share of PRs closed without ever merging, abandoned, superseded, or stale work. | ↓ lower |
AI Intelligence: adoption and effectiveness
Needs an AI tool integration (Cursor, GitHub Copilot, Claude Code, and others). Shown on the AI intelligence page.
Metric | What it measures | Direction |
AI-assisted % | Share of code lines written with AI assistance. Adoption and leverage. | ↑ higher |
AI cost | Total spend on AI coding tools, and per-engineer. Read against output and quality gains. | context |
Tokens per delivery | AI tokens consumed per delivery point. Usage efficiency. | ↓ lower |
Delivery lift | Output multiplier vs. a prior period (e.g. 1.4×). Productivity change alongside AI use. | ↑ higher |
User adoption | Share of active engineers using AI tools at all. Org-wide reach. | ↑ higher |
Collaboration (Reviews): review quality and knowledge sharing
Metric | What it measures | Direction |
Collaboration ratio | Share of delivery on enablement (reviews, pairing, mentoring). Balance of output vs helping. | balance |
PR review ratio | Review effort vs creation effort. Review coverage across the team. | balance |
PR review usefulness | Share of review comments authors mark useful. Feedback quality. | ↑ higher |
PR addressed rate | Share of review comments that led to changes. Iterative development. | ↑ higher |
PR pairing rate | Share of PRs with multiple authors. Pairing and knowledge sharing. | ↑ higher |
Scope of Work: investment mix
Needs a ticketing integration with work-type labels. Targets change a lot by company stage.
Metric | What it measures | Direction |
Roadmap alignment | Share of delivery mapped to roadmap/epics. Strategic vs ad-hoc execution. | ↑ higher |
New stuff % | Share of delivery on net-new features. Innovation investment. | context |
Improvement % | Share of delivery on enhancements to existing features. Product polish. | balance |
KTLO % | Keeping-the-lights-on: maintenance, bug fixes, tech debt. | ↓ lower |
Performance % | Share of delivery on performance and scalability. | context |
Code Rank: composition and knowledge risk
Metric | What it measures | Direction |
Code Rank density | Share of team who are high performers (“get things done”). Top 20% often drive 50%+ of output. | ↑ higher |
Knowledge gaps | Share of codebase with only 1–2 experts. Bus-factor and succession risk. | ↓ lower |
Financial: accounting and compliance
Metric | What it measures | Direction |
Share of delivery qualifying for capitalization under accounting rules. New development typically qualifies; maintenance doesn’t. | neutral |
Reading metrics together
No single metric tells the whole story. The combinations below are the ones to recognize at a glance.
Healthy patterns
High delivery per head + low defect rate: productive and careful about quality.
High collaboration ratio + high delivery: strong support for others without losing output.
High roadmap alignment + balanced scope mix: work follows the plan and the portfolio is healthy.
High AI adoption + high delivery lift: AI tools are really improving productivity.
Warning patterns
Low cycle time + high waste rate: merges are fast, but a lot of work is started and then dropped. This points to planning or requirements problems.
High KTLO % + low new stuff %: maintenance is crowding out new work.
Low collaboration ratio + high knowledge gaps: people work in silos, which creates succession risk.
High defect rate + low code coverage: testing discipline has dropped.
High delivery + high revert rate: speed is winning over stability.
Categories at a glance
Category | Metrics | Primary use | Data required |
Delivery | 3 | Productivity, capacity planning | Git data |
Quality | 5 | Code quality, technical excellence | Git + optional coverage |
Efficiency | 5 | Process speed, bottlenecks | Git + optional ticketing |
AI Intelligence | 5 | AI adoption and effectiveness | AI tool integration |
Collaboration | 5 | Knowledge sharing, reviews | Git data |
Scope of Work | 5 | Investment mix, strategic alignment | Ticketing integration |
Code Rank | 2 | Composition, knowledge risk | Git data |
Financial | 1 | Accounting and compliance | Ticketing integration |
How Delivery Is Scored
Delivery is the metric most conversations depend on, so it helps to know how it is built. In short: every contribution is scored as magnitude × complexity. Boilerplate is removed first. Penalties lower the score for work that was undone or duplicated. Scores add up from person to organization. Knowing this helps you read a score correctly and explain it to an engineer.
The formula
Delivery score = Magnitude × Complexity Magnitude = how big the change is (size). Complexity = how hard it was (difficulty). |
Complexity measures the task, not the person. A junior engineer doing an architecture change gets high complexity. A staff engineer fixing a typo gets low complexity. The system credits what was built, not who built it.
Magnitude: size of the change
A T-shirt scale that is not linear. A Medium is not twice a Small, because bigger changes have a much bigger effect. The maximum is XL, no matter how many lines changed. This stops people from inflating scores with volume.
Size | Value | Typical meaning |
Zero | 0.0 | No relevant changes (closed PR, machine-generated code) |
XXS | 0.15 | Bug fix |
XS | 0.5 | Documentation typo |
S | 1.0 | Minor fix or comment updates |
M | 3.0 | Small feature or refactor |
L | 5.0 | Significant feature |
XL | 8.0 | Major feature or architecture change |
Complexity: difficulty of the work
Scored from 1.0 to 3.0. It reflects the difficulty shown in the work itself.
Level | Range | Meaning |
Level 1 | 0.5–1.5 | Junior work: follows existing patterns, basic implementation |
Level 2 | 1.5–2.5 | Mid–senior work: independent feature, cross-team coordination |
Level 3 | 2.0–3.0 | Staff+ work: architectural decisions, org-wide impact |
The five skills: Code (quality of the implementation), System Design (architecture and scale), Delivery (execution and coordination), Communications (collaboration and docs), and Ownership (decisions and initiative). Not every skill applies to every item. A document has no Code score. The AI sets a level for each skill that applies, weights them by the person's career level, and adds them into one complexity number.
Worked examples
Same formula, very different kinds of work:
Architectural decision, small code, high difficulty Changing DB connection pooling, org-wide blast radius. Magnitude S (1.0) × Complexity 3.0 = 3.0 points |
Boilerplate-heavy feature, large code, low difficulty Adding endpoints across 20 files, following an established pattern. Magnitude L (5.0) × Complexity 1.0 = 5.0 points |
Critical refactor, large and hard Refactoring authentication, security implications, multi-team rollout. Magnitude XL (8.0) × Complexity 3.0 = 24.0 points |
What counts as work
Pensero scores more than code, so work that helps others is visible. Each item gets its own magnitude × complexity score.
Artifact | What it captures |
Pull request | Code contributions, features, fixes, refactors. Scored only when merged. |
Trunk-based commits | Direct commits to main, grouped by author + day + repo into one unit. |
PR review | Code-review effort: comments, approvals, review discussion. |
Ticket / ticket review | Jira / Linear work items and reviews on them. |
Document / doc review | Design docs, RFCs, architecture writing and feedback. |
Communication | Technical discussion and Q&A (requires integration). |
Why it matters: senior engineers get credit for reviews and mentoring, technical writers show up in delivery, and collaboration is rewarded, not punished. Open PRs are tracked but do not score until they are merged.
How each artifact is scored
Each artifact follows its own scoring and attribution rules:
Pull requests: Scored when merged or closed. Credit is split between contributors based on their code contribution.
Documents: Scored as they evolve, with rules to avoid counting repeated small edits as new delivery. Credit can be split between authors.
Tickets: Credit the person who creates and defines the work, separately from the person who implements it.
Communications: Score meaningful technical contributions connected to real work. Message volume, reactions, and casual conversations are not counted.
Boilerplate filtering
Raw line counts can be inflated by code nobody really wrote. So Pensero removes boilerplate before sizing a change: lock files, generated schemas (Protobuf, GraphQL, OpenAPI), migrations and compiled assets, large test fixtures, and whitespace-only changes. Magnitude is then measured on the cleaned diff.
Example A 770-line PR = 450 lines of package-lock.json + 200 lines of test fixtures + 120 lines of real feature code. After filtering, only the 120 feature lines count, so it scores as Medium, not Large. |
This works both ways. It stops padding with volume, and it also means needed boilerplate committed together with real work does not inflate or distort the score.
Penalties
Penalties adjust scores so they show delivered value, not raw activity. They add up but are capped at 100% (a score cannot go below zero). They are recalculated as new PRs merge, which is why scores can change after the fact.
Penalty | Amount | When it applies |
Revert | 100% | PR reverts a previous PR. Both the original and the revert lose credit. |
Duplicate code | 100% | PR duplicates another PR’s code (e.g. same branch to a second target). |
Duplicate document | 100% | Document duplicates another’s content. |
Release merge | Partial | PR merges other PRs; only net-new work counts, credit stays with the originals. |
Rework and bugfix deductions exist in the system but are currently turned off for customers.
Reading penalties as a manager
Revert points to a quality or testing gap. It is a coaching moment, not a scoring quirk: “what testing would have caught this before production?”
Release merge and duplicate code penalties are normal in feature-branch and multi-environment (dev → staging → prod) workflows. They only prevent double counting. No need to worry.
Duplicate document penalties catch copy-paste gaming and also accidental duplicate design docs.
False positives can be flagged for review, and admins can override them in rare cases.
How scores roll up
Adding up is simple, with no tricks. Item scores (after penalties) add up to the individual total, and individuals add up to the organization total. Work goes to its original owner. Reviews go to the reviewer. Shared work like multi-author PRs is split automatically. Only completed items in the date range count: merge date for PRs, completion date for tickets, publish date for documents.
Timing, so you can set expectations with your team: an item scores within minutes of merging. Then a background job checks for penalties over the next day or two and may adjust the score. Daily numbers move around. Weekly is the right pace to read delivery.
Reading delivery well
The healthy and warning signs below matter more than any single number.
Healthy signs
Steady week-to-week rather than erratic.
A balanced mix of items: not 100% PRs, but reviews and docs too.
Complexity that follows seniority (a senior's work shows higher complexity than a junior's).
Warning signs
Big differences within a team (some people carrying far more than others): unbalanced load or a skill gap.
A steady decline over time: process friction or growing tech debt.
Very little review work across the team: silos.
Context first: compare engineers at similar levels, allow for team type (platform teams deliver fewer, more complex points than product teams), and remember that some real work is not captured at all, such as verbal mentoring, on-call, and dev tooling.
How Quality Is Measured
The five quality metrics are defined in What Pensero Measures. This section explains what is behind them: how Pensero finds a bug, links it to the PR that caused it, and spots rework, reverts, and duplicates automatically. Knowing how detection works helps you trust the numbers and explain them without sounding like you are blaming someone.
Detection runs automatically on every merge
When a PR merges, Pensero looks at it within minutes. It decides if the PR fixes a bug, links a bugfix back to the change that caused it, compares the diff with recent PRs to find rework, reverts, and duplicates, and pulls in test coverage data if a coverage tool is connected. Results show in the quality metrics within hours. Past data is updated if a bugfix links back to an older PR.
Consistent and fair: the same analysis runs for everyone. Engineers can see their own metrics and why they changed. The focus is the trend over time. Bugs happen. What matters is the direction.
How bugs are traced to their source
The hard question is: when a bugfix lands, how does Pensero know which PR caused the bug? It finds the files and lines the fix touches, goes through git history to find which PR last changed those lines, and links the two: bug introduced ↔ bug fixed.
Example March 1: Jane merges PR #100 adding authentication logic. March 15: Tom merges PR #150, “Fix login crash,” touching lines 45–60 of auth.py. Those lines trace back to Jane’s PR #100. Jane’s defect rate reflects the source; Tom gets credit for the fix. |
So engineers get credit for fixing bugs, see the effect when their own code causes one, and the organization sees the total cost of defects. No manager has to judge it by hand.
How each metric is detected
Defect rate: bugs the org introduced
AI reads each PR's title, description, and diff to decide if it fixes broken behavior: a crash, a wrong result, a logic error, or something that used to fail. It is careful on purpose. Feature work, refactoring, performance tuning, security hardening, config changes, and dependency updates do not count. It then separates the lines that are the real fix from unrelated cleanup, so only the bug-fixing part of the PR counts toward the defect signal.
Rework rate: rewriting recent code
Pensero compares each PR line by line with PRs merged in about the last month in the same repo. When many lines added by a recent PR are now being removed or replaced, that is rework. Lower overlap is just normal change. The useful difference for conversations: rework means replacing code because the first try was wrong (often due to unclear requirements). Refactoring improves structure but keeps behavior, and shows up as improvement work, not rework. The metric is built to tell them apart.
Revert rate: rolled-back PRs
Reverts are found from revert or rollback patterns in titles and git metadata, then checked by comparing the revert PR's removed lines with the original's added lines. This is a count (a PR was rolled back or it was not), so size does not matter. A revert is the most serious quality signal: code that could not be fixed step by step and had to be undone.
Duplicate code rate: copy-pasted code
The content of each changed line is normalized (whitespace and formatting ignored), hashed, and matched against earlier PRs. Many identical lines mean duplication. Because it matches exact content, renaming variables or reformatting will not hide it, and formatting alone will not cause false matches.
Code coverage: test coverage
Coverage comes from an external tool (Codecov today, others planned). On merge, Pensero fetches the diff coverage, that is, the share of new lines that are tested. It adds this up weighted by PR size, so a large PR with little testing moves the number more than a tiny fully tested one. Without a coverage integration this metric shows N/A. One point worth repeating to teams: coverage shows what is tested, not whether the tests are good.
Reading the metrics together
As with delivery, combinations tell the story. The pairs below link a metric pattern to its likely cause. Use them to read a team at a glance.
Pattern | Likely meaning |
Low defect + high coverage | Testing discipline is working. |
High defect + low coverage | Insufficient testing; bugs slipping through. |
High rework + high duplicate | Poor code organization, messy, hard to maintain, frequently rewritten. |
High rework + high revert | Unstable requirements or rushed code. |
Quality sliding across consecutive periods | Compounding pressure: ramping hires, deadlines, or accruing tech debt. |
Context still matters. High rework is normal for an early product that is still finding its fit, and expected for a short time after a big refactor. It is a concern in a mature product. A blanket coverage rule tends to backfire: engineers write empty tests to hit the number. It is better to require coverage on critical paths (auth, payments, data integrity) and watch the trend.
The delivery-vs-quality quadrant
Pensero shows a quadrant chart with delivery on the x-axis and a quality score (the inverse of defect rate) on the y-axis. It shows the trade-off between shipping and stability at a glance. The four quadrants:
Quadrant | What it suggests |
High delivery, high quality | Sustainable practices, shipping steadily without introducing bugs. |
High delivery, low quality | Shipping fast but introducing defects, often a sign of pressure to ship without enough testing. |
Low delivery, high quality | Careful but slow, may reflect genuinely complex work or an overly cautious culture. |
Low delivery, low quality | Where support is most likely needed, worth understanding why before concluding anything. |
Read the quadrant at the group level, not as a judgment of one person. A cluster in “high delivery, low quality” points to a system that rewards shipping without testing. A cluster in “low delivery, high quality” may mean the work is harder than it looks, or that the culture is too cautious. Use a person's position to open a conversation (“what is the work in this area like right now?”), not as a label. The same caveats as delivery scoring apply: complex and enablement work moves where someone lands.
Using quality data in conversations
Start with curiosity, not a verdict. “Your defect rate is higher than usual. What is going on, and where do you need support?” finds the real cause. “Your defect rate is unacceptable” only creates defensiveness. Common causes to check: an unfamiliar or complex area (pair with a senior), no time for tests (revisit priorities and deadlines), an edge case the engineer did not see as a bug (testing guidance), or changing requirements (ask product for clarity).
In the product, organization-wide rates and trend charts are on the Quality page. Per-person quality is in the Quality → Contributors tab, which you can sort and drill into to prepare for 1-on-1s. Use it to find who might need support, not to rank people.
Can engineers game quality metrics?
Mostly no. Defect detection reads the PR content, so calling a bugfix a “refactor” does not fool it. Rework and duplication are line-level content matches, so renaming or reformatting does not hide them. Reverts come straight from git history.
The one weak spot is coverage: you can write tests that run code without really checking anything. The defense is review. Check test quality, not just the percentage.
How Reviews & Collaboration Are Measured
The five review metrics are defined in What Pensero Measures. This section covers how they work underneath: how a review comment is judged useful or not, how Pensero knows if feedback was acted on, how it detects pairing, and the collaboration patterns to recognize. Reviews are where knowledge sharing and peer quality control happen. These signals tell you if your organization is learning together or working in silos.
What counts as review work
Review (or enablement) work is anything that helps someone else ship: reviewing PRs, documents, or tickets; co-authoring through pairing or mobbing; and technical help in messages and Q&A. Approving your own PR and creating your own work do not count. The metrics compare this enablement effort with creation effort. The point is not to maximize either one, but to see the balance.
How the harder metrics are detected
Review usefulness: is the feedback meaningful?
AI classifies every review comment. It counts as useful when it points out a defect (bug, bad logic, security issue), a weakness (performance problem, missing edge case), or a concrete suggestion for a better approach. It counts as noise when it is only “LGTM”, a style nitpick, or a non-technical remark. The metric is the useful share. It measures whether reviews catch real issues or just approve.
Addressed rate: was the feedback acted on?
Pensero checks if the author pushed changes after a comment: fully addressed, partly addressed, ignored, or unknown (for example a comment right before merge). Only useful comments count. Noise is left out. It is weighted by severity, so fixing a critical bug matters more than a style note. The telling combination is usefulness and addressed rate together: useful feedback that is often ignored is a culture signal, not a tooling one.
Pairing rate: collaborative authorship
A PR counts as paired when it has several commit authors, Co-authored-by tags, or other signs of joint work. Pairing is expensive on purpose (two people on one task), so the goal is not a high number but the right uses: onboarding, unfamiliar or complex problems, and spreading knowledge in critical systems.
Collaboration patterns to recognize
The patterns matter more than any single rate. Each one below is a combination you can spot quickly, and what it usually means.
Pattern | What it looks like and means |
Healthy collaboration | Balanced creation and review, useful feedback that gets acted on, some pairing. Knowledge moves and quality holds. |
Siloing risk | Very low review and pairing, too few reviews to even measure quality. Knowledge concentrates, bus-factor and coverage risk if someone leaves or takes leave. |
Review theater | Lots of review activity but low usefulness and low addressed rate. The motions of review without the value, “LGTM” culture, feedback ignored, a false sense of safety. |
Over-collaboration | Review far exceeds creation, constant pairing. Feedback quality may be high, but delivery drags under process overhead; can be right during big refactors or heavy onboarding. |
Review theater is the one to watch for most. High activity looks healthy on a dashboard, but low usefulness plus a low addressed rate means the review process is a checkbox, not a quality gate.
The delivery-vs-reviews quadrant
This mirrors the delivery vs. quality quadrant: delivery on the x-axis, enablement (review) activity on the y-axis. It shows collaboration style at a glance.
Quadrant | What it suggests |
High delivery, high reviews | Multipliers, shipping their own work and lifting others. Seniors often land here. |
High delivery, low reviews | Solo contributors, shipping independently; watch for knowledge-sharing risk. |
Low delivery, high reviews | Enablers, focused on helping others; appropriate for some roles, worth a check for juniors. |
Low delivery, low reviews | May be blocked or need support, understand why before concluding. |
Read it at the group level. A senior in “solo contributor” may need a reminder about mentoring. A junior in “enabler” may be reviewing when they should be building. A wide, healthy spread means a balanced culture. As always, position opens a conversation. It does not settle one. Role and project phase move where someone lands.
Reading review ratios in context
There is no single “right” review ratio. It depends on role, seniority, and project phase. Seniors run higher (they multiply others by mentoring and unblocking). Juniors run lower while they learn and build. Onboarding pushes everyone's ratios up. Instead of setting a number, watch the quality signals (usefulness, addressed rate) and the trend. A low pairing rate is not a problem by itself. But low pairing plus a low review ratio plus high knowledge gaps is a real silo signal.
In the product, organization-wide review metrics and trends are on the Reviews page. Per-person collaboration is in the Reviews → Contributors tab, useful for spotting who is isolated or overloaded before a 1-on-1.
How Scope of Work Is Categorized
The six scope-of-work metrics are defined in What Pensero Measures. This section covers how each piece of work is sorted into a category, and how to read the resulting mix. Think of scope of work as your engineering investment portfolio: where time really goes, which is often not where you think it goes.
How work gets classified
When work completes (a PR merges, a ticket closes, a document is published), AI reads its content (title, description, commit messages, ticket labels) and sorts it into one of five categories. It looks at keywords (“add”, “fix”, “optimize”, “refactor”), structure (new files vs. changed files), ticket labels (bug, feature, enhancement), and the purpose of the change.
Category | What lands here |
New stuff | Building something that didn’t exist, new features, modules, endpoints, services. |
Improvement | Enhancing something that already exists, adding to or refining a current feature. |
KTLO | Keeping the lights on, bug fixes (any severity), security patches, tech debt, dependency updates, infra maintenance. |
Performance | Scalability and optimization, caching, query tuning, sharding, architecture for scale. |
Other | Work that doesn’t fit the above. |
New stuff and improvement are the easy pair to mix up: “add user search” is new stuff (it did not exist before); “add filters to user search” is improvement (making something better). The five shipping categories add up to 100% of delivery.
Capitalizable % is derived, not classified. It is simply new stuff plus improvement: the development that creates new assets (CapEx), as opposed to maintenance and operations (OpEx). There is no separate tagging. It comes straight from the categories above, which makes it easy to audit.
Roadmap alignment is configured, not guessed
Unlike the five categories, roadmap alignment is not an AI judgment about the type of work. It is about whether work connects to your strategic tracking. Each organization defines what “aligned” means: specific epic IDs, projects or initiatives, custom labels like “roadmap”, certain ticket states, or custom query logic that matches your planning process. Ad-hoc bug fixes, unplanned experiments, incidents, and untagged work fall outside it. So roadmap alignment measures planning discipline (following the plan vs. firefighting), not what kind of code was written.
Reading the investment mix
The mix is a portfolio. The patterns matter more than any single percentage. Three to recognize:
Pattern | What it looks like and means |
Balanced portfolio | Most work tied to roadmap, real innovation capacity, steady polish, manageable maintenance. Executing strategy while holding quality. |
Tech-debt crisis | Maintenance consuming a large share, roadmap alignment and new-stuff both depressed. Firefighting mode, little room to innovate, and capitalizable % drops with it. |
Premature optimization | Heavy performance investment with low innovation for the stage. If user volume is low, scale work is stealing from features you actually need. |
Stage decides what “good” looks like. See the stage notes in the glossary. In general, the share of new stuff starts high at MVP and goes down as a product matures, while KTLO and improvement go up. Performance work should be low before scale (do not optimize too early) but goes up when you are really scaling. Low performance investment during fast growth is the warning sign, not low investment early on. The same number means opposite things at different stages.
Using scope data well
The most valuable move is to compare your actual mix with your intended one, then act on the gap: “we are at 38% KTLO, so let's reserve 20% of next sprint for debt reduction.” Trends beat snapshots. A KTLO number going up month after month, or new stuff steadily shrinking, says more than a single reading. Organization-wide mix and trends are on the Work page. Per-person category breakdowns are in the Work → Contributors tab.
Roadmap alignment as an early signal: a sharp drop usually has a concrete cause, such as a major incident, a wave of customer escalations, urgent tech debt, or planning that slipped. It deserves a “what pulled us off plan?” conversation, not a verdict on the team.
Common questions
How accurate is the AI categorization?
Generally high. It is trained on many examples, uses several signals (title, description, labels, code), and stays careful when unsure. You can click any work item to see its category and the reasoning. Admins can override edge cases (the original category is kept for audit). Overrides should stay rare. Routine re-tagging weakens the consistency that makes the data trustworthy.
What about mixed or ambiguous work?
Edge cases like “refactor for performance” (KTLO or performance?) or “fix a bug and add a feature” (KTLO or new stuff?) are decided by main intent. The AI picks the category that covers most of the effort. It is a reasonable default, and the per-item view lets you check anything that looks off.
Is a high KTLO number bad?
It depends on whether it is a spike or a pattern. A one-month jump from incident response, security patching, or a planned debt sprint will return to normal next period. The same level over several months points to a systemic quality or debt problem. Cross-check the defect rate, and consider a dedicated debt-reduction sprint plus better testing upstream to stop bugs at the source.
How Efficiency Is Measured
The five efficiency metrics are defined in What Pensero Measures. This section covers two things: the statistical views Pensero offers for the time metrics (P90, P80, median, and average) and why the default is P90; and, more useful, how to read the four time metrics in order to find exactly where delivery gets stuck. Efficiency is about finding where time leaks, not pushing people to go faster.
Why P90, and the other statistical views
All four time metrics default to the 90th percentile: 90% of work finishes within the stated time. But you can switch the view (P90, P80, Median (P50), or Average) from the dropdown on efficiency charts. Each one answers a slightly different question.
Averages are ruined by a single stuck PR. Four PRs that merge in 2 to 5 hours plus one stuck for a week average about 36 hours, which describes none of them. The median is 4 hours and P90 is 5, both much closer to what engineers actually experience. P90 removes the real outliers (truly complex edge cases) and shows the systemic delay you can act on. It is also the DORA standard and much more stable when comparing teams or periods.
When to reach for each:
P90 (default): finding systemic bottlenecks, comparing teams or periods, setting SLAs (“90% of PRs merge within a day”).
P80: a slightly tighter view when P90 feels too lenient.
Median (P50): the typical engineer's experience, when outliers are truly rare rather than systemic.
Average: overall trend. Most useful next to P90 to understand the spread.
The gap between views is itself a signal. A large gap between average and P90 means a long tail: most work is fast, a few items drag. To tell “a few outliers” from “everything is slow”, compare median with P90. If the median is close to the average, a handful of outliers skew the spread. If the median is close to P90, most PRs really are slow. That difference changes the fix: chase the few stuck PRs, or change the process itself.
A large gap between P90 and average is not a problem. It only means there is a long tail. That is exactly what P90 is designed to show, so you can fix the tail.
The four time metrics, and what sits between them
Each metric measures from PR creation (or ticket start) to a later moment. Read from start to end, the gaps between them tell you where the time goes.
Metric | Measures from → to | The gap before it represents |
Time to comment | PR created → first comment | How long until anyone looks at it (attention). |
Time to approve | PR created → first approval | Comment → approve: how long the actual review takes. |
Time to merge | PR created → merge | Approve → merge: CI, merge queue, deployment steps. |
Cycle time | Ticket assigned → PR merge | The full end-to-end, including time-to-start-coding before the PR existed. |
Cycle time needs ticketing (Jira, Linear) with start timestamps and PRs linked to tickets. The three time-to-* metrics and waste rate come from git data alone, so they work even without a ticketing integration.
Finding a bottleneck from the sequence
Because the metrics build on each other, the shape of the sequence tells you where to look. Match your numbers to one of these:
Where the jump is | Diagnosis | What to try |
Fast comment, slow approve | Review itself is the bottleneck, PRs get noticed but approval drags. | Add approved reviewers; revisit strict multi-approval policies; check for a single-approver chokepoint. |
Slow comment and approve | PRs sit idle, nobody looks until late. | Review rotation, PR notifications, assign reviewers at creation, budget review time into sprint capacity. |
Fast approve, slow merge | A big approve→merge gap points outside review, CI, merge queue, or deployment. | Profile and parallelize CI, check merge-queue policy, automate manual deploy steps. |
Fast throughout, low waste | Healthy process, maintain it. | No action needed. |
Waste rate: work that never shipped
Waste rate is the share of PRs closed without merging: code written but never delivered. This includes dropped experiments, work in the wrong direction after requirements changed, duplicates, and old PRs nobody reviewed. Open PRs (still in progress) and merged PRs do not count. Merged and then rewritten work is rework, which is part of the quality metrics, not this one.
It measures planning quality, not code quality. Some waste is healthy. Experiments fail, prototypes are dropped, startups change direction based on feedback. The question to ask is why PRs are closing: intentional learning, or dysfunction (starting too early on unclear requirements, duplicate work from poor communication, forgotten old PRs)? An innovation team running many experiments will and should waste more than a mature product team with clear requirements.
Reading efficiency in context
Compare like with like. Team type sets the baseline. Product teams move fastest and waste more (simple changes, lots of experiments). Platform teams are slower by nature (complex changes that affect many teams need careful review) and waste less (more planned work). Infrastructure / SRE sits in between, with careful review because production impact is high. Comparing a platform team's cycle time with a product team's and calling the platform team “slow” is the classic mistake.
Organization-wide P90s with trend indicators are on the Efficiency page. Per-person times are in the Efficiency → Contributors tab. As always, the trend matters more than the absolute number. Pick one bottleneck, change one thing, and watch if the number moves.
Common questions
Can we improve speed without hurting quality?
Yes. The two are not opposites. The efficiency gains that help quality are exactly the ones to go after: faster reviews catch issues sooner, clear requirements cut rework, parallel CI speeds up merges without risk, and better communication prevents duplicate work. What you should not do is buy speed by skipping review, cutting coverage, or merging without approval. That only moves the cost into the quality metrics.
What if we don’t use tickets?
You lose cycle time (it needs ticket start timestamps), but the three time-to-* metrics and waste rate still work from git data alone. GitHub Issues or Projects can work if set up. Ad-hoc spreadsheets cannot be connected.
How AI Intelligence Is Measured
The AI intelligence metrics are defined in What Pensero Measures. This section covers where the data comes from and how to read adoption against effectiveness. The core idea: track both whether people use AI tools and whether that use really helps. High adoption with poor results means training is needed, not celebration.
Which tools are tracked, and how
Pensero connects to the major AI coding assistants (Cursor, GitHub Copilot, Claude Code, Gemini Code Assist, OpenAI Codex, Cline, and AWS Bedrock) through OAuth or API keys. It pulls usage daily (lines generated, tokens, cost), links that usage to specific commits and PRs, and adds it up. This needs admin access to the AI tool accounts, credentials set up in Pensero, and each engineer's AI tool account linked to their Pensero profile. Without those links, usage cannot be assigned to anyone.
Adoption vs. effectiveness
The metrics answer two questions. Adoption (are people using AI?) is measured by user adoption (the share of active engineers using AI at all) and AI-assisted % (the share of merged lines that were AI-generated). Effectiveness and cost (is it worth it?) is measured by tokens per delivery (efficiency, lower is better), AI cost (total spend), and delivery lift (change in output vs. an earlier period). Together they tell you if spend turns into reach and into output.
Reading | What it suggests |
High adoption, rising delivery lift | Tools are landing, reach and output both moving. |
High cost, low AI-assisted % | Paying for capacity that isn’t being used, check access, training, unused licenses. |
High AI-assisted %, climbing tokens per delivery | Usage is getting less efficient, over-long prompts, repeated retries, or AI used where it isn’t needed. |
Low user adoption after many months | Barriers, not early days, access, awareness, or cultural resistance. |
How the key figures are derived
Each AI tool reports lines generated per user per day. Pensero links those to commits, then PRs, then the organization. AI-assisted % is AI-generated lines divided by total merged lines. Tokens per delivery is total tokens used divided by delivery points. A token is about four characters of prompt or code, so longer prompts and more retries cost more. User adoption counts active engineers (those who merged at least one PR) with any AI usage. Delivery lift compares this period's delivery with an earlier period.
Treat delivery lift with care: correlation is not causation. It shows output changed while AI was adopted, but headcount changes, simpler work, or seasonal effects can all play a part. Use it as a signal to look closer, not as proof on its own, and compare like with like (same team, same kind of work) before drawing conclusions.
Reading AI metrics in context
AI-assisted % varies for good reasons by work type and person. UI and new code tend to be higher; complex algorithms and legacy refactors lower. Juniors often use it more than seniors. Adoption also takes time, so a low number in the first months is early days, while the same number a year in points to a barrier worth investigating. Organization-wide AI metrics and trends are on the AI intelligence page (AI adoption group in the sidebar). Per-person figures are in the AI intelligence → Contributors tab.
Common questions
Is high AI usage always good?
Not on its own. Usage is a means, not the goal. What matters is whether it produces quality output efficiently. Pair AI-assisted % with the efficiency and quality signals. High usage together with rising tokens per delivery or more defects means someone relies on AI without getting clean, efficient results. The answer there is coaching, not a usage target.
Can we cut cost without losing value?
Usually yes. Tokens per delivery shows inefficient usage. Someone using far more than the team norm is often getting poor suggestions and retrying, which coaching on prompts can fix. AI is not the right tool for every task (simple edits and boilerplate rarely need it), and user adoption shows paid but unused licenses you can take back. High total usage is also leverage for volume pricing.
What if AI makes engineers over-reliant?
Watch for high usage together with worse quality signals, such as more reverts or rework. Healthy AI use looks like an engineer who still understands and can explain their code, reviews suggestions critically, and uses AI to speed up the obvious rather than skip the thinking. The framing that works: AI speeds up routine work so you can spend your judgment on the hard parts. It does not replace judgment.
Comparison Tools: Benchmark & Calibrate
Two comparison pages in the Company group answer different questions. Benchmark is for executives only. Calibrate is available to executives and managers (managers see their reporting line). Benchmark asks “how are we doing compared with the industry, over time?” Calibrate asks “how do these groups compare with each other, right now?” The quick rule: Benchmark for trends and board context, Calibrate for talent and resourcing decisions.
Both show metrics already defined in What Pensero Measures. One naming note: Benchmark and Calibrate call new-feature investment “Innovation rate”. It is the same metric the glossary calls New stuff %.
Benchmark: you vs. the industry
Benchmark (Company › Benchmark) plots your organization against the anonymous median of all Pensero organizations, on a 0 to 100 percentile scale. 50th is the industry median. 75th means you are ahead of three quarters of organizations. 25th puts you in the bottom quarter. It shows 26 weeks of history per metric as a trend line, so you see direction as well as position. Higher is better for most metrics, but defect rate, cycle time, and knowledge gaps are reversed (lower ranks higher). The summary view lists all ten metrics with current percentiles. Click any one to open its detail page.
It covers ten metrics: delivery per headcount, defect rate, AI-assisted code, collaboration, innovation rate (new stuff %), roadmap alignment, cycle time, capitalizable %, code rank density, and knowledge gaps. The comparison is against all Pensero organizations, whatever their size, industry, or stage. There is no segment filter yet. Keep that in mind before reading too much into a single percentile.
How to read it: a percentile is a question, not a verdict. Being below the median can be completely right: foundation work that lowers short-term delivery, a team that is ramping up, or a deliberate choice of quality over speed. Use it to ask what is driving the position.
Calibrate: groups side by side
Calibrate (Company › Calibrate) is a matrix: metrics in the rows, comparison groups in the columns, each cell color coded. Two columns are always there: Company (your whole organization) and Industry (the Pensero median). You can add up to ten custom columns for any mix of people, teams, saved cohorts, or filters (for example “backend developers hired in 2025”). It is built for talent calibration sessions and resourcing decisions.
The color coding is relative to both baselines at once, which is what makes it easy to read:
Cell color | Meaning |
Dark green | Better than both company and industry, excelling; worth understanding what they do differently. |
Light green | Better than company but below industry, strong internally, room to reach industry-leading. |
Light red | Below company but above industry, your org is high-performing overall; this group lags it. |
Dark red | Below both, investigate: blockers, priorities, or training need. |
Gray | Not enough data to score (low activity, empty group, or metric not applicable). |
Calibrate shows eleven metrics: the ten from Benchmark plus active headcount (shown for context, not color coded), with AI user adoption in place of AI-assisted code. Add columns with “Add cohorts”. Hover over any cell to see the value and highlight its row and column. Remove columns from their header. There is a limit of ten columns. To compare more groups, swap columns between sessions or compare cohorts instead of individuals. The matrix cannot be exported yet. Take a screenshot or copy the values to share.
Color is relative, so context still matters. A dark red cell can be exactly right: a new team ramping up, or a group working on harder problems on purpose. Read it as “where should I look”, not “who is failing”.
Which to use
If you need to… | Use |
Track org trends over time | Benchmark |
Compare to the industry | Benchmark (Calibrate has it as a reference column) |
Compare teams or individuals side by side | Calibrate |
Run a code rank calibration session | Calibrate |
Build a board presentation | Benchmark |
Make a resource-allocation decision | Calibrate |
Do historical analysis | Benchmark |
Applied Scenarios
Use this section when you face a real situation. Every scenario below makes the same point in a different way: no single metric tells you what is going on. A number that looks alarming on its own usually makes sense, or means the opposite, once you read it next to two or three others. Each case shows the misleading single-metric reading first, then the fuller picture, then what to do.
The “low performer” who is really the team's glue
Situation: An engineer's delivery is well below the team's. On a delivery-only view they look like an underperformer, and a quarterly ranking would put them at the bottom.
The single-metric trap: Reading delivery alone, you would coach them to “ship more”, or worse, flag them in a review.
The fuller picture: Put their collaboration and quality signals next to delivery. Their collaboration ratio is high, their PR review usefulness is among the best in the team, and knowledge gaps in the areas they touch are low. On the delivery vs. reviews quadrant they sit in enabler: low delivery, high enablement. They are the person unblocking everyone else, reviewing the hardest PRs, and spreading knowledge that keeps the bus factor down. Their low personal delivery is the cost of lifting everyone else.
Action: Recognize the enablement openly. It is real work the delivery number does not capture. The risk is not underperformance. The risk is that this person is invisible to a review that only looks at metrics, and may burn out or leave. If anything, check the opposite: is the team relying on them so much that their own growth has stalled?
The high deliverer who is quietly a risk
Situation: An engineer tops the delivery chart, quarter after quarter. The obvious reading is “star, promote and copy”.
The single-metric trap: Delivery alone says star performer. But output is only one dimension.
The fuller picture: Their collaboration ratio is near zero, their PR pairing rate is tiny, and the knowledge gaps metric shows they are the only contributor across a large part of the codebase. They are a solo contributor on the quadrant: shipping fast, sharing nothing. If their defect or revert rate is also going up, the speed may be costing quality. The high delivery is real, but it concentrates critical knowledge in one person and builds succession risk.
Action: Do not punish the output. Redirect some of it. Pair them with others, route reviews through them so knowledge spreads, and make enablement a clear expectation if they are senior. Celebrating the single metric would have deepened exactly the risk you most want to avoid.
The team working hard but not shipping
Situation: A team works long hours but output is flat, and leadership asks if they are productive enough.
The single-metric trap: Low delivery per head looks like a performance or staffing problem. The instinct is to push harder or question headcount.
The fuller picture: Efficiency tells a different story. Cycle time and time to merge are high, and the bottleneck sequence points to review: fast time to comment but slow time to approve. The PR review ratio shows two people doing almost all the reviews.
The team is not unproductive. Its work is stuck in a review queue because review load sits on a couple of engineers. Delivery is low because finished code is waiting, not because people are not working.
Action: Fix the process, not the people. Spread the reviewer load, set an expectation for review response time, and watch cycle time and time to approve fall. Pushing the team to “work harder” would have made the real bottleneck worse.
A quality regression with a hidden cause
Situation: Customers report more bugs and support is escalating. You need to know the scope and the cause quickly.
The single-metric trap: Defect rate is up, so the easy conclusion is “engineers got sloppy”. That invites blame and fixes nothing.
The fuller picture: Look at several signals. Defect rate has jumped, but split it by team and it is concentrated in one group. There, code coverage has dropped and PR review usefulness has fallen: reviews stopped catching issues. Check scope of work: that team's KTLO % spiked and roadmap alignment dropped, so they were buried in reactive work. And the Contributors view shows two senior engineers were out, leaving less experienced members merging unreviewed code. The regression is not sloppiness. It is a team overwhelmed by maintenance with its safety nets temporarily down.
Action: Treat the cause, not the symptom. Pause feature pressure on that team, bring in review support while the seniors are out, and restore a coverage expectation. The defect rate was the smoke. The cause only appeared when you read coverage, reviews, scope, and staffing together.
Defending (or questioning) the AI tool spend
Situation: Finance wants to know if the AI coding tools are worth the money.
The single-metric trap: Pointing only at a high AI-assisted % proves usage, not value. Pointing only at delivery lift invites the fair objection that other things changed too.
The fuller picture: Read adoption, effectiveness, and quality together. User adoption and AI-assisted % show the tools are really used, not shelfware. Tokens per delivery shows if that usage is efficient. Most important, check whether defect and revert rates stayed stable as AI usage rose. Adoption with stable quality is the real sign that the tools help, rather than just creating more code to clean up. Delivery lift is supporting evidence, treated as correlation and compared like for like.
Action: Bring the combination, not one headline number, and be honest about the caveat on delivery lift. If usage is high but tokens per delivery is poor or quality slipped, that is a coaching and optimization finding, not a reason to cancel the tools. (Note: this guide avoids a single ROI multiplier on purpose. The combination of metrics is more honest and lasts longer than a made-up payback figure.)
The through-line
In every case, the single metric pointed one way and the truth was in the combination. Build the habit: when a number surprises you, do not act on it. Ask which two or three other metrics would confirm or overturn the obvious reading. Delivery next to collaboration. Defect rate next to coverage, reviews, scope, and staffing. AI usage next to quality. The metric tells you where to look. The combination tells you what is true. And the conversation with the person tells you why.
Common questions
Can engineers game their score?
It is hard and it backfires. Boilerplate filtering removes padding. Magnitude is capped at XL no matter the line count. Complexity is assessed, not counted. Penalties remove duplicated or reverted work. Because the system covers many metrics, gaming delivery usually hurts quality or efficiency. Spam of tiny changes stays at XXS × Level 1, which is worth almost nothing. The real defense is culture: when scores lead to support instead of punishment, there is no reason to game them.
What about glue work that does not show up?
Reviews, RFCs, design docs, and technical discussion are scored, and complex unblocking work scores well. But verbal mentoring, process improvements like CI/CD, and on-call firefighting are not captured yet. So when delivery looks low, ask: “are you doing glue work or mentoring that is not captured?” Treat the metric as the start of the conversation.
Why did the score change retroactively?
Penalties run after the merge, in the background. A score can appear right away and then drop a day later when an overlap with another PR, a revert, or a duplicate is found. This is expected. Check weekly, not daily.
How do I explain a low score to an engineer?
Get the context first. Check the mix of items (mostly reviews?), recent penalties (reverts or duplicates?), complexity (strategic or architecture work takes longer), timing (starting on a new project?), and outside factors (time off, on-call, incidents). Then ask with curiosity, not blame: “I see your delivery is lower than usual. What is taking up your time, and how can I help?” Metrics show symptoms. The conversation finds the cause.
Should I mandate code reviews?
Most organizations already require an approval before merge. The real question is whether review is effective. Effective review means meaningful feedback that is acted on, without taking too much of people's time. Ineffective review is rubber-stamping, ignored feedback, or so much review activity that it becomes pure overhead. Measuring usefulness and addressed rate, and recognizing people who give valuable reviews, works better than a quota.
Is a low pairing rate bad?
Not by itself. Pairing is worth its cost for onboarding, unfamiliar or complex problems, and spreading knowledge of critical systems. Solo work is fine for well understood tasks and clearly owned parallel work. It only becomes a red flag together with a low review ratio and high knowledge gaps.
Review usefulness looks low: what now?
Look into it before reacting. Low usefulness usually means people rush reviews, an “LGTM” habit has set in, reviewers lack the context to add value, or PRs are too large to review well. Reading a sample of the low-usefulness comments shows which one. The fix follows the cause: coach on what a good review catches (logic, not just style), encourage smaller focused PRs, and send PRs to reviewers who know the area.

