Measure Engineering Productivity Without Harming Team Trust
Engineering leaders face a persistent challenge: tracking productivity metrics without creating surveillance cultures that erode team morale. This article brings together expert perspectives on measurement frameworks that balance accountability with autonomy, showing how teams can demonstrate value while preserving psychological safety. The strategies outlined here help organizations build transparent systems where engineers understand how their work connects to business outcomes and trust that metrics serve improvement rather than punishment.
Quantify Customer Results Quarterly
One lesson I learned the hard way: the moment you turn a metric into a target, people optimize for the metric, not necessarily for the outcome you intended.
As the Global Lead of Business Intelligence, I was leading a team of about 35 engineers supporting thousands of customers. Some customer issues took too long to troubleshoot because we lacked diagnosability in the product. Senior leadership was asking why fixes were taking so long, and I was tasked with speeding things up.
My first move was to create "diagnosability" bugs/enhancements and make creating them part of engineers' goals. The intention was to systematically improve our ability to diagnose problems.
Every engineer hit the goal. Yet there was no clear evidence customers were getting better outcomes. Engineers were creating bugs because we were measuring bug creation, but not whether those bugs actually helped customers. I quickly realized my mistake and the wrong incentive.
Rather than defend the metric, I changed it. We stopped asking, "How many diagnosability issues did you create?" and started asking, "What changed for the customer?" We started tracking how long diagnosing a similar issue took before versus after the fix. Did resolution time actually improve? How many recurring issues did an enhancement eliminate?
The shift wasn't popular at first. The new process was harder, since engineers now had to track impact, not just output. I convinced them to trust it for a couple of cycles. To test whether the new metric held up, I ran a survey each quarter. That rhythm mattered as much as the metric — it gave engineers a way to flag if the measurement felt unfair and gave me real evidence to bring to leadership instead of a defensive claim that "the new way is better."
By the second and third cycles, engineers could see the correlation between the work and fewer recurring issues. When I presented findings to senior leadership, their question was, "Is the core issue being addressed?" One enhancement eliminated roughly 50 customer bugs a month, about 80% of issues in that area. That's a far more meaningful signal than "engineers closed 50 tickets." That metric and quarterly rhythm have held for three years.
If people can hit the metric while the customer experience stays the same, you've measured the wrong thing. Good productivity metrics should make the system clearer to leaders without creating fear or inviting engineers to game it.

Connect Cycle Time to Business Value
I stopped letting my teams report tickets closed per sprint. Every time leadership saw that number climb, they assumed we were getting faster. My engineers were splitting work into smaller tickets to inflate the count, and nobody was talking about whether the features we shipped moved revenue or reduced churn. The metric was clean on a dashboard and disconnected from what mattered.
So I moved my reporting to team-level cycle time and tied each release to a business outcome. I started asking how long it took a squad to move a feature from commit to production, and whether that feature hit the usage or retention target we scoped it against. My weekly review with leadership became a short narrative around those two numbers, with context about what slowed us down or sped us up. I kept individual scorecards out of it.
The first thing that changed was the conversation itself. My VPs stopped asking who was productive. My engineers stopped gaming the backlog because there was nothing to game.
When we spotted a team consistently lagging on cycle time, we dug into whether it was a tooling problem, a dependency bottleneck, or a staffing gap. We stopped finger-pointing at individuals, and the reporting became useful for decisions.

Tell Impact Stories Biweekly
The core tension here is that the metrics that are easiest to count are also the easiest to game, and engineers know it. Lines of code, tickets closed, PRs merged—these numbers go up the moment you start measuring them, and they stop meaning anything useful almost immediately. The problem isn't that leadership wants visibility; that's completely reasonable. The problem is reaching for output metrics when what you actually want to understand is outcomes.
The shift that made the biggest difference for us was moving from activity metrics to impact narratives in our leadership reviews. Instead of showing a dashboard of velocity numbers, we started framing engineering progress around three questions: what shipped, what did it change for the customer or the system, and what did we learn that changes what we build next? That format gives leadership the clarity they're looking for without creating an environment where engineers feel like they're being measured on throughput. The secondary benefit was that it forced the engineering team to stay connected to outcomes themselves, which improved prioritization decisions. When you have to explain every two weeks what actually moved because of your work, you get much more deliberate about what you choose to work on.
Surface Delays in Weekly Delivery Reviews
Measure flow, not individual output.
I do not use lines of code, ticket counts, or hours online as proof of engineering productivity. Those numbers are easy to game and punish people who handle hard problems, review other people's work, or prevent incidents.
I use a short weekly delivery review instead. We look at what reached customers, what is blocked, what created rework, and which decision would improve flow next week. The discussion stays at team level. It is about the system, not ranking developers.
The most useful signal is time spent blocked. Long waits usually expose unclear ownership, slow approvals, unstable requirements, or technical dependencies. That gives leadership something concrete to fix without asking engineers to work faster around a broken process.
I pair the signal with a plain narrative: what changed for the customer, what slowed us down, and what we are changing next. Leaders get clarity because progress connects to outcomes. Engineers get fairness because complex work is not reduced to activity counts.
My rule is simple. Use metrics to improve the environment around the team, never to frighten individuals into producing prettier numbers.

Define Shared Success at 30 60 90 Days
I changed our review rhythm so every engineering project starts with a shared 30/60/90-day definition of success and a cross-functional cadence to review progress. Rather than tracking activity counts, we review the agreed leading and lagging metrics, telemetry, user behavior, and risk posture together. Holding product, sales, data, and compliance accountable in the same forum prevents one-number scorecards that create fear or incentive to game results. That rhythm gives leaders clear, business-oriented signals of progress while staying fair to engineering work by balancing short-term telemetry with milestone-based reviews.

Establish Tiered Decision Cadence
When leadership asks for proof of productivity, I try to shift the conversation from activity counts to decision-ready progress tied to outcomes, so teams do not feel pushed to game a number. One change that worked well for us was moving away from weekly status meetings and adopting a tiered rhythm: a short weekly sync focused only on blockers and decisions, a monthly performance review that connects work to business outcomes, and a quarterly reset to revisit priorities. That cadence gives leaders consistent visibility without turning every week into a scorecard. It also keeps the discussion grounded in what is being learned, what is changing, and what support the team needs to deliver. The result is clearer accountability with less fear, because the focus stays on progress and impact rather than performative metrics.

Show Capability Growth Within Constraints
I changed the question leaders received from how much was finished to what can the organization now do safely that it could not do before. Every update answered that question with one observable capability and one boundary that still remained. The pairing prevented polished success stories from hiding unresolved complexity.
The review rhythm was biweekly, but comparison happened over a rolling six-week window. That avoided overreacting to a difficult cycle while still making stalled progress visible. Teams were encouraged to name boundaries early, because an explicit limit was treated as a planning input. Leaders gained a grounded view of capability growth, and engineers avoided pressure to convert unfinished work into artificial completion. The metric was capability gained under known constraints, which is difficult to inflate and easy to discuss fairly.

Demonstrate Live Builds Weekly
The metric I killed was tickets closed per week. It measures who typed the most, not what got better for the person using the product. What replaced it is one demo a week. Every person on the team shows something running instead of giving a status update. If there's nothing to show, that's the finding. Story points and lines of code both invite gaming, since people slice tasks smaller or pad estimates just to look busy. The review rhythm changed too. We stopped reporting progress on a slide and started walking through the actual build together, with the screen in front of everyone. Leaders see the friction firsthand instead of reading a summary written to hide it. A demo also protects the people doing the work, because it shows what's hard about the problem instead of a percentage that flattens it. On a small team, you can't afford a metric that rewards looking productive over doing the actual work.

Combine Pipeline Data With Sprint Context
The teams that get this right stop reporting the person and start reporting the pipeline. Individual output metrics — commits, story points, PR counts — get gamed the moment someone realizes they're being watched, and the gaming makes the actual codebase worse: smaller PRs padded for count, safer tickets picked over the ones that matter. What survives scrutiny without distorting behavior is delivery-flow data — lead time from commit to production, how often a change needs to be rolled back — reported at the team level, never tied to a name. Nobody can game "how long it takes a change to reach production" without genuinely fixing something real.
The change that mattered most for us wasn't a new dashboard; it was pairing that data with a short narrative every sprint: what shipped, what we deliberately slowed down for, and why. A number alone invites leadership to read a dip as a problem; a sentence of context turns it into a visible tradeoff instead of a mystery. Across 450+ engineers on enterprise delivery, that combination is what keeps the reporting honest instead of adversarial.

Examine Claim Approval Waits
I do not count tickets closed. When leadership wants proof of engineering work, I talk about whether claims move through the product without extra handoffs.
The 2025 Expense Trends Report looked at 371,381 claims across 460 organisations. 2.6 percent were approved immediately. 27 percent were approved after 30 or more days. That gap is a system result, not a personal score. A review that looks at those delays is harder to game than a velocity league table.

Credit Prevented Incidents and Trade-Offs
We stopped reporting output and started reporting incidents avoided, and it changed the conversation more than any velocity chart ever did.
The problem with productivity metrics is not that engineers game them. It is that the metrics measure the wrong direction. Commits, story points, tickets closed — all of them count things produced. Leadership does not actually want things produced; it wants the business not to break and to move faster than competitors. Those are outcomes, and none of the standard numbers touch them.
So the report we give has three parts. First, what shipped that a customer could notice, in one line each, with the customer-facing effect rather than the ticket title. Second, what did not break: incidents this period against the same period last year, and specifically the ones prevented by work done earlier, which is the part that never gets credit. Third, the one thing we chose not to do and why, because a team with no visible trade-offs is a team that is not being honest about capacity.
The measurement that earned trust was one we were nervous to share: how often our own checks caught something before a customer did. We publish it internally every month. If that number were zero, the honest reading would be that the checks are ceremony and should be deleted. It is not zero, so the checks stay, and leadership can see why the time spent on them is not overhead.
Two things I would push back on. Never report a number you cannot explain the mechanism behind — "deploy frequency doubled" invites the question "so what," and if the answer is not ready, the number becomes a liability. And do not compare teams to each other. It produces gaming within a quarter, every time, and the comparison was never informative anyway because the teams do different work.
The test: could the person receiving the report act on it without asking a follow-up question? If not, it is a status update, not a measurement.
Richard Meadows, Head of Content, Streamrise

Use Plan Drift as Early Signal
Before you ask how to measure engineering productivity, you have to answer a harder question: what does productive actually mean here? Because whatever leadership measures, the team will optimize for, regardless of whether optimizing for it produces the outcome anyone actually wants. That isn't gaming the system. It's a rational response to the signal you're sending.
Lines of code measure output. Tickets closed measure activity. Velocity measures throughput. None of those measure whether the team is solving the right problems, learning fast enough to catch mistakes before they become expensive, or building something that will hold under real use. When those are the metrics, a team can score well on all of them and still ship the wrong thing, on time.
So the first change I'd make isn't the metric. It's the definition. Productive means the team is making decisions that lead to good outcomes, catching problems early enough to adjust, and learning from what doesn't work faster than the cost of those failures compounds. That definition changes what you measure, because now you're looking for signals of good judgment and early learning, not signals of motion.
At Microsoft, I spent years re-engineering how large product groups understood their own work. The shift that changed everything was moving the question from "how much did we do" to "what did we learn, and did it change what we're building?" Practically, we stopped using sprint metrics to evaluate people and started using them to find friction. The question in every review was not "did you hit the number" but "what slowed you down, and what does that tell us about the next sprint?" Velocity became a diagnostic, not a grade. And we separated the cadence for learning from the cadence for reporting: teams reviewed their own work for what to adjust, and leaders got a monthly view of trajectory. That meant engineers weren't performing for leadership at every stand-up, and leadership still had what they needed.
The metric I'd keep: the ratio of planned work completed to unplanned work absorbed. When that ratio degrades, something is wrong upstream—in prioritization, alignment, or technical debt—and you know early enough to address it. It is almost impossible to game because gaming it requires fixing the actual problem.

Normalize Defects by Release Volume
The metric I changed was defects per delivery volume. At Ronas IT, I don't use individual output counts for this discussion because they push engineers toward visible activity: more tickets, more comments, more closed items, even when the work that protects quality is review, design clarification, or removing a source of defects. The useful question is whether the system is shipping usable work at a healthy quality level.
A raw defect count has weak meaning because a team that ships more can naturally expose more defects. We count defects from the tracker, including closed ones, and discuss them against the amount of work delivered. In the review, put defect load beside delivery volume so the raw number does not carry the whole story.
I pair that metric with our developer experience survey. Each wave is converted into a form our BI tool can chart, so repeated themes become a signal the company can act on. If engineers name review queues, technical debt, or handoff problems, those themes sit beside the delivery and defect data.
The review is a team-level system review. It is never a person-by-person ranking. Once people feel individually scored, they protect themselves and the work stops improving. Show delivery volume, defect load, and recurring friction together, then ask what in the system needs to change.

Audit Durable Code Changes
The metric that took the fear out was the one that got harder to satisfy, not easier.
Most productivity measures count output: commits, points, tickets closed. Everyone knows they're countable, so everyone games them, and leadership quietly learns to discount the number it asked for. I went the other way and built a verifier that asks whether a merged change was still standing later. Merging is the easy part. Survival is the claim worth making.
The interesting work turned out to be throwing results away. The first version counted 19 entries. Then I added a ubiquity guard: if a file path shows up in more than a quarter of the change sets, it can't be the thing confirming anything; it's just a file everyone touches. Demoting nine README-only pairs took the count from 19 to 13. I added a fan-out cap too, because one fix claiming more than five originals is a batch-close signature rather than five separate wins, and that rule alone demoted all 16 implications of a single merge request.
Thirteen verified beats nineteen asserted, and the engineers could see exactly which rule killed which entry. That's what changed the temperature of the conversation. Nobody is afraid of a number they can audit and argue with.
The limit, and I'd say this to any leader adopting it: this works because the checks are public and run against the repository, not against people. Point the same machinery at individuals and you rebuild the thing you were trying to avoid.


