We Scored 53,000 PRs. Counting Them Is a Bad Way to Measure Engineers.
PR count isn't just an incomplete measure of engineers. It's an inaccurate one, and AI coding tools are making it worse. Here's what 53,000 open-source PRs from 1,444 engineers show.
Mark Raasveldt co-created DuckDB. Out of 1,444 open-source engineers we looked at, he's #215 by PRs merged a month, but #79 by how much engineering is in those PRs.
A lot of engineering dashboards still count pull requests to measure people. So we scored 53,193 merged PRs from 248 open-source projects to see how often that count gets engineers wrong. Turns out it's wrong often enough that you shouldn't use it on individuals, and AI coding tools are making it worse.
A few of the 248 open-source orgs in this analysis
We spend a lot of time looking at open source, since a lot of the best engineering happens in public. GitVelocity gives every PR a score in GV points (0 to 100) based on how much engineering is in it. A one-line docs fix gets 1 or 2, a typical bug fix about 13, and a big feature can hit 80. To be clear, GV points measure engineering, not business value. A clever one-line fix can matter more than a feature and still score low. For reference, a typical engineer here merges about 6 PRs a month and earns about 120 GV points.
1. A docs fix counts the same as a new feature
On a PR-count dashboard these three Deno PRs count the same, even though one is a docs tweak and one is a whole new subcommand. Across all 53,000, half of PRs score 13 GV points or less, and one in ten scores 58 or more.
2. The same PR count can hide 4x the engineering
Here are nine engineers who each merge about 10 PRs a month. On a PR-count dashboard you can't tell them apart. Their GV points run from 337 a month down to 87.
Some people just ship bigger changes. A typical engineer's PRs average 20 GV points and the top 10% average 43, and merging more PRs doesn't make you any more likely to be in that top group.
3. Between teammates, it picks the wrong person one time in five
Comparing two people on the same team should be the easy case. Same codebase, same reviewers. We checked 10,781 pairs of teammates, and the one who merged more PRs a month earned fewer GV points 20.4% of the time. The catch is you can't tell which of your own comparisons are the wrong ones, and some are wrong by a lot.
This is an extreme pair. Engineer A merges twice as many PRs a month, and Engineer B earns seven times the GV points. That doesn't make A a weak engineer (they may be doing docs, release or review work that neither metric captures well). But PR count would still rank them first.
4. Among the top 20 by PR count, more than a third of comparisons are wrong
The people at the top of a leaderboard are the ones you actually care about. Across everyone, PR count gets 20% of comparisons wrong. Among the top 20, it gets 35% wrong.
Peter Steinberger, who builds OpenClaw, is #1 on both lists. Of the other 19, 7 drop out of the top 20 by GV points, one of them to #167.
If you like the stats version: Spearman's rank correlation runs from 0 (random) to 1 (same ranking). Across all 1,444 engineers, PR count scores 0.79. For the top 20 it drops to 0.45.
5. It hides engineers who ship fewer, bigger changes
That's how the co-creator of DuckDB ends up at #215.
Mark Raasveldt merges about 17 PRs a month, and 214 engineers merge more. But his PRs average 41 GV points, about twice the typical engineer's. Nick Fitzgerald, who works on the Wasmtime runtime, goes from #51 by PR count to #28 by GV points.
AI is making it worse
Coding agents make a PR cheap to open. With Claude Code or Codex, an engineer can open far more PRs than before, and an agent will happily split a change into as many PRs as you ask for. PR count goes up whether or not more engineering happened, and every problem above gets bigger.
Splitting a change doesn't add up to more GV points, because the score follows what the code does. No score is immune to AI (generated code can look busier than it is), but GV points at least read the code, and a PR count doesn't.
So why does anyone use it?
To be fair, across a whole org PR count does move with the amount of engineering done. As a rough org-level signal it's fine. If your org's PR count drops by half, ask why. Where it falls apart is when you use it on people. It's wrong too often on individuals, and splitting work into more PRs is the easiest way to game it. If you want to compare engineers, you have to look at what's in the PRs. That's what GitVelocity does.
See it on your own team
Connect your GitHub org and see your engineers ranked by PR count next to GV points. GitVelocity is free. You bring your own AI key, so scoring runs on your account.
Score your org for freeHow we measured
- Data: merged PRs from 248 public GitHub orgs in GitVelocity's community program, April 4 to October 4, 2026. Engineers with at least 10 scored PRs. Bot accounts excluded by name.
- Monthly averages: orgs joined our dataset at different times, so they have different amounts of history in the window (3.8 months on average). Each engineer's PRs and GV points are divided by the months of data their org has, so someone in a newer org isn't ranked lower for having less history.
- Scoring: every PR scored by the same model (Kimi K2.6) against the same rubric. GV points measure the engineering complexity of a change as judged by a model, not its business value. See how that model compares to a reference.
- The nine engineers: picked evenly from the middle 80% of the 184 engineers who average 8 to 12 PRs a month, so the gap doesn't depend on outliers.
- Statistics: Spearman rank correlations. Across all engineers, PRs a month and GV points a month correlate at 0.79 (Kendall's tau 0.60). PRs a month and average GV points per PR correlate at 0.00. "Pairs in the wrong order" is derived from Kendall's tau. Team comparisons use pairs in the same org who each had at least 10 PRs.
- Limits: this is open source. Review norms and PR habits differ from a company codebase, and orgs differ in how much of their history we've scored.
Conrad is CTO and Partner at Headline, where he leads data-driven investment across early stage and growth funds with over $4B in AUM. Before becoming an investor, he founded Munchery (raised $130M+) and held engineering and product leadership roles at IAC and Convio (IPO 2010). He and the Headline engineering team built GitVelocity to help engineering organizations roll out agentic coding and measure its impact.