AI Maturity Scoring Model
Data-driven maturity evaluation: transforms weekly AI-usage signals into 0-100 scores and five developer levels measuring relative AI maturity.
EARLY ACCESS
Turning weekly AI-usage signals into a 0–100 maturity score and a single maturity level for every developer.
Data-Driven Maturity Evaluation
Basic adoption dashboards answer whether a developer uses AI and can even give insights into how much.
Maturity answers how well.
Every developer gets one row per week, from the week they first used an AI tool onward. This includes weeks in which they used nothing, because a silent week is also signal, not a missing one.
Each week, fourteen behavioral signals are measured, normalized to a 0–1 scale, averaged into three group sub-scores, and combined into a weighted composite score. The score is based on a four week rolling average, and that value places the developer on the maturity ladder.
1. Absolute - the signal has a natural ceiling, so it is scored against a configured target and capped at 100%. Hitting the target every week during the window is a full score.
2. Relative - the signal unbounded and only meaningful in context, so it is ranked against everyone else active in the same week. Where ties occur, the middle of the band is used, so a large group tied at the top scores high and a large group tied at zero scores low.
As always in TargetBoard, the scoring model for your organization can be customized extensively to give weight to and set targets on, what matters most to you. All examples shown below are sample only and based on default calculation logic.
Signal Groups
Signal Group | Sample Weight | What the Signal is Tracking |
Adoption Score | 35% | Is this a habit? How often is the developer using AI tools and for how long? How intensely are they working on those days? |
Breadth Score | 20% | How widely is AI applied and used? Distinct tools, programming languages, named skills, and distinct product surfaces touched within a single day and across the week. |
AI Output Score | 45% | Does the work reach production, and does it stick? AI-authored commits and pull requests, the AI share of merged PRs, the rate at which suggested code is kept, and the cost of getting there. |
Signals Being Tracked
Within a group, each signal is weighted equally, so a signal's effective weight is the group weight divided by the number os signals in the group.
Signal | Group | Period | Measures | Normalization |
Active Days Ratio | Adoption | Rolling, 4 weeks | Active days across the window measured against the target (4 days per week) | Target (4 days/week over 4 weeks > 100%) |
Sessions per Active Day | Adoption | Week | How intensely the developer is using AI tools. | Peer Rank |
Weeks Since First Activity | Adoption | All time | Tenure with AI tools, how ramped up should they be? | Peer Rank |
Count of Tools Used | Breadth | Week | Distinct AI tools used in the week. | Peer Rank |
Count of Products Used | Breadth | Week | Distinct surfaces used within those tools. | Peer Rank |
Language Breadth | Breadth | Week | Distinct programming languages touched. | Peer Rank |
Count of Skills Used | Breadth | Week | Distinct named AI agent skills used (null on accounts whose tools do not report skill usage). | Peer Rank |
Count of Surfaces Used | Breadth | Week | Distinct Claude surfaces touched in a day (null on accounts whose tools do not include Claude). | Peer Rank |
Commits with AI | AI Output | Week | Commits authored with AI assistance. | Peer Rank |
PRs with AI | AI Output | Week | Pull requests opened with AI assistance. | Peer Rank |
AI-Assisted Merge Share | AI Output | Week | AI assisted PRs / All PRs, capped at 100%. Only factors in PRs merged to the main branch. | Target |
Line Acceptance Rate | AI Output | Week | AI suggested lines kept / AI lines suggested. | Target |
AI Cost per Merged PR | AI Output | Week | AI spend / PRs merged to the main branch. | Peer Rank |
AI Cost per Merged Line | AI Output | Week | AI spend / AI lines accepted by the developer. | Peer Rank |
Not every tool reports every signal. Copilot carries no per-user cost or token data, Cursor carries no session counts, and surfaces_used_count is reported only by the Claude Enterprise analytics feed.
Any unavailable signal is dropped from its group's average, and a group with nothing available is dropped from the composite, with the remaining weights renormalized to fill the gap. Consequently, a Copilot-only team will still get a meaningful 0–100 rather than an artificially depressed one.
The same rule covers a signal for which every peer shows an identical score in a given week. As the signal does not separate individuals, it is set aside rather than diluting every developer's score equally.
The Maturity Ladder
Each level represents a different level of adoption as described below. While the core calculations that feed the scoring on based on individual usage, it is important to note that you can also see these maturity scores on Team, Role and other levels.
* scoring cutoffs can be calibrated per account.
Maturity Level | Score | Description |
Leader | Over 70 | Top of the organization, using AI in a consistent and sustained manner. The people worth studying and pairing others with. |
Advanced | Over 59 | Frequent use that consistently reaches merged code. |
Adopter | Over 47 | Regular established use, but likely more assistive and within a narrower set of tools. |
Explorer | 0-46 | Occasional shallow use. Active, but not yet regularly integrating AI tools into work. |
Not Adopted | No activity | Has registered AI activity in the past, but not in the current scoring window. Should be monitored carefully and acted upon as needed. |
A developer who stops using AI drops to Not Adopted rather than disappearing from the chart: once someone has been seen, every later week is reported, so churn shows up instead of quietly vanishing.
Direction of Travel
Alongside the level, each developer-week carries a trend, comparing the current four-week mean against the preceding four-week window: Improving (up more than 2 points), Declining (down more than 2), Stable, or New for anyone without enough history to compare.
Additional Notes
Most signals are ranked against peers within the same week, so the average score across the organization stays roughly flat by construction, even while absolute AI usage grows substantially.
That is the model working as designed: it measures relative maturity, not volume.
To show progress over time, track the share of developers at Advanced or Leader, the shrinking Not Adopted group, or the raw absolute signals. Do not chart the mean composite and expect a slope.
How did we do?
Calculated Metrics