Skip to main content

AI Maturity Scoring Model

Data-driven maturity evaluation: transforms weekly AI-usage signals into 0-100 scores and five developer levels measuring relative AI maturity.

EARLY ACCESS

Turning weekly AI-usage signals into a 0–100 maturity score and a single maturity level for every developer.

Data-Driven Maturity Evaluation

Basic adoption dashboards answer whether a developer uses AI and can even give insights into how much.

Maturity answers how well.

Every developer gets one row per week, from the week they first used an AI tool onward. This includes weeks in which they used nothing, because a silent week is also signal, not a missing one.

Each week, fourteen behavioral signals are measured, normalized to a 0–1 scale, averaged into three group sub-scores, and combined into a weighted composite score. The score is based on a four week rolling average, and that value places the developer on the maturity ladder.

Two Ways a Signal is Scored:
1. Absolute - the signal has a natural ceiling, so it is scored against a configured target and capped at 100%. Hitting the target every week during the window is a full score.
2. Relative - the signal unbounded and only meaningful in context, so it is ranked against everyone else active in the same week. Where ties occur, the middle of the band is used, so a large group tied at the top scores high and a large group tied at zero scores low.

As always in TargetBoard, the scoring model for your organization can be customized extensively to give weight to and set targets on, what matters most to you. All examples shown below are sample only and based on default calculation logic.


Signal Groups

Signal Group

Sample Weight

What the Signal is Tracking

Adoption Score

35%

Is this a habit? How often is the developer using AI tools and for how long? How intensely are they working on those days?

Breadth Score

20%

How widely is AI applied and used? Distinct tools, programming languages, named skills, and distinct product surfaces touched within a single day and across the week.

AI Output Score

45%

Does the work reach production, and does it stick? AI-authored commits and pull requests, the AI share of merged PRs, the rate at which suggested code is kept, and the cost of getting there.

Signals Being Tracked

Within a group, each signal is weighted equally, so a signal's effective weight is the group weight divided by the number os signals in the group.

Signal

Group

Period

Measures

Normalization

Active Days Ratio

Adoption

Rolling, 4 weeks

Active days across the window measured against the target (4 days per week)

Target (4 days/week over 4 weeks > 100%)

Sessions per Active Day

Adoption

Week

How intensely the developer is using AI tools.

Peer Rank

Weeks Since First Activity

Adoption

All time

Tenure with AI tools, how ramped up should they be?

Peer Rank

Count of Tools Used

Breadth

Week

Distinct AI tools used in the week.

Peer Rank

Count of Products Used

Breadth

Week

Distinct surfaces used within those tools.

Peer Rank

Language Breadth

Breadth

Week

Distinct programming languages touched.

Peer Rank

Count of Skills Used

Breadth

Week

Distinct named AI agent skills used (null on accounts whose tools do not report skill usage).

Peer Rank

Count of Surfaces Used

Breadth

Week

Distinct Claude surfaces touched in a day (null on accounts whose tools do not include Claude).

Peer Rank

Commits with AI

AI Output

Week

Commits authored with AI assistance.

Peer Rank

PRs with AI

AI Output

Week

Pull requests opened with AI assistance.

Peer Rank

AI-Assisted Merge Share

AI Output

Week

AI assisted PRs / All PRs, capped at 100%. Only factors in PRs merged to the main branch.

Target

Line Acceptance Rate

AI Output

Week

AI suggested lines kept / AI lines suggested.

Target

AI Cost per Merged PR

AI Output

Week

AI spend / PRs merged to the main branch.

Peer Rank

AI Cost per Merged Line

AI Output

Week

AI spend / AI lines accepted by the developer.

Peer Rank

Because the maturity level comes from a four-week rolling average, a developer who starts or stops using AI takes a few weeks to move levels. The weekly score reacts immediately; the level deliberately does not.
On Missing Data Points:
Not every tool reports every signal. Copilot carries no per-user cost or token data, Cursor carries no session counts, and surfaces_used_count is reported only by the Claude Enterprise analytics feed.
Any unavailable signal is dropped from its group's average, and a group with nothing available is dropped from the composite, with the remaining weights renormalized to fill the gap. Consequently, a Copilot-only team will still get a meaningful 0–100 rather than an artificially depressed one.
The same rule covers a signal for which every peer shows an identical score in a given week. As the signal does not separate individuals, it is set aside rather than diluting every developer's score equally.


The Maturity Ladder

Each level represents a different level of adoption as described below. While the core calculations that feed the scoring on based on individual usage, it is important to note that you can also see these maturity scores on Team, Role and other levels.

* scoring cutoffs can be calibrated per account.

Maturity Level

Score

Description

Leader

Over 70

Top of the organization, using AI in a consistent and sustained manner. The people worth studying and pairing others with.

Advanced

Over 59

Frequent use that consistently reaches merged code.

Adopter

Over 47

Regular established use, but likely more assistive and within a narrower set of tools.

Explorer

0-46

Occasional shallow use. Active, but not yet regularly integrating AI tools into work.

Not Adopted

No activity

Has registered AI activity in the past, but not in the current scoring window. Should be monitored carefully and acted upon as needed.

A developer who stops using AI drops to Not Adopted rather than disappearing from the chart: once someone has been seen, every later week is reported, so churn shows up instead of quietly vanishing.

Direction of Travel

Alongside the level, each developer-week carries a trend, comparing the current four-week mean against the preceding four-week window: Improving (up more than 2 points), Declining (down more than 2), Stable, or New for anyone without enough history to compare.

Additional Notes

Most signals are ranked against peers within the same week, so the average score across the organization stays roughly flat by construction, even while absolute AI usage grows substantially.

That is the model working as designed: it measures relative maturity, not volume.

To show progress over time, track the share of developers at Advanced or Leader, the shrinking Not Adopted group, or the raw absolute signals. Do not chart the mean composite and expect a slope.

How did we do?

Calculated Metrics

Contact