Design System Maturity Model
Introduction
Ask a design system lead how mature their system is and youâll get a number, a shrug, or both. âWeâre probably a three out of five?â A three at what, though?
Itâs a question Iâve been asking teams for years, and the honest answer is that nobody has had a good way to respond. Teams can feel whether their system is healthy or struggling, but they canât accurately talk about it it. That gap was survivable while design systems were new enough that enthusiasm carried them. It isnât survivable now, with budgets under scrutiny, headcount challenged, and the whole discipline being asked to justify itself.
So I built something to fix it: a maturity model that charts a system across six independent axes rather than collapsing it into a single score, a book that argues the case and works through each axis in depth, and a free assessment tool that turns the whole thing into eight minutes and a shareable diagnosis.
This case study covers all three: the argument the model makes, the writing that gave it substance, and the product work that made it usable.
Why a single number wasnât enough
Sparkboxâs Design System Maturity Model, developed by Ben Callahan and his team, has been shaping how we talk about system growth for years. Its four stages describe a systemâs life like a productâs life, and part of their genius is how instantly recognisable they are. Every practitioner who reads them has the same reaction: âoh, weâre teenage.â
Weâd been recommending it for years. But the limitation of any single-track model is baked into the format⌠it assumes a system matures as one coherent thing. In practice, the systems I see through zeroheight mature all over the place.
A grassroots system built by craftspeople can hit Stage 4 on component quality while its governance is one person doing it alongside their real job. A well-funded, mandate-driven programme can have immaculate process ceremony wrapped around a component library nobody trusts. Average either into a single score and you get âa three,â which hides the actual diagnosis. The gap between your strongest and weakest axis is where your next move lives, and a single score conceals it by design.
There was also a more pressing reason⌠the maturity models the community grew up on were written before AI-assisted design and development were part of anyoneâs daily workflow. Whether your design system is the context AI tools build from, or the thing they route around, is now a live maturity question. A system that was exemplary by every 2023 standard can be bypassed daily by every AI-assisted workflow in the building.
So I kept Sparkboxâs four stages entirely unchanged and altered what they measure, running them across six independent axes: Foundations, Documentation & Knowledge, Governance & Team, Adoption, Measurement & Impact, and AI Readiness. So, instead of a score at the end, you get a radar diagram that shows you the shape of your system.
The FigJam plan, where I mapped out the entire model and assessment before building anything.
The website design before starting development, created using the existing zeroheight marketing design system.
Behaviour, not artefacts
Almost every maturity conversation falls into the same trap: counting artefacts. Do you have tokens? A component library? A docs site? A contribution model? Tick enough boxes and youâre mature. Itâs appealing because artefacts are easy to see, easy to audit, and easy to buy.
Iâve watched the failure mode play out repeatedly. A team sets everything up perfectly â layered tokens, live docs site, written processes â on a build-it-and-they-will-come assumption, and nine times out of ten they donât come. The teams whose systems genuinely power their organisations treat the system as infrastructure, with community and communication at the centre. Their maturity isnât in what theyâve built. Itâs in how the organisation behaves around what theyâve built.
That became the guiding light of the maturity model: maturity is measured in organisational behaviours and culture first, and artefacts second. Every stage definition in the model is written in those terms. Not âdo you have tokens?â but what happens when you change one. Behaviours tell you what an organisation trusts.
Writing the book
The model needed somewhere to live properly. Alongside the tool, I wanted the reasoning to be availa ble in full, so Measuring Design System Maturity became the substance behind the assessment.
What started as a downloadable PDF became a 15,000 word book. I wanted to go indepth on the thinking behind the axes, the archetypes, the recommendations, and the future of design systems when it comes to maturity.
As well as writing the book, I designed and typeset the whole thing, leaning on my past lives as a graphic designer and type nerd. I also illustrated the entire book, creating visual representations and icons to represent every axes and arrchetype. Representing something as messy as âgovernanceâ was a challenge!
Building the assessment
A book makes the argument, but a tool gets it into peopleâs hands. The assessment had to do real diagnostic work in the time someone will actually give you.
The questionnaire was the hard part, and the constraint was brutal. The fewer questions the better, but too few and the diagnosis is worthless. I settled on 29 questions and about eight minutes. Iâd have preferred fewer, but that was the minimum that still produced a defensible read. The diagnostic core is four questions per axis, each with four graduated options mapping to the four stages.
The framing rule was the same one that governs the model. Wherever possible, questions are behavioural rather than artefactual â not âdo you have contribution docs?â but âwhat actually happens when someone wants to contribute?â Artefacts are easy to have and easy to ignore. Behaviour is where maturity actually lives. Writing four graduated options per question, each recognisable enough that people self-select honestly rather than aspirationally, took far longer than writing the stage definitions did.
Alongside the diagnostic questions, the tool collects a small amount of context (organisation size, governance model, and so on). None of it touches your scores, but all of it tunes the recommendations. âFormalise your governanceâ means something completely different at a 40-person startup than at an enterprise, and the advice should know the difference. A startup matching the Green Bud shouldnât over-formalise governance at all, but an enterprise with the same score is dangerously under-resourced.
Building it inside zeroheightâs existing marketing site was a deliberate trade-off. It put the assessment where people already arrive rather than behind a separate domain, and it meant no auth, no accounts, and no barrier between reading the argument and testing yourself against it. It also meant working within what that platform allows, and within the existing design system, which meant creating something that felt bespoke enough, but still felt like zeroheight.
I designed the entire experience, and then coded it using Next.js and Tailwind. I also created a custom scoring algorithm that would give me the six-axis shape and the matched archetype.
The final landing page for the assessment, with some illustrations created by our amazing in-house brand designer.
The book, a labour of love entirely created by me, currently in production.
Conclusion
The model launched in July 2026 alongside the blog post making its case, and the book follows it. What I keep coming back to is how much of the work was writing rather than designing. The six axes are a product decision, but the reason theyâre useful is the 15,000 words underneath them defining what each stage concretely means and what behaviour reveals it.
It also sharpened something Iâd half-believed for years. Measuring behaviours and communicating value arenât two separate jobs. A maturity model that only tells you where you are is an interesting diagnostic, but I wanted to create one that would help teams justify themselves to the people who control their funding.
Design systems are sliding into the trough of disillusionment, and the way out isnât another conference talk about craft. Itâs the unglamorous work: earning engineeringâs trust, externalising the why, funding the team like infrastructure, making the system the easiest path, measuring in the askerâs language, and serving it all to the machines. This project is my attempt to give teams a map for that, and a vocabulary they can use in the rooms where the decisions actually get made.