Curate Take the assessment

A framework for taste

Production became free.
Judgment did not.

Your organisation can now make far more than it can meaningfully evaluate, and the gap widens every quarter. What used to be a pleasant quality in a senior designer has become the only thing separating your work from everyone else’s.

Score your team 18 questions · four minutes · free

What actually changed

The bottleneck moved

For as long as anyone has made things professionally, making was the expensive part, and everything else arranged itself around that fact: teams were sized for production, budgets were written for production, and quality control sat as a thin layer at the end, because the cost of building anything guaranteed that not much got built.

Judgment was never the constraint under those conditions. One creative director could review everything, because everything was not very much.

A team that produced twelve concepts a quarter now produces twelve before lunch, and the twelve are competent. What did not change is the number of people who can look at any of them and say which one is worth building, or why the other eleven are not.

2019 2021 2023 2025 now Coding agents cross into daily use MID-2025 · THE LINES SWAP WHAT YOU CAN PRODUCE WHAT YOU CAN JUDGE
Judgment to spare Ships unjudged

The inversion · shape is illustrative; the crossover follows agent capability data

Organisations used to have more judgment than they had output to apply it to. Now they have more output than judgment. That reversal explains something most teams have noticed without quite being able to name it: work got faster and quality got flatter. Not worse. Flatter. More things arrive at a competent middle, and fewer arrive anywhere else.

Why the usual answers fail

You have probably tried four things

Hiring more reviewers

So you hire two more senior designers. Output has gone up eightfold and your review capacity has gone up by perhaps a third, and both new hires now spend their afternoons catching things a written standard would have caught before the work reached them. Headcount buys a linear improvement against a problem that stopped being linear some time last year, and it is the most expensive option on this list.

Better prompting

Prompting raises the floor, and that is worth having. A team that writes careful prompts gets competent output more reliably than one that writes lazy prompts, and competent is a long way from distinctive. The model returns the middle of what it has seen, and a sharper instruction returns that middle more precisely. Teams who treat this as a prompting problem usually plateau within two quarters and conclude the tools are not ready yet.

Design systems and brand guidelines

Ask what your design system is actually for. It stops twelve people building twelve different buttons, and it does that job well. It will also ensure that everything you ship is coherent, on-brand, and very hard to tell apart from what your competitor ships, because they bought a system too and it solved the same problem the same way. Consistency was the right answer when the risk was fragmentation.

Waiting for the tools to improve

They will. They will get better at producing the middle, which raises the floor for every company at once and lowers the value of standing on it. Faster generation does not relieve a bottleneck in judgment. It widens it.

The framework

Six dimensions of taste

Six behaviours you can observe in a finished piece of work, score against a rubric, and write into a brief or a hiring bar. They describe how a decision was arrived at, and say nothing about what the decision should have been.

They read the same way whether a person or a model produced the draft, which is why the failure modes below are worth naming. The assessment scores how your team decides, not which tools it uses.

CUR ATE

Scored 1–5 on each axis · the shape says more than the total

C

Confidence

Does the work trust itself?

Hedged work is rarely wrong, which is why it survives review. The extra label, the second call to action, the tooltip that exists because nobody was quite sure the interface was clear: each addition was defensible on its own, and the accumulation reads as anxiety.

In AI workflowsA model satisfies every clause of a prompt, so its default output is hedged. The constraint has to be negative and explicit.

U

Unique

Could only you have made it?

Nothing is wrong with the work. Nothing identifies it either. Replace the logo and nobody notices, including the team that made it.

In AI workflowsConvergence on the centre of the distribution is the mechanism, not a defect in it. Distinctiveness is supplied as input, never requested.

R

Refusal

What did you decide not to make?

When production costs approach zero, declining is the only scarce act. Its absence is invisible, because nothing that was never built appears in any review.

In AI workflowsGenerating another variant now costs nothing, which removes the friction that used to enforce restraint.

A

Accountability

Who owned the decision?

Compromise rarely comes from disagreement about quality. It comes from the absence of an owner. The work gets worse and no single person made it worse, which is why it is so hard to argue against.

In AI workflowsWhen the draft was generated rather than authored, nobody made the first choice, and ownership becomes easy to avoid entirely.

T

Tension

Does anything surprise?

The deliberate break that turns out to be right, as opposed to the animation added late because something felt flat. Work without it is clean, correct, usable, and forgotten within the hour.

In AI workflowsA model asked to be surprising returns the most conventional available form of surprise.

E

Endurance

Will it still be good in three years?

Whether the work depends on conditions that will not persist. Anything optimised for immediate reception is systematically vulnerable here.

In AI workflowsGenerated work weights recent conventions, a precise description of what dates first.

What it contains

Specificity is the product

Most frameworks name a quality and leave you to recognise it, which holds up until two people disagree and neither can explain the disagreement. Taste has been described that way for roughly a century. The rubrics here describe behaviour in enough detail that a disagreement has somewhere to land: you can point at the clause you think is wrong instead of trading impressions. Here is Accountability at Level 3, unedited:

Accountability · Level 3 · Articulated

Decisions have named owners and the team can identify them. Ownership is real but conditional: it survives peer disagreement and does not survive seniority. When a sufficiently senior person objects, the decision moves, and everyone understands this as how things work. The owner is accountable for the outcome without holding the authority that would make that fair.

Thirty of those, five levels across six dimensions. Each one paired with a single bounded exercise for moving up one level, and an honest note on what that move costs. Some are cheap. Accountability is not:

Moving from 3 to 4 · Accountability

Pick one project. Give its owner explicit authority to decline input from anyone, including their own management chain, for its duration. Communicate that authority upward before it is tested rather than after. This transition fails almost exclusively at the moment it is first exercised.

One question from the assessment

Try it on your own team

Every question asks about something that happened rather than something you believe. Ask a team to rate its own judgment and you measure its self-image. Ask when it last killed a project that was going fine, and you measure the thing itself.

For the last contested decision on your team, who made the final call?

The maturity model

Five levels, and only one of them is the target

Level 4 is where taste stops depending on who happens to be in the room, and most teams can get there within a year on at least three of the six dimensions. Level 5 is not a reasonable near-term objective for anyone, and aiming at it tends to produce theatre, not progress.

1AccidentalQuality occurs by chance, if at all.
2ReactiveRecognised after the fact, usually at review, by particular people.
3ArticulatedCan be named, taught, and defended. Most teams rest here.
4SystematicTargetEnforced by process rather than presence. Survives absence and turnover.
5InstinctiveAssumed, and its absence treated as a defect. Rare, and easily lost.

Most organisations perform adequately at Level 3. Whether that is sufficient depends on whether your market rewards adequacy.

From the assessment results page

Reasonable objections

Questions you should be asking

Preference is subjective. Judgment is not, or not entirely. Two experienced people rarely agree on which option is best and very often agree on which options are bad, and that shared negative space is what the six dimensions describe. The framework never tells you what to make. It describes how decisions were arrived at.
Yes, if someone answers to produce a score rather than to describe what happened. The questions are built to make that harder: they ask about a specific instance rather than a general impression, and several have an obviously flattering answer that a careless respondent picks by default. It will not catch someone determined to lie. Nothing self-reported does, which is exactly why the honest version of this costs more than four minutes.
One person, in four minutes, is enough for a useful first read. But it is worth having two people answer independently, because where their answers diverge tells you more than either score on its own. A gap usually means the standard lives in a person, not in how the team works, and that is a lower score on Accountability than either of them would have guessed.
A design system governs what things look like. This governs how decisions get made, including the decision to make something at all. They operate at different layers, and a mature design system will happily enforce a stale judgment for years without alerting anyone.

Find out where your taste actually sits

Eighteen questions, about four minutes, no sign-up. You get a score on each of the six dimensions, the shape they make together, and one exercise for whichever dimension needs it most. The exercise is the whole thing rather than a summary of one.

Score your team free · four minutes · nothing stored

Read the free sample chapter (PDF)