Skip to content

213 posts tagged with “process”

Patrick Neeman, writing in UX Collective revisits the Double Diamond and design thinking. By going back into history, he argues why the Double Diamond is no longer relevant today.

The Design Council launched the Double Diamond model in 2004: two diamonds, problem then solution, both of equal size. Four years later IDEO’s CEO Tim Brown made the case in Harvard Business Review for Design Thinking as the process, putting designers and researchers at the start of innovation and at the start of the investment cycle, not the end.

Despite Fortune’s claim that IDEO invented human centered design — I even hesistate to link to it from here for obvious reasons—the movement started with Herbert Simon, whose 1969 The Sciences of the Artificial proposed the science of design thinking, and it was built out at Stanford, where Rolf Faste ran the design program before David Kelley took it to IDEO. The Design Council drew the shape for the other.

[…]

The lineage holds more names than Simon and Faste. Robert McKim published Experiences in Visual Thinking in 1972 while teaching at Stanford, Faste expanded that work into the 1990s, and Kelley founded IDEO in 1991. The first prominent book carrying the term was Peter Rowe’s Design Thinking in 1987, about architects and urban planners.

Notice what those people had in common:

  • Simon wrote about engineering and administration
  • Faste taught mechanical engineers
  • Rowe was writing about buildings and urban planning
  • None of them were building software

Every founding discipline made things that were slow, costly, and impossible to take back, and the method inherited their constraints along with their rigor.

That’s always been the crux: designers and designing is cheap; engineers and engineering is expensive. So spend more time on research and design before committing to coding the thing. But as Neeman writes, prototyping three things in an afternoon with AI is now trivial:

At that price [an afternoon] you stop debating which of three directions is right and build all three, because the debate costs more than the answer.

This gets at something I think a lot about from my time at Apple—creative selection, or exploring all the possibilities to discover the right solution. It’s now made a lot easier with AI.

Double Diamond design process diagram showing four phases: Discover, Define, Develop, and Deliver, represented as two red diamond shapes with directional arrows.

The UX Double Diamond is dead, and in AI only one survives for software

Both the double diamond and design thinking priced building as the risky step, and that price collapsed. What’s left is a one-page brief…

uxdesign.cc icon uxdesign.cc

As a young designer who was really into music, designing album covers was always something I really wanted to do. I had always admired the surrealistic work of Storm Thorgerson for Pink Floyd or the moodiness of Vaughn Oliver’s work for 4AD. I had fun in art school designing mixtape covers, but that was the closest I’d ever gotten to creating cover art.

Fast-forward to 2001 when Wilco released their now-classic album, Yankee Hotel Foxtrot. You know the cover, a mostly empty beige square with twin scalloped-sided towers rising from the bottom in black and white. That photo by Sam Jones of the Marina City apartment buildings in Chicago—the band’s hometown—combined with the lyrics from “Ashes of American Flags” indelibly captured the mood of America still reeling from the 9/11 attacks that happened two weeks prior. The band nor the album cover designer could have anticipated the coincidence.

Through twists and turns, the release of the album was digital at first. Wilco’s label rejected the album so the band left the label, negotiated to get the rights, and released the album themselves. They streamed it first on their website before being signed onto another label and having a retail release the following spring.

When I got the album, I pored through the credits and spotted the name of the designer: Lawrence Azerrad. I was pleasantly surprised and shared it with my colleagues at the time, for Azerrad, like me and most of my designer friends, was an alumnus of California College of the Arts. He was in my graphic design class.

Azerrad and I reconnected years later. For a design conference in Chicago that I was co-producing, we held a session at Wilco’s Loft, their home base and studio. Certainly one of my dreams come true!

What surprises me and what I admire about Azerrad and his work with Jeff Tweedy and Wilco specifically, is how much it’s not about him. I remember him telling me about sessions where he’d just set up his computer as the band was rehearsing or recording, and they’d jam while he designed. Tweedy would come in, give some feedback, and they’d jam on the design together. Rinse and repeat. The way Azerrad designs these covers, it’s more like he’s channeling the artist’s creative energy, and their collaboration is the output.

Rachel Cabitt recently interviewed Azerrad for the 25th anniversary of the album’s release. He says:

…as our connection to culture becomes more digitized, as AI plays a bigger role in our lives, and as discovering music and art through our phones gets almost effortless, it’s on creatives to find new ways to stretch toward deeper connection, to provoke more awe and wonder, not less. Because art, in all its forms, can get us to think more deeply about who we are, what we’re doing with our time here, and what we want to leave behind.

Right on.

Marina City's twin corncob towers in Chicago, photographed from below in sepia tone, showing their distinctive stacked circular balconies against a tan sky.

A 25 Year Collaboration: Wilco & Lawrence Azerrad

Chatting with the Grammy award-winning designer ahead of our Los Angeles panel

theartofcoverart.substack.com icon theartofcoverart.substack.com
Black-and-white ink illustration of a crowd of faces on an orange background, each speaking through a blank white speech bubble.

AI Discourse, Writing, and Me

I’m tired of the discourse. The more I read about how AI disrupts the way we design and build software, the less I want to read and write about it. (Yes, ironically, this is about AI.)

It’s not that there aren’t interesting things happening—there are. Jev, Muse, Grok Bot, Siri AI, continuing revelations about rogue OpenAI swarms. There’s a lot.

However, I think we’ve turned the corner. AI is here to stay, regardless of whether the frontier gets paced or not, or whether data centers continue to proliferate. The models we have today have forever changed how software is made. At this point, I believe the sea changes that have been happening every two to three months have slowed. With each new model release, they might be getting faster, more intelligent, and able to perform longer-running tasks. But the major shifts have already happened. We have agents writing code, creating mockups and prototypes, doing research, and autonomously fixing bugs. Improvements will be more of the same.

For designers and product managers, it’ll be dispatching agents to take on work and then reviewing their work. I already see this on my company’s Slack using the Claude Tag. Everyone in EPD just @mentions Claude, “Fix this” and the agent dutifully does it. How fast a company can build depends on the token budget and human review bottleneck.

Lightspark’s chief design officer Geoff Teehan, writing in Design Field Notes, traces the work that comes with adding a UI control:

By details, I don’t just mean the small visual decisions. Every feature, control, mode, state, exception, and bit of behavior becomes something a team has to design and someone has to understand. There’s only so much attention to go around. Same for time, taste, and patience.

A button seems harmless enough, but someone has to decide where it goes, what it says, what it looks like, what happens when you press it, what happens when you can’t, and what happens after you do. Add another and now you don’t just have two buttons. You have the relationship between them.

Teehan’s examples distinguish reducing that work from simply hiding controls. Teenage Engineering’s gorgeous OP-1 combines a synthesizer, sampler, sequencer, recorder, and mixer. It still takes practice to use. But its four colored knobs correspond to colors on the screen, giving the musician a consistent way to make adjustments across different functions. The colors do a job here. Removing them to make the instrument look simpler would undermine the relationship Teehan describes.

He closes with what a team can do when it has fewer details to maintain:

Removing something gives you time back. Time to reconsider the typography, rewrite three words, make an interaction feel right, move something two pixels, tune the sound, or notice what becomes annoying after the twentieth use.

Most people won’t notice any of it. Fine. They shouldn’t have to.

Minimalism isn’t the point. You can make something sparse and still make it bad. The goal isn’t to have the fewest things. It’s to have few enough things that you can care deeply about every one of them.

Cover graphic for Geoff Teehan's Design Field Notes essay on limiting interface details.

Limit the number of details

Every detail asks for attention. Limiting how many we add is what gives us room to make the ones that remain better.

design.lightspark.com icon design.lightspark.com

Peter Merholz, who has urged design leaders to connect their work to what the business values, is reconsidering that advice. A company can track the screens a team produces while overlooking the exploration that helped it choose what to build. That’s what he means by design’s “legibility gap.” He explains why predictable work gets recognized:

This is not some fault of design, but instead reflects the values of business operations. Organizations value predictability, so that they can plan, budget, and communicate likely outcomes to their stakeholders. Predictability requires determinacy, where the nature of the work (its processes and outputs) are understood ahead of time—engineering will take 6 weeks to work through these Jira tickets; marketing will spend $50,000 to draw this amount of traffic. Determinacy in turn defines legibility, the qualities of work that can be seen, appreciated, valued, and accounted for.

Much of UX/Design work is indeterminate, and thus illegible. Examples include:

  • You explore four concepts and choose one. The three abandoned efforts are invisible, and you’re asked why you didn’t just start with the workable one.
  • Research reveals the requirements were misguided, so you reframe them around actual customer need. The roadmap retains the same name, so it appears nothing has changed.
  • You draft the experience principles that will inform hundreds of downstream decisions, but none of which will be traced back to them.

I’ve supported connecting design work to an existing metric, with the qualification that a number can’t capture everything design work contributes. Here, he’s reconsidering whether the business’s existing criteria are sufficient.

His warning comes from a forestry example: in Merholz’s account, clearing a forest for timber production destroyed the undergrowth and other life that sustained it. The first generation flourished, but the next collapsed. He applies that warning to design leadership:

This legibility frame illuminates a potent detail that had escaped me. Where in the past I argued that design leadership is about translation and connection (I have used the phrase “connect design to what the business values” countless times), the real work is subtler and actually much more difficult.

If we want UX/Design to avoid the fate of the forest’s second generation, the real work of leadership is to evolve what businesses see as ‘legible.’ We need them to widen their aperture to appreciate the value of indeterminate work, which is open-ended, exploratory, emotional, meaningful, even beautiful.

Merholz leaves the practical question open: how does a leader persuade the business to support exploration whose result can’t be specified in advance? An existing metric gives that conversation a starting point, but it can’t settle which unmeasured work deserves support.

And it’s a question that I personally continue to grapple with.

Double Diamond design-process diagram illustrating Peter Merholz's essay on design's legibility gap.

Design’s Legibility Gap

Peter Merholz argues that businesses make predictable, determinate outputs legible while overlooking the exploratory work that enables good design.

petermerholz.com icon petermerholz.com

Design educator Michael Buckley, writing for UX Collective, asks design schools to reconsider how they recruit students. He starts with what happens when those students get to class:

I see the mismatch most clearly when students reach the parts of a project that precede the visual work. They may be eager to design screens but impatient with research, resistant to technical constraints, or inclined to treat strategy as documentation to complete before the real design begins. Technology is sometimes learned well enough to operate but not well enough to question.

These reactions are not simply failures of effort. Many students entered the field expecting a professional outlet for artistic creativity. A discipline increasingly concerned with behavior, systems, evidence, and consequences can feel like a change in the terms of that promise.

An interest in making beautiful work is a good reason to study design. But students also need to learn strategic and systems thinking. The school owes them a description of the profession that makes those expectations clear. They shouldn’t have to discover the difference after they’ve enrolled.

Buckley brings the question back to recruitment and admissions:

Changing the curriculum without changing the recruitment pitch leaves that mismatch intact. Programs that present design primarily as a practical outlet for creativity will attract students who expect creative expression to organize the profession. Their frustration is understandable when the education they encounter reflects the realities of contemporary design rather than the false expectations that brought them to the field.

Design programs should therefore reconsider what they advertise and what they evaluate during admissions. A visual portfolio can remain part of the process without serving as the only meaningful evidence of potential. Applicants could also explain how they understood a problem, considered conflicting needs, or changed a decision after encountering new evidence. An imperfect artifact supported by thoughtful reasoning may reveal more than a polished portfolio organized around personal style.

Cover illustration for Michael Buckley's UX Collective essay on design education and recruitment.

Design has outgrown the traditional designer

Why design education must reconsider the kind of student it imagines.

uxdesign.cc icon uxdesign.cc

When does a repeated element deserve its own component? Anton Sten is using an AI agent to build his design system. He measures its value by the decisions it saves people. Here’s his filter for what belongs:

Lately I’ve been using a simple filter before adding anything to a system. Does it show up often enough to deserve a shared solution? Does it have behavior or states that need to stay consistent? Is there an actual rule here that we want to encode rather than leave to individual judgment? And, just as importantly, is this something someone is willing to own over time?

The third question is the one I keep coming back to. Repetition alone isn’t enough. Something becomes valuable as part of a system when it captures a decision we don’t want humans or agents to keep making independently.

I’d treat the owner question as part of that decision, too. Someone has to remain responsible for the rule as the product changes. Otherwise, the team has standardized something without deciding who can resolve the next exception.

Once a rule is shared, people and agents still need to recognize when it applies. Sten uses token names to show how:

Take naming. green-600 tells you what something looks like. text/success tells you what it’s for. A designer with years of context may know which green is appropriate, and an engineer may remember which one the team normally uses, but an agent only has the system in front of it.

The more intent the system contains, the less any of its users have to infer. Semantic naming is not just a preference in that context; it carries the decision across design, code and the agents working between them. That clarity helps people too, but agents expose just how much ambiguity we’ve historically allowed humans to compensate for.

Hero graphic for Anton Sten's essay on design systems as decision-capturing tools.

The point of a system is fewer decisions

A good design system captures decisions so people and agents do not have to keep making them independently.

antonsten.com icon antonsten.com

In a YouTube walkthrough, Michael Shimeles demonstrates how a file of instructions coordinates his coding agents. His “software factory” uses reusable skills, each describing how to do a particular job:

I built my own software factory. Now, when you hear that phrase, you’re probably thinking of a giant custom harness, and that could be far from the truth. To be frank, a proper software factory is more about your workflow, skills, and domain knowledge more than it is the harness. So, a software factory should be harness agnostic and model agnostic.

The harness is the software that runs the coding agent and gives it access to tools.

Shimeles’s instruction file, AGENTS.md, specifies four stages: isolate, build, prove, and ship. First, each agent gets separate working files and a branch for its task, using Git worktrees. That prevents agents from overwriting each other’s working files, though shared resources still need coordination. During the build, a skill specifies how to organize the code. Shimeles explains:

This is just basically a way I like code to be structured because if it ever come a time where I have to manually review this code like it’s 2022 or something then it’s written in a way that is easy for me to digest and get used to. Right?

So for the developers you might disagree with a service architecture. You might want it written in a different way. The point for this is get this skill written in such a way where the code written makes sense to you. For me, a service layer architecture makes sense for me.

Then the agent runs checks and captures the changed behavior. For a visual change, before-and-after images accompany the request to merge the code, called a pull request or PR. An automated code reviewer, Greptile, supplies feedback for the agent to address before passing the PR to Shimeles:

This is not getting merged by me until I get a five out of five. So, my software factory will continue to read the feedback from Greile and then will only give me the PR URL once it’s got a five out of five. Once we’ve reached five out of five status, it’s my job to look at before and after. It’s my job to read the description and then I just merge.

The score doesn’t guarantee correct code. I’d treat the before-and-after evidence as something to inspect separately from the reviewer’s confidence. His shared skills limit the number of review iterations and spell out what’s needed to capture evidence, so there’s setup and adaptation involved in copying this workflow.

I have a software factory, it’s just all skills (you can copy them)

Ras Mic demonstrates a software-development workflow coordinated by an AGENTS.md file and reusable skills: isolate each task in a Git worktree, apply consistent code structure, gather runtime evidence, and iterate on automated review before human approval.

youtube.com icon youtube.com

A harness is the environment around agentic work: the models, instructions, skills, tools, and files that let agents do the job. A graph describes how work moves through that environment. Anatoli Kopadze, writing an article on X, starts with two familiar parts:

A graph is just a plan for your AI work, drawn out so you can see it. It answers two questions: which jobs need to happen, and which job has to wait for which.

There are only two parts, and getting them straight fixes most of the confusion.

A box is called a node. It’s one job: one agent doing one task, with one thing going in and one thing coming out. Researching a competitor. Writing a draft. Checking a claim.

An arrow is called an edge. It just means one job needs what another job produced, so it has to wait for it. And the arrow only counts when something real actually passes along it.

That last condition changes how you design the workflow. A list written as “do A, then B” looks sequential even when B never uses A’s result. Kopadze’s test is to remove those false dependencies so independent jobs can run at the same time:

Look at the AI workflow you run today and walk it step by step. At each step, ask one thing: does this step actually need the result of the one before it?

If yes, the edge is real. Keep the order. If no, there is no edge, and the wait is wasted. Those two jobs can run at the same time.

Take a simple one: “review file A for bugs, then review file B for bugs.” It reads like a sequence, but the check on file B never looks at what file A returned. They only run one after another because that is the order you typed them in. Run them side by side and the whole thing finishes in the time of the slower single file, not the two added together.

This resembles building an automation in n8n: each node has a job, and the connections determine what data it receives and when it can run. AI adds another requirement, though. The agent checking a result should not inherit the worker’s reasoning and simply agree with it:

So you never let the agent that did the work check the work.

You put a separate node on the edge. Its only job is to try to kill the finding before it moves on. If it survives, it passes. If not, it dies right there.

Here is the catch nobody names: that checker needs a clean context.

Give it the same chat the worker had and it is not checking anything, it is nodding along to itself in a different font. A graph of agents sharing one context is just a single loop in a costume, and it breaks the same way, only later and pricier.

So make the verifier fresh. Own context. Checking a real signal, not “did the agent say it is done” but “does the test actually pass.”

Cover illustration for Anatoli Kopadze's X article on graph engineering, showing nodes and edges representing AI agent workflows.

Graph Engineering explained: what it is, when to use it and when not to

Graph engineering treats AI workflows as graphs — nodes are bounded tasks, edges are real data dependencies — which lets independent work run in parallel instead of one long chain. The core test: remove edges that don’t carry data, then verify each finding with a separate, fresh-context checker.

x.com icon x.com

Scott Fryxell, writing on his blog, has built his workday around a harness: a shared setup for instructions, tools, files, and scripts that he can use from Cursor, Claude, or the Pi coding agent. The durable workflow matters more than whichever AI tool happens to be attached to it.

Single developer projects can build to the caliber and consistency of large development teams. You can and should build bespoke applications and you don’t have to sweat onboarding experienced engineers; if they know their stack backwards and forwards they’ll quickly know how to contribute to yours. But most of all, I have learned that the harness is the thing; the fulcrum from which my expectations meet the LLM’s capabilities.

At the moment my rig is supported by two subscriptions (Cursor, Claude) that I can augment with Pi as needed. All three share my skills and AGENTS.md. Though I am using three TUIs, I have a unified experience. This has commodified the models for me; there is no magic sauce or special experience in Claude or Cursor that I need in order to be productive. I have zero anxiety about the transition from Cursor to Codex at the end of this month.

Fryxell also assigns different kinds of work to different models. He uses frontier models for planning, the first task, and parts of review, while cheaper models handle routine work. Here, “frontier” means the most capable models available:

I can run Deepseek on most maintenance and simple tasks. It’s when I am exploring a serious feature or large refactor with lots of moving parts that I reach for the frontier. Recently I learned about prewalk2, Can Bölük’s technique that uses frontier for the planning phase and first task, then hands off once the pattern is set. I paired it with the planner/worker/critic split from Building an Advanced Agentic Harness3 - a single prompt that plans, executes, and critiques itself confuses its own objectives, so each role gets isolated instead. I built both into a skill, with a supporting Pi extension that can take over at any stage of work.

Exploration leads to a plan formalized into an explicit DAG (directed acyclic graph) task list. Then a worker takes over, focusing on implementing the DAG one node at a time. Once complete, I bring in the critic to simplify and question what was implemented. Often this phase will push back enough that the worker phase is revisited. But once satisfied, the critic gives way to a promoter, which is my reminder that a job is not complete until you’ve properly communicated it to others.

The task plan is only one part of the harness. It also keeps the plans, files, scripts, and notes in one place where Fryxell can inspect and reuse them. He calls the command-line apps TUIs, short for text-based user interfaces. Those apps can change while the work stays put:

The harness is self-contained to support more than a home directory (sandboxing, a web interface, File System Access API, Docker, Deno executable, etc).

These concepts are still forming in my mind so I’ve been referencing npm start, cursor-agent, claude as TUIs to keep the concept of a harness clear. All TUIs share the harness.

Remaining auditable is important enough that the TUIs are instructed to keep things inside the artifacts/ directory. Cursor uses .gitignore to ignore files, which I think is smart, so a git-less root is required. I have a skill that syncs the harness with the repo in the work directory. Skills, extensions, and AGENTS.md are first-class citizens at the root, waiting to be modified and built upon. TUIs have to toe the line.

Generic Open Graph share image for Scott Fryxell's blog post on building a shared AI coding harness across tools.

The Harness Is the Thing

Models are commodifying. The harness - skills, extensions, a scriptable app - is where the work and the leverage now live.

scott-fryxell.github.io icon scott-fryxell.github.io

Organizations often miss work their metrics were not built to see. AI agents face a similar problem when they encounter something their rules cannot handle. Service designer and University of Michigan urban-technology professor Ron Bronson turns to Stafford Beer’s idea of algedonic signaling: a way to report an exception outside the usual channels.

One reply to the thread referenced algedonic signalling, which sent me back to Stafford Beer.

In the Viable System Model, an algedonic signal is an exception channel: information that can move outside the normal reporting structure when actual conditions have departed badly enough from what the system expects. Beer wasn’t thinking about “agent welfare” and to be honest, neither am I. At least it relates to anthropamorizing an agent’s welfare, when in reality what i care about exclusively is what gets done well, what’s gets done successfully and ensuring an agent stays within the boundaries of its remit. So, the usefulness of the idea here doesn’t depend on deciding whether a model experiences anything like pain because that’s stupid and makes me angry, even as a suggestion.

But a system that operates with some autonomy needs a way to tell the rest of the system when ordinary control is no longer adequate.

Bronson separates two jobs. A judgment router decides whether an agent is allowed to act. An exception channel gives the agent somewhere to report a situation the router does not understand.

An agent can be completely authorized to perform the task in front of it and still encounter something it doesn’t know what to do with. The API returns 200, but the answers contradict each other. A tool works exactly as documented while exposing some behavior that looks dangerous, or the task can be completed but the agent finds something else that seems wrong and goes on a goose hunt.. Nothing has necessarily crossed an authorization threshold because none of this is within the original parameters, but it’s also not outside of them.

That makes the exit door a different piece of infrastructure from the gate.

An alert matters only if the system responds. Bronson’s diagram shows that it can let the work continue, stop the task, or change its rules for next time. Otherwise, it has only recorded the problem.

The author of the thread built a small pipe from an agent to Teams and called it distress_call. and the key is this, there should be somewhere for the agent to go when it needs to alert something is off.

The judgment router still matters because somebody eventually has to decide what an agent is authorized to do, when its discretion ends and where accountability sits, but the exit door takes that concept a step further and helps us track the inconsistencies, and perhaps the errors to better spot problems that might consequential.

For Agent Experience, that expands the design surface considerably. Making an environment easy for an agent to operate is only part of the job. We also have to decide how an agent stops.

The system learns only if the alert changes what happens next time. The next agent should stop sooner or call the right person.

Screenshot of the article page at blog.ronbronson.com.

Agent Experience Needs Failure Affordances

AI agents can be fully authorized for a task and still hit something no rule anticipated. Ron Bronson borrows Stafford Beer’s algedonic signaling to argue agents need a separate exit door for reporting that, not just a judgment router deciding what they’re allowed to do.

blog.ronbronson.com icon blog.ronbronson.com

Service designer Daniel T Santos uses British cybernetician Stafford Beer’s Viable System Model to explain why strategic design work can disappear in plain sight:

System 1 is operations: these are the units that produce the countable output. The screen ships, the order moves, the revenue lands. System 1 is legible by construction, because the organisation was built to count what it produces.

System 4 is strategy and coherence. This is the outward and forward-facing work: sensing the environment, holding coherence across the parts, deciding what the service should become. This is where research that reframes a requirement lives. Where the journey gets held together across channels that each have their own backlog. Where the principles that shape a hundred downstream decisions are set.

System 4 work produces no countable artefact in the quarter it is done. So the organisation, looking through a System 1 lens, does not see it. Peter Merholz recently called this design’s legibility gap, and it is worth reading. Beer gives it a mechanism: the work is not illegible because it is done badly or explained poorly. It is illegible because it sits in a tier the measurement system does not read.

The organization cannot value work that it doesn’t measure. Santos’s answer is “a shared unit of account”:

Indeterminate work becomes legible when it is anchored to a number the organisation already counts.

Anchored instead of translated or championed!

You take the System 4 work and you attach it to the System 1 metric it was quietly moving all along. You decide, up front, which existing operational or financial number this work should shift if it succeeds.

Santos knows a number cannot capture everything design contributes. But organizations already manage by numbers. Linking design work to one it affects at least makes the work visible.

Screenshot of the article page at danieltsantos.substack.com.

Design work isn’t vague. Our measurement is.

The work your best people do is invisible to your own metrics. This article aims at explaining why.

danieltsantos.substack.com icon danieltsantos.substack.com

Ageism in design is real. In my last job search a couple of years ago, I got zero callbacks. I landed a new role via connections.

What’s kept me sane and relevant has been curiosity. Tom May, writing in Creative Bloom:

One thing does need to stay sharp, though. As creative director Olivia Downing explains: “The secret sauce is you have to stay curious. About culture, about platforms, about all of it. You don’t have to be on social media, but I believe it’s our job to understand the importance and relevance of the medium.”

Editorial image for a Creative Boom piece on aging well as a creative professional.

Can you actually grow old as a creative, successfully?

Creatives can grow old successfully—but only by trading the race for speed and novelty for judgment, reliability, and deep client relationships. The one thing that has to stay sharp is curiosity.

creativeboom.com icon creativeboom.com

Dan Maccarone, product strategist and co-founder of the design studio Charming Robot, draws a line between what expertise can be documented and what must still be learned through first-hand experience:

The move nobody makes: run both plays on purpose. Structure what you can, and still teach it to a person.

The trap is thinking you have to choose between them like my co-founder clients. But you don’t. The better move is to run both plays at once, on purpose: write down everything about your judgment that will survive being written down, and build the apprenticeship for everything that won’t.

Start with the writing-down, because there’s a catch everyone hits. Getting your judgment out of your head isn’t the same as dumping your files into Claude. Do that and you get mush. The model grabs the loudest thing you’ve ever said and hands it back with total confidence.

In design, AI guardrails can store the reasoning, conditions, and exceptions behind a product decision. Apprenticeship gives people decisions in context through critique and design reviews, where correction becomes pattern recognition. Reading the expert’s answer is not the same as learning how to arrive at one.

Trust it like you’d trust a very sharp new hire who’s read everything and lived nothing: useful, fast, and occasionally, cheerfully, wrong.

Maccarone’s prescription is to document the method and pair it with apprenticeship, so people learn when to use their own judgement.

Cartoon of a person passing a glowing stream of ideas to a smiling teal robot, illustrating judgment transferred to AI.

Bottle your judgment and make it outlive you

The most valuable thing you own dies with you, unless you get it out of your head. Let a machine and the person next to you finally use it.

uxdesign.cc icon uxdesign.cc

Matt Shumer, writer of the viral “Something Big Is Happening” essay, argues that most agent workflows try to prevent mistakes by specifying every move. His Gauntlet Loop takes a different approach: define what good looks like, then let the agents decide how to reach it.

You give a lead agent a goal and a real example of what great looks like. The lead agent decides how to break the goal into the smallest pieces that can be improved separately. Each piece gets its own builder and a separate critic with fresh context.

The builder makes something. The critic compares it against the reference example. If the reference example wins, the critic explains the biggest remaining gap and sends the work back to the builder. The builder fixes it. Then another round begins.

That continues until the result reaches the bar (or, more likely, you decide it is ready).

The agents get specificity about the outcome, not a recipe for producing it. On a design team, one agent could decide how to build a prototype while a fresh critique agent—Shumer calls them critics—compares the result with a reference and an explicit quality bar. Shumer separates the critic so freedom in execution does not weaken the judgment:

The builder and critic should be separate agents.

The builder has seen every decision it made. It remembers why it made them. That makes it very good at explaining why its work is reasonable.

You do not want reasonable. You want an independent judgment.

Spawn a fresh critic and give it the goal, the bar, the relevant rules, and the actual artifact. Do not give it the builder’s history or explanation.

Smarter agents can work with fewer procedural rules when their output remains inspectable. Shumer puts it into practice with a goal, reference, artifact, and independent review.

Diagram-style graphic for 'How to Run a Gauntlet Loop,' illustrating the builder-critic iteration cycle described in the post.

How to Run a Gauntlet Loop

Matt Shumer’s Gauntlet Loop gives agents a clear outcome and an example of what good looks like, then lets them determine how to get there. Decomposed work and independent critics keep that freedom inspectable without prescribing every move.

somethingbig.ai icon somethingbig.ai

Creativity researcher Keith Sawyer described art and design school as a place where students learn to see their own work: they bring unfinished work into critique, hear what others see in it, and practice identifying the gap between intention and result.

David Hoang turns that studio lesson into exercises early-career designers can continue long after school:

To train the eye, spend time doing three things: Capture, Collect, and Curate. To Capture, bring the tools that work best for you: a physical notebook, camera, recorder, or phone. As you walk around, travel, read a book, or surf the internet, Collect what matters. The point isn’t volume but discernment. I keep folders on my Mac for interesting textures, color palettes, fonts applied in the wild, and shapes found in nature. Finally, Curate what you captured. Don’t bury it in a Photo Library with tens of thousands of images. Put your observations somewhere you can revisit daily. Print them out or make a mood board. Whatever you choose, give your trained eye the proper attention.

Hoang then moves from observation to making. He recommends rebuilding strong interfaces from scratch, with the original on the canvas beside your work:

For interface designers, my recommendation is to copy from the masters. By copy, I don’t mean taking a screenshot and putting it in Codex or cloning from a UI library. Look at the work and reproduce it in your UI drawing tool. Put the screenshot on the canvas and reconstruct it from scratch. As you do this, note the margins, spacing, and font sizes. Ask why they work. What deliberate decisions produced this output?

And designers have to explain those decisions. Hoang treats critiques, presentations, and written rationales as more practice because each forces intuition into words. All of it feeds the same loop:

Experience isn’t time spent; it’s evidence of what you can do. The proof of work comes first. Design intuition can feel mysterious because the reasoning behind it becomes difficult to see. An experienced designer may recognize that something is wrong before anyone can explain why.

Edward de Bono argued that creativity isn’t an innate talent but a deliberate skill anyone can develop. I believe this, but it requires rigorous training.

This is why the repetitions matter. Every time you notice a pattern, reconstruct an interface, study an aesthetic movement, or defend a design decision, you give your intuition more evidence to draw from. Over time, the distance between seeing, understanding, and making gets shorter. Your eye begins to match your hand, and your hand begins to keep pace with your judgment.

Header image for David Hoang's newsletter essay on deliberately training design taste and judgment.

Training design senses

Taste and judgment get called the differentiators of the AI era, but nobody explains how to build them. David Hoang on the deliberate practice behind design intuition.

proofofconcept.pub icon proofofconcept.pub

The startup grind is often the same: ship fast, but maintain a high bar of quality. At least that’s the sentiment. Those things are in opposition and the system incentivizes people to just move on.

Miguel Fernandez, writing for Slack Design, describes why shared standards matter:

Craft is subjective, until it isn’t. Disagreement surfaces the moment you have to define it. Is this interaction good enough? Is this component polished? Without shared standards, these discussions become battles of personal taste rather than principled decisions. You can’t argue craft in a document. You can only experience it.

That makes the quality bar a cross-functional working agreement rather than a designer’s private standard. Fernandez continues:

It seems like craft competes with velocity. Engineering and Product are often rewarded for shipping, not refining. Craft requires slowing down at key moments, which feels costly when measured against sprint commitments. The incentive systems work against craft, but this is a false trade-off: craft debt compounds just like technical debt, and eventually you pay for it in churn, reputation, and the slow erosion of what made your product special.

Fernandez then quotes Linear co-founder and CEO Karri Saarinen on the conditions beyond the individual:

“Craft is the mindset that creates quality. But it’s not enough. You need to have the right skills and ideas. You need individuals who take their profession and craft seriously, then build teams that work this way together, and have a company that creates the conditions for it. Not only incentivizing with deadlines and metrics, but also caring if the experience is good enough.”

Illustration for Slack Design's essay on craft, product quality, and shared standards in the age of AI.

Product Quality: A Shared Commitment to Craft in the wake of AI

Craft is subjective until you have to define it — and it seems to compete with velocity. Miguel Fernandez on why product quality needs shared standards, not just individual taste.

slack.design icon slack.design

At Apple, my boss Hiroki Asai, then the VP of Graphic Design, often gave me the same feedback: “This needs to be more considered.” Meaning, spend more time with it. Keep your quality bar high.

That memory came back while listening to Michael Riddering interview Charlie Deets, a former Safari designer now working on Dia, on Dive Club. Riddering:

It’s that extra loop and practice of consideration that I find often is the only way that I can get to simplicity.

I might have the right idea as my knee-jerk reaction. Like we put in a lot of reps, like that’s not uncommon for me. Sometimes I return back to the idea.

But then knowing that this is directionally correct, but then going through it one more time and say, “Okay, how what can I distill or cut or combine?”

Deets connects Riddering’s extra loop to Apple’s willingness to interrupt its own momentum:

Yeah, this is something I think Apple just does better than anyone else, and I think it’s partially because of sort of this one-year release cycle thing. At any point in that cycle you can say, “Is this really the right way?” And you can go back to the beginning.

And people almost are expect that disruption as part of the process. Like it’s very much part of the process to just say, “Actually we’re on the wrong path. Let’s start over from this position.”

Reopening the work also asks something of the team. Deets again:

And I think that requires a lot of confidence and it requires a lot of low ego or something.

That was Hiroki’s lesson, too. “More considered” meant that the work could be further iterated upon.

Quotes lightly edited for clarity.

Charlie Deets - Designing the internet computer

Charlie Deets on why an extra loop of consideration — going back through a directionally-right idea one more time — is often the only way to reach real simplicity.

youtube.com icon youtube.com

Paul Bakaus says he helped turn Impeccable from what he calls a “vibed personal skill” into a design toolkit. He says hundreds of thousands of designers and developers now use it. Three of his techniques show how much of that improvement came from engineering around model behavior instead of polishing a prompt.

The first is “Make them argue · two blind reviewers beat one confident guess”:

For a concrete example, /impeccable critique, Impeccable’s design review command, runs both deterministic checks against your code (to detect stuff like low contrast quickly and decisively) and reviews the design like a human would. But early versions had a major flaw that produced bad critiques: hand the LLM the detector output up front and it skews BOTH ways:

  1. detector noisy → condemns a strong page over fixable nits (false alarm)
  2. detector silent → rubber-stamps a generic page (meaningless clean bill)

To debias, critique now spawns two sub-agents that never see each other’s work: A = an LLM design director (hierarchy, slop, heuristics), B = the deterministic detector + browser evidence. When the sub-agents finish, the main agent synthesizes (weave, never concatenate) the results from both for a balanced critique. Two blind opinions beat one confident guess.

I agree with the architecture. Review independence matters as much for agents as it does for human teams; the reviewer needs a perspective the maker didn’t already shape.

Bakaus’s second technique is “Force divergence · escape the cluster, don’t chase it with bans”:

Bans just move the model from one region of latent space into another. The escape is sideways, not a longer denylist.

Here are three techniques that produce higher divergence (cheapest but weakest first):

  1. Shave the safe picks: make the model name its top 3 fonts, then discard all three. Out-rank yourself.
  2. Generate ~50, a blind sub-agent keeps the 5 most distinct. (Lives in Radiant’s codebase, not Impeccable yet.)
  3. Seed from outside the model: palette.mjs --from 8f2a starts in a region the model never picks on its own.

This is a better response to repetitive AI aesthetics than accumulating another denylist of fonts and visual effects. The intervention changes the candidate set instead of pleading for a better choice.

The third is “Give them memory · runs that compound instead of restarting”:

Most skill authors treat skills as stateless prompts, but there’s no rule that says a skill can’t have state. Keep in mind that skills can be bundled with scripts that can be run anywhere during the lifecycle of the skill!

Impeccable’s critique and polish commands use this technique quite effectively:

  1. /critique writes a snapshot per target (score, P0/P1, markdown + frontmatter)
  2. /polish reads the latest snapshot as its fix backlog

One non-obvious detail that makes this work across a team: the snapshot slug comes from the resolved file path, not from how you phrased the request. A teammate pointing at the same file inherits the same memory, and the score trend (24, then 28, then 32) survives across sessions.

This one is immediately practical: save the work on disk, then let the next command pick it up.

Profile photo of Paul Bakaus, whose X post explains techniques for engineering more reliable AI agent skills.

The Dark Arts of Skill Engineering

Turning a vibed personal skill into a battle-hardened harness extension: adversarial review to counter self-bias, forced divergence to escape the model’s median, and memory that lets skill runs compound instead of restarting.

x.com icon x.com

The loop engineering theme explained how to replace turn-by-turn prompting with a system that discovers work, assigns it, checks the result, and remembers what comes next. AI engineering writer and former Google engineering leader Addy Osmani scales that idea into a software factory and adds the governing constraint: the work can move only as fast as humans can review it.

A software factory is many harnessed loops running at once, fed by a queue of work and drained through a review gate into production, with humans owning the whole thing from above. It is not a bigger agent; it is an org chart made of loops.

Osmani follows that definition through the factory’s wiring diagram:

By and large, every box in this diagram is almost zero cost: generation, tests, scanning. They all run at scale for negligible cost. There is only one expensive box that proves stubbornly resistant to scaling, and that’s the review gate. That shiny amber box is “judgment”, and where the crux of the argument about whether we can make development faster and more frequent resides.

For Osmani, the review gate sets the pace. More generation only helps when verification scales with it; otherwise the factory manufactures a queue of code no one has the attention to understand. Imagine a conveyor belt of products being assembled together, only to pile up at the end, waiting for a poor human quality checker.

Osmani turns that constraint into architecture:

You might be thinking that all sounds unglamorous. You’re right. The safety net is made up of perfectly ordinary architectural practices we’ve always known about and mostly ignored: good types and method signatures so that mistakes are caught by the compiler instead of in production; test seams where we can pin behavior and make change observable; laying out the code so the next reader, human or model, knows where to find the thing they care about; keeping call stacks short and legible; keeping component boundaries well defined so a change doesn’t have a huge blast radius; and dependency injection so we can swap out one piece for another. None of it is new. We’ve always said we care about good architecture. But now that we’re using automated coding agents, that architecture is finally doing a second job as a cheap and hard-to-fake safety net against the mistakes the agent will make.

Put plainly, structure the work so mistakes are easy to spot and changing one thing doesn’t break everything else.

For designers, that means clear rules for each component, documented states, testable prototypes, and explicit review criteria. Build loops that can keep moving without you, but let them move only as far as those checks can prove the work is sound.

Diagram illustrating loop, harness, and factory layers in an agentic software production pipeline.

Software Factories, Light and Dark

A software factory is many harnessed loops running at once — the system that builds your software instead of you. You can keep humans in the loop, or take them out entirely.

addyosmani.com icon addyosmani.com

I’ve started using AI this way in a lot of my daily work. In my day job, after diving into a problem, I give Claude the right context and iterate with it to get to a problem brief. In my freelance work, I set the parameters—the concept, style, or general functionality—then iterate with the AI.

AI product manager and builder Karo Zieminski, writing in Product with Attitude, calls this “AI-assisted craft”:

AI-assisted craft means the human sets the intention and the standard for the work, then directs how it gets made. AI gets a defined supporting role. Supporting, as in: it does not get to make the decisions. The practice applies across knowledge work and digital creation, from writing and research to coding and design.

Zieminski separates assistance from direction:

AI-assisted. The human is the primary maker. The choices, mistakes, revisions, and final form are theirs. AI is one of the tools they use.

AI-directed. The model produces the work; the human directs it. That direction must be consequential enough to shape the result. One prompt followed by a shrug is generation with supervision theatre.

Both require the human to make consequential choices. In my workflows, that happens throughout the iteration, not only when I write the first prompt.

Zieminski’s bounded-task rule makes that concrete:

“Improve this” is not a bounded task.

Give the model a job with sharp edges: do this, not that.

Research for me, BUT bring me facts, not conclusions. Challenge my assumptions, BUT do that through Socratic questions so I have to do the thinking. Explain this code block, BUT test whether I understood it. Suggest design fixes BUT don’t bleach my personality out of it. Whatever it is, leave space for my judgment.

I agree, with one small addition: setting the boundary starts the work. AI can’t answer the question at the other end for me: Would another change improve it, or is it time to stop asking?

Karo Zieminski’s AI-assisted craft framework chart contrasting deliberate human work with AI slop, alongside a 100% human writing detection result.

AI-assisted Craft: A Manifesto

Karo Zieminski names the missing category between human-made work and AI slop: AI-assisted craft, where a human sets the intention, gives AI a bounded job, and keeps every consequential decision.

karozieminski.substack.com icon karozieminski.substack.com

Yennie Jun, writing for Art Fish Intelligence, asks which parts of thinking we surrender along with the task. Her example shows the difference between asking AI to answer a question and asking it to test thinking we’ve already begun:

I suggested (with only a little bit of initial resistance) that we pause and think about why this might be. I suggested a few theories. Perhaps it was Portugal’s relative homogeneity and religiousness, compared to the US’s diversity of immigrants. Perhaps Portugal clung on to so-called “Age of Exploration” as one of the most prominent chapters in its national story. We wondered, postulated, made wild guesses, backtracked, connected our ideas, disagreed, and remembered historical details we learned in high school many years ago. We drew on our collective memories, knowledge, understanding of the world, and critical thinking skills. We knew we were speculating, and some of our theories were probably wrong; that was part of the exercise.

Eventually, we asked the same question to AI. Its response corroborated many of our theories and supplied several explanations we had missed. It also omitted a few possibilities we still found plausible. We had begun with a question, generated hypotheses, and only then used AI to test and extend our thinking. I relished the exercise.

The backtracking is the point. A finished answer can save time, but repeatedly skipping the work of forming and testing a hypothesis also skips the practice that builds judgment. For designers, that’s the work we should keep, even as it accelerates production.

We still need to use trial and error to learn what to ask or try next. Otherwise, faster production leaves us less able to tell whether what we made is any good.

Jun turns from productivity to autonomy:

Am I any different from the Microphone Man? Perhaps what differentiates me is that I still collected and curated the data, formulated the questions I wanted answered, and evaluated the end results? Or that the data was my own, instead of recording other people’s conversations? There will always have to be some balance between automating menial tasks to free up time for rewarding endeavors, and doing the work yourself as a learning experience.

Jenny, another character in Ken Liu’s story, aims to counterpoint the main character’s over-reliance on his AI assistant. She exclaims, “Tilly doesn’t just tell you what you want! She tells you what to think. Do you even know what you really want anymore?” Our autonomy depends, at least in part, on continuing to participate in forming our own desires. But when we offload thinking about what we want (What music should I listen to? What movies should I watch? What food should I eat? What shoes should I wear?), who do we become?

What are we automating? Human work or human agency? Human tasks or human thinking?

Designers still have to decide what deserves to exist before asking AI to make it.

Ken Liu’s The Paper Menagerie beside a handwritten notebook, pen, and headphones.

Are we offloading too much of our thinking to AI?

AI can test and extend a line of thought, but using it before we form a hypothesis skips the practice that builds judgment and weakens our agency.

artfish.ai icon artfish.ai

Artist Katya Ross, in a talk published by CreativeMornings, the global creative community and lecture series, argues that content abundance has made intentional curation the problem:

We don’t have a content problem. We have a curation problem as a culture, as people in this overloaded culture. We don’t want more content. We want more meaningful content. […] The problem with trying to seek out meaningful content: your brain wants you to stay alive. Your brain curates according to survival, not necessarily meaning. And social media algorithms, which is a lot of where we get all of the content that we view, curate according to ad revenue. They curate according to what you pay attention to, whether you like something or not, whether you agree with it or not. It just wants your ad revenue, and it wants you scrolling.

Ross means something more specific by curation than choosing the best items from a pile. For her, it is the act of constructing relationships that create meaning:

To define curation for the sake of this talk, I would say that curation is constructing and carefully considering the relationships that make meaning possible. Curation is, in other words, the fundamental mechanism through which we can create and communicate meaning. Now, it’s just a mechanism. There’s no guarantee that your curation is going to be any good.

That shifts originality away from producing something with no precedent. The contribution is the relationship a creator sees among inherited ideas, materials, and experiences:

We are in relationship with reality. We are not truly inventing anything. We are working with what’s already there, whether that’s materials or ideas. Every idea has a lineage. Even if it is an original idea, everything that’s gone into your mind has laid the groundwork for this idea. The ideas are related to one another. Every work that you create exists in this messy, beautiful, complicated web of work that has come before, work that sits beside, work that will come after.

That is a more useful standard for originality than novelty. The materials can be familiar and the work can still be original when the relationship it reveals carries the creator’s particular way of seeing.

Katya Ross: We Have a Curation Problem

In an overloaded culture, the creative work is not making more material. It is constructing the relationships among ideas, experiences, and artifacts that make meaning.

youtube.com icon youtube.com

Ben Callahan on what a design system can and can’t guarantee:

A design system can only raise the quality floor. It sets the baseline below which nothing should ship. An accessible-by-default button, a holistic and thoughtful approach to spacing, a template that starts a consuming team ten steps ahead.

But a design system alone can’t raise the quality ceiling. That’s not something you can do by delivering assets. The worst product teams can make awful experiences with the best design systems. That’s because the quality ceiling is set by the choices product teams make with what you give them. It’s their restraint, it’s where they push, and it’s knowing when to deviate from the standard because the standard isn’t serving the end user.

Callahan’s title invokes AI, though the essay only touches it indirectly. The connection follows from his distinction: faster generation and stronger defaults can produce more acceptable work, but neither can decide when the standard is failing the user. That decision still requires careful judgment.

Callahan on the loop:

And, of course, this loop just continues to run. Over time, the quality floor and the quality ceiling are raised.

The most important step here isn’t the shipping of a new component. It’s the time in conversation that results in alignment on a definition of quality.

Your system sets the floor. The way your system is used sets the ceiling. If you’ve poured everything into the first and bowed out of the second, it’s time to step back into ring.

Screenshot of the article page at bencallahan.com.

What is product craft in the age of AI and design systems?

Design systems can raise the quality floor, but product teams set the ceiling. Raising both requires shared standards, judgment, and an ongoing practice of craft.

bencallahan.com icon bencallahan.com

Patrick Neeman, writing for UX Collective, compares today’s AI interfaces to the browser wars:

[Jeffrey] Zeldman did not invent the specifications, he did something harder: He convinced an entire industry that shared conventions were worth fighting for, and he won. Zeldman changed the world with a stance, not a specification and we should thank him for it.

We are living through that moment again, this time for the interfaces we wrap around models, the skills we scaffold on top of them and representations they mean.

The browsers have new names: ChatGPT, Claude, Gemini, and Copilot each handle the same task their own way, with their own conventions for parsing content, showing reasoning, citing a source, and asking permission before they act.

The connection to design systems is structural. Browser standards gave different products a shared foundation without forcing them to look identical. Neeman doesn’t claim that the conventions for AI interfaces are settled; he proposes design systems as the way practitioners can develop and share them:

The core move was to pull structure, presentation, and behavior into distinct layers so each could change without breaking the others. That one idea outlived every specific technology it was built on.

It is why a design system works at all.

When Brad Frost introduced atomic design, he was extending the same instinct: stop shipping pages, start composing interfaces from small, shared, recombinable parts. Design systems are the standards movement’s direct descendant, and they are the closest thing we have to a working model for AI interface conventions.

That working model is already appearing in the Markdown files agents use as project-level contracts:

Agents increasingly take their instructions from plain text files that sit beside the work — AGENTS.md for how an agent should behave in a project, SKILL.md for what a capability can do, README.md for the context around both.

This is the new semantic layer. It is markup again, written in Markdown and read by a model instead of a browser.

[…]

A design.md that carries your design system’s patterns, tokens, and rules into every agent that touches the product. An accessibility.md that states the non-negotiables in language a model can follow. A content.md that fixes voice, terminology, and the content model.

Neeman also points to a broader protocol stack taking shape:

You are not waiting for this to begin. It has begun. A partial map of the standards taking shape right now:

  • Model Context Protocol — a shared way for a model to reach tools, data, and context, already adopted across rival platforms and now stewarded by a neutral foundation.
  • A2UI — a declarative protocol for agents to describe interfaces that render natively across web, mobile, and desktop, keeping what the interface is separate from how each client draws it.
  • Agent2Agent — an open protocol for agents to discover one another and collaborate across frameworks and vendors, launched by Google and handed to the Linux Foundation.
  • The W3C AI Agent Protocol Community Group — a grassroots group drafting open rules for a trustworthy web of agents.
  • Agent identity work — cross-body efforts, at the W3C and beyond, to verify who an agent is and what it is allowed to do before it acts.

None of these is finished, and that is the opening. The conventions are still soft enough to shape, which is exactly where Zeldman’s coalition made its difference.

Web standards diagram connecting AI interfaces, protocols, and design systems.

Designing with web standards: The playbook for this AI moment

AI interfaces are in their browser-wars moment. Shared patterns, readable contracts, and protocols can create consistency without making every product identical.

uxdesign.cc icon uxdesign.cc