Skip to content

Yennie Jun, writing for Art Fish Intelligence, asks which parts of thinking we surrender along with the task. Her example shows the difference between asking AI to answer a question and asking it to test thinking we’ve already begun:

I suggested (with only a little bit of initial resistance) that we pause and think about why this might be. I suggested a few theories. Perhaps it was Portugal’s relative homogeneity and religiousness, compared to the US’s diversity of immigrants. Perhaps Portugal clung on to so-called “Age of Exploration” as one of the most prominent chapters in its national story. We wondered, postulated, made wild guesses, backtracked, connected our ideas, disagreed, and remembered historical details we learned in high school many years ago. We drew on our collective memories, knowledge, understanding of the world, and critical thinking skills. We knew we were speculating, and some of our theories were probably wrong; that was part of the exercise.

Eventually, we asked the same question to AI. Its response corroborated many of our theories and supplied several explanations we had missed. It also omitted a few possibilities we still found plausible. We had begun with a question, generated hypotheses, and only then used AI to test and extend our thinking. I relished the exercise.

The backtracking is the point. A finished answer can save time, but repeatedly skipping the work of forming and testing a hypothesis also skips the practice that builds judgment. For designers, that’s the work we should keep, even as it accelerates production.

We still need to use trial and error to learn what to ask or try next. Otherwise, faster production leaves us less able to tell whether what we made is any good.

Jun turns from productivity to autonomy:

Am I any different from the Microphone Man? Perhaps what differentiates me is that I still collected and curated the data, formulated the questions I wanted answered, and evaluated the end results? Or that the data was my own, instead of recording other people’s conversations? There will always have to be some balance between automating menial tasks to free up time for rewarding endeavors, and doing the work yourself as a learning experience.

Jenny, another character in Ken Liu’s story, aims to counterpoint the main character’s over-reliance on his AI assistant. She exclaims, “Tilly doesn’t just tell you what you want! She tells you what to think. Do you even know what you really want anymore?” Our autonomy depends, at least in part, on continuing to participate in forming our own desires. But when we offload thinking about what we want (What music should I listen to? What movies should I watch? What food should I eat? What shoes should I wear?), who do we become?

What are we automating? Human work or human agency? Human tasks or human thinking?

Designers still have to decide what deserves to exist before asking AI to make it.

Ken Liu’s The Paper Menagerie beside a handwritten notebook, pen, and headphones.

Are we offloading too much of our thinking to AI?

AI can test and extend a line of thought, but using it before we form a hypothesis skips the practice that builds judgment and weakens our agency.

artfish.ai iconartfish.ai

Om Prakash, writing for UX Collective, argues that chat should handle ambiguous intent rather than replace graphical interfaces:

This is a distinction Erika Hall draws precisely in Conversational Design (A Book Apart, 2018) — arguably the sharpest book written on this subject. Hall argues that conversation is the right design choice only when the system needs to negotiate meaning with the user. When the user already knows what they want — “add to cart,” “filter by price,” “submit the form” — conversation is overhead. Structure is faster, more accurate, and less exhausting. Conversation earns its place only when intent is genuinely ambiguous: when the user is exploring, when their goal is fuzzy, when they need the system to meet them partway.

This framework changes how you evaluate chat implementations in the wild. The products that figured it out early didn’t go all-in on chat. They built hybrid interfaces , structured UI for known, predictable tasks; conversational AI for open-ended, exploratory ones.

For designers, the practical choice comes before any screen or prompt: how clear is the user’s intent?

Choosing the interaction mode becomes more consequential when software can act. A form submission can usually be corrected; an agent may send an email or book a flight before the user realizes it misunderstood. Prakash shifts from intent to oversight:

The UX problem this creates is the most interesting one in our field right now: how do you design an interface for an agent that doesn’t need an interface to do its job, but whose users absolutely need one to trust it?

Ben Shneiderman, in Human-Centered AI (MIT Press, 2022), has been asking a version of this question for years. He argues that the dominant framing of AI as an autonomous agent replacing human judgment is both technically premature and ethically dangerous — and proposes instead a framework of high human control combined with high automation. Not one at the expense of the other. The goal is not to minimize human involvement; it’s to make human oversight legible, accessible, and non-burdensome. Shneiderman’s model is the right north star for agentic UX: not autonomy versus control, but autonomy with control — designed into the system from the start.

This is the transparency paradox of agentic design. An agent that operates silently is efficient and terrifying. An agent that narrates every action is trustworthy and exhausting. The design challenge is finding the right level of visibility — enough that users feel in control, not so much that they’re overwhelmed by a stream of system-generated activity logs.

Conceptual interface showing agent activity, controls, and user oversight.

The interface has left the building

As interfaces recede into chat, voice, and agents, the design problem is not removing controls. It is giving people enough visibility and override to trust autonomous action.

uxdesign.cc iconuxdesign.cc

Jakob Nielsen’s AI-generated illustrations do his work no favors. His visual taste is, to put it kindly, not mine. But Jakob Nielsen has done something few technology forecasters ever do: he pulled out papers from 1993 and 1996, scored 23 predictions against computing in 2026, and published the misses alongside the hits.

The 71% headline is Nielsen’s own grade, and he acknowledges the obvious conflict:

Two clarifications. First, timing gets no separate penalty because the delay affects nearly every row. Neither paper promised a full system within a decade; Noncommand explicitly said such a system was unlikely “within the next ten years, which is about as far as one can predict in the computer field with a minimum of credibility.” The vision nonetheless took 30 years to become a widely shipped product pattern, a delay I return to in the lessons.

Second, I’m grading my own homework, which is an obvious conflict of interest. I’ve tried to score against what shipped and stuck, not against what demos well, and I show every score so you can regrade me.

So I wouldn’t get hung up on whether 71% is exactly right. Nielsen and his late co-author, computer scientist Don Gentner, still saw an intent-driven interface decades before machine learning made one practical. Some of the details are uncanny; others reveal how a sound principle can survive the failure of its original mechanism.

Nielsen on one of those switcheroos:

In 1993, I described moving an on-screen object “by selecting it by looking at it and then pressing a selection button (to prevent accidental selection).” In February 2024, Apple shipped Vision Pro with the same core selection pattern: look at a target, then pinch to commit. The prediction reappeared three decades later with its confirmation step intact. Of course, the Vision Pro remains a niche product, which is why the bandwidth and interaction-stream rows score in the middle of the scale: the high-bandwidth immersive future arrived, but as a sideshow. The main stage went to the lowest-bandwidth input device imaginable: an empty text field.

The physical channel narrowed while the semantic channel widened. Our mistake was measuring the interface by how much raw data crossed it rather than by how much work each user token could trigger. The better metric for AI is intent leverage: useful output divided by the effort required to specify the goal.

And this 1996 sentence comes remarkably close to describing the interface designers should be building beyond the prompt box:

We assumed users already knew what they wanted and just needed a better way to say it. Conversational AI revealed that human intent is rarely a pre-formed cognitive object just waiting to be translated; rather, intent is fluid and often discovers itself through the act of conversation. Current AI user interfaces do little to help users figure out their intent, but even primitive AI UX already supports some degree of iterative co-articulation.

One sentence from 1996 aged better than everything else Don and I wrote: “Real expressive power comes from the combination of language, examples, and pointing.” That’s a working definition of multimodal prompting: type your intent, paste an example of what you want, and point by uploading an image or selecting a region. If I could grade a single sentence at 100%, this is the one.

Language is best for leaping across a large solution space: “make this calmer,” “compare these contracts,” or “plan a week in Kyoto.” Pointing is best for local correction: this paragraph, that number, the face in the upper-right corner. Examples communicate qualities users can’t easily name. The winning AI interface will therefore let language propose, examples constrain, and pointing repair.

Nielsen and Gentner’s prediction worked because the underlying idea was stronger than the technologies they had available. They got the vehicle and timing wrong. They also expected expert users to benefit most. Yet language, examples, and pointing still sound like a better creative interface than an empty text field.

Two aged papers titled Noncommand User Interfaces and The Anti-Mac Interface connected by glowing lines.

Predicting the AI Interface 30 Years Ago: I Was 71% Right

Jakob Nielsen’s old predictions got the mechanism and timing wrong, but their core insight holds: language, examples, and pointing help people express intent together.

jakobnielsenphd.substack.com iconjakobnielsenphd.substack.com

Artist Katya Ross, in a talk published by CreativeMornings, the global creative community and lecture series, argues that content abundance has made intentional curation the problem:

We don’t have a content problem. We have a curation problem as a culture, as people in this overloaded culture. We don’t want more content. We want more meaningful content. […] The problem with trying to seek out meaningful content: your brain wants you to stay alive. Your brain curates according to survival, not necessarily meaning. And social media algorithms, which is a lot of where we get all of the content that we view, curate according to ad revenue. They curate according to what you pay attention to, whether you like something or not, whether you agree with it or not. It just wants your ad revenue, and it wants you scrolling.

Ross means something more specific by curation than choosing the best items from a pile. For her, it is the act of constructing relationships that create meaning:

To define curation for the sake of this talk, I would say that curation is constructing and carefully considering the relationships that make meaning possible. Curation is, in other words, the fundamental mechanism through which we can create and communicate meaning. Now, it’s just a mechanism. There’s no guarantee that your curation is going to be any good.

That shifts originality away from producing something with no precedent. The contribution is the relationship a creator sees among inherited ideas, materials, and experiences:

We are in relationship with reality. We are not truly inventing anything. We are working with what’s already there, whether that’s materials or ideas. Every idea has a lineage. Even if it is an original idea, everything that’s gone into your mind has laid the groundwork for this idea. The ideas are related to one another. Every work that you create exists in this messy, beautiful, complicated web of work that has come before, work that sits beside, work that will come after.

That is a more useful standard for originality than novelty. The materials can be familiar and the work can still be original when the relationship it reveals carries the creator’s particular way of seeing.

Katya Ross: We Have a Curation Problem

In an overloaded culture, the creative work is not making more material. It is constructing the relationships among ideas, experiences, and artifacts that make meaning.

youtube.com iconyoutube.com

Ben Callahan on what a design system can and can’t guarantee:

A design system can only raise the quality floor. It sets the baseline below which nothing should ship. An accessible-by-default button, a holistic and thoughtful approach to spacing, a template that starts a consuming team ten steps ahead.

But a design system alone can’t raise the quality ceiling. That’s not something you can do by delivering assets. The worst product teams can make awful experiences with the best design systems. That’s because the quality ceiling is set by the choices product teams make with what you give them. It’s their restraint, it’s where they push, and it’s knowing when to deviate from the standard because the standard isn’t serving the end user.

Callahan’s title invokes AI, though the essay only touches it indirectly. The connection follows from his distinction: faster generation and stronger defaults can produce more acceptable work, but neither can decide when the standard is failing the user. That decision still requires careful judgment.

Callahan on the loop:

And, of course, this loop just continues to run. Over time, the quality floor and the quality ceiling are raised.

The most important step here isn’t the shipping of a new component. It’s the time in conversation that results in alignment on a definition of quality.

Your system sets the floor. The way your system is used sets the ceiling. If you’ve poured everything into the first and bowed out of the second, it’s time to step back into ring.

Screenshot of the article page at bencallahan.com.

What is product craft in the age of AI and design systems?

Design systems can raise the quality floor, but product teams set the ceiling. Raising both requires shared standards, judgment, and an ongoing practice of craft.

bencallahan.com iconbencallahan.com

Modernism’s typographic principles have survived so long that clarity and legibility can feel less like choices than natural law. New New Typography asks what changes when designers treat those principles as products of their time rather than permanent rules.

Paul Moore, writing for It’s Nice That:

Are we stuck with the established typographic principles of the 20th century? This is what German design studio Matter Of is asking in its design book New New Typography (published by Sorry Press). The title is in direct reference to the proclamation of Jan Tschichold’s original “new typography” in Central Europe 100 years ago. Now Matter Of (and designer Wiegand von Hartmann) are at the spearhead of an even newer typographic revolution: the ‘new new,’ searching to unsettle established paradigms, expand vocabularies and innovate type design. That is, until it’s time for the ‘new new new’, but let’s not worry about that right now.

The book makes that challenge through form, not just argument. It switches among six sans-serif typefaces and marks its black-and-white pages with aggressive cyan, turning the act of reading into a test of what counts as coherent:

New New Typography makes itself known immediately with its minimalist aesthetic, using black-and-white print with a searing cerulean cutting through the page in disruptive squiggles or sometimes blocks of highlighted text (a conscious effort to make the book look like a textbook blasted with felt-tip markers). The book is designed using six different sans-serif typefaces that alternate across chapters, raising questions in ironic and subtle ways, such as ‘when is a typeface considered new or different?’, playfully reflecting on the evolution of Helvetica, Helvetica Neue, Helvetica Now. To find ways to bring proliferating typographies into relation, the book sketches engagement between the typeface’s different dependencies, all with a tongue planted firmly in cheek.

That makes the book less a replacement doctrine than a device for questioning doctrine:

Described as a messy and unstable anti-manual for typography, 243 figures and six theoretical proposals present an overview of the impossibility of an overview on typography – it’s adventurously speculative, blending high design theory with a modern irony and cheekiness. What if the modern typographer isn’t just a typographer, but an active negotiator of political systems? Nevertheless, it isn’t a book about delivering an answer. Nowadays, just daring to imagine the future is enough.

But refusing a single answer does not mean refusing standards altogether. If typography helps negotiate political systems, its experiments still need to account for who can read them and what power they reinforce. Imagination opens the field; judgment decides what deserves to remain.

Black, white, and cyan experimental typography from Matter Of's New New Typography.

Are we stuck in the typographic principles of the 20th century?

Matter Of’s anti-manual treats modernist typographic rules as historical choices rather than natural law, using form to reopen the question of what type can become.

itsnicethat.com iconitsnicethat.com

Patrick Neeman, writing for UX Collective, compares today’s AI interfaces to the browser wars:

[Jeffrey] Zeldman did not invent the specifications, he did something harder: He convinced an entire industry that shared conventions were worth fighting for, and he won. Zeldman changed the world with a stance, not a specification and we should thank him for it.

We are living through that moment again, this time for the interfaces we wrap around models, the skills we scaffold on top of them and representations they mean.

The browsers have new names: ChatGPT, Claude, Gemini, and Copilot each handle the same task their own way, with their own conventions for parsing content, showing reasoning, citing a source, and asking permission before they act.

The connection to design systems is structural. Browser standards gave different products a shared foundation without forcing them to look identical. Neeman doesn’t claim that the conventions for AI interfaces are settled; he proposes design systems as the way practitioners can develop and share them:

The core move was to pull structure, presentation, and behavior into distinct layers so each could change without breaking the others. That one idea outlived every specific technology it was built on.

It is why a design system works at all.

When Brad Frost introduced atomic design, he was extending the same instinct: stop shipping pages, start composing interfaces from small, shared, recombinable parts. Design systems are the standards movement’s direct descendant, and they are the closest thing we have to a working model for AI interface conventions.

That working model is already appearing in the Markdown files agents use as project-level contracts:

Agents increasingly take their instructions from plain text files that sit beside the work — AGENTS.md for how an agent should behave in a project, SKILL.md for what a capability can do, README.md for the context around both.

This is the new semantic layer. It is markup again, written in Markdown and read by a model instead of a browser.

[…]

A design.md that carries your design system’s patterns, tokens, and rules into every agent that touches the product. An accessibility.md that states the non-negotiables in language a model can follow. A content.md that fixes voice, terminology, and the content model.

Neeman also points to a broader protocol stack taking shape:

You are not waiting for this to begin. It has begun. A partial map of the standards taking shape right now:

  • Model Context Protocol — a shared way for a model to reach tools, data, and context, already adopted across rival platforms and now stewarded by a neutral foundation.
  • A2UI — a declarative protocol for agents to describe interfaces that render natively across web, mobile, and desktop, keeping what the interface is separate from how each client draws it.
  • Agent2Agent — an open protocol for agents to discover one another and collaborate across frameworks and vendors, launched by Google and handed to the Linux Foundation.
  • The W3C AI Agent Protocol Community Group — a grassroots group drafting open rules for a trustworthy web of agents.
  • Agent identity work — cross-body efforts, at the W3C and beyond, to verify who an agent is and what it is allowed to do before it acts.

None of these is finished, and that is the opening. The conventions are still soft enough to shape, which is exactly where Zeldman’s coalition made its difference.

Web standards diagram connecting AI interfaces, protocols, and design systems.

Designing with web standards: The playbook for this AI moment

AI interfaces are in their browser-wars moment. Shared patterns, readable contracts, and protocols can create consistency without making every product identical.

uxdesign.cc iconuxdesign.cc

Pick one metric on your team’s dashboard. What decision would change if it moved?

Ant Murphy offers two questions for finding out:

Two of my favourite questions to ask about any product metric are:

  1. What are you trying to learn?

  2. What change will you make as a result?

There are thousands of things you can measure, but measuring for measurement’s sake isn’t helpful. It needs to drive decisions.

The second question puts a team on the hook before the data arrives. It forces the decision rule into the open, when there’s less temptation to rationalize whatever the chart happens to show.

Murphy also argues that difficulty assembling a metric can be evidence that the team has stopped settling for whatever its analytics tool already exposes:

And that’s the trap that Ridgway was describing all those years ago and it’s still relevant today.

So I see friction as a good sign - it means the metric is:

  • Unique: specific to your product and your problem, not pulled off the analytics shelf
  • Meaningful: tied to the behaviour change you actually care about
  • And thought through: you’ve actually unpacked what you’re trying to learn

So it’s worth pushing through the friction. I’ll take a meaningful proxy, or a smaller sample of something high-quality, over a generic metric every time.

Product analytics dashboard showing metrics connected to decisions.

Measuring ≠ Learning

A dashboard becomes useful only when each metric answers a specific question and changes a decision. Otherwise it is measurement without learning.

antmurphy.me iconantmurphy.me

Kai Wong asked 32 design leaders what they do when companies mistake faster production for faster design. He begins with a plumber:

Imagine a plumber walks into your house, looks at the pipes for a few minutes, tightens one valve, and hands you a bill for $300. Your first reaction might be, “Anybody could have turned that valve.”

Wong borrows a distinction from neuroscientist and Tiny Experiments author Anne-Laure Le Cunff: “the ancient Greeks had two words for time, not one.”

The first is chronos: time as quantity. The number of hours in a day, the number of weeks in a year. This is the time your projects are based on.

The second is kairos: time as quality. Not how much, but how good. Le Cunff frames it not just as better quality time: it’s having the time to recognize a pattern from everything you’ve seen before and know it’s the right time to act.

[…]

Design runs on Kairos. AI might have made things faster, but businesses don’t need 500 screens by lunchtime.

They need the right solution to their problem. And that comes from the quality of thinking that happens along the way.

Wong closes:

When AI is your competition, the temptation is to compete on its terms. Faster. Cheaper. More. You’ll lose that race. And you’ll produce worse work while you lose it.

Compete on the thing AI doesn’t have. Judgment.

Generation is chronos, and chronos is cheap now. Judgment is kairos, and kairos is the whole job. Protect it. Make it visible. When someone asks you to cut the timeline in half, be ready to explain clearly what they’d actually be cutting.

Design leader reviewing work at a desk amid rapid interface production.

What 32 design leaders do when told to move faster

AI makes production cheaper, but it does not make judgment cheaper. The time to recognize the right solution is still the work design leaders need to defend.

uxdesign.cc iconuxdesign.cc

Lola Famulegun, writing for Nielsen Norman Group, separates the UX metrics that describe design performance from the business metrics executives use to judge an investment:

Upstream metrics tell you how the design performed. Examples include task success rates, error rates, SUS scores.

Downstream metrics capture what changed in the business as a result. Examples include support contact volume, conversion rates, and churn.

Downstream metrics tell you what the work was worth. You don’t need to abandon the metrics you already collect, but you do need to build a bridge from them to the ones leadership tracks. The table below maps common upstream metrics to the business priorities they most directly connect to, with a suggested framing for each.

Famulegun also points out that the translation depends on access to data outside the UX team. I recall that my team had to fight with my company’s revenue operations team to get access to Salesforce data so we could make prioritization decisions with ARR as an input.

Famulegun:

The data you need already exists inside your organization. Partner with finance, product analytics, customer support, or marketing to understand what they track and to get access to before-and-after data for flows you’ve redesigned. Even directional data is persuasive when it’s honest: “Contacts about [feature] dropped 30% in the quarter following the navigation change.”

A practical note: this translation only works if you have access to downstream data. If your team isn’t currently connected to product analytics, customer support, or finance reporting, that’s the first conversation to have.

The goal isn’t to overstate what UX delivers. It’s to surface the connection that already exists between the work your team does and the numbers the business is tracking.

Diagram linking UX metrics to revenue, retention, and support outcomes.

Stop Reporting UX Activity and Report Business Outcomes

UX metrics matter most when teams connect them to the business outcomes leadership tracks, from revenue and retention to support costs and risk.

nngroup.com iconnngroup.com

Laura Summers, writing for the Pydantic blog, describes the strange math of coding agents: the work can run in parallel, but our attention can’t.

Marcelo, another Pydantic colleague, when asked about his Claude Code session freezing said: “just open 5 claude sessions. You’ll never notice because you’re busy giving feedback to the others.” He was joking. I think. But it captures something true about the current moment. The parallelism is exhilarating and kind of feral. The number of things you can start has dramatically increased. The number of things you can thoughtfully finish hasn’t changed at all, because that part still requires the one resource we can’t parallelise: your brain.

The design version is easy to recognize: an agent can generate dozens of screens and states while one brain still has to judge the product intent and every edge case. You may spend less time drawing the interface, but every state still demands a decision.

Summers calls the emotional cost “the human reward function problem”:

Here’s a term for what I think is happening: the human reward function problem. In machine learning, a reward function tells an agent what good looks like. Writing code by hand was never easy, but it was full of small rewards. Solving a problem in your head. Understanding a gnarly bit of logic. Watching the code compile. The feeling of control. LLM-assisted programming has automated much of the work that generated those dopamine hits and replaced it with the cognitive load of review and supervision. The satisfying part shrank. The exhausting part grew. And there are no new rewards to fill the gap.

If you’re feeling like your work is simultaneously more productive and less satisfying, you’re not broken. The feedback loop is broken. And I think we need to start treating that as an engineering problem in its own right, not a personal failure.

Person monitoring multiple coding-agent sessions on a computer screen.

The Human-in-the-Loop is Tired

Coding agents can multiply the amount of work started, but not the attention required to judge intent, review output, and finish work thoughtfully.

pydantic.dev iconpydantic.dev

I’ve intentionally covered loops a lot this week. It’s been the talk of the virtual town of late, so it’s an important concept to understand as AI-assisted software design and development continues to mature.

MC Dean pulls it all together practically for designers in this piece. She reminds us that designers need to stay in the conversation; be in the room where it happens:

The conversation about AI and design tends to run in one direction: here are tools you can use to do your job faster. That framing keeps designers in the task loop. Tools help you execute. The loop stack shows you where the real work is.

The real work is authoring the system loop. Writing the specifications that constrain how AI behaves. Encoding the quality standards that define what good output looks like. Making the judgment calls that no loop below the oversight layer can make on its own.

Taking a “loop stack” built by engineers, she adapts it for designers:

At the base: the execution loop. The model fires, tool calls happen, tokens generate. This is genuinely not your concern. You don’t need to understand transformer architecture to design well with AI, any more than you need to understand TCP/IP to design a good website.

One layer up: the task loop. The agent works on a specific thing until a condition is met, then stops. Who defines that condition? You do. “The task is complete when the output meets the brief” is a design decision. What counts as meeting the brief is yours to specify.

Then the product loop. This is the experience layer, the thing a person actually encounters. Does the flow hold together? Does the output feel like it belongs to a coherent system? Does it match the quality bar? Every heuristic you’ve ever learned, every design principle you’ve internalized, lives here.

Then the system loop. This is where the AI gets better, or doesn’t. The patterns it learns, the constraints it operates within, the values it embodies. Design systems, behavioral specifications, brand guidelines, content principles, tone of voice. Everything you’ve encoded about what good looks like. This is the layer that trains the loops below it.

At the top: the oversight loop. This is where human judgment lives permanently. Not as a checkpoint at the end. Not as a review gate before launch. As a continuous presence that can redirect anything in the stack at any time.

This is what we call “craft”.

In the full post, she shares a prompt for you to try that illustrates how the loop actually works. Go try it.

A glossy white sphere centered among flowing blue concentric waves and curved lines, suggesting gravitational pull or ripple effects.

/loop

Every AI conference this year was about loops. Here’s what that means for design, and a loop you can run in the next ten minutes.

marieclairedean.substack.com iconmarieclairedean.substack.com

The interesting part of agent loops isn’t how to keep an agent running. It’s knowing when the work is understood well enough to let it run.

Peter Yang, interviewing Claude Code’s Thariq Shihipar on Behind the Craft, asks about Claude Code’s new loop, goal, and workflow features. Shihipar describes /goal less as an autonomy switch than as a signal that the uncertain work has already been done:

/goal is great when you have a complicated task and need to make sure it is done at the end. It’s the user indicating, “I’ve done enough specification and exploration. I understand the problem space. Just go execute on it, and if you run into something, fill it in.”

That puts a useful boundary around loop engineering: don’t ask the agent to power through ambiguity you haven’t investigated. Shihipar treats planning as the work of reducing that ambiguity:

We often talk about plans as one shot: you plan, then you do something, and that’s it. But planning is an iterative process of exploring, investigating, and finding out what you don’t know and what you want.

A few minutes later, he gives that process a better name:

I like to call it getting rid of your unknowns. With almost any task, there’s a lot you don’t know—either how things work or what you want. It’s very iterative. You don’t write it all down once and then implement it. There are many steps and different passes.

Implementation doesn’t end that process. Shihipar asks the agent to keep notes about what it discovers while building, then feeds those discoveries back into the specification:

The model can find things that it—or you—didn’t anticipate during implementation. I ask it to keep implementation notes as it goes: what did we not expect about this implementation? Once we have that, we can respec if needed. It’s much less one handoff from specification to implementation and more a back-and-forth process.

That’s the more useful loop: explore until “done” is concrete, build the smallest version that can expose what the plan missed, then revise the plan from what the build teaches you. Longer-running agents are the consequence, not the point.

Quotes lightly edited for clarity.

How I Plan, Build, and Run Loops with Claude Code in 40 Minutes | Thariq Shihipar

Thariq works on the Claude Code team, and I’ve wanted to see how he builds for a long time. In our episode, he showed how to use /goal to keep Claude working, how he plans with Claude to remove unknowns before building, and how he runs a team of agents in Slack. He also shared why his team cut…

youtube.com iconyoutube.com

Wayland is a display protocol many Linux desktops use to coordinate how applications draw to the screen. Its “every frame is perfect” goal means users should see a complete, coherent image rather than partially updated output. Nikita Prokopov argues interface designers should hold transitions to the same standard:

A stated goal of Wayland is “every frame is perfect “.

And I think this is a goal we should all aspire to. Wayland is talking about the technical side of things (modern GPU stacks are very complex and Wayland is trying to take control back) but it could be applied to UI too.

His screenshot test belongs in design review. Stop the transition at arbitrary points and ask why each component is where it is.

Prokopov shows how a small mismatch becomes a larger trust problem:

Not the end of the world by any means, but it does create a feeling that these two components are not in sync with each other. Next thought: maybe they weren’t designed together? If so, then they might not work well together. That’s how trust is lost.

Motion can also tell the user something happened when it didn’t:

This creates a false feeling that something subtly changes when you switch between modes. And you know what? I don’t want my UI to give me false feelings. I want it to be a precise instrument, not an animated toy.

Interface animation frames demonstrating alignment and coordinated motion.

Every Frame Perfect

Interface quality lives in the states between screens. Each visible frame should be coherent enough that users can understand what changed and why.

tonsky.me icontonsky.me

The click-target problem is why I’m fine with macOS’s mandatory squircle. Older icons sometimes made negative space look clickable when it wasn’t. Apple’s uniform shape makes the target predictable—a small accessibility improvement worth the trade, IMHO.

Louie Mantia, writing on the Parakeet blog, explains the icon designer’s tradeoff:

Masking all of these app icons to a squircle, and even applying Liquid Glass effects to them, aims to solve this problem. And this follows the same principle of iOS 7, which is to make it easier for all apps to fit in on the platform, especially apps built by designers and developers who aren’t familiar with how to make an icon that looks great next to first-party icons.

Just so I’m clear about my preference, I would love if Apple provided a way for designers to poke outside that squircle boundary. Some of my favorite app icons did that. But also some of my least-favorite app icons ignored this shape entirely, when it was used for every system icon in the last five years. Whenever those apps showed up in my Dock, it was like a stain on my shirt I couldn’t get out.

Despite the genuine loss associated with the squircle restriction, there’s more than one way to design with it.

Mantia is right that Apple handled some older icons poorly. Designs made specifically for the pre-Tahoe squircle were shrunk and placed inside the new one, punishing designers who had followed the previous rules. But Mantia’s broader point holds. The Dock also has to accommodate neglected utilities, unmodified corporate logos, and apps whose makers never learned the platform’s conventions. The squircle makes more of them look correctly sized next to one another.

Mantia goes beyond visual consistency and describes the shape’s semantic job:

Seventeen years later, now that app icons on macOS do have a uniform rounded-square appearance, they do look better adjacent to each other in the Dock. What’s even better is that over the last 19 years, apps on Apple platforms have distinguished themselves with a very recognizable visual identity.

The shape of apps is a squircle. And it has been proven to work for everyone. Companies can use their logo as an app icon. Designers can create something specifically for the platform. And both of these get to look like an app. Whether people consider the squircle a container or a canvas, this uniform appearance communicates its function: a squircle represents an app, just like how a piece of paper represents a digital document, or a folder represents, well, a folder.

Mac app icons arranged in rounded-square containers in a Dock.

The Shape of Apps

macOS’s mandatory squircle trades some icon individuality for a more predictable, accessible target and a clearer visual signal that something is an app.

parakeet.co iconparakeet.co

Jihoon Jeong asks the question: if the model starts fresh each cycle, what actually compounds? He uses software developer Geoffrey Huntley’s Ralph Wiggum loop to illustrate:

If the loop’s power came from looping — from persistence of effort, from the agent grinding away at the problem — then the longer you could keep one agent going, the better it should get. The opposite is true, and every practitioner knows it. A long agentic session curdles as its window fills with dead ends and stale state; the agent gets worse with continuity, not better. The winning configuration, rediscovered by everyone who runs loops at any scale, is maximum discontinuity: kill the agent every iteration, resurrect it blank, and let it inherit nothing except what the last iteration wrote to disk. Huntley’s design wasn’t naive. It was surgical. Discard the mind, keep the files.

Fresh context is only useful if the loop can tell progress from activity. Jeong puts that burden on the verifier:

Second, one of the five decisions is load-bearing in a way the others aren’t. The verifier is the wall the whole structure hangs on. A loop repeats whatever its verifier accepts; if the verifier is strong — tests, compilers, benchmarks, anything with teeth — the loop compounds progress, and if the verifier is weak, the loop compounds output. Every experienced loop practitioner converges on the same rule: the loop is exactly as good as its stopping test. A loop with a weak verifier isn’t an autonomous engineer. It’s an expensive random walk with excellent posture.

Jeong’s answer is that the model starts each cycle from scratch, while plans, tests, commits, and code preserve progress for the next one:

The loop works, and the skeptics are right about why its working is strange. It adds no intelligence. It makes no model smarter. It rents the same brilliance every cycle at full price, extracts what it can, and throws the brilliant thing away — keeping only the residue on disk, because the residue is the only part that compounds. It works better than it has any right to, exactly as I said last time. And its characteristic failure mode is now visible at scale too, and it is not a crash. It’s a flatline. The loop keeps turning, the tokens keep burning, and the density of correct answers stays wherever the verifier pinned it — because nothing inside the system learns from one cycle to the next. The agent that finishes iteration forty is precisely as capable as the one that started iteration one. Only the pile of files has grown.

Illustration for an article about verification and durable state in agent loops.

The Year of the Loop

Agent loops do not improve because a model remembers. They improve when durable plans, tests, commits, and verification preserve the right residue between fresh sessions.

medium.com iconmedium.com

After Addy Osmani’s introduction to loop engineering, Robert Ross, writing at The Thought Drop, opens up the machinery. What looks like one agent loop is really three nested loops:

Agent loops are often oversimplified. They’re presented as a single loop, when really it’s three loops in a trench coat that make up an “agentic” experience for a customer. I’m here to write (yes, I wrote this, insane right?) yet-another-blog about agent loops. The example code blocks are also pseudo-code and for illustrating these ideas. Also I’ve omitted streaming, which complicates the post but the shape of these stays the same.

Those are the inference loop, which manages model calls and conversation history; the tool loop, which turns model output into actions; and the human loop, which approves, rejects, or redirects consequential work.

Ross’s “brain in a jar” analogy explains why the tool loop changes a model into an agent:

LLMs are brains in a jar. They provide no functional value on their own. The tools you give an LLM are what make it an agent.

When you tell a model “here are the tools you have” in your outer inference loop, the model may try to “use” them in its inference (response). This is the same thing as a brain sending an electrical signal telling your index finger to hover over the enter key of the email you desperately want to send Laney. Tom, we need to set boundaries my man.

The separate tool definitions you include in your API request are usually serialized into the system prompt field of the token stream the model processes. And it may infer the usage of multiple tools in one turn. (Hence: Tool Loop).

The human loop is the final layer—and the hardest to build:

The Human Loop is arguably the hardest part to implement in agentic systems. You can’t have a piece of code block for hours. What if the server restarts? What if you have thousands of other requests coming in you need to respond to? The first two loops (inference and tool) are simple enough. The human loop ups the ante of difficulty. This is why durable execution frameworks exist, like Temporal.

But the human loop is necessary, because it’s the only thing stopping Tom from actually sending that message to Laney. IT WAS TWO YEARS AGO TOM, MOVE ON!

Three nested circles labeled inference loop, tool loop, and human loop beside the article title.

The Agentic Loop: Three loops in a trench coat

Agentic systems are not one loop but three: inference, tool use, and human oversight. The last is the hardest—and the one that keeps consequential work accountable.

bobbytables.io iconbobbytables.io

We’ve heard about prompt engineering and then context engineering, and now it’s loop engineering. Googler Addy Osmani offers a clear introduction to it, beginning with a simple definition:

Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead. A loop here can be thought of a recursive goal where you define a purpose and the AI iterates until complete.

What changes is who keeps the work moving:

For like two years the way you got something out of a coding agent was you wrote a good prompt and shared enough context. You type a thing, you read what came back, you type the next thing. The agent is a tool and you are holding it the entire time, one turn after the other. That part is kind of over, or at least some think it’s going to be.

Now you build a small system that finds the work, hands it out, checks it, writes down what is done and then decides the next thing, and you let that system poke the agents instead of you. I wrote before about the cousin of this, agent harness engineering, which is making the environment one single agent runs inside and the factory model - the system that builds the software. Loop engineering sits one floor above the harness. The harness but it runs on a timer, it spawns little helpers, and it feeds itself.

And the loop itself has a recognizable anatomy:

A loop needs five things and then one place to remember stuff. Let me list it first and then map it.

  1. Automations that go off on a schedule and do discovery and triage by themselves.
  2. Worktrees so two agents working in paralell dont step on each other.
  3. Skills to write down the project knowledge the agent would otherwise just guess.
  4. Plugins and connectors to plug the agent into the tools you already use.
  5. Sub-agents so one of them has the idea and a different one checks it.

Then the sixth thing, the memory. A markdown file, or a Linear board, anything that lives outside the single conversation and holds what’s done and what is next. Sounds too dumb to matter. But it’s the same trick every long running agent depends on and I went into it in long-running agents, the model forgets everything between runs so the memory has to be on disk and not in the context. The agent forgets, the repo doesnt.

Illustration accompanying Addy Osmani's article on building recursive coding-agent loops.

Loop Engineering

Loop engineering moves the work from one-off prompts to durable systems that discover, delegate, verify, and remember work while humans remain accountable for the result.

addyosmani.com iconaddyosmani.com

Karo Zieminski and Dheeraj Sharma recommend starting with a critic: one recurring job, one explicit standard, and a loop that stops before autonomy outruns our ability to inspect it. Their example reviews PRDs, but the pattern fits any creative work whose quality we can describe clearly enough to test. Sharma grounds that advice in the 30-plus agents he has built for his content operation:

I have built 30+ agents that now keep a real content operation running across my newsletter and YouTube channels. You’d be surprised how modest the useful ones look. If you start with an agent that “runs your whole business”, you’ll most likely build something fragile. OpenAI’s advice is to maximize a single agent’s capabilities first before even thinking about multiple agents. Anthropic’s rule is even stricter: add complexity only when it demonstrably improves outcomes. One agent, one job, one loop.

The rubric is the consequential design artifact. It turns tacit judgment into criteria the agent can apply consistently and the human can challenge. The retry limit matters for the same reason: repeated failure becomes evidence that the product thinking needs work, rather than an invitation to let the loop run forever.

A real critic checks whether the doc can do its job after engineering pokes holes in it. It needs to be forced to review every PRD through the same fixed format (every single time), and come back with a score, a diagnosis, and a concrete fix list. It also needs a retry limit. For PRDs, 2–3 rounds is usually enough. If it still fails after 3 loops, revisit the product thinking. And keep notes about every failure. Anthropic’s evals guidance treats every bug as a test case. The PRD your critic scored wrong last week is the exact document you re-test it against after every change.

For designers, this is a practical way to keep judgment inside the system. The agent can expose weak reasoning and carry the review process forward; deciding what deserves to ship remains a human responsibility.

I always come back to the same rule: agentize the tasks, not the craft.

Use your agents to move the PRDs to GitHub, but review them first.

Keep human decision gates at the moments where judgement matters.

Keep using the parts of your brain that make the work yours. Keep the joy you find in creating it.

Visual-guide cover for building your first AI agent as a PRD critic.

How to Build Your First Agent. One That Works.

Your first AI agent should be a critic: one recurring job, one explicit rubric, and a loop that stops before autonomy outruns your ability to inspect what it produces.

karozieminski.substack.com iconkarozieminski.substack.com

Claire Vo, who built a bug-triage harness for her company ChatPRD, offers a usefully plain definition of an AI harness. The important part is that the intelligence does not live only in the model. Some of it lives in the surrounding code that prepares the work, limits what the agent can do, and decides what it must leave behind.

A harness is some code around an AI agent. Yes, you heard it here first. A harness is just code around an AI agent that makes it more effective. Can that code have AI in it? Sure. Does that code have to have AI in it? Not necessarily. What is the goal of a harness? To make the AI better. It is so simple, and I feel like the way that people have been talking about this has made it such a mystery that I wanted to make it very clear to you all. It is just writing more code around your AI to make it more useful for a specific use case.

Vo’s threshold for building one is equally practical: look for work where the setup and expected result recur.

So what are the parts of a harness? Well, a harness is going to have specific context. It’s going to be able to take specific actions, and it’s going to have a goal of specific outcomes. It’s just as simple as that. And I want to talk about when it makes sense to build a harness and when it doesn’t. I think you’ll want to build a harness when the same workflow needs the same setup and the same outcomes. It’s really when there is a combination of deterministic and non-deterministic workflow, step-by-step process, tools, and use cases you want your AI to follow to do a specific job.

That turns harness-building into a design problem. The work is choosing the job, shaping the workflow, narrowing the tools, specifying the artifacts, and creating an interface through which a person can direct and inspect the system.

I identified a specific workflow. I determined what the run against the task would look like. I made very opinionated calls to tools or data sources. I didn’t just say, “Use an MCP,” although that could be part of your harness. What I did is make adapters that made the calls to these external APIs and tools very specific. I thought about what the structured artifacts out of that workflow might be. I decided what rules and permissions I wanted to give this harness and which ones I didn’t. I decided whether I wanted to use Claude Code or Codex or a model router to actually run these things. And then I built a surface to interact with this agent. It could be a TUI. It could be a CLI. It could be a web app. But I built some way to interact with this.

The model supplies capability. The harness makes a repeatable workflow legible and enforceable.

What is a harness and how to build one with Claude Agent SDK

A plain definition of an AI harness: the code around an agent that prepares its work, limits what it can do, and decides what it must leave behind. Built around a live bug-triage example.

youtu.be iconyoutu.be

Christine Vallaure, a UI designer who teaches Figma and AI workflows, offers designers a map of the hidden infrastructure between a convincing Figma-to-code demo and a production workflow:

The demos show you one clean layer working under perfect conditions. Your actual work needs three or four layers stacked together, and nobody shows you the stack, because the stack is where it gets messy and half-solved. The confusion comes from not knowing they are separate things at different stages that need different skills.

A useful distinction is between context and connection. Figma’s Model Context Protocol (MCP) connection can expose the file, markdown can preserve working rules, and skills can make repeated tasks more consistent. None of those guarantees that a generated button is the button already maintained in the product. That takes an explicit mapping to the codebase, plus someone responsible for keeping it current.

Look at the whole stack. Each layer covers the hole under it. The pipe lets Claude see your design. The note carries your rules. The recipe keeps your repeated jobs consistent. And the mapping, the top layer, is the only one that truly welds design and code together so they never drift.

But that top layer needs a codebase, developers, and constant upkeep. Most people do not have that, and should not pretend to. So for almost everyone, the design and the code will drift the moment either side changes, and that is not a sign you set it up wrong. It is simply what these tools are without the expensive wire: they take a snapshot and build from it. When things drift, you regenerate. You do not try to hand-repair a connection that was never really there.

That makes the stack an ownership map as much as a technology map. Each added layer creates another artifact that can become stale. The right setup is the most infrastructure a team can actually maintain, rather than the most complete diagram it can assemble.

And the flip side holds: if you are not this team, do not build like this team, or you will spend your life maintaining a machine you never needed and cannot keep up with.

Diagram-style cover mapping the layers between a Figma-to-code demo and a real production workflow.

You design it. Then what? A clear map of the Figma-to-code AI mess

A map of the layers between a convincing Figma-to-code demo and a real production workflow—the pipe, the note, the recipe, and the mapping—and how much of that stack a team can actually maintain.

uxdesign.cc iconuxdesign.cc

Slack Design Ops practitioner Sheila Kazan begins with a question designers were asking privately:

“How do I use AI?”

That question, typed in a DM rather than asked out loud, told us everything. Our design team was not short on curiosity. What was missing was somewhere to be a beginner. A space where not knowing wasn’t a liability, but the whole point. So we built one. In true Slack fashion, our AI origin story starts with a Slack channel and a lot of enthusiasm.

Kazan on why Slack built its own program:

The problem wasn’t a shortage of learning opportunities. It was the opposite. Tool enablement sessions started flooding our calendars from every direction, and almost none of them were built with designers in mind. Most of these sessions were designed for engineers, and we were just along for the ride. There was another wrinkle: things were moving so fast that a setup guide from Monday was outdated by Friday.

So we decided to build our own AI enablement programming. For design, by design.

That phrase became our north star. Hearing what AI tools can do from a fellow designer lands completely differently than hearing it from an engineer. It wasn’t evangelism. It was permission, and for the designers on our team who were still waiting to be convinced, that distinction mattered more than we expected.

Kazan on an outcome that doesn’t show up in a prototype:

And you know what? That’s okay. That’s also a really good finding. Not every designer walked away with a working prototype or a merged PR. Some walked away with something harder to measure and more important: a clearer sense of where they currently stand with this technology, what excites them, what makes them uneasy, and what questions they still need to answer for themselves.

That’s the thing about building a learning culture: the value isn’t always in the output. Sometimes it’s in the container. When people know there’s a place to bring their confusion, they bring it. And when confusion is visible, it becomes something the whole team can work on together.

Design leaders should budget for the learning environment alongside the licenses.

Cover for Slack Design's Builder Days, where designers learned to build with AI together.

We Didn’t Teach Our Designers AI: We Built a Place Where They Could Learn It Together

Slack Design didn’t run another tool-enablement session. They built a place to be a beginner—for design, by design—where not knowing was the whole point.

slack.design iconslack.design

Phil Morton, who writes about product design, research, and AI, argues that design teams can’t adopt AI one designer at a time. The process only changes when design and engineering change it together:

New ways of working mean that designers and engineers have to work closely together. The ideal is that you’re in the same code repository as the engineers, not throwing a Figma file over the wall.

For most teams that’s a long way from today. Designers and developers might sit in the same squad, but they’re still siloed, with a big handover in the middle.

So when you’re trying to work out what your AI design process is going to look like, you’re not just choosing new tools for yourself. Your engineering partners have to adopt the same approach, because it makes no sense for you to work one way and them another.

The handoff is the stubborn part. A designer can learn Claude Code and still end up handing a different artifact across the same organizational boundary. Without a shared repository, components, and review process, the team has improved one person’s output while leaving the production system alone. Tool fluency helps, but it doesn’t create a shared way of working.

Morton also shows why there can’t be one standard AI workflow for every design project:

In the old way of working, the output was roughly the same whatever the project: a high-fidelity Figma file.

Now it varies wildly. If you’re assembling a feature on an established product with a mature design system, you barely need Figma at all. AI is good at using existing components and you’re not asking it to do any visual design, which it’s bad at. You can sketch the rough idea, have it assemble it and iterate. There’s little point building a pixel-perfect version by hand first.

But if you’re working on a new product which needs its own visual design or brand, it’s a different job. AI can get to a rough wireframe using generic components, then it stalls. Creative and original visual design still needs a human.

Illustration for an essay on why design teams struggle to adopt AI without changing how they work together.

Five reasons design teams are struggling to adopt AI

Design teams can’t adopt AI one designer at a time. The process only changes when design and engineering change it together, sharing a repository instead of a handoff.

philmorton.co iconphilmorton.co

In their December 2025 survey, Noam Segal and Lenny Rachitsky found that designers were getting less from AI than their peers. I wondered whether designers were failing to make the shift from production to strategy. Their 2026 survey points to a harsher answer: AI is raising output expectations faster than organizations are redesigning work around human judgment.

But then we looked closer at what “better at my job” means. When we asked people to describe in their own words how AI had changed their work, “better” turned out to mean producing more and faster, but not higher quality. The productivity gains are coupled with deep unease about the costs of leveraging AI.

“I can do more, faster, but not better.”

“Amplified and destabilized at the same time. We just set a new denominator for the job. And it moves higher and higher every month.”

This is what happens when leaders add AI without redesigning jobs. Every saved hour becomes capacity to fill. People still carry the judgment calls and quality bar, only now at machine speed. The tool creates leverage; management decides whether workers experience that leverage as agency or pressure.

Design and research show what happens when that redesign lags:

Among researchers, 51% are “anxious about my job security,” versus 15% of founders. Among designers, 63% feel “overwhelmed by the pace of change” and 61% feel “tired,” the highest of any role. Researchers are among the most likely to fear “losing my job to AI” (36%, just behind Data/Analytics at 38%), and designers are the most likely to feel the comp squeeze (61% selected “expected to do more for the same compensation”). Both report the lowest willingness to recommend their field of any role, and designers, as we’ll see, report the worst-rated managers in the survey.

Last year, designers and researchers showed the largest negative sentiment shift of any group. A year later, they’re the most negative on nearly every measure we have.

The damage also reaches the people who have not entered the field yet. Experienced workers can augment judgment built over years, while companies automate the tasks through which juniors would have developed that judgment. The industry must rebuild the apprenticeship pipeline, not merely reopen entry-level requisitions. Segal and Rachitsky’s respondents can already see the break:

More than half of working tech professionals would actively steer a newcomer away from the path they chose. That translates to an average NPS score of –39. Moreover, a third of the people who call themselves optimistic still wouldn’t recommend their own field.

The cleanest way to say it: “The water’s fine; don’t come in.” People have largely made peace with their own trajectory. They’ve got the skills, the relationships, and the seniority to ride it out. But they’ve lost faith that the on-ramp still works for someone behind them.

“I’m lucky I’m later in my career … AI can augment what I’ve built. I think I won’t be in a position to hire and mentor new PMs, but I’ll be safe. Which feels really crappy to say.”

The real split is between people with enough accumulated agency to turn speed into leverage and people forced to absorb speed as a higher baseline. Companies pocketing the gains while neglecting job design and apprenticeship are spending down the human systems that made those gains possible.

Chart-driven cover for a 2026 survey of tech workers showing a workforce splitting in two.

How tech workers are feeling in 2026: a workforce splitting in two

A survey of 5,920 tech workers finds a field splitting in two: rising burnout, a productivity squeeze that trades quality for speed, and the sharpest anxiety among designers and early-career workers.

lennysnewsletter.com iconlennysnewsletter.com

Joe Wilkins reports a sharp reversal in what one finance firm wants from new graduates:

As one New York financier told Financial Times journalist Gillian Tett, new hires who were seen as “AI natives” are turning out to have alarmingly shallow ideas. So much so, the anonymous finance worker admitted, that his firm now actively avoids seeking out AI-literate STEM graduates, and opts to comb through humanities students instead.

“We want critical thinking, not just AI,” the financier told the FT.

The risk begins when familiarity with a tool gets mistaken for the ability to think through the work the tool produces. Wilkins points to what students lose when the shortcut replaces the practice:

Over the past few years, a veritable tidal wave of headlines, studies, and think pieces have flooded the internet with horror stories about the decline in literacy rates, social skills, and critical thinking abilities of the country’s college students. While there’s a kernel of truth that these factors had already been slowly dwindling prior to the widespread adoption of AI, the tech only seems to be accelerating the drop-off in real-life abilities, particularly among young people for whom it can serve as a cognitive crutch.

The state of higher education is so bad that many of today’s higher ed students are not only offloading their coursework to AI chatbots like ChatGPT — a shortcut, educators say, that’s even impacting their ability to participate in face-to-face discussions.

AI fluency can expand what someone is capable of producing. It cannot supply the close reading and judgment needed to recognize whether that production is any good, or the communication skills needed to explain why. Early in a career, doing the work is how people learn to judge that output and explain their decisions.

While plenty of thought leaders have waxed lyrical about the importance of “AI literacy” — an understanding of how to effectively use AI tools, basically — the businesses these future students are heading toward are still heavily reliant on literacy literacy. For all its revolutionary potential, there’s ample evidence that AI has yet to meaningfully impact productivity in the US, meaning that students who go all-in on AI at the expense of other skills will likely find themselves ill-prepared for the actual demands of life after college.

Illustration for a report on 'AI native' graduates arriving in the workplace with shallow critical thinking.

Bosses Horrified as “AI Native” College Graduates Hit the Workplace

One finance firm now avoids AI-literate STEM grads and combs through humanities students instead. They want critical thinking, not just AI fluency.

futurism.com iconfuturism.com