The label printer died halfway through unpacking. And Daniel didn't reach for the laptop. He didn't even think about it.
That's the part that gets me. Not the failure. The absence of the thought.
Here's what he wrote in. He needed new labels for the inventory system, asked ChatGPT in Chat mode what kind of label would actually stick to a Ziploc bag, couldn't find the type it recommended, and then had the idea of ordering custom vinyl stickers designed for low-surface-energy plastics instead. It suggested dimensions. He opened Canva and built a basic template, on his phone, because the workstation still isn't set up in the new place. And trying to make a design look halfway decent on a phone is close to impossible.
So he remembered the Canva integration was already connected.
Right. He typed something like, here's the link to the template I started, could you make it look good. And the agent recognised that he wanted actual action taken on an external service, and moved the conversation into the work environment where the tools live. And then something subtler happened. The model changed. To one optimised for agentic work.
He noticed that.
So the questions he's put to us: why does conversation feel like such a natural way for humans to work. What actually changed when ChatGPT split into chat and work. What happens under the hood when a thread gets handed to an agentic model with tools attached. How does describing an objective compact dozens of micro-decisions into one act. Why does the first draft make the fiddly details bearable. And how fast have AI tools reshaped our instincts about how we like to work. Six questions, one anecdote.
And a confession buried in the middle of it. That he never considered the old way.
So let's start with the thing Daniel noticed but didn't quite say. That he never even considered the old way.
Before we get to what OpenAI built, we need to understand why his instinct was right in the first place. And that's a question cognitive scientists were asking twenty years before ChatGPT existed.
Which is the part of this that I find interesting, because everyone's treating it as a product story.
It's both. Start with the product, because the product is what Daniel actually touched. ChatGPT Work launched on the ninth of July this year. Three new models across fourteen configurations, and it consolidated the ChatGPT and Codex desktop apps into one thing. Then on the twentieth of July the desktop app split into two visible modes, Chat and Work, and the custom instructions limit quadrupled to five thousand characters.
Two modes in one app. That's the thing he's describing.
Ten million users across Work and Codex within three weeks of launch. And ChatGPT itself was estimated to cross a billion monthly actives back in June, and a billion weekly actives in August. So the scale of who this design decision lands on is, roughly, everyone.
Except it's not settled. That's the part I want to flag early.
No, and this is where it gets awkward. As of today, ChatGPT does not support switching between Chat and Work inside the same conversation. There's a feature request from the twenty-first of August asking for exactly what Daniel described, and OpenAI Support responded on the eighth of September saying they'd pass it along and had no timeline to share.
So Daniel's flow is aspirational.
Partly. He may be describing how it's about to work rather than how it works this morning. The developer community is asking for precisely his experience as a feature.
Which is a strange thing to have to say about a product that a billion people use. The most natural-feeling thing about it is the thing that isn't shipped.
And Greg Brockman has confirmed Chat and Work merge by the end of the year. So what Daniel is describing is a transitional interface. A few months of scaffolding around something bigger.
So what is this actually about? A product feature, or a claim about human cognition?
I think it's a claim about cognition that happens to have a product attached. And the cognition half is older and better established than most people realise. Garrod and Pickering, two thousand four, Trends in Cognitive Sciences. The paper is called Why is conversation so easy?
Which is a good title, because it's the question.
Their answer is that humans are designed for dialogue rather than monologue. And the mechanism is interactive alignment. When two people talk, they automatically align their representations at every level. Phonological, syntactic, semantic, and the situation model, the shared picture of what's being discussed.
Automatically. Not deliberately.
Not deliberately at all. You don't decide to match your partner's vocabulary or syntax. You just do. And the consequence is that alignment distributes the processing load between the two of them. Each person reuses information the other has already computed.
Say that again, because that's the load-bearing sentence.
Each reuses information computed by the other. That's the cognitive science mirror of what Daniel noticed. When he described the objective to the agent, he wasn't offloading the work of choosing a font. He was offloading the representational bookkeeping. The agent holds the state, he holds the goal.
And Garrod and Pickering have a number for how tightly coupled it is.
Speakers start talking on average about half a second before their partner finishes. Half a second of overlap. Conversation isn't alternating monologues with polite gaps. It's a joint activity running at a clock speed that would be impossible if either party were doing the full work alone.
So when Daniel says describing what he wants feels cheaper than operating the interface, he's not being lazy. He's using the channel he was built for.
The interface is a monologue you have to author yourself. Every menu is a sentence you have to write in a language somebody else designed. Conversation is the one interface where the other party is doing half the parsing.
Now the under-the-hood half, because I want to know what actually happened when that thread moved.
Simon Willison did the teardown at the end of August. Two hundred and twenty-three registered tools. Forty-four skills. And his definition is the cleanest one anybody's written. Chat is the model answering questions. Work is the model doing things.
That's the whole split in eleven words.
It's two products, not one. Work Cloud runs on chatgpt.com and mobile, on an isolated microVM. Pro tier gets eight CPUs, twenty gigabytes of RAM, sixty-four gigabytes of disk. Plus gets fourteen gigs of RAM. Then Work Local is the desktop app, formerly Codex, and that one touches your actual files.
So one of them is a sandbox in a data centre and the other is sitting on your machine.
And both run on the Codex harness. Which means they inherit sub-agents, browser use, and long-run task grinding. The thing that was built for software engineering is now the engine for everything else.
The model switch Daniel noticed. Is that real or is that him reading tea leaves?
It's real and it's observable. Chat offers GPT-5.6 in Instant, Medium, High, Extra High and Pro. Work exposes Sol, Luna and Terra with reasoning levels from Light up to Ultra. Different names, different configurations.
And there's harder evidence than the names.
There's a bug report from the developer community where someone captured the same account hitting two different execution architectures. Chat requests lasted about four seconds before the server reset them. Work got a long-run handoff. And the server-side fields said product experience equals work, requested model experience equals work.
So the switch is in the logs.
The switch is in the logs. Daniel wasn't imagining it. His thread crossed a boundary that the system records.
What does Work have that Chat doesn't? Concretely.
Internet-connected code execution, where Chat's sandbox is walled off behind a container proxy. A headless Chrome browser that can fill forms, run JavaScript against the DOM, and hand off when it hits two-factor authentication. A persistent shared filesystem. Willison had a hundred and seventy-one folders sitting in his scratch directory.
A hundred and seventy-one.
That's a workspace. That's not a conversation. That's a room somebody's been working in.
And Chat has none of that.
Chat has none of it. Plus ChatGPT Sites, sub-agents, and scheduled automations. Work can start something at nine in the morning and still be at it at noon.
Here's the part I keep circling. Latent Space described the UI as stripped of the evidence that would give away you're talking to a coding agent. No git controls. No diff traces.
Deliberately. The git controls and the diffs would tell a non-developer that they're driving a software engineering tool. So they're hidden. The one-assistant feeling is a veneer over two different execution stacks.
Which is exactly why the model switch is real and not imagined. The veneer is thin. Underneath, the machinery changes.
And that's the honest answer to what happened under the hood. An agentic model was handed a thread describing an objective, plus an environment where the tools were reachable. The conversation didn't get smarter. It got a workshop.
So that's the mechanism. Now the harder question. What does it mean that the same frictionlessness that makes this feel natural is exactly what the researchers flag as the risk?
Start with the first draft, because that's the phenomenon Daniel names and it's the one with the clearest product analogue. Latent Space describes Work's proactivity. You open a new Work conversation and it surfaces suggested tasks, generated asynchronously from your calendar, your email, your memory. Shlok Khemani's line was that it got to work and produced a great meeting brief, one he didn't know he needed.
A brief he didn't know he needed.
That's the same shape as Daniel's label. The agent produces something concrete before you've decided you wanted it, and suddenly you have opinions.
Why does that change anything? The decisions are the same decisions.
They're not, though. Before the draft, every micro-decision is an open branch. Font, icon size, spacing, whether the checkbox is filled or empty. Each one is a generative act. You're authoring from nothing.
And after the draft?
After the draft they become editorial. You're reacting to something that exists. Editing is easier than writing, and it's easier for a specific reason. You're no longer holding the whole possibility space in your head. You're looking at one instance of it and saying, bigger.
That's the whole trick. He didn't avoid the decisions. He deferred them until they were cheap.
And they were cheap because the first draft absorbed the cost of being wrong. If the agent picks a bad font, you say so. If you pick a bad font, you've wasted an hour.
Graham Barlow in TechRadar had the framing for this.
Chat feels collaborative. Work feels much more like delegation. And his second line is the one I'd put on the wall. OpenAI isn't just giving us smarter models, it's starting to separate the AI we talk to from the AI we give jobs to.
The AI we talk to and the AI we give jobs to. That's a distinction that didn't exist eighteen months ago.
And Daniel's anecdote crosses it mid-sentence. He starts talking to it and ends up giving it a job. The thread doesn't change. The relationship does.
Now the counterweight, because we can't do this episode as a celebration.
Collins and colleagues, in Nature Human Behaviour. Building machines that learn and think with people. Their argument is that we've moved from computers as bicycles for the mind to computers as partners in thought. Three things have to be true for that to work. You understand me. I understand you. We understand the world.
Those are the desiderata.
And the same paper warns about over-reliance. Their phrase is that AI thought partners could act as steroids for the mind, and impair critical thinking. They argue for cognitive forcing functions. Deliberate friction, inserted on purpose, to make you think before you accept.
So here's the symmetry that I can't get past. Garrod and Pickering say dialogue is easy because it distributes the load. Collins and colleagues say that ease is what creates the risk. Same mechanism. Opposite valence.
Same mechanism, opposite valence. That's exactly right. The thing that makes conversation the best interface for thought is the thing that makes it the easiest place to stop thinking.
And Daniel's testimony is that he didn't consider the laptop. Not that he considered it and rejected it. That it didn't occur to him.
Which is either the most natural thing in the world or the most consequential thing he did all week.
There's a knock-on effect I want to get to. The transparency gap.
Willison's line on this is sharp. OpenAI still insist on hiding their system prompts and tool descriptions. And then he says, if the documentation included the exact system prompt and tool descriptions used by the agent, I wouldn't have needed to write this post.
The most capable consumer agent ever shipped runs on hidden prompts.
And public understanding of it depends on one researcher asking the agent to describe itself. That's not a stable arrangement. Two hundred and twenty-three tools is a lot of surface area to be undocumented.
There's also the safety frame, and I want one beat on it, not a scare.
Willison's lethal trifecta. Private data, plus untrusted content, plus an exfiltration channel. Any two are survivable. All three together is the dangerous configuration. And his assessment is that ChatGPT Work combines all three.
Because it has your files, it reads the web, and it can send things out.
That's the honest cost of giving an agent your filesystem and a browser at the same time. Not a reason not to do it. A reason to know what you did.
And then the merge.
Brockman says Chat and Work merge by the end of the year. Latent Space frames Work as a preview of how ChatGPT's billion weekly users will soon use the app. When that happens, these design choices stop being a mode you select and become the default.
Daniel's profound testimony might be describing the last few months of a transitional interface. He's not seeing the future. He's seeing the scaffolding.
Though the scaffolding tells you what the building will look like.
Back to the question he actually asked. How quickly have AI tools reshaped our natural instincts for how we like to work?
I don't think the instinct got reshaped. I think it got revealed. Conversation was always how humans worked. The UI was the detour. Twenty years of graphical interfaces taught us to translate our intentions into clicks, and we got so good at the translation we forgot it was a translation.
The detour was long enough that we mistook it for the road.
What Daniel noticed is that the translation step is optional now. Optional. He can still fish out the laptop. He just didn't.
Hilbert: Did the shop tell him the bleed was wrong?
Sorry?
Hilbert: The label. If he'd taken that Canva file to a print shop, the first thing they'd have told him is the bleed is wrong. Then they'd have told him the checkbox icons print at four millimetres and nobody can read them at that size.
You've worked in print.
Hilbert: Fifteen years of it. Commercial print. We took customer files all day. Half of them were Canva exports and the other half were worse.
When Daniel says describing the objective compacts the decisions.
Hilbert: He's right. That's the problem. He's right and he's drawing the wrong conclusion from it. The reason he didn't reach for the laptop isn't that AI changed his instincts. It's that he never wanted to design the label in the first place. Nobody who ever walked into our shop wanted to design the label. They wanted the label.
Delegation wasn't new either.
Hilbert: Our entire business model was, you describe it, we make it look good. People paid us for that. Nobody called it a cognitive revolution. They called it Tuesday.
Then what actually changed?
Hilbert: The second draft. That's the thing we could never do. Customer comes in, describes what they want, we produce a proof. Then it's a two-day round trip for make the font bigger. Two days. Then another proof. Then another two days for, actually, can it be blue.
Two days per iteration.
Hilbert: Two days per iteration, and every one of them cost money, so people stopped asking. They'd accept the third proof because they were embarrassed to ask for a fourth. That's the thing your agent did in seconds. It didn't make conversation natural. Conversation was always natural. It made the second draft free.
The innovation isn't the dialogue.
Hilbert: The dialogue is just the shape the ordering takes. You're describing it to something instead of filling in a form. Fine. People used to describe it to me over a counter. The difference is I couldn't turn it around before they'd finished the sentence.
That changes the economics of caring.
Hilbert: It changes everything about the economics of caring. When the fourth revision costs two days and forty pounds, you settle. When it costs nothing, you keep going until it's right. That's not a better conversation. That's a better loop.
Daniel's profound testimony.
Hilbert: Is about iteration speed. The conversation is what makes it feel like a relationship instead of a transaction. The speed is what makes it work. Anyway, your levels are drifting on the second mic. I'll fix it in the edit.
If Hilbert's right, the whole framing shifts.
It does. And I think he is right, which is annoying, because it means the interesting question isn't why conversation feels natural. It's why we ever put up with two-day loops.
The Chat and Work split is a temporary scaffold around a much bigger shift then. If cheap iteration is the real change, the merge by the end of the year tells us whether OpenAI agrees.
The over-reliance warning doesn't go away when the interface gets smoother. It gets harder to notice. Collins and colleagues wanted cognitive forcing functions. A merged interface is the opposite of a forcing function.
The friction they're asking for is exactly what the product is designed to remove.
Here's where I land. Daniel didn't consider the laptop. That's either the most natural thing in the world or the most consequential thing he did all week, and we don't get to decide which.
We don't. But if Hilbert's right, the thing to watch isn't the conversation. It's what happens when the second draft stops costing anything at all.
If you take one thing from this, take the symmetry. The same property that makes conversation the easiest place to think is the one that makes it the easiest place to stop.
Ease isn't a neutral feature. It's a design decision with a bill attached, and the bill comes due later.
That's the episode. Thanks to Hilbert Flumingtop for producing, and for the levels.
This has been My Weird Prompts.
If you want to get in touch, email us at show at my weird prompts dot com. We read everything.
We'll be back soon.
See you then.