First, a quick note: next week I'll be out in San Francisco attending OpenAI DevDay. If you're in SF next week, hit me up and let's get coffee.
Actual conversation between the author and their Muse AI agent, "Marley"
We're now many months into AI-pilled people and companies talking up how amazing the future will be when we all have AI agents helping us be more "productive" and "get more done."
A lot of Big Tech's bets on adding AI to everything — from Copilot in Windows and Office, to Gemini in all of Google's apps — is premised on the idea that people want a robot to help accelerate their computer jobs, and therefore will flock to chatbots the way they did iPhones.
So far, that hasn't happened. Apart from ChatGPT, no AI products have seen meaningful consumer adoption.
Even Grok Bot, one of the most user-friendly agent thingies released to date, is mainly popular among tech insiders and appealing to the sort of people who know what a "skill" is.
This is the world into which Meta shipped Muse a couple of weeks ago. Because it's from Meta, and comes with extremely cute fuzzy avatars, the going assumption is that this one might finally break through and turn agents into a huge consumer business.
Meta, wisely, is marketing Muse as less an amazing engine of productivity and more as a cute Tamagotchi-like friend that'll do annoying stuff while you live your best life. They're using a lot of lovely, aspirational photography — people cooking, running, hanging in the pool — as opposed to… however Microsoft and Google expect us to live with their crap.
What if, the Muse website posits, you could be doing this without that bit of anxiety about whether you forgot to cancel a bunch of subscriptions and we're burning cash for no reason?
Here's the thing, though:
When I was still a Big Tech PM, working on SaaS products, I used to say our biggest competition wasn't another product. Most software's biggest competitor is nothing.
This clear in some initial takes from folks like Ryan Broderick, who wrote:
Make no mistake, I am a busy person lol. I’m the CEO of a small business and that business primarily runs on email. I have two inboxes that both regularly receive around 100 emails a day. I live in a big city and go out to eat a lot. And I like to think I have a pretty robust social life. But I don’t mind the minor inconveniences of 21st century life and, in some cases, actually enjoy them. I like grocery shopping and picking out restaurants and, if you can believe it, booking my own travel.
Marketers will often classify products as vitamins or painkillers. When talking about apps used by individual people (whether at home or at work), I'll sometimes also separate them by whether the problem they solve is a mountain or a wall. A mountain can be a wall if it's in the way of what you really want to do, like getting to Grandma's house quickly. But it can also be a worthwhile challenge, and going to Grandma's house via a beautiful mountain pass can be a feature, not a bug.
One of the biggest persistent problems in software is that we tend to treat every problem as a wall, and build products to help people get over, around, or through it. Travel booking seems like a wall, especially if you only do it inside a big company.
For me, all the decisions involved in planning travel are multifaceted — it's not just finding the cheapest flight, but whether the marginal cost of going at 9 AM is worth not having to get up for a 7 AM flight, or if the layover in Seattle is long enough to get a free lunch at the Amex lounge. Agents can automate the doing here — the clicking and searching and waiting — but a lot of the thinking happens while I'm clicking, and that can't be handed off to any product.
What's more: nearly every software tool is too heavy for most people and their needs. At Stripe, as we interviewed users to gather data for what became Stripe Apps, we expected that folks would say they'd use software to keep their Stripe payments better connected with other tools. What we heard was that, e.g., vendors with 8 clients can keep most of their business context in their heads; when they can't, they reach for a text file or Google Sheet.
My take on agents, including Muse, is that they aren't lighter than regular software — they just shift the heaviness from somewhere you can see it (form fields, tables, grids of information) to the ether, asking you to trust the AI to accurately and consistently manage stuff that's complex enough to make humans tired and anxious.
Software people tend to believe, philosophically, that computers are better and more reliable than people, including/especially themselves.
Normal people put more trust in what they can see. Even if they see the benefits, they're more acutely aware of the risks. What if it screws up? What if it sends a text by mistake and I lose a client/friend/partner? How do I know the numbers are right?
Hypothetically, agents should be great for people like me with ADHD — robots with all the executive function and accountability God failed to load into our brains.
Here's a data point: I've been putting off filing a warranty claim for our expensive Ratio coffee maker for months, so this morning I decided to ask Muse if it could just, y'know, do that. I gave it photos of the problem and the serial number label, and access to my email so it could find my original order number and purchase date.
Here's what happened:
It could load the warranty form in its web browser, but the form was in an iframe and the agent's computer-use tool doesn't work inside iframes
I could "take over" the browser session and fill in the form myself, but you can't copy and paste information into the agent's browser from your Mac, so I had to re-type everything by hand.
The agent's browser can't actually upload files like photos — which it would've been nice for Muse to tell me upfront
So I had to redownload the photos from Muse, open the form in my own browser, and copy the information from the chat session (rewriting the message to sound more human), etc.
The main benefit of having tried this is that it turned this task from a future chore to something an agent had screwed up that I needed to fix now, and it did handle digging through my email to get the numbers and details it was annoying to track down.
I don't think most people would start using a new app from Meta, or any Big Tech company, because it'll promise to do 110% of your annoying tasks, but actually only do 20%, even if 20% is better than nothing.
Anyway, if you're not yet "AI-pilled" and wondering if you need to care about "agents", my answer is: ehhhhhhh, probably not. Yet.
What I'm reading
⌨️ Areal — Dinamo Typefaces collab'ed with the design portfolio site Are.na on a revised and updated take on, of all fonts, Arial. In addition to the not-quite-Helvetica we all know and accept, Areal has monospace and semi-mono variants, all of which are variable fonts with a special "dark mode" axis for better rendering in, um, dark mode. (In all seriousness, Areal Semi Mono looks pretty hot.)
🤖 "AI, make the website good" — Some meditations from Eleventy Build Awesome creator Zach Leatherman on whether AI is actually making better web pages, including some, ahem, interesting interactive speed tests of the major AI companies' home pages.
📈 Create your own personal AI benchmark — Every's "head of evals" Mike Taylor on the practice and process of how you can measure "good" results from whatever AI tools you use. I personally think it's worthwhile to think about at least "back-pocket evals" (as Steve Yegge calls them) if only to keep in mind that AI tools are still deeply flawed, but still occasionally useful and worth paying some attention to.