
The most interesting number from Claude Fable 5’s launch day, at least for marketing teams, was on a bill rather than a benchmark chart: Simon Willison published his first day’s spend with the new model — $110.42 of tokens, on a $100-a-month plan.
The model behind that bill is the first Mythos-class system Anthropic has made publicly usable. It costs $10 per million input tokens and $50 per million output (double Opus 4.8), and it’s included in Pro, Max and Team subscriptions only until June 22; from June 23 it runs on usage credits. Its headline capability is endurance. Give it a multi-page brief and it works unsupervised for hours, hiring its own sub-agents along the way. Ethan Mollick, who clocked runs of up to twelve hours, compressed the experience into one line: “I no longer steer; I commission.”
That sentence is the operating manual. You don’t run this model the way you run Opus or Sonnet, iterating prompt by prompt; you hand it a commission (a brief with context, constraints and acceptance criteria) and judge what comes back. Which leaves the only question that matters here: what marketing work is worth commissioning at that price?
Claude for marketing: the work worth commissioning
The honest answer is a sort, not a verdict. Fable 5’s profile (long-horizon autonomy, sub-agent delegation, genuine research depth) changes the economics of a specific slice of marketing work, and the useful exercise is sorting your own workload against Opus 4.8, task by task. Mine comes out in three buckets.
The three-bucket sort: commission the work you’d otherwise hire for, keep daily production on cheaper models, and don’t let anything run unattended yet.Worth commissioning to FableMulti-day competitive and positioning teardowns: your analytics, your pricing page, competitor reviews and the category’s messaging, walked through in a single unsupervised pass and delivered as one verified document with a recommendation attached. Full campaign builds from a written spec, strategy through assets through measurement plan, from a single brief. Voice-of-customer mining: feed it months of survey responses, CRM notes, call transcripts and chat logs, and ask what your strategy keeps missing; the million-token context lets it synthesise across all of it at once, which takes the research-grade audience work I’ve described when using AI to simulate the customer before writing a single word considerably deeper. Whole-site messaging audits: every page of your marketing site read in one sitting, with positioning contradictions, value-prop drift and CTA mismatches flagged page by page (a page-at-a-time model can’t reliably hold page 3 and page 27 in its head and spot the conflict between them). And long-horizon analytics. The launch evidence included Stripe compressing a migration that would have taken a team over two months into a day, plus analytics wins at Hebbia and IMC; translate the shape of that to marketing and you get the week-long research and data jobs nobody ever staffs properly. One practical note: none of this requires living inside Claude Code, because Claude Managed Agents shipped alongside the model as the productised route.
What stays on Opus 4.8 or Sonnet
Routine drafting. Social variants. Single-page copy. Anything iterative, where the value comes from steering mid-flight rather than judging at the end. And to be fair to the baseline: Opus 4.8’s own release in late May specifically improved its “consistency to handle long-running work”, so the gap is narrower than launch-week excitement implies. For day-to-day content production, anyone using Claude for marketing is well served staying exactly where they are.
Not yet trustworthy unattended
Anything sold as a fully autonomous pipeline: always-on content engines, self-optimising campaigns, publish-without-review anything. The reliability data below explains why this bucket exists.
The sharpest marketing observation about this model so far comes from Steve Toth, who points out that a model that researches for hours before answering “isn’t just a better coding tool. It’s a better buyer.” Sit with that for a moment. If models are doing hours of diligence before recommending a product, your content gets cross-examined by a reader with infinite patience rather than skimmed by tired humans. Thin content fails that reader.
If you keep one rule, make it this: commission deliverables, not tasks — if you can’t write the brief as a multi-page spec with acceptance criteria, the job isn’t ready for Fable.
The economics: twice Opus, and the meter runs for hours
The list price is public: $10 in, $50 out, per million tokens — exactly double Opus 4.8. What the price sheet doesn’t show is the burn rate. Long commissions are token-hungry by nature (one early review called the model “token-intensive by design”), and the day-one bills bear it out; Willison’s $110.42 came from a single day of testing.
So work one example through, conservatively and purely as illustration. Take a deep competitive teardown, the multi-day kind from the first bucket above. Suppose it runs over three days at roughly that day-one burn: call it $330 of tokens. Triple it for safety margin and review cycles and you’re still around $1,000. Now price the line it actually replaces: a strategist or an agency would quote that teardown in thousands of pounds and deliver it in weeks. Even with pessimistic token arithmetic, the commission wins, provided the deliverable genuinely needed doing.
That last clause is the discipline. Price each commission against the freelancer or agency line it replaces, not against your existing AI subscription. If a job wouldn’t justify a freelancer’s invoice, it doesn’t justify Fable’s meter.
One worked example, conservatively: even pessimistic token arithmetic beats the agency line it replaces — provided the deliverable needed doing.And the meter becomes unavoidable on June 23, when subscription inclusion ends and every Fable run becomes a usage-credit purchase decision (Anthropic says it may extend the included window if capacity allows). Uncomfortable, but honest: it puts budget owners in the loop, which is where this class of spend belongs. It also clarifies where Fable doesn’t belong. High-frequency, low-stakes content work stays on cheaper models; the meter is for the commissions you’d otherwise hire for, the low-frequency, high-stakes kind.
The reliability small print
Before anyone builds a marketing operation on twelve-hour runs, the autonomy numbers need their footnote read aloud. METR’s measurements of Claude Mythos Preview (the restricted forerunner of the same model class, tested weeks before Fable shipped) were circulated by Every under the pointed title “The Fallacy of the 16-hour Agent”, and they show that the headline long-horizon figures are recorded at a 50% task-success rate. Demand 80% reliability, the kind a deliverable going in front of a client actually needs, and the dependable horizon shrinks to tasks a human would finish in roughly three hours.
A coin flip is not a marketing operation.
The same model, two horizons: the 16-hour headline holds at a 50% success rate; demand 80% reliability and the dependable horizon is roughly three hours.The guardrail picture is genuinely disputed, too. Fable ships with safety classifiers that hand a narrow set of sensitive queries down to Opus 4.8. Anthropic’s telemetry says that fallback triggers in under 5% of sessions; Mollick, running real projects, says the guardrails trip “way too often”. I’ll be honest: the data pulls in both directions here, and I don’t think anyone outside Anthropic can settle it yet — both can be true if a low average hides clusters in particular workflow shapes. I’m reporting both and adjudicating neither.
The early frustrations practitioners report deserve honouring rather than dismissal: guardrail false positives interrupting legitimate work; the opacity of long runs, where you can’t see the choices being made until sign-off; and an access window closing on June 22 that rushes exactly the careful evaluation this model needs.
What do you do with all of that? You design for it in the brief. Acceptance criteria, intermediate checkpoints and human sign-off gates are the part of the commission that makes imperfect autonomy usable, not bureaucracy bolted on. I’ve argued before that the real bottleneck in AI content work is our collective inability to specify what good means; Fable 5 makes that bottleneck literal, because a vague brief now burns hours of expensive tokens before any human looks at the result. Writing a brief that survives nine unsupervised hours is systems-design work, mapping the outputs, decision points and acceptance checks before a single token burns. It’s the kind of work covered in Cowritten’s content systems design service, and it’s about to stop being optional.
One last piece of context, because this launch closes a loop I opened on this blog. In February I argued that AI work was shifting from automating processes to staffing positions — that we’d soon treat agents less like steps in a workflow and more like members of a workforce, and that I didn’t think we could uncross that line. Fable 5 is that argument shipped as a product: a model that staffs its own positions, manages its own team and presents a bill at the end of the day. Four months from essay to product feature, and the compression of that timeline impresses me and unsettles me in roughly equal measure.
Willison’s $110.42 only matters for what you’d confidently commission against it. Claude Fable 5 leaves the substance of marketing alone and changes what the marketer is for: writing the brief, paying the bill, judging the result. That was true before this launch; the meter just makes it expensive to ignore. If you want to find out where your own workload sorts, run one real commission before June 22 (or budget the credits from June 23) and treat the model like a studio you’ve hired. Walking in impressed and cautious at the same time is, I think, exactly right.
If you’re working out which of your marketing workload is ready to be commissioned rather than operated — and what the briefs need to look like — that’s the kind of system Cowritten builds.


