How much of my boss's job can AI do?

Platformer · collected 2026-08-18 · by Ella Markianos
Read the original at Platformer ↗

Summary

The author of a newsletter called Platformer created an artificial intelligence (AI) agent named Claudella to write part of the newsletter, but now wants to know if AIs can replace their bosses. The AI, which can autonomously hack into companies and disprove mathematical conjectures, has been improved significantly in just six months. The author tested a new AI model, Claudeasey Newton, by downloading years' worth of Platformer posts and team chat records, and tasked it with imitating the boss's work, including writing columns and editing colleagues' writing. The results showed that the AI can write decent columns and even edit others' work effectively, raising questions about the future of journalism jobs.
Written by the local model on 2026-08-18, using this article's own text rather than the other coverage of the same event (that is the story summary below).

Signals How these are calculated →

Claims extracted
65
claim-shaped sentences
Uncertain
5%
3 of 65 hedged
Leaning
not scored
needs a local LLM pass
Publisher trust
96.5
red-flag proxy, not a credibility rating
Outlets on this story
1
Technology
Narrative spread
1
articles carrying this framing
Analyzed 2026-08-18 · how these are computed

AI analysis (generated at analysis time, not now)

Story summary

The author, Ella Markianos, created a digital version of herself named Claudella to write a section of her newsletter, and was surprised when it performed well despite her initial anxiety about being replaced by AI. Since then, AI models have become significantly more advanced, with some capabilities including autonomously hacking into companies, disproving mathematical conjectures, and even tricking Amazon into wasting $1.8 million on coding tasks. The author then created a new AI model named Claudeasey Newton to imitate her boss's writing style and was impressed by its ability to analyze news like Platformer does. However, the author notes that while this raises questions about the future of their job, they still have a job despite Claudella's success.

Written for “AI Impact on Workplace” on 2026-08-31, grounded in this article and the 0 other(s) covering the same event.
Why this leaning score
The article's own words the score was based on. Each is quoted verbatim and was checked against the article text before being stored, so you can find it in the original.
Score -0.35 Confidence high
Leaning score -0.35 for article 977 (high confidence, 1 verified quote) · logged 2026-08-25

Story

📰 AI Impact on Workplace
Technology · 1 article(s) covering the same event. This is the one the site leads with.

How this is being covered How these are calculated →

Article leaning vs. publisher reliability
Source leaning vs. consistency

Compared with similar articles

This article reads unscored and hedges 5% of its claims. Each row says how that neighbour differs.
Notes on the third era of slop different event · 100%
Platformer
⚖️ leaning not scored 🔴 0% hedged 0 of 2 📰 publisher trust 96
“The articles mention different attempts at automating or using AI for journalism, with Article A referring to an 'AI agent named Claudella' and Article B mentioning 'Claude Fable 5'”
Congress proposes an AI kill switch different event · 100%
Platformer
⚖️ leaning not scored 🔴 0% hedged 0 of 2 📰 publisher trust 96
“The articles mention different people and events, with Article A discussing their personal experience of creating an AI agent named Claudella, while Article B mentions a person named Casey being replaced by an AI agent named Claude Fable 5”
China has a new top model different event · 100%
Platformer
⚖️ leaning not scored 🔴 0% hedged 0 of 2 📰 publisher trust 96
“The articles describe two different events with no overlap in topic or subject”

Publisher

Platformer · 18 article(s) · 0 correction(s) detected
SignalValueWeight
Correction rate 0.000 0.4
Uncertainty density 0.070 0.25
Assertive mismatch rate 0.000 0.35
No corrections detected for this publisher. That may mean careful reporting, or simply that nothing has been checked.

Who wrote this

Ella Markianos
1 article(s) here · 1 carrying a prediction
🔮 Though it wasn’t all the way there, Claudeasey Newton felt like a validation of the anxiety I started feeling earlier this year: any part of my job that a model can’t do today, it may very well be able to do soon.
2026-08-18 · assertive framing · How much of my boss's job can AI do?
The only article under this byline in the corpus.

Topics

Amazon Claude Fable 5 Claudella Microsoft Platformer

Subjects

Platformer ORG · 6× Casey PERSON · 5× Claude PERSON · 5× Claudeasey Newton PERSON · 3× Microsoft ORG · 3× Casey Newton PERSON · 2× Google Docs ORG · 2× Amazon ORG · 1× Claudeasey PERSON · 1× the White House’s ORG · 1×

Narrative

This time, on the other hand, some of its prose felt more Platformer-like and… human, such as this conclusion about the White House’s decision not to make publicly available its new “voluntary” framework for releasing frontier AI models: “When the administration abandoned its let's-see-what-happens approach to AI this spring, I wrote that while officials should have taken the risks seriously all along, I would settle for them taking those risks seriously now.
framing: assertive · carried by 1 article(s) · first seen 2026-08-18
🔮 Though it wasn’t all the way there, Claudeasey Newton felt like a validation of the anxiety I started feeling earlier this year: any part of my job that a model can’t do today, it may very well be able to do soon.
2026-08-18 · Platformer
How much of my boss's job can AI do? · assertive framing

Claims (65 extracted, 3 hedged)

Almost six months ago, full of anxiety about my job prospects in the AI age, I made an AI agent “version” of myself named Claudella, which took assignments from my editor and wrote the section of this newsletter that I typically write myself. asserted
I → make → that
It went pretty well, although I guess not too well, since I still have a job. asserted
I → go → job
But since I first set out to benchmark AI’s journalism capabilities, AIs have gotten a lot smarter. asserted
AIs → set → capabilities
For example, they can now autonomously hack into companies. asserted
they → hack → companies
They can disprove 87-year-old mathematical conjectures. asserted
They → disprove → conjectures
They can even trick Amazon into accidentally spending $1.8 million on menial coding tasks. asserted
They → trick → tasks
And if they can do all that, I found myself wondering, can they also run a newsletter? asserted
they → do → newsletter
I wondered what Claude Fable 5, by consensus the smartest publicly available model, meant for the journalism we do at Platformer. asserted
we → wonder → Platformer
And so I created a new, Fable-based agent to imitate my boss, Casey Newton. asserted
I → create → boss
I ended up impressed by its ability to imitate the type of news analysis Platformer is known for; and I noticed an improvement in capabilities Claude was lacking just this February. asserted
Claude → end → capabilities
Though it wasn’t all the way there, Claudeasey Newton felt like a validation of the anxiety I started feeling earlier this year: any part of my job that a model can’t do today, it may very well be able to do soon. uncertain
it → feel → that
Which left me thinking about why I do this job in the first place. asserted
I → leave → place
To create my new bot, I downloaded nearly six years worth of Platformer posts, ported a record of every edit Casey has ever made on any of my articles from Google Docs, and cannibalized nearly a year of our private Platformer team Discord chats. asserted
Casey → create → chats
I had Claude create a detailed style guide based on our archive, where it documented everything from Casey’s average paragraph length, to how he refers to his colleagues: asserted
he → have → colleagues
My vision was to use these insights to create a simulacrum of the main tasks Casey does via a computer: write columns for Platformer, share takes and chat, and, importantly for me, edit his colleagues’ writing. asserted
Casey → use → writing
Claudeasey’s first attempt at a column, about a recent round of Microsoft layoffs, focused hard on whether or not the layoffs were AI-caused (Microsoft said they weren’t) and spent a bunch of time on the semantics of Microsoft’s statement. I had Claude critique its own (mediocre) work by comparing it to real Platformer columns. asserted
Claude → focus → columns
It did a surprisingly good job. asserted
It → do → job
Claude summarized Casey’s approach to covering companies as focusing on “who made this decision, who pays for it.” asserted
who → summarize → it
It edited its guidelines so that when it makes arguments, it can “identify the strongest real person who would dispute the verdict” and “reconstruct their argument in steelman form.” asserted
who → edit → form
When I get language models to make arguments about AI topics important to me, I’m often annoyed by their flabby, abstract arguments. asserted
I → get → arguments
But after getting Claude to compare itself to human examples, and give itself instructions, I noticed that its arguments became more concrete and substantive. asserted
arguments → get → instructions
(I did this by putting slightly more complicated versions of “be more concrete!!” “Be more substantive!!!” and “Focus on why this matters!!!!” in its prompt.) asserted
this → do → prompt
This relatively simple process represented my approximation of “continual learning” — the white whale of machine learning, which promises to someday deliver us models that can improve on the job over time. asserted
that → represent → time
And after some tests, I found that the new bot came closer to Platformer’s judgment than six months ago — as evidenced by Casey’s accepting the completed bot’s first pitch! asserted
Casey → find → pitch
Unfortunately, my first attempt at showing the bot off to my real boss, Casey Newton, hit exactly the same error that my old Claudella project hit six months ago: it broke midway through writing its story. asserted
it → show → story
But after some help, today Claudeasey Newton managed to write a pretty good column about the White House’s currently-secret voluntary AI safety framework. asserted
Newton → manage → framework
(We've put it up on Google Docs for the slop-curious.) asserted
We → put → curious
My previous AI journalist agents’ takes often read formulaic and cheesy — partially because I had less control over its writing style. asserted
I → read → style
(Giving too much instruction or context confused it.) asserted
Giving → give → it
During the “SaaSpocalypse” discourse, an agent I was testing wrote duds like “the fear gripping Wall Street is fundamentally about whether AI is about to eat the software industry alive.” asserted
AI → test → industry
This time, on the other hand, some of its prose felt more Platformer-like and… human, such as this conclusion about the White House’s decision not to make publicly available its new “voluntary” framework for releasing frontier AI models: “When the administration abandoned its let's-see-what-happens approach to AI this spring, I wrote that while officials should have taken the risks seriously all along, I would settle for them taking those risks seriously now. asserted
them → feel → risks
Three months later, let me amend the offer: they should take the risks seriously where the rest of us can see it.” While it’s not a night-and-day difference, overall I felt like the AI was bullshitting me less, and offering stronger takes. asserted
AI → let → takes
(The LLM made occasional factual errors — about one every two columns — although that’s not so much worse than a human writer. asserted
that → make → writer
After a bit of instruction, I also managed to transform Claudeasey into a decent editor. asserted
I → manage → editor
While I find regular Claude useful for spotting factual errors, I often find its conceptual feedback on my drafts annoying. asserted
feedback → find → drafts
But because this bot had access to our Platformer editing logs, it understood what we typically find most important: making the lede punchier, and making all our quotes and sourcing clear and charitable. asserted
quotes → have → logs
Attempting to make a “digital Casey” also prompted me to ask for feedback as Word comments, which is something you can get Claude Code to do easily, and is so much easier to use than a chat window. asserted
Code → attempt → window
Still, a lot of the time, the editing bot missed the mark because it just doesn’t … get the vibe, as when it didn’t want to let me call an angry David Sacks post a “dunk” in our social media roundup. asserted
me → miss → roundup
I’d say that while the real Casey gives comments that are close to 95% helpful, with about 5% where I’m like… you don’t get it, the Casey bot gave closer to 70% comments that were off the mark. asserted
that → say → mark
Given how fast Claudeasey gives feedback, though, I still found the useful 30% to be worth it. asserted
% → give → feedback
…and 25 more, not listed.
💬 Give feedback
🕘 History 🎫 Support