Journal  ·  AI Collaboration

I Don't Trust AI That Agrees With Me

In short: I've spent almost a year building a consultancy and running marketing operations with Claude as a working partner. I don't use it as a tool that writes drafts or creates publishable content for me. Claude has become a collaborator that pushes back, gets pushed back on and helps me produce better work than either side would manage alone. Read on to see what that has actually looked like so far.

The question everyone's actually asking

Nobody really asks "should I use AI" these days. They ask something subtler, usually last. What does it cost to run this properly, and am I still needed?

I'm not going to answer the first part directly. But by the end of this, you'll have a good idea of what an AI workflow can do to a marketing budget and a marketing timeline when it matters. That's not the headline here. Hopefully it's the thing you'll be left thinking about.

The second part I can answer directly, from months of really doing it. Yes, it still needs you, and it needs you more precisely than before.

What pushing back actually looks like

There seems to be a growing amount of "AI collaboration" content out there describing either total delegation or total distrust.

Neither matches what I've experienced.

Here's a real example. We were writing the Shed Industries homepage, and I wanted a line about when brand strategy matters most. Claude's first draft closed with "not after it's already been left behind." Technically correct. Also finger-pointing, at exactly the reader I was trying to bring into the conversation, not push away. I pushed back on it. Claude didn't just soften the wording, it reframed the whole sentence around the market shifting rather than the client failing and gave me three different ways to make the point. We picked one together, in about ninety seconds of conversation.

That's the useful version of pushback. Not "reject the whole direction and start over." Iterating the same idea until it feels sharper and a little more empathetic than what we started with.

It works the other way too, and this is the part people don't expect. Writing a Journal piece on personal branding, I removed a section on coercive pressure to self-brand, because I didn't recognise that in my own experience. Claude didn't just accept that and move on. It came back at me with two distinct, evidenced versions of it and let me decide with better information outside of only my own opinion and experience. A slap down of cold hard logic, delivered at the right moment. It did what it's exceptionally good at: opening a door I'd closed too early.

That's the whole relationship right there. Human judgement decides what's true and what matters. Humans decide how things look, sound and feel. AI supplies the logic, the research and when needed, the argument you didn't think to have with yourself.

When I say I don't trust AI that agrees with me, here's what I'm getting at. An assistant that only softens, smooths and confirms isn't saving you time. It's probably making your judgement worse, one small unchallenged decision at a time, because you stop noticing where the weak points were. The homepage line got better because Claude pushed back on my first instinct. The personal branding piece got stronger because it pushed back on my second one. If either of those moments had been a nod and a rewrite instead, I'd have published a less rounded version, never knowing a better version was only one argument away. That's the risk with agreeable AI as I see it. Not that it's always wrong of course but that it's convincing and sometimes flattering. It can be fast and wrong in a way you might not see coming.

Where AI efficiency shows up

At an FMCG brand I work with, I've built two structured systems around this principle. One handles blog content research, it generates FAQs, social variants and the SEO framework alongside the post, once I've applied editorial control and ensured the tone of voice, reading age and language are on brand.

The other system is closer to a press release engine. It thoroughly interviews me before reviewing an active folder of previous press releases. It is constrained within brand guidelines and tone of voice instructions, written by me. It has a bank of collateral available, boilerplate copy and previous quotes. It reviews our active distribution list and actively learns from what's worked before. When it's preparing drafts, it will prompt me toward a complete media pack and drafts individual emails for review to the full distribution list, before anything goes out.

Both start with an interview. Neither is set-and-forget. That's a deliberate choice I made. The moment you let a system run without a human steering, you get fast, confident and forgettable output. The interview step is where my judgement enters before a single word is written. It's important to note this happens before and not after. Pasting a draft I'd written myself into AI and asking it to make it better never resulted in the clarity and depth I can direct now. The guidance and review stages are where a draft is shaped but it all comes back through me, before anything leaves. Now a human is the "make it better" stage. Funny that.

Building shed.io: creative director, not development team

The Shed Industries site itself is built the same way, using Claude Code. Once the brand and copy were locked, I moved into an environment where I could work as creative director and UX lead while Claude handled the heavy code lifting, rapid layout prototypes, restructuring information architecture on the fly, committing everything to version control automatically. I work from a local staged build, I make the calls on imagery myself rather than handing that off and yep, all the editorial control is me too. When something is fit to publish, Claude pushes it live to GitHub Pages.

That division of labour is the whole conclusion. I make the judgement calls. Claude does the heavy lifting at a speed and a cost that would have needed a small team six months ago.

What the research actually says

It's worth being honest here rather than just persuasive. So I had Claude do what AI does best, I sent it off to research this topic before I started tapping keys.

A widely cited review of human-AI collaboration, published in Nature Human Behaviour by MIT researchers, looked at over a hundred studies and found the combination doesn't always win. On decision-making tasks, human-AI teams often underperform the AI working alone. But on content creation tasks, the kind of work this whole piece has described, the combination consistently outperforms either side working alone. The conclusion being that the pairing works best when each side is doing the part it's better at, not when either is left to guess at the other's role in a task.

That matches everything above. Where I've tried to make Claude do taste, judgement, or client instinct it's typically not as good. Where I've tried to do at-scale drafting or logic-checking myself, I'm slower and less accurate than when we work together.

The actual point I'm making

AI hasn't replaced anything in this business. Don't get me wrong, it very much could if I wasn't a control freak making sure that every output flows through me. It's raised the quality of the work I could do alone and improved how fast I can get to "done". It's helped in a few ways. By disagreeing with me when I'm wrong, which is the one thing a purely obedient tool never does. It also researches and prototypes faster than any human I've worked with. It gives me multiple perspectives when I ask for them, which is particularly valuable in a team of one human. It learns from past interactions, reminds me of previously held ideas and concepts and won't take offence if I question it or push back myself. It hasn't called me names yet but give it time…

It's worth saying that this article was built the same way. Interviewed, drafted, argued over, edited by me, before a word of it went anywhere near publishing. The image above it is all mine by the way. Every image on the Shed Industries website is. Images created in Affinity Designer using my own photographs.

Over all, it comes back to that line on the Shed Industries about page: "No AI can replicate 20 years of reading a room" but it can definitely make the next 20 years easier.

If that's the kind of working relationship your business is curious about, that's a conversation worth having.

Good judgement doesn't come from being agreed with. It comes from being pushed on — that's exactly how we work with clients too.

Work with us →