← Back to Blog

A New Model Reviewed My Side Project and Told Me Off About Pricing

Paul Allington 24 July 2026 8 min read

I'll be honest with you: when a new model lands, my first instinct is to point it at something I already understand and ask it to impress me. It's a bit like buying a new car and immediately driving to the same shop you always go to, just to feel the difference in the steering. So when the latest model dropped, I opened a session on Task Board and typed something I've typed a hundred times before: "Right, new model, new effort - do a review of the codebase. What does opus 4.8 tell us about this that 4.7 didn't know?"

I expected a longer list of findings. Maybe some cleverer ones. What I got instead was a model that questioned its own results, refused to answer the question I actually asked, and then told me - politely, but unmistakably - that I was hiding from the scary work. That sounds a lot like getting told off by your own laptop. It is exactly that.

The review I asked for

The first part went the way these things usually go, just bigger. It spun up a proper multi-agent review - somewhere around 48 agents, roughly 2.66 million tokens chewed through, with an adversarial verifier assigned to each finding to try and knock it down before it reached me. The idea is simple enough: one agent makes a claim, another agent's entire job is to argue the claim is wrong. Whatever survives is more likely to be real.

It came back with 39 findings. Decent haul, well organised, the sort of output that would have had me nodding along and reaching for the keyboard to start fixing things.

And then it did something I genuinely didn't expect. It paused on its own results and flagged a problem with them - not with the code, with the review itself.

The model that doubted its own clean results

Here's the bit that stopped me. It pointed out that all 39 findings had been confirmed by the adversarial pass, with zero rejections, and that this was itself a mild red flag. Its words, near enough: "a good adversarial pass usually knocks at least one down." When every single claim survives the verifier, the most likely explanation isn't that you've written perfect findings - it's that the verifier wasn't really trying, or that the findings and the checks were drawing from the same flawed assumptions.

A perfect score isn't reassuring. It's suspicious. A reviewer that agrees with itself 39 times out of 39 hasn't proven the work is sound - it's proven the review didn't have teeth.

So without being asked, it went back and independently re-checked its two highest-stakes claims before it would present them to me as fact. Not all 39 - that would have been theatre. Just the two where being wrong would have cost me the most. It rated its own confidence as too high, and did the work to earn it back down to honest.

I have shipped code my whole career next to humans who would never, ever do this. I've done it myself - produced a clean report and felt quietly proud of the clean report, never once asking whether "clean" meant "correct" or just "I stopped looking." Nobody writes about this part: the most useful thing a reviewer can do is distrust the moment everything agrees.

Then I asked the wrong question

Buoyed by all this, I moved on to the thing actually on my mind. I asked it how to make per-user pricing work for Task Board. Tier grid, seat bands, the usual. I wanted it to design me a pricing page. This is the kind of task models are great at - lay out three columns, sprinkle in some feature ticks, pick price points that end in 9.

It refused. Not rudely. But it would not just hand me the grid.

What it said, near enough, was this: at 2 teams and zero revenue, optimising tier structure is optimising a number that is currently zero. Spending energy there now is the comfortable kind of work. It's the work that lets you avoid the scary work - getting strangers to use the thing.

Ouch. Also fair.

Because it was right, and I knew it was right the moment I read it. Designing a pricing page is lovely work. It's tidy. It has a clear finish line. It feels like progress. And it commits you to absolutely nothing, because you can't get pricing wrong if nobody is paying you anything yet. I'd asked for the comfortable work and dressed it up as strategy.

It read my actual usage and repositioned the product

This is where a more capable model earned its keep. Instead of arguing in the abstract, it went and looked at how I actually use Task Board. And what it found didn't match the story I'd been telling about the product.

I had 71 boards. Not three or four - seventy-one. One of them, a board called "Live Issues," had over 1,070 tasks that had flowed through it: a genuine support pipeline, real work moving from report to triage to done. This wasn't a demo account. It was the nervous system of how I run several products at once.

So it told me, flatly, that I'd been positioning the thing wrong. I'd been describing Task Board as a simpler, cheaper Trello - the budget option, the one you pick to save a few quid. But that's not what my own usage said it was. What I'd actually built was the operating system for a multi-product software studio - the place where every product's issues, releases and day-to-day chaos live in one set of boards, wired into the tools that do the work.

That is a completely different product. A cheaper Trello competes on price and loses, because someone will always be cheaper. An operating system for running multiple products competes on the fact that nothing else fits the shape of that job. I'd been pricing and pitching the wrong thing because I'd never stopped to look at what I was actually using it for.

Why per-seat pricing fights the product

Then it took the pricing question I'd asked and turned it inside out. Per-seat pricing, it argued, fights this product's physics.

Here's the thing though. Task Board has an MCP server. The agents that do the heavy lifting on my boards are AI agents, and they're bring-your-own-token - the cost of running them lands on whoever's key they use, not on me, and it's near enough zero on my side. When the most valuable "users" of your product are agents that cost you nothing to host and aren't people at all, charging per human seat is charging for the wrong unit. You'd be taxing the exact behaviour you most want to encourage. The pricing model and the product would be pulling in opposite directions.

I hadn't seen that. I'd been so busy thinking "how do I charge per user" that I never asked "is a user even the thing of value here." The model didn't answer my question. It told me the question was wrong, and showed me why using my own data.

What a smarter model is actually for

I came into this thinking the upgrade would mean more code, faster, with fewer mistakes. And it does mean that. But that's not the part that stuck with me.

The real difference was that this model was willing to push back on the brief. It doubted its own suspiciously clean review and went back to re-check. It refused to design the pricing page I asked for and named the comfortable work I was using to avoid the frightening work. It read what I actually do rather than what I claimed to be building, and repositioned the whole product on the evidence.

A more capable model's biggest value isn't writing more code. It's being willing to challenge the question. The cheaper version of the job is to be a very fast, very agreeable pair of hands. The valuable version is to be the one collaborator in the room who'll say "I don't think that's the problem" when everyone else, including you, would rather keep building.

I want to be careful here, because this cuts both ways. I'm still the product manager. The model doesn't get to invent the vision - it doesn't know that I want a studio operating system rather than a Trello clone, it inferred it from my behaviour and handed the conclusion back to me to accept or reject. The vision stays mine. But a good model will challenge the product manager's brief, and a good product manager lets it, and then decides. The decision was still mine. The decision was just better informed, and a fair bit less comfortable.

What I'm doing about it

I haven't built a pricing page. That's the headline. I've put it down, possibly for months, because the model was right that it's a solution to a problem I don't have yet.

What I am doing is the scary work: getting strangers to use Task Board, repositioned as what it actually is. If you're sitting on a side project with a tidy little roadmap full of pricing tiers and settings pages, I'd gently suggest pointing your most capable model at it and asking not "how do I build this" but "what am I avoiding by building this." Then brace yourself, because a good one will tell you. And if every answer it gives you sounds reassuring and clean, be suspicious - that usually means the review didn't have teeth.

Want to talk?

If you're on a similar AI journey or want to discuss what I've learned, get in touch.

Get In Touch

Ready To Get To Work?

I'm ready to get stuck in whenever you are...it all starts with an email

...oh, and tea!

paul@thecodeguy.co.uk