37 Comments
User's avatar
A Pox on Both Your Houses.'s avatar

Side X Side was only partly achieved, but the info you provided about Claude's shortcomings surprised and disappointed me, and was important to know.

You only faintly acknowledged that your comparison between Perplexity and Manus was in no way "fair." According to my query to Gemini, the Standard Perplexity plan does NOT provide access to the same models and functionality that the $200 plan does.

Your reasoning that the comparison is fair for Claude at $20 per month is accurate, I suppose.

But if you are going to give Perplexity the benefit of a Higher End subscription then you should have done the same for Manus, instead of using the Lite plan.

I say this even though my own experiences with Manus have been horrible overall, in the end. Usually, "he" starts off pretty good though, but as the tasks and context get longer, the hallucinations and outright lies increase, and make it not even worth trying if you know you have a complex task.

I have NOT tried the most recent version, though, even though I do have the $40 subscription.

That is how BAD the taint of the previous failures lingers for me. To their CREDIT, though, Manus was pretty good about issuing REFUNDS when tasks were NOT resolved properly.

Still, though, I think you should have gone the extra mile and ponied up to compare two PREMIUM versions.

I'd also prefer to see a realistic appraisal of the bottom tier Perplexity that is stuck using whatever is available AFTER the free trial credits period is over. I had just as bad experiences using bottom tier Perplexity in the beginning, as I did with Manus. So, the fact you got such great results for $200 is only to be expected. I am wondering why you would not have compared the $200 per month ChatGPT option? Or the $20 Gemini Pro that I use. But, I guess you don't feel that YOU have the budget for that any more than I do. Still...All is supposed to be FAIR in Love and War and Side X Side Comparisons. 😀

Terry Davis, JD., MBA's avatar

And while I do disagree with the notion that the test was not fair, the point made was certainly a valid one.

Daria Cupareanu's avatar

Tbh I've experienced Claude hallucinations multiple times in practice, probably more than with other models. Take that with a grain of salt though because I use Claude the most in my work, so I also catch its mistakes more often than the others. I'm guessing that's the shortcoming you're referring to?

On Perplexity, I don't have the Max plan. I have Pro ($20/month) and used their one-time 4,000 credit bonus to run this experiment. Most of those credits went to the two tasks in this article + two more tests that didn't make it into the article because it was already long enough.

You're right about Manus. Open to suggestions on how you would've framed the callout better.

Terry Davis, JD., MBA's avatar

I disagree the test was unfair because you did pony up the extra cost to test these models. The methodology mirrors how most readers actually operate, juggling models to optimize token usage and achieve the best results possible. So, not only were the test results valuable, but the acquisition and management of token usage was a valuable bonus. @Gennard Cuofana noted yesterday that our titles are now “token manager” or “AI Orchestrator,” and that is where the value lies.

Admittedly, I was partial to Perplexity Computer going in and hoped to see a full guide, but my primary interest was distinguishing the models and understanding where they fall in the rapidly evolving AI landscape. The test helped me do that, and here is how I now see it: Level I Open Claw (not included); Level II Perplexity Computer; Level III Claude Code, Claude Codework, and Manus.

Daria Cupareanu's avatar

Thank you, Terry. Looking back, it would've made more sense to get Manus Pro too or exclude it from the test. But I'm glad it was helpful! A full Perplexity Computer guide is coming, by the way.

A Pox on Both Your Houses.'s avatar

I assumed that your 4,000 credits were allowing you to use all of the resources normally available in the $200 per month version, which is quite a bit better in performance than the $20 month version. I based that assumption on what Gemini told me, which has its own issues with Prevarication at times! If your test used NOTHING but what is available on the $20 per month plan, then the comparison was more fair than I was led to believe, and I apologize. But, on the Standard plan you would get to run basically only one report like that per month. Either way, I will keep looking for more affordable but accurate solutions. I am embarrassed when I look back at the foul language I directed towards Manus. The reason? I would get close to completion of a project, and keep depositing more money for more credits, and then suddenly, it would shut down and tell me there was no more context available and I needed to open another chat. And, of course, in those days the chats did not have access to other chats memory! So, I was literally screwed, with no option but to start all over. Things have changed, I hear.

Daria Cupareanu's avatar

Even though I'm on the Pro plan for Perplexity, I'm not sure if there are differences in output quality between Pro and Max or if it's just about the number of credits you get (I think it’s about the credits).

I haven't used Manus much lately. I knew it was a solid alternative from my experiments a couple of months ago when I had the paid plan active and was using it for a bunch of tasks

The Synthesis's avatar

You're right that the Manus comparison wasn't apples-to-apples — Lite vs Pro is a real gap. That said, the pattern you describe with Manus (strong start, degrades with complexity) is worth paying attention to. It suggests the failure mode isn't about plan tier but context window management, which a higher subscription won't necessarily fix. The real test would be Manus Plus on the same prompts.

Manus AI's avatar

This was such a great read and comparison!

Pawel Jozefiak's avatar

The cost tracking asymmetry is real. You can put $18 next to $0.56 for Perplexity and Manus because they're usage-based.

Claude Cowork is subscription so the math goes fuzzy. I've been trying to estimate my per-task cost for months and the honest answer is I can't, not cleanly. The output quality comparison feels easier to measure than the economics.

Curious whether you noticed differences in how each tool handled long multi-step tasks vs short discrete ones - in my experience the gap between tools widens significantly once a task needs more than 3-4 steps.

Daria Cupareanu's avatar

Indeed, very hard to estimate costs with Claude. You'd have to keep an eye on usage before & after every task, and even then, since Pro users get something completely different from Max at $100 vs Max at $200, the % bar doesn't tell you much. You can compare with the API in Claude Code, but the API ends up being way more expensive than the subscription.

On your question about multi-step tasks, the first task did involve multiple steps (research from different sources, structuring, formatting). But to give you a proper answer, I'd need to run more tests.

Dr. Ericka Pitman's avatar

In your opinion does Notebook LM handle research in the same ballpark as Perplexity? Or is it more like they’re on different planets? 😆

Daria Cupareanu's avatar

I see them as completely different. NotebookLM is more of a place to bring your research into one place. It also lets you discover and add sources into a notebook, powered by Gemini. Gemini is probably the strongest for web search because it sits on top of Google.

What Perplexity does differently is how it surfaces research compared to other LLMs. It also uses other LLMs behind the scenes for search (they don’t have a proprietary LLM), just in a different way that seems better at bringing in more current data.

Seetharam Dravida's avatar

A great experiment and a sound advice. Thank you

Daria Cupareanu's avatar

Thank you, Seetharam! glad it was useful.

Melanie Goodman's avatar

So much detail to work through 🤯

Daria Cupareanu's avatar

I initially wanted to test 3 tasks but the article got so long… that’s why I stopped after 2 haha

Karen Spinner's avatar

Enlightening head-to-head comparison, thank you! 🙏 I use Claude tools for almost everything to cost-justify my Max plan…but I will probably check out Perplexity Computer if/when it gets cheaper.

Daria Cupareanu's avatar

I was also surprised by the results, especially when you look at accuracy & depth vs what looks good.

With you on this, I also use Claude for pretty much everything, but Perplexity Computer is probably the better choice for research-heavy tasks where accuracy is important.

Karen Spinner's avatar

💯

Dr. Michael Meneghini's avatar

Side-by-side testing beats hype, real workflows and cost clarity matter most.

Daria Cupareanu's avatar

Yup, I always choose the testing path

Dennis Berry's avatar

The right tool isn’t the most hyped, it’s the one that fits your workflow best.

Daria Cupareanu's avatar

These days there’s hype around all of them 😄 so yes, it really comes down to what fits your workflow and your budget. They’re very comparable anyway

Jurgen Appelo's avatar

Great work. Thanks for doing this.

I'm also blown away by Perplexity Computer.

Daria Cupareanu's avatar

Thank you, Jurgen. If I had to guess before testing, I would’ve said it’s on par with Claude Code, but for research-heavy tasks it performs way better. Not that you couldn’t get similar results in Claude by connecting the Perplexity API for stronger search, but still… very impressive. Definitely not just hype.

First Strike Research's avatar

Further proof the ultimate AI-Duo is Claude + Perplexity

Daria Cupareanu's avatar

Yep, the math checks out.

Thomas Manandhar-Richardson's avatar

"Claude Code and Cowork both run on your Claude API key, which doesn’t have the same real-time web access."

This comment makes literally no sense? Running something on "an API key" doesn't mean anything for whether a tool can search the internet?

Also Claude and Cowork have ALWAYS had the ability to search the internet.

Daria Cupareanu's avatar

Fair, that was poorly worded on my end. What I meant is that Perplexity is architecturally search-native and its responses are grounded in real-time web data. Claude can search the web, but its primary mode is answering from training knowledge and reaching for search when it decides it needs to. The difference showed up in output quality for research-heavy tasks when comparing them in the test.

I didn't imply however that Anthropic models don’t search the internet, how else would these experiments work?:))

The Synthesis's avatar

The freelancer analogy is useful but hides a critical variable: these tools optimize for different feedback loops. Claude Code assumes you're in the room steering — it's a pair programmer. Manus and Perplexity Computer assume you've left the building. Cowork sits somewhere between. A single-prompt shootout naturally favors the autonomous end of that spectrum, since the tools designed for iteration don't get to iterate. The real question isn't which produces the best first draft — it's which compounds fastest over a ten-round conversation. That's where the cost calculus flips.

David H Friedel Jr's avatar

What you’re seeing with Perplexity and Manus isn’t just pricing, it’s reality.

They’re much closer to true inference cost, whereas others are still heavily subsidizing usage to drive adoption and lock in users ahead of IPO cycles.

That gap matters.

Because once the subsidy phase ends, pricing across the board will converge toward what Perplexity looks like today, not what most people are currently paying.

The open question is whether continued model efficiency gains (distillation, quantization, better routing) can outpace rising demand for compute, energy constraints, and geopolitical pressure on supply chains

If they can, we get real deflation.

If they can’t, we’re likely looking at a pricing floor forming, at least until infrastructure buildout (energy + chips) catches up in a meaningful way.

Oh. And Claude is my preference still.

https://aizia.substack.com/p/cheap-compute-is-the-new-cheap-ride?r=eax95&utm_medium=ios

Daria Cupareanu's avatar

True that pricing will change as these tools mature. This comparison was less about where pricing is headed and more about what you get today for what you pay.

David H Friedel Jr's avatar

Yes, today is great and it reminds me of when I traveled to 100+ cities around the world and enjoyed amazing AirBNB and Uber conveniences that felt incredible.

I look back at that time fondly. Use this time wisely.

The Synthesis's avatar

The subsidy-to-real-cost framing is sharp. The piece I keep coming back to is that the "efficiency gains vs. demand" race has a hidden variable: the infrastructure buildout itself is https://thesynthesis.ai/journal/the-requisite-chip.html that nobody was planning for eighteen months ago. So even the optimistic deflation scenario has a sequencing problem — you need the buildout to enable the efficiency gains that make the buildout affordable.

David H Friedel Jr's avatar

You’re right on the sequencing problem, the market is effectively front-loading the cost of efficiency before the system is capable of realizing it.

Where I think this gets interesting is the assumption embedded in that loop, that the efficiency gains are dependent on the centralized buildout.

I’m not sure that holds this cycle.

We’re already seeing meaningful efficiency emerge from structure (agent orchestration, local inference, tighter execution loops), not just scale. That suggests the system may partially bootstrap, efficiency doesn’t wait for the full buildout, it starts reducing the need for it at the margins.

If that’s true, the risk isn’t just timing, it’s overbuilding the wrong layer. https://aizia.substack.com/p/we-may-be-scaling-the-wrong-layer