An AI assistant comparison synthesis

Over 20 years ago, I read Guy Kawasaki‘s book Rules for Revolutionaries, and one of his pieces of advice that has stuck with me is “Eat like a hummingbird, poop like an elephant.”
Allow me to explain.
Hummingbirds have one of the highest metabolic rates in the animal kingdom, so they must eat about 50% of their body weight every day to support their level of activity. He wasn’t suggesting that we eat that much, but rather that it is good to consume information at a very high rate, preferably from a wide variety of sources, much like hummingbirds consume nectar.
For the second part of that statement, it is good to know that Elephant dung is a highly fibrous, nutrient-rich waste produced in massive quantities – up to 100 kg (220 lbs) daily per adult. Because elephants digest only 40–50% of their food, their dung is rich in cellulose, making it excellent for sustainable paper, fuel, and fertilizer. It also plays a key role in ecosystems by dispersing seeds and providing habitats for insects. His parallel is that we should process the information we consume at such a high rate so that we leave a trail of knowledge bombs behind us, which can have beneficial effects on our environment.
With that in mind, I’m choosing to share artifacts I create to satisfy my occasionally overactive curiosity, in the hope that others might find them useful. This is one such artifact.
I certainly don’t use A.I. for everything, but when I do, I want to be sure I’m using the best tool in the best way possible. To that end, I have access to a variety of tools, and I periodically compare them against subject matter I know to assess the quality and completeness of their answers and to see what kinds of interesting hallucinations different tools can produce under different circumstances.
To help evaluate the relative strengths and weaknesses of the tools available to me, I sent the following prompt to each of them: “Tell me about the relative strengths and weaknesses of the following A.I. services: – Copilot – ChatGPT – Gemini – Perplexity – Claude.”
I then copied the answers from each tool, except Copilot (as I don’t have it available on my personal computer), and uploaded them to Claude with the following prompt: “Analyze the results of this document and these pages, and combine them into a synthesized summary.”
The following was the result. Take it for what it is, and make use of it as you will.
*****
AI Assistant Comparison
Synthesized Analysis from Four Independent Assessments
Sources: Perplexity, ChatGPT (two responses), Claude, and Gemini
March 2026
Overview
This document synthesizes four independent assessments of the five leading AI assistants—ChatGPT, Claude, Gemini, Copilot, and Perplexity—produced by Perplexity, ChatGPT, Claude, and Gemini. Each model was asked the same question about the relative strengths and weaknesses of all five services, meaning every model evaluated both itself and its competitors.
Despite the inherent conflict of interest, a remarkably consistent picture emerges: the AI assistant landscape in 2026 has shifted from general-purpose chatbots toward specialized tools, each optimized for a different use case. As the ChatGPT assessment frames it, these tools are not trying to win the same game. Rather than asking which is best in absolute terms, the practical question is which is best for a given task and workflow.
Consensus at a Glance
| Service | Best For | Key Strength | Main Weakness |
|---|---|---|---|
| ChatGPT | General productivity and versatility | Largest ecosystem; strongest all-rounder with deep research mode | Can default to a generic "AI voice"; sometimes overconfident; factual answers need checking |
| Claude | Writing, coding, and deep reasoning | Superior prose; leads coding benchmarks; 200K context; strong structured thinking | Overly cautious safety filters; smaller ecosystem; can be verbose |
| Gemini | Google Workspace users and multimodal tasks | 1M+ token context; deep Google integration; strong multimodal capabilities | Less consistent output quality; UX criticized as clunky; generic tone |
| Copilot | Microsoft 365 and enterprise workflows | Native integration with Word, Excel, Teams, Outlook; strong GitHub coding | Limiting outside Microsoft ecosystem; weaker reasoning vs. ChatGPT/Claude |
| Perplexity | Research, fact-checking, and live information | Every claim cited with sources; real-time web data; deep research engine | Weak at creative tasks and coding; not a full workspace or long-term assistant |
Service-by-Service Synthesis
ChatGPT (OpenAI)
All four sources position ChatGPT as the market leader and strongest generalist. Claude’s analysis cites roughly 60% market share among AI chatbots, and all four highlight its unmatched plugin and Custom GPT ecosystem. The deep research mode is praised across the board. ChatGPT’s second response adds emphasis on strong UX, persistent memory, and balanced performance—calling it the “default choice” AI. Perplexity’s assessment, notably well-sourced with citations, describes ChatGPT as a “strong all-rounder for brainstorming, writing, coding help, and general conversation.”
Consensus strengths: Versatility across writing, brainstorming, coding, and conversation; the largest third-party ecosystem and custom GPT marketplace; strong voice and multimodal features including DALL-E 4 integration; good memory and customization capabilities. All four sources agree it is the safest single-tool choice for everyday use.
Consensus weaknesses: A tendency to produce confident-sounding answers that may not be accurate—raised in all four assessments. Gemini describes the default tone as “corporate,” Claude notes it “optimizes for confident-sounding answers over accurate ones,” and ChatGPT’s own second response concedes it is “not always the deepest thinker.” Perplexity notes it is “not as tightly embedded in Google or Microsoft workflows” as ecosystem-native competitors. Gemini adds that persistent memory can struggle with consistency over long projects.
Claude (Anthropic)
Claude receives the strongest coding and writing praise across all four sources. Perplexity’s sourced assessment highlights Claude for “long documents, structured reasoning, careful drafting, and coding/automation-oriented work.” Claude’s own assessment cites leading SWE-bench and ARC-AGI-2 benchmarks. Gemini calls it “widely considered the best for natural, nuanced prose.” ChatGPT’s second response specifically notes Claude “often wins” over ChatGPT for deep thinking.
Consensus strengths: Superior long-form writing that avoids typical AI patterns; top-tier coding and structured reasoning performance; the 200K context window for processing long documents; strong instruction-following; and a more natural, human-like writing style. ChatGPT’s second response adds that Claude excels specifically at strategy, ethical topics, and nuanced analysis. Enterprises in finance and cybersecurity have adopted it for cautious, honest outputs.
Consensus weaknesses: Overly cautious safety filters that sometimes refuse legitimate requests—conceded by all four sources, including Claude’s own assessment. A smaller plugin ecosystem compared to ChatGPT; fewer integrations with external tools. Perplexity adds that Claude is “often positioned as less playful/creative,” ChatGPT’s second response notes it can be “overly cautious or verbose,” and Gemini flags that its internal search capabilities lag behind competitors.
Gemini (Google)
All four sources converge on Gemini’s core value proposition: Google ecosystem integration and industry-leading context length. Gemini’s own assessment highlights the 1-million-token context window and recent reasoning improvements in 3.1 Pro. Perplexity describes it as strong for “multimodal work such as text, images, audio, and video.” ChatGPT’s second assessment adds that Gemini benefits from Google’s infrastructure for speed and scalability.
Consensus strengths: Deep integration with Gmail, Docs, and Drive; the largest context window available (1M+ tokens); strong multimodal processing; fast performance on Google infrastructure; and good research capabilities tied to Google Search. All four agree it’s the obvious choice for heavy Google Workspace users.
Consensus weaknesses: Less consistent output quality compared to ChatGPT and Claude. All four sources note a less distinctive conversational style; Gemini itself calls this an “identity crisis,” ChatGPT’s second response calls the tone “generic,” and the UX is criticized as “clunky or fragmented.” Perplexity notes it “can be less appealing if you are not in the Google ecosystem” and “more workflow-oriented than creatively polished.” Gemini uniquely flags privacy friction with personal Google data integration.
Copilot (Microsoft)
All four analyses agree that Copilot’s value is almost entirely tied to its Microsoft 365 integration. Claude’s analysis notes 85% of Fortune 500 companies use Microsoft’s generative AI platforms. ChatGPT’s second response highlights GitHub Copilot as a standout coding tool and the advantage of free access to high-end OpenAI models for Windows users. Perplexity’s assessment frames Copilot’s advantage as being about “where it sits, not just how it chats.”
Consensus strengths: Native, friction-free integration with Word, Excel, Teams, Outlook, and Windows; context-aware tasks like summarizing missed meetings or drafting slide decks; strong enterprise workflow automation; and excellent coding assistance through GitHub Copilot.
Consensus weaknesses: Limited appeal outside the Microsoft ecosystem—the most unanimous finding across all four sources. ChatGPT’s second response notes weaker “creativity and long-form reasoning” compared to Claude and ChatGPT. Gemini describes it as feeling like a “wrapper” over ChatGPT. Claude notes Microsoft has focused on improving the existing tool rather than releasing major upgrades, and Gemini adds that pervasive Windows integration can feel intrusive. Perplexity concurs it is “less flexible” for “open-ended creativity or cross-platform use.”
Perplexity
The strongest consensus across all four analyses is around Perplexity’s unique positioning as a research and citation tool. Claude’s analysis calls it “arguably better than any chatbot” for sourced research; Gemini highlights its deep research engine; and ChatGPT’s second response frames it as “AI-powered Google.” Perplexity’s own assessment is notably modest, describing itself as “best for live, source-backed research and quick fact-finding” without overselling general capabilities.
Consensus strengths: Every claim comes with clickable source citations; strong live web data retrieval and real-time search; the deep research engine for comprehensive investigation. Gemini adds the ability to toggle between underlying AI models (Claude, GPT) for different perspectives. ChatGPT’s second response highlights specific use cases: market research, news monitoring, and competitive analysis. Perplexity’s own sourced assessment notes that its “citations make verification easy.”
Consensus weaknesses: Poor at coding, creative writing, and open-ended ideation—unanimous across all four sources, including Perplexity’s own assessment. Not designed for long-term project management or ongoing conversational threads. ChatGPT’s second response notes it is not a full “AI workspace.” Perplexity itself concedes it is “less strong for creative writing, open-ended ideation, or deeply nuanced drafting.”
Shared Weaknesses Across All Services
ChatGPT’s second assessment uniquely highlights weaknesses shared by all five services—a useful corrective to the tendency of each source to frame limitations as belonging only to competitors.
Hallucination risk: All five services can still generate plausible-sounding but incorrect information, especially without citation mechanisms. This risk is lowest with Perplexity (which cites sources by default) and highest in open-ended creative or reasoning tasks across any tool.
Reasoning limitations: All services still struggle with certain types of reasoning, particularly spatial and visual tasks. No single tool has fully solved complex multi-step logical reasoning, though Claude and recent Gemini updates score highest on benchmarks.
Prompt sensitivity: Performance across all five services varies significantly depending on prompt quality. A well-crafted prompt can dramatically improve output from any tool, and a vague prompt will produce mediocre results from even the best model.
Notable Biases and Divergences
While the four assessments are broadly aligned, patterns of self-favoritism and emphasis are worth noting.
Perplexity’s response (the uploaded document) is the most carefully sourced, with numbered citations for nearly every claim. It is also the most modest in self-assessment, positioning itself simply as “best for live, source-backed research” without overclaiming general capabilities. Perplexity is the only source that does not cite benchmarks or market data in its own favor. Its competitor assessments are fair and balanced, making it arguably the most neutral of the four evaluations.
ChatGPT’s response is the most decision-oriented and practical, framing the comparison as a workflow question rather than a capabilities ranking. It is the only source to articulate a shared-weaknesses section, which adds balance. However, it is also more generous to itself (“most versatile”) and more pointed in criticizing competitors’ UX and consistency.
Claude’s response is transparent about potential bias (“take my self-assessment with appropriate skepticism”) but is the most data-heavy in its own favor, citing specific benchmark leads (SWE-bench, ARC-AGI-2) and enterprise adoption figures. It is notably the only source to cite ChatGPT’s 60% market share—acknowledging a competitor’s dominance lends credibility.
Gemini’s response emphasizes its 1M+ context window and recent reasoning improvements (Gemini 3.1 Pro) more prominently than other sources. It is the most self-critical about personality and tone (“identity crisis”) but uniquely highlights features others understate, like privacy concerns and the ability to upload entire books or hour-long videos for analysis.
Practical Decision Framework
Drawing from the consensus across all four analyses, the practical recommendation reduces to three questions:
1. Where do you work? If your day is built around Microsoft 365, choose Copilot. If you live in Google Workspace, choose Gemini. Both tools derive their primary value from ecosystem integration rather than standalone capability.
2. Do you need thinking or answers? For deep reasoning, analysis, strategy, and polished writing, Claude is the consensus pick. For fast, sourced answers to factual questions, Perplexity is unanimously recommended.
3. Do you want one tool for everything? ChatGPT remains the safest single-tool choice, with the broadest versatility and largest ecosystem. However, all four analyses converge on the same practical insight: most serious users end up using two or three of these tools for different tasks rather than committing to one.
| Task | Best Choice |
|---|---|
| Everyday AI assistant | ChatGPT |
| Writing, strategy, and deep analysis | Claude |
| Coding and development | Claude or GitHub Copilot |
| Research with sourced citations | Perplexity |
| Office productivity (Excel, PowerPoint, Teams) | Copilot |
| Google Docs, Gmail, and Drive workflows | Gemini |
| Long document processing | Gemini (1M context) or Claude (200K context) |
Methodology
Four of the five AI services—Perplexity, ChatGPT, Claude, and Gemini—were each asked the same question about the relative strengths and weaknesses of all five services (ChatGPT provided two separate responses). The resulting assessments were compared for areas of consensus, disagreement, and self-bias. This synthesis prioritizes claims that appear in at least two of the four assessments and flags single-source claims accordingly. Notably, all five services were evaluated by at least one competitor, and four of the five also evaluated themselves, providing a useful cross-check on self-serving claims.
This article was first published on Substack on March 29, 2026.