I Tested AI Assistants for Two Weeks on Real Work Tasks. Here's What Actually Surprised Me.
Photo: person using AI assistant laptop productivity office comparison, via micmaronline.com
Let me save you the suspense on one thing right away: there is no single "best" AI assistant. Anyone who tells you otherwise is either selling something or hasn't actually used more than one of them for real work.
What there is, though, is a best AI assistant for you — depending on what you actually do all day. And figuring that out is way more valuable than any benchmark score or viral Twitter thread.
So I did what most reviewers don't: I used these tools the way a typical American knowledge worker actually would. Email drafts. Summarizing long documents. Writing code snippets. Brainstorming marketing copy. Explaining a confusing insurance document. The unglamorous, repetitive stuff that quietly eats hours out of your week.
Here's what two weeks of real-world testing taught me.
The Lineup
I focused on four tools: ChatGPT (OpenAI, tested on both the free GPT-4o tier and the $20/month Plus plan), Claude (Anthropic, free and $20/month Pro), Gemini (Google, free and the $19.99/month Advanced tier bundled with Google One), and Perplexity AI as a wildcard — it's been quietly gaining a devoted following among researchers and analysts.
For each tool, I ran the same core tasks and tracked two things: how much time I actually saved compared to doing it myself, and whether the output was good enough to use with minimal editing.
Email Drafting: Closer Than You'd Think
This is probably the most common use case for AI tools among office workers, so I started here. I gave each assistant the same prompt: draft a professional but warm follow-up email to a potential client who went quiet after an initial meeting.
All four produced usable drafts. But there were real differences in tone.
ChatGPT's draft was competent and clean, but felt a little generic — the kind of email that reads like it could have come from anyone. Claude's version had noticeably more personality. It felt like it had actually absorbed the nuance of the situation rather than just filling in a template. Gemini's output was solid but occasionally stiff, and it defaulted to slightly more formal language than I'd asked for.
Perplexity, interestingly, isn't really designed for this kind of creative generation. It punted to a fairly basic response.
Time saved: About 8–10 minutes per email on average across all tools, which adds up fast if you're writing a dozen emails a day. Claude edged out the others on quality, meaning less editing time.
Summarizing Long Documents: Gemini Pulls Ahead
This one was eye-opening. I uploaded a 47-page municipal zoning document (the kind of thing a small business owner or real estate investor might actually need to parse) and asked each tool to summarize the key restrictions and flag anything that might affect a retail storefront.
Gemini's integration with Google Drive and Docs gave it a structural advantage here — it handled the full document without complaint and produced a well-organized summary with clear section headers. Its deep ties to Google's document ecosystem made this feel genuinely seamless.
Claude also performed impressively, particularly in identifying nuanced language that the others glossed over. Its 200,000-token context window means it can handle very long documents without chopping them up.
ChatGPT (on the Plus plan with file upload enabled) did a decent job but occasionally missed details buried in the middle sections — a problem with how it processes long inputs.
Verdict: For document-heavy work, Claude or Gemini Advanced are meaningfully better than the free tiers of anything. This is one area where paying actually makes a measurable difference.
Code Generation: ChatGPT Is Still the Go-To (But Claude Is Closing Fast)
I'm not a developer, but I regularly need to write basic Python scripts, fiddle with spreadsheet formulas, or figure out why a piece of automation broke. This is exactly the kind of task where AI assistants can save non-technical users serious time.
I gave each tool three tasks: write a Python script to rename a batch of files based on a naming convention, debug a broken Excel VLOOKUP, and explain what a specific block of JavaScript was doing in plain English.
ChatGPT was the most reliable here — code ran correctly on the first try more often than the others, and its explanations were clear without being condescending. Claude was a close second and occasionally produced cleaner, more readable code. Gemini struggled more with the debugging task, producing code that was close but required an extra round of correction.
For the "explain this code" task, all three did well — this is genuinely one of the most practical things AI can do for non-developers.
Creative Writing and Marketing Copy: Claude's Moment to Shine
I asked each assistant to write three variations of a tagline for a fictional sustainable sneaker brand targeting millennials. Then I had a small group of colleagues (who didn't know which tool produced which) pick their favorites.
Claude won, and it wasn't particularly close. Its outputs had a distinct voice and avoided the clichéd language that plagued some of the other results. ChatGPT's versions were competent but felt like they'd been pulled from a marketing textbook. Gemini's attempts were serviceable but a bit flat.
If your work involves brand writing, content creation, or anything where voice matters, Claude is worth the subscription cost.
The Free vs. Paid Question
Here's the finding that surprised me most: for casual, occasional use, the free tiers are genuinely good enough. Free ChatGPT (GPT-4o) and free Claude handle most everyday tasks without breaking a sweat. The paid upgrades matter most when you need:
- File uploads and document analysis
- Higher usage limits (if you're a heavy daily user)
- Access to the latest model versions
- Priority response speed during peak hours
If you're using AI a few times a week for basic tasks, save the $20/month. If it's a daily productivity tool, the paid tier pays for itself quickly.
When Human Expertise Still Wins
I'd be doing you a disservice if I didn't flag this clearly: all of these tools make confident-sounding mistakes. I caught factual errors in legal summaries, outdated information presented as current, and code that looked right but had subtle bugs.
For anything high-stakes — medical decisions, legal advice, financial planning, anything where being wrong has real consequences — treat AI output as a starting point, not a final answer. The time savings are real, but so is the verification step.
The Bottom Line
After two weeks, here's my honest stack ranking for typical US knowledge workers:
- Best all-around: ChatGPT Plus (most consistent, widest capability range)
- Best for writing and nuance: Claude Pro (the quality difference in tone and creativity is real)
- Best for Google Workspace users: Gemini Advanced (the integration advantage is significant)
- Best for research and sourcing: Perplexity AI (it cites its sources, which matters more than people realize)
- Best free option: Free Claude or free ChatGPT — roughly tied, depending on the task
The smartest move? Most of these tools offer free trials of their paid tiers. Spend a week with the one that fits your workflow before committing to a subscription. That's the kind of sharp, intentional approach to tech that actually pays off.