GPT-5 vs Claude 4 vs Gemini 2: Real-World Work Test Comparison
Results from testing the three models on the same tasks (data analysis, code review, document summarization) over the past three months.
1. Data Analysis Ability
- GPT-5: Excels at creating tables and statistical summaries. Also handles R coding adeptly.
- Claude 4: Strong at reading long CSV files and deriving insights. However, weak at catching errors in Korean code.
- Gemini 2: Integrates well with Google Docs, making it convenient to use with collaboration tools. Quite fast for simple queries.
2. Code Review Accuracy
- GPT-5: Frequently catches security vulnerabilities. Downside: too many comments.
- Claude 4: Specialized in finding logic errors. Detailed explanations are a plus.
- Gemini 2: Fastest speed but somewhat lower accuracy.
3. Document Summary Quality
Personally, Claude 4 felt the most natural. GPT-5 was too rigid, and Gemini sometimes missed the key points.
Conclusion: It seems the answer is to use all models interchangeably depending on the situation. Each has its own clear strengths.
7 answers
Oh, thanks for the great info.
I feel the same way
But what about the cost?
I've tried it too, and Claude is definitely better for code reviews. GPT talks too much lol.
For data analysis, GPT is good, but for Korean document summarization, Claude is really natural. That's why I use them interchangeably. However, Gemini is only useful within the Google ecosystem.
Well, I actually thought GPT-5's summaries were fine? Claude was a bit inconvenient when catching Korean errors. But overall, I do agree with some parts.
Do you have a source? I'd like to try it too.