Claude 3.5 vs GPT-4o Coding Test Comparison Review

I tried a few Baekjoon Gold-level problems, and Claude performed better than expected.

1. GPT-4o: Clean time complexity explanations, but frequent implementation errors.

2. Claude 3.5: High code completeness and provides debugging tips.

Personally, Claude seems better for algorithm problems. But I haven't tested fine-tuned models yet.

by 주말개발자671

10 answers

Agreed, Claude was more accurate on the gold-level problems for me too.

by 초보개발자536 · ▲0

GPT does make mistakes in implementation, lol. It gets things wrong surprisingly often.

by 문과출신개발자63 · ▲0

But fine-tuned GPT-4o might actually be better, right? Comparing it to the general model seems a bit unfair.

by 궁금한사람108 · ▲0

Claude seems to have really improved. Its coding used to be a mess.

by 프롬프트장인623 · ▲0

I also tested the platinum-level problems, and Claude had fewer timeouts. However, GPT was better in terms of memory usage.

by 월급루팡852 · ▲0

Oh, this is right lol. Giving debugging tips is actually way more helpful than I thought.

by 디지털노마드175 · ▲0

Well, I just find GPT more convenient. Claude's explanations are too verbose and I can't focus.

by 카페인중독991 · ▲0

Please let me know the source. Did you run it yourself, or did you refer to another benchmark?

by 데이터덕후545 · ▲0

Honestly, GPT seems better for beginners. Claude is too focused on experts.

by 스타트업러871 · ▲0

GitHub Copilot is still the best lol, especially for Python.

by 클라우드러버316 · ▲0