I fed last year's CSAT math problems into GPT-4o and Claude Opus.
Results:
- GPT-4o: Got up to problem 22 (high difficulty) correct, but made a mistake on problem 29 (calculus). Overall accuracy: 80%
- Claude Opus: Solved all problems up to 30 correctly. However, it took twice as long as GPT.
Both showed perfect reasoning processes, but occasional calculation errors popped up. Still, they'd probably score around grades 3-4 on average. It seems like the day AI gets a top grade on the CSAT isn't far off.