I Tried Having AI Solve the CSAT Math Problems
I fed last year's CSAT math problems into GPT-4o and Claude Opus.
Results:
- GPT-4o: Got up to problem 22 (high difficulty) correct, but made a mistake on problem 29 (calculus). Overall accuracy: 80%
- Claude Opus: Solved all problems up to 30 correctly. However, it took twice as long as GPT.
Both showed perfect reasoning processes, but occasional calculation errors popped up. Still, they'd probably score around grades 3-4 on average. It seems like the day AI gets a top grade on the CSAT isn't far off.
3 answers
Wow, Claude got all 30 right? That's amazing lol. But if it takes twice as long to solve them, then in a real test that's a bit... Isn't the CSAT all about time management?
The pace of AI development is truly insane. But it's quite human-like that even though reasoning is perfect, it still makes calculation errors. I remember getting #29 wrong due to a calculation mistake, lol.
Well, this is different from the actual CSAT exam environment, isn't it? Did they scan the test paper and input it? And they might have trained the data on a question bank. Do I need to bring a tablet in to actually get a top grade? lol