GPT and Claude failed Bridgewater's finance tests because the right answers were never public
Bridgewater and Thinking Machines Lab found that off-the-shelf AI models like GPT and Claude struggle with specialized financial document analysis because the necessary data is not part of their public training sets. The researchers demonstrated that a smaller, fine-tuned open-source model performed more accurately and at a lower cost than these general-purpose systems. This suggests that domain-specific training remains essential for high-stakes financial applications, as general models cannot reliably infer proprietary or non-public market information.
Covered by 1 source
- TThe Decoder↗Maximilian SchreinerJul 3