08/12/2026
In a new study from Nature, researchers used AI models to predict the outcomes of social science studies. As illustrated in the 2nd image, the researchers examined 70 survey studies and almost 120,000 participants. They found that ChatGPT-4's predicted results and the observed results were highly correlated (r = .85). This means that ChatGPT was good at identifying the strongest effects, weakest, etc. (see 3nd image). However, it systematically overestimated the magnitude of effects, predicting that effects would be twice as strong as they really were. To be fair, humans overestimated effects by about the same amount.
This means that AI models are good at simulating human responses to surveys. They are less capable of predicting humans' responses to experimental studies that expose people to different interventions (r = .39), such as studies designed to change attitudes. AI models also overestimated the magnitude of the effects in these types of studies. But that's also a problem with human judgments of intervention studies, too.
This study has important implications for how science will be conducted in the future. AI can be used to run or augment pilot studies, which reduces cost and can help full studies reach fruition earlier. AI can also be used to identify studies that are likely to replicate (or not), which will help scientists know where to concentrate their resources.
Link to the full study in the comments.