Testing GPT-4 with Wolfram Alpha and Code Interpreter plug-ins on math and science problems

  • 2023-08-10 18:22:28
  • Ernest Davis, Scott Aaronson
  • 0

Abstract

This report describes a test of the large language model GPT-4 with theWolfram Alpha and the Code Interpreter plug-ins on 105 original problems inscience and math, at the high school and college levels, carried out inJune-August 2023. Our tests suggest that the plug-ins significantly enhanceGPT's ability to solve these problems. Having said that, there are still often"interface" failures; that is, GPT often has trouble formulating problems in away that elicits useful answers from the plug-ins. Fixing these interfacefailures seems like a central challenge in making GPT a reliable tool forcollege-level calculation problems.

 

Quick Read (beta)

loading the full paper ...