OpenAI GPT-6 – Back on Pole

OpenAI retakes the lead.

  • Although the tests are unverified, GPT-6 Astra looks very promising, with big improvements on some of the most difficult tests, and crucially it also appears to be much cheaper to operate, although this has made the model more difficult to track and monitor.
  • OpenAI has announced GPT-6, which is currently in soft release to select customers but will become more widely available to paid users of ChatGPT pretty soon.
  • It is important to remember that the release document (see here) is a marketing press release and not a scientific paper, and so the conclusions that it draws need verification by independent 3rd parties.
  • So confident is OpenAI about its technical superiority that it chose both extremely difficult benchmarks and the same ones that Anthropic chose when it released Fable-5.
  • Looking at the tests, GPT-6 Astra’s performance superiority and the lower cost to run it are the two standout conclusions.
  • This is significant because the cost of AI is becoming important as token spend by enterprises has gone through the roof and the shortage of data centre capacity has meant that the cost of compute is also at supernormal levels.
  • This has led finance departments to take a much closer look at AI usage and, in many cases, to cap the spend.
  • The cost advantage is not because OpenAI has lower token prices, but because GPT-6 uses fewer tokens to arrive at its output.
  • However, this advantage has come at the cost of visibility.
  • A lot of the processing that goes on in the model is unreported, which is one reason why it is more efficient, but this means that it is more difficult to work out how the model is coming up with its answers.
  • From a reliability and safety perspective, this is problematic, but OpenAI claims to have managed to get the hallucination rate down, which is in line with the recent trends of frontier models.
  • This is something I have generally observed when asking models to gather data on a certain topic and double-checking for hallucinated data or references.
  • As one would expect, when it comes to coding, there is very little between Fable-5 and GPT-6 other than GPT-6 using fewer tokens to produce the same result.
  • If OpenAI can hang onto this efficiency lead, then it is in a very good position, as I think it is also sourcing compute from its providers at a much lower price.
  • This is because it struck its compute deals far earlier than Anthropic and did so when one would buy compute at 20% of the current going rate.
  • If OpenAI can prevent its suppliers from wiggling out of their contracts and putting their prices up, it could enter 2027 with a very large cost advantage over its competitor.
  • This is how it may be able to wrest the initiative that it has lost to Anthropic and start growing revenues much more quickly.
  • Up until last week, the general view was that Anthropic was running away with the AI race.
  • This appears to be no longer the case, which is the last thing Anthropic wants to hear right before its IPO.

RICHARD WINDSOR

Richard is founder, owner of research company, Radio Free Mobile. He has 16 years of experience working in sell side equity research. During his 11 year tenure at Nomura Securities, he focused on the equity coverage of the Global Technology sector.

Leave a Comment