Is ChatGPT losing its edge? - Stanford-Berkeley study investigates

Published 20/07/2023, 03:24 pm

OpenAI's ChatGPT is reportedly deteriorating in capability and researchers are yet to determine the cause, according to a recent study conducted by Stanford and UC Berkeley.

The recent study demonstrated that newer versions of ChatGPT provided significantly less accurate answers to the same set of questions within a span of a few months, with researchers unable to explain this deterioration in performance.

Researchers Lingjiao Chen, Matei Zaharia and James Zou put ChatGPT-3.5 and ChatGPT-4 models through a series of tasks involving solving math problems, answering sensitive questions, writing new lines of code and conducting spatial reasoning from prompts to gauge the reliability of the different versions.

Highlighting the potential for substantial change in LLM behaviour over relatively short periods, the researchers stressed the importance of continuous monitoring of AI model quality.

They recommend that users and companies relying on LLM services in their workflows implement a form of monitoring analysis to ensure consistent performance.

We evaluated #ChatGPT's behavior over time and found substantial diffs in its responses to the *same questions* between the June version of GPT4 and GPT3.5 and the March versions. The newer versions got worse on some tasks. w/ Lingjiao Chen @matei_zaharia https://t.co/TGeN4T18Fd https://t.co/36mjnejERy pic.twitter.com/FEiqrUVbg6
— James Zou (@james_y_zou) July 19, 2023

Shift in models

ChatGPT's responses to sensitive queries, particularly those relating to ethnicity and gender, also evolved to become more concise and avoidant.

The researchers observed a shift in the models' approach to dealing with sensitive questions.

While earlier versions offered extensive reasoning for refusing to answer certain sensitive queries, the June versions simply issued an apology and refused to respond.

In the tests, the March version of ChatGPT-4 could identify prime numbers with an impressive 97.6% accuracy.

However, by June, the same model's accuracy had sharply declined to a mere 2.4%.

This study follows OpenAI's announcement on June 6 about plans to create a team dedicated to managing potential risks associated with superintelligent AI systems, which the organisation anticipates emerging within the decade.

Latest comments

S&P/ASX 200

8,285.20

+61.20

+0.74%

ASX 200 Futures

8,275.50

-44.0

-0.53%

ASX All Ordinaries

8,539.00

+59.10

+0.70%

US 500

5,874.40

-74.8

-1.26%

Dow Jones

43,444.99

-305.87

-0.70%

China A50 Futures

13,318.00

-149.0

-1.11%

Dollar Index

106.62

+0.020

+0.02%

Most Popular Articles

News

Analysis

Tesla stock target lifted at RBC on increased confidence in AVs

By Investing.co...

15 Nov 2024

Musk, RFK Jr side with Howard Lutnick in Trump search for Treasury secretary

By Reuters

16 Nov 2024

Russia cuts gas to Austria in payment dispute, keeps EU flows

By Reuters

16 Nov 2024

Lilly's tirzepatide reduced the risk of worsening heart failure events by 38% in adults with heart failure with preserved ejection fraction (HFpEF) and obesity

By Investing.co...

16 Nov 2024

Palantir CEO Alexander Karp sells shares worth $399 million

By Investing.co...

15 Nov 2024

More News

Market Movers

Name	Last	Chg. %	Vol.
BHP Group Ltd	40.070	+0.15%	7.19M
ANZ Holdings	32.450	+2.59%	6.68M
Westpac Banking	33.060	+1.91%	5.29M
National Australia Bank	39.220	+1.34%	4.33M
Commonwealth Bank Australia	155.130	+1.50%	2.49M
CSL	277.01	-2.48%	921.02K
Xero	172.61	+0.94%	778.87K

Name	Last	Chg. %	Vol.
PYC Therapeutics Ltd	1.850	+902.70%	276.31K
Noviqtech	0.056	+200.00%	77.65M
Excite Tech Services	0.012	+142.86%	200.00K
Bentley Capital Ltd	0.017	+90.91%	262.59K
Cyclone Metals	0.030	+52.63%	31.98M
Tambourah Metals	0.03	+50.00%	100.00K
Tungsten Mining NL	0.075	+37.04%	280.38K

Name	Last	Chg. %	Vol.
Koonenberry Gold	0.02	-50.00%	2.23M
Pro Pac Packaging	0.02	-50.00%	1.15M
Aumega Metals	0.042	-33.90%	596.60K
Juno Minerals	0.02	-33.33%	10.00K
Flynn Gold	0.03	-25.00%	737.68K
Trigg Mining	0.03	-25.00%	20.54M
M3 Mining	0.03	-25.00%	27.33K

Trending Stocks

Name	Last	Chg. %	Vol.
Commonwealth Bank Australia	155.130	+1.50%	2.49M
CSL	277.01	-2.48%	921.02K
Woodside Energy	24.000	+1.48%	5.29M
Magellan Financial Group	10.49	+2.04%	477.74K
BHP Group Ltd	40.070	+0.15%	7.19M

Install Our AppScan QR code to install app

Risk Disclosure: Trading in financial instruments and/or cryptocurrencies involves high risks including the risk of losing some, or all, of your investment amount, and may not be suitable for all investors. Prices of cryptocurrencies are extremely volatile and may be affected by external factors such as financial, regulatory or political events. Trading on margin increases the financial risks.
Before deciding to trade in financial instrument or cryptocurrencies you should be fully informed of the risks and costs associated with trading the financial markets, carefully consider your investment objectives, level of experience, and risk appetite, and seek professional advice where needed.
Fusion Media would like to remind you that the data contained in this website is not necessarily real-time nor accurate. The data and prices on the website are not necessarily provided by any market or exchange, but may be provided by market makers, and so prices may not be accurate and may differ from the actual price at any given market, meaning prices are indicative and not appropriate for trading purposes. Fusion Media and any provider of the data contained in this website will not accept liability for any loss or damage as a result of your trading, or your reliance on the information contained within this website.
It is prohibited to use, store, reproduce, display, modify, transmit or distribute the data contained in this website without the explicit prior written permission of Fusion Media and/or the data provider. All intellectual property rights are reserved by the providers and/or the exchange providing the data contained in this website.
Fusion Media may be compensated by the advertisers that appear on the website, based on your interaction with the advertisements or advertisers.

Popular Searches

Please try another search

Is ChatGPT losing its edge? - Stanford-Berkeley study investigates

Latest comments

Trending Stocks