3 pointsby vicnovalmost 2 years ago

2 comments

vicnovalmost 2 years ago

"What are the main findings? In a nutshell, there are many interesting performance shifts over time. For example, GPT-4 (March 2023) was very good at identifying prime numbers (accuracy 97.6%) but GPT-4 (June 2023) was very poor on these same questions (accuracy 2.4%). Interestingly GPT-3.5 (June 2023) was much better than GPT-3.5 (March 2023) in this task. We hope releasing the datasets and generations can help the community to understand how LLM services drift better."

jamesmurdzaalmost 2 years ago

How is it possible to for the success rate to go from 98% to 2%? What is the author's explanation for this?

LLM Drifts: How Is ChatGPT’s Behavior Changing over Time?

2 comments

LLM Drifts: How Is ChatGPT’s Behavior Changing over Time?

2 comments