3 点作者 vicnov将近 2 年前

2 条评论

vicnov将近 2 年前

"What are the main findings? In a nutshell, there are many interesting performance shifts over time. For example, GPT-4 (March 2023) was very good at identifying prime numbers (accuracy 97.6%) but GPT-4 (June 2023) was very poor on these same questions (accuracy 2.4%). Interestingly GPT-3.5 (June 2023) was much better than GPT-3.5 (March 2023) in this task. We hope releasing the datasets and generations can help the community to understand how LLM services drift better."

jamesmurdza将近 2 年前

How is it possible to for the success rate to go from 98% to 2%? What is the author's explanation for this?

LLM Drifts: How Is ChatGPT’s Behavior Changing over Time?

2 条评论

LLM Drifts: How Is ChatGPT’s Behavior Changing over Time?

2 条评论