| Abstract | Deep learning models including large language models (LLMs) suffer from concept drift for various reasons and causes. A drift can occur either during data collection due to changes in data distribution and data imbalance, or due to external factors such as shifts in the economy, policies, or technological advancements. The drift can occur in one or more key data components such as language, context, and the web. Unlike traditional machine learning models, LLMs are increasingly used as agents to generate responses and content that their users use in their daily lives. Numerous papers of research have been published on the impact of LLMs on user behaviour, content generation and consumption, and contextual understanding. For instance, there has been a noticeable decline in the use of public knowledge-sharing platforms such as Stack Overflow, or how the spoken language of individuals is influenced by words or patterns resembling those generated and used by LLMs, particularly by models like GPT. In this paper, we study concept drift in LLMs, how it occurs, and the consequences of its (re-)occurrence if it remains unhandled, particularly at times of model convergence when the model might suffer from a sudden or gradual concept drift, class imbalance, or data (sparsity) imbalance. If we assume that a hypothetical LLM has reached convergence at a certain point in time, and then the data encounters a class imbalance or concept drift after such a point in time, a model may not understand new concepts and/or properly generate responses to new concepts since it was trained on outdated concepts. In particular, this paper looks at concept drift in LLMs related to one or more of the foundational building blocks of LLM training data: context, language, and the web as a whole. The web is a fundamental component of data as it reflects human behaviours, finite (from the perspective of new data) human-generated content, and patterns, while also being the primary source of various shifts and drifts. |
|---|