상세 보기
The Evolution of Large Language Models
초록
This paper explores the evolution of large language models (LLMs), tracing their development from statistical methods to advanced neural architectures. By examining milestones such as n-grams, Hidden Markov Models, recurrent neural networks (RNNs), and transformer-based frameworks, it highlights key innovations addressing long-range dependencies and model scalability. Attention mechanisms and autoregressive frameworks like GPT transformed natural language processing (NLP), enabling breakthroughs in tasks such as translation and adaptive education. The study also evaluates commercial versus open-source LLMs, considering their pedagogical applications, ethical concerns, and resource demands. The paper concludes by advocating for sustainable, inclusive NLP practices that balance technological innovation with transparency and equity.