Descriptif
This course introduces the foundations and recent advances in Large Language Models (LLMs), covering their design, training, evaluation, and deployment. It combines core theoretical principles with practical insights drawn from current research and industry applications. The course follows the full lifecycle of LLMs, from large-scale pretraining on web data to post-training and alignment techniques, as well as efficient adaptation to downstream tasks. It also addresses key challenges in scaling, including data cura- tion, distributed training, and computational efficiency. In addition, the course explores how to evaluate and understand LLMs, covering a range of evaluation methodologies and interpretability approaches. It further introduces state-of-the-art research topics, including advanced capabilities such as reasoning and multilingualism. Overall, the course provides a comprehensive view of modern LLM development, equipping students with both a strong conceptual foundation and hands-on experience to pursue research or industry roles in this rapidly evolving field. Topics to be covered • Architectures: recap on transformers and attention, sparse attention, grouped query attention, multi-query attention, sliding window and chunked attention, mixture of experts • Pretraining: scaling law, data sources and web-scale corpora, data filtering and deduplication, pretraining objectives and variations of next-token prediction • Post-Training: supervised fine-tuning, reinforcement learning from human feedback, preference learning, parameter-efficient tuning • Inference and Efficiency: quantization, decoding methods, parallelism and distributed training, test-time scaling, memory bottlenecks and KV cache • Evaluation and Interpretability: overview of state-of-the-art LLMs and benchmarks, paradigms of automatic text evaluation, LLM-as-a-judge, human evaluation best practices, attention visualiza- tion, representation analysis, mechanistic interpretability • Reasoning: chain-of-thought prompting, self-consistency and test-time reasoning, tool use and in- termediate reasoning, planning and decomposition, benchmarks for reasoning • Multilingualism: tokenization challenges across languages, data disparity and representation im- balance, evaluation challenges in multilingual settings, social and cultural bias in NLP, machine translation as a case study
24 heures en présentiel
Diplôme(s) concerné(s)
Format des notes
Numérique sur 20Pour les étudiants du diplôme M2 DS4Health - Digital Skills for Healthcare Transformation
Le rattrapage est autorisé (Note de rattrapage conservée)- le rattrapage est obligatoire si :
- Note initiale < 7
- le rattrapage peut être demandé par l'étudiant si :
- 7 ≤ note initiale < 10