Didattica integrativa
Gizem Gezici
Modalità d'esame
L'esame consisterà in un progetto con presentazione e discussione.
Prerequisiti
Il corso non prevede prerequisiti.
Il corso è rivolto agli studenti del Corso ordinario di II livello della Scuola Normale Superiore (SNS) e ai dottorandi del Dottorato Nazionale in Intelligenza Artificiale (Università di Pisa)
e del Dottorato in Metodi computazionali e modelli matematici per le scienze e la finanza (SNS). Sono inoltre invitati a iscriversi anche i dottorandi degli altri corsi di dottorato della SNS
interessati agli argomenti trattati nel corso.
Programma insegnamento
Il corso è articolato in tre moduli:
Modulo I – L'evoluzione dell'intelligenza artificiale generativa
Evoluzione dell'elaborazione del linguaggio naturale (NLP): dai sistemi basati su regole ai Large Language Models (LLM).
Modulo II – Fondamenti tecnici dell'intelligenza artificiale generativa
Architettura Transformer, transfer learning e calcolo su larga scala.
Modulo III – Intelligenza artificiale generativa responsabile
Sfide etiche, sociali e tecniche, tra cui bias, allucinazioni e interpretabilità.
Strategie di valutazione e mitigazione.
Seminari, casi di studio e attività pratiche.
Riferimenti bibliografici
Reference bibliography
- Vaswani, Ashish, et al. "Attention is all you need." Advances in neural information processing systems 30 (2017).
- Paaß, G., & Giesselbach, S. (2023). Foundation models for natural language processing: Pre-trained language models integrating media (p. 436). Springer Nature.
- Devlin, Jacob, et al. "Bert: Pre-training of deep bidirectional transformers for language understanding." Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 2019.
- Radford, Alec, et al. "Improving language understanding by generative pre-training." (2018).
- Brown, Tom, et al. "Language models are few-shot learners." Advances in neural information processing systems 33 (2020): 1877-1901.
- Wei, Jason, et al. "Chain-of-thought prompting elicits reasoning in large language models." Advances in neural information processing systems 35 (2022): 24824-24837.
- Kojima, Takeshi, et al. "Large language models are zero-shot reasoners." Advances in neural information processing systems 35 (2022): 22199-22213.
- Ouyang, Long, et al. "Training language models to follow instructions with human feedback." Advances in neural information processing systems 35 (2022): 27730-27744.
- Karpukhin, Vladimir, et al. "Dense Passage Retrieval for Open-Domain Question Answering." EMNLP (1). 2020.
- Dettmers, Tim, et al. "Qlora: Efficient finetuning of quantized llms." Advances in neural information processing systems 36 (2023): 10088-10115.
- Gallegos, Isabel O., et al. "Bias and fairness in large language models: A survey." Computational Linguistics 50.3 (2024): 1097-1179.
- Shuster, Kurt, et al. "Retrieval augmentation reduces hallucination in conversation." Findings of the Association for Computational Linguistics: EMNLP. 2021.
- Goodfellow, Ian, et al. "Generative adversarial networks." Communications of the ACM 63.11 (2020): 139-144.
- Bommasani, Rishi, et al. "On the opportunities and risks of foundation models." arXiv preprint arXiv:2108.07258 (2021).
- EMNLP 2024 Tutorial: Language Agents: Foundations, Prospects, and Risks https://language-agent-tutorial.github.io/