Course Materials & Downloads
Searchable hub for lecture slide PDFs, landmark literature, interactive Jupyter notebooks, and evaluation datasets.
History of NLP & Modern LLM Landscape
Word Vectors & Learned Representations
Backpropagation & Neural Computing for NLP
Autoregressive Language Modeling
Recurrent Neural Networks (RNN, LSTM, GRU)
Attention Mechanisms & The Transformer Architecture
Tokenization & Multilingual Language Modeling
LLM Pretraining: Data, Systems & Architectures
Scaling Laws, In-Context Learning & Prompting
Fine-Tuning, Parameter-Efficient Adaptation & LoRA
Reinforcement Learning Fundamentals in NLP
Post-Training with Human Preferences: RLHF & DPO
Model Quantization & Distributed Training Systems
Mixture of Experts (MoE) & Long-Context Scaling
Decoding Algorithms, Chain-of-Thought & Test-Time Scaling
Retrieval-Augmented Generation (RAG)
Language Model-Based Agents & Tool Use
Evaluation Techniques, Benchmarking & Leaderboards
Interpretability, Safety, Bias & Privacy in LLMs
Multimodal LLMs: Vision-Language Models (VLM)
Diffusion Models & Flow Matching for NLP
Graph Neural Networks for Natural Language Processing
Research Skills, Experimental Design & Future Frontiers
Training language models to follow instructions with human feedback
Ouyang et al. (InstructGPT, 2022)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafailov et al. (2023)
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Lewis et al. (2020)
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei et al. (2022)
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer et al. (2017)
Graph Neural Networks for Natural Language Processing: A Survey
Wu et al. (2021)
Lab 1: Implementing Word2Vec & CBOW from Scratch (PyTorch)
Hands-on coding tutorial with PyTorch and Hugging Face Transformers.
Lab 2: Scaled Dot-Product & Multi-Head Self-Attention Transformer
Hands-on coding tutorial with PyTorch and Hugging Face Transformers.
Lab 3: Parameter-Efficient Fine-Tuning with LoRA & QLoRA
Hands-on coding tutorial with PyTorch and Hugging Face Transformers.
Lab 4: Building an Agentic Dense-Retrieval RAG Pipeline with LangChain
Hands-on coding tutorial with PyTorch and Hugging Face Transformers.
Lab 5: Post-Training Alignment with DPO (Direct Preference Optimization)
Hands-on coding tutorial with PyTorch and Hugging Face Transformers.
PersianNLP Benchmark Suite (PQuAD, Digikala Reviews, ParsNLU)
Standard datasets for Persian reading comprehension, sentiment, and NLI.
MMLU & MMLU-Pro
Massive Multitask Language Understanding benchmark for reasoning and knowledge evaluation.
GSM8K & MATH
Grade school math and competitive reasoning datasets for chain-of-thought evaluation.
Alpaca & UltraFeedback
High-quality instruction tuning and human preference preference datasets.