
This intermediate Google DeepMind course demystifies the inner workings of the transformer architecture. Through hands-on activities, learners dissect prompt processing, visualize attention weights, and explore multi-head and masked attention while learning community-centered AI design practices.
Data scientists, ML engineers, and researchers wanting a deep mechanical understanding of transformer architectures and attention mechanisms.
In this Google DeepMind course you will discover the mechanisms of the transformer architecture. You will investigate how transformer language models process prompts to make context-sensitive next-token predictions. Through practical activities you will explore the attention mechanism, visualize attention weights, and encounter advanced concepts like masked attention and multi-head attention. You will also learn other techniques that are necessary to build neural networks that are well-suited to be used as language models. Finally, through activities on values, stakeholder mapping and community engagement, you will practice concrete tools for ensuring AI projects are developed with communities, not just for them.
Price
Advertisement
This course is free to enrol.
Enrol free on DataCamp →
More in Data Science & AI