
This intermediate Google DeepMind course addresses text data representation, tokenization, and vector embeddings. Learners analyze how language models capture semantic meaning through vector spaces while learning to mitigate data bias and design transparent, ethical datasets using Data Cards.
NLP engineers, data scientists, and AI developers focusing on text pre-processing, vector embeddings, bias reduction, and ethical dataset design.
In this Google DeepMind course you will learn how to prepare text data for language models to process. You will investigate the tools and techniques used to prepare, structure, and represent text data for language models, with a focus on tokenization and embeddings. You will be encouraged to think critically about the decisions behind data preparation, and what biases within the data may be introduced into models. You will analyze trade-offs, learn how to work with vectors and matrices, how meaning is represented in language models. Finally, you will practice designing a dataset ethically using the Data Cards process, ensuring transparency, accountability, and respect for community values in AI development.
Price
Advertisement
This course is free to enrol.
Enrol free on DataCamp →
More in Data Science & AI