EC9D8: Foundations of Data Science in Economics
Introduction
Analyses in all fields of Economics nowadays make frequent use of large and detailed datasets ("big data"). The explosion in data access and availability opens many opportunities for applied research, as well as new challenges on how to handle, process, and extract meaningful conclusions from the data. The aim of the module is to introduce students to the R and Python programming languages and basic concepts of data science; and to provide and "hands-on" experience with economic data. The module lays the foundation to more advanced materials.
Principal Aims
The primary aim of the module is to introduce students to the analytical tools of data science provided by the R and Python computing languages. It builds computer coding skills from the basic principles, and covers methods to acquire, process and manipulate large volumes of data, often obtained from the web. Data visualisation methods are also presented.
By the end of the module, students should feel comfortable with collecting, visualizing and presenting
both structured and unstructured datasets. This should include using multiple programs/languages,
depending on their application of interest, including R and Python. They should be aware
of the advantages/disadvantages of each package and applications of each one.
Principal Learning Outcomes
Subject Knowledge and Understanding: Be able to process and work efficiently with large datasets.
Subject Knowledge and Understanding: Develop and enhance computer skills in the R language, including the writing of clear and reproducible R codes.
Subject Knowledge and Understanding: Be able to use R to process data and apply data-science methods.
Subject Knowledge and Understanding: Develop and enhance computer skills in the Python language, including the writing of clear and reproducible Python codes.
Subject Knowledge and Understanding: Be able to use Python to process data and apply data-science methods.
Syllabus
The module has two distinct components: data science in R and Python.
Data science in R:
• Overview of R: data types and vectors;
• Programming in R: operators, conditional statements, loops, the apply family, user-defined functions, and scoping rules;
• Reading and writing data in R;
• Data visualization in R;
• Organising, merging, and managing data; data transformation using tidyverse;
• Data tidying, string operations, and working with unstructured data;
• Data extraction and acquisition in R: web scraping and querying databases;
• Simulations and econometrics in R;
• Using R for data analysis in economic research, with an emphasis on reproducibility.
Data science in Python:
• Introduction to Python: setting up the environment (Anaconda, Jupyter), data types, control flow, and writing functions;
• Reading and writing data in Python;
• Data analysis with pandas: DataFrames, indexing, filtering, and summary statistics;
• Advanced DataFrame operations: merging, reshaping, group-by aggregation, and handling time series;
• Data cleaning and preparation of economic datasets;
• Natural language processing in Python: text preprocessing, tokenisation, and basic linguistic features;
• Vector representations of text: bag-of-words, TF-IDF, and word embeddings;
• Applied NLP in economics.
Context
- Core Module
- L1I1 - Year 1
Assessment
- Assessment Method
- Centrally-timetabled examination (On-campus) (100%)
- Coursework Details
- Centrally-timetabled examination (On-campus) (100%)
- Exam Timing
- January