¹û¶³´«Ã½

Skip to main content Skip to navigation

EC9D8: Foundations of Data Science in Economics

  • Mateusz Stalinski

    Module Leader
30 CATS - Department of Economics

Introduction

Analyses in all fields of Economics nowadays make frequent use of large and detailed datasets ("big data"). The explosion in data access and availability opens many opportunities for applied research, as well as new challenges on how to handle, process, and extract meaningful conclusions from the data. The aim of the module is to introduce students to the R and Python programming languages and basic concepts of data science; and to provide and "hands-on" experience with economic data. The module lays the foundation to more advanced materials.

Principal Aims

The primary aim of the module is to introduce students to the analytical tools of data science provided by the R and Python computing languages. It builds computer coding skills from the basic principles, and covers methods to acquire, process and manipulate large volumes of data, often obtained from the web. Data visualisation methods are also presented.

By the end of the module, students should feel comfortable with collecting, visualizing and presenting

both structured and unstructured datasets. This should include using multiple programs/languages,

depending on their application of interest, including R and Python. They should be aware

of the advantages/disadvantages of each package and applications of each one.

Principal Learning Outcomes

Subject Knowledge and Understanding: Be able to process and work efficiently with large datasets.

Subject Knowledge and Understanding: Develop and enhance computer skills in the R language, including the writing of clear and reproducible R codes.

Subject Knowledge and Understanding: Be able to use R to process data and apply data-science methods.

Subject Knowledge and Understanding: Develop and enhance computer skills in the Python language, including the writing of clear and reproducible Python codes.

Subject Knowledge and Understanding: Be able to use Python to process data and apply data-science methods.

Syllabus

The module has two distinct components: data science in R and Python.

Data science in R:

• Overview of R: data types and vectors;

• Programming in R: operators, conditional statements, loops, the apply family, user-defined functions, and scoping rules;

• Reading and writing data in R;

• Data visualization in R;

• Organising, merging, and managing data; data transformation using tidyverse;

• Data tidying, string operations, and working with unstructured data;

• Data extraction and acquisition in R: web scraping and querying databases;

• Simulations and econometrics in R;

• Using R for data analysis in economic research, with an emphasis on reproducibility.

Data science in Python:

• Introduction to Python: setting up the environment (Anaconda, Jupyter), data types, control flow, and writing functions;

• Reading and writing data in Python;

• Data analysis with pandas: DataFrames, indexing, filtering, and summary statistics;

• Advanced DataFrame operations: merging, reshaping, group-by aggregation, and handling time series;

• Data cleaning and preparation of economic datasets;

• Natural language processing in Python: text preprocessing, tokenisation, and basic linguistic features;

• Vector representations of text: bag-of-words, TF-IDF, and word embeddings;

• Applied NLP in economics.

Context

Core Module
L1I1 - Year 1

Assessment

Assessment Method
Centrally-timetabled examination (On-campus) (100%)
Coursework Details
Centrally-timetabled examination (On-campus) (100%)
Exam Timing
January

Let us know you agree to cookies