Магистратура
2026/2027



Теория вероятностей и математическая статистика
Статус:
Курс по выбору (Науки о данных (Data Science))
Где читается:
Факультет компьютерных наук
Когда читается:
1-й курс, 1 модуль
Охват аудитории:
для всех кампусов НИУ ВШЭ
Язык:
английский
Кредиты:
3
Контактные часы:
28
Course Syllabus
Abstract
The overwhelming majority of courses on data-mining and artificial intelligence implies a strong background in probability and statistics. The goal of this course is to provide those students who are not so keen on the subject matter with its fundamentals.
Learning Objectives
- To introduce the theoretical foundations of Probability theory.
- To introduce the theoretical foundations of Mathematical statistics.
- To provide the students with practical skills of modelling real-world in the framework of probability and statistics.
Expected Learning Outcomes
- Student will be able to derive Markov’s and Chebyshev’s inequalities and explain with numbers the degree of their conservativeness
- Student will be able to apply the Chernoff method: build a tail bound by optimizing over the parameter of the moment generating function (MGF)
- Student will be able to give the definition of a sub-Gaussian quantity and prove Hoeffding’s lemma for a bounded random variable
- Student will be able to choose between Hoeffding and Bernstein based on the variance and confirm the choice with a numerical comparison on rare events
- Student will be able to compute the sample size n guaranteeing accuracy ε with confidence 1−δ, and construct a confidence interval for the mean/proportion
- We'll learn how to formalize an A/B test as a randomized experiment: choose the unit of randomization and the metric, state H0 and H1, derive the z/t-statistic of the difference.
- We'll learn how to link the four knobs of an experiment — the level α, the type II error β, the effect size, and the sample size n— and derive the sample size formula for the difference of means and the difference of proportions.
- We'll learn how to correctly estimate ratio metrics (CTR, etc.) when the unit of analysis is not equal to the unit of randomization: the delta method for the variance of a ratio and the cluster bootstrap.
- We'll learn how to design an OEC and guardrail metrics so that the test measures what the business needs, not what is easy to compute
- We will learn how to formulate the peeking problem rigorously and understand why a fixed α-threshold breaks under optional stopping
- We'll learn how to build confidence sequences (anytime-valid CI) via the mixture method and compare their width with a fixed confidence interval
- Formalize the exploration–exploitation dilemma and define regret as a metric of strategy quality
- Understand the idea of contextual bandits (LinUCB) and Bayesian optimization of hyper- parameters (GP + EI/UCB).
- Recognize the risk of biased estimates under adaptive data collection and adaptive stop- ping
- Explain by example why correlation does not imply causation (confounding, reverse causa- tion, selection, collider), and pose a causal question through Rubin’s potential outcomes Y(1),Y(0);
- Prove unbiasedness of the difference in means in a randomized experiment and build a confidence interval for the effect;
- Recognize situations for instrumental variables and difference-in-differences and under- stand why a collider must not be controlled for.
- Formulate OPE as a counterfactual problem and write policy value through bandit feed- back logs; understand the role of the overlap assumption.
- Build the direct method and doubly robust; prove double robustness and understand the reward model as a control variate that reduces variance.
- Apply OPE as an offline screening of dozens of policies before an online A/B and see how off-policy learning grows out of the same mathematics
Course Contents
- Сoncentration of measure
- A/B Tests
- Anytime-valid statistics
- Multi-Armed Bandits
- Сausal inference
- Off-policy evaluation
Bibliography
Recommended Core Bibliography
- 9780262257053 - Sutton, Richard S.; Barto, Andrew G. - Reinforcement Learning : An Introduction - 1998 - A Bradford Book - http://search.ebscohost.com/login.aspx?direct=true&db=nlebk&AN=1094 - nlebk - 1094
- 9781838820046 - Lapan, Maxim - Deep Reinforcement Learning Hands-On : Apply Modern RL Methods to Practical Problems of Chatbots, Robotics, Discrete Optimization, Web Automation, and More, 2nd Edition - 2020 - Packt Publishing - http://search.ebscohost.com/login.aspx?direct=true&db=nlebk&AN=2366458 - nlebk - 2366458
- A basic course in measure and probability : theory for applications, Leadbetter, R., 2014
- Burk, S. (2006). A Better Statistical Method for A/B Testing in Marketing Campaigns. Marketing Bulletin, 17, 1.
- Imbens, G. W., & Rubin, D. B. (2015). Causal Inference for Statistics, Social, and Biomedical Sciences. Cambridge University Press. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsrep&AN=edsrep.b.cup.cbooks.9780521885881
- Introduction to Statistics and Data Analysis, With Exercises, Solutions and Applications in R, Christian Heumann, Michael Schomaker, Shalabh, Springer Nature Switzerland AG 2022, 978-3-031-11833-3, published: 30 January 2023
- Kohavi, R., & Thomke, S. (2017). The Surprising Power of Online Experiments: Getting the Most out of A/B and Other Controlled Tests. Harvard Business Review, 95(5), 74–82.
- Larsen, R. J., & Marx, M. L. (2015). An introduction to mathematical statistics and its applications. Slovenia, Europe: Prentice Hall. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.19D77756
- Mao, H., Venkatakrishnan, S. B., Schwarzkopf, M., & Alizadeh, M. (2018). Variance Reduction for Reinforcement Learning in Input-Driven Environments. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsbas&AN=edsbas.BAB65515
- Rohatgi, V. K., & Saleh, A. K. M. E. (2015). An Introduction to Probability and Statistics (Vol. 3rd edition). Hoboken, New Jersey: Wiley. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=1050364
- Shepherd, B. E., Jarrett, R., & Fu, L. (2016). Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction. Biometrics, 72(4), 1387–1388. https://doi.org/10.1111/biom.12615
- Wiering, M., & Otterlo, M. van. (2012). Reinforcement Learning : State-of-the-Art. Berlin: Springer. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=537744