Tarastats Statistical Consultancy https://www.tarastats.com/ Tue, 21 Apr 2026 02:59:21 +0000 en-GB hourly 1 https://wordpress.org/?v=7.1.2 https://www.tarastats.com/wp-content/uploads/2018/04/icon-big.png Tarastats Statistical Consultancy https://www.tarastats.com/ 32 32 A New Approach To Statistical Thinking https://www.tarastats.com/new-approach-statistical-thinking/ Sat, 15 Oct 2022 12:00:27 +0000 https://www.tarastats.com/?p=7922 Statistical thinking requires a new approach due to recent developments of the modern statistical science. This new approach puts causal thinking at the heart of the key statistical thinking concepts, which reflects new developments of modern statistical science in the field of causal inference. This new approach is based on the key concepts of modern […]

The post A New Approach To Statistical Thinking appeared first on Tarastats Statistical Consultancy.

]]>

Statistical thinking requires a new approach due to recent developments of the modern statistical science.

This new approach puts causal thinking at the heart of the key statistical thinking concepts, which reflects new developments of modern statistical science in the field of causal inference.

This new approach is based on the key concepts of modern statistical science — understanding modern sampling theories, missing data mechanisms and their impact on usefulness of data interpretations obtained from Descriptive Statistics.

Furthermore, understanding of sampling theories and missing data mechanisms is critical for performing Inferential Statistics in a scientifically objective way.

This new approach to statistical thinking is based on advances of modern statistical science and is accordingly called modern statistical thinking.

This article has not been released yet. Be the first to know when it comes out.

Leave your email below.

The post A New Approach To Statistical Thinking appeared first on Tarastats Statistical Consultancy.

]]>
Do you wonder how statistics lies? https://www.tarastats.com/how-statistics-lies/ Fri, 14 Oct 2022 20:02:27 +0000 https://www.tarastats.com/?p=7853 Statistics lies in presence of ignorance. By definition, ignorance means lack of knowledge, understanding, or information about something. Statistics lies when a person presenting statistical data lacks knowledge, understanding, or information about statistical-methodological techniques which enables us to analyse data in a scientifically objective way. Although we live in an evidence-based world, our societies have not […]

The post Do you wonder how statistics lies? appeared first on Tarastats Statistical Consultancy.

]]>

Statistics lies in presence of ignorance.

ignorance

By definition, ignorance means lack of knowledge, understanding, or information about something.

Statistics lies when a person presenting statistical data lacks knowledge, understanding, or information about statistical-methodological techniques which enables us to analyse data in a scientifically objective way.

Although we live in an evidence-based world, our societies have not done enough to enable citizens development of statistical literacy – the key skill of the digital era.

In the past, being able to read and write enabled citizens to have more prosperous lives. Today’s data-driven world requires statistical literacy.

Does this mean that if a presentation of statistics is done by those who have studied statistics we do not need to worry about statistics lying?

No. One can study statistics with a focus on developing new theories (mathematical statistics) or learning how to apply statistical methodologies to real-world data (applied statistics). Applied statisticians are much more familiar with modern statistical thinking — the key skill to analyse data in a scientifically objective way, while mathematical statisticians are frequently not equipped with such understandings. 

Why is it so?

While things work beautifully in a theory, this is not the case with real world data. We are not living in a linear world, yet, ‘linearity’ is the most common assumption that needs to be satisfied when analysing data. 

Not living in a linear world means that collected data are filled with non-linearities. To correctly address complications arising due to non-linearities, a high level of modern statistical thinking is required. 

The modern statistical thinking connects understandings of the key concepts of modern statistical science with causal thinking. Such connection is of immense importance because we live in a cause-and-effect world — the world where most questions of interest are causal in their nature.

What is new?

In the early 20th century, a theory in physics called quantum mechanics was developed. Quantum mechanics enables calculations of properties and behaviours of physical systems.

The quantum mechanics showed us that our world is all about cause-and-effect relationships — an action (a cause) manipulates behaviour of an object or a subject, and the effect is the impact of the action (of the cause/manipulation) that we observe.

Connecting this information with the fact that most questions of interest are causal in their nature leads us to the following conclusion: those who analyse data and present its outcomes must understand how to analyse causal relationships in a scientifically objective way.

Why are modern statistical thinking and causal analysis still absent or underrepresented from most statistics courses?

The science of the 20th century was not responsible only for development of quantum mechanics, but also for development of statistical methods that enable us to analyse causal relationships in a scientifically objective way.

Briefly said, it started with William S. Gosset and Student t-test in the beginning of the 20th century, continued with development of randomisation machinery by Sir Ronald Fisher, and development of modern sampling approaches and confidence interval by Jerzy Neyman.

In the second half of the 20th century Donald B. Rubin developed a causal model – broadly known as the Rubin Causal Model (Holland 1986). This causal model is a foundation for cause-and-effect studies in varieties of fields, from medicine and public health to economics, environment, biology, law and business. The model enables analysis of causal relationships also with data collected outside of an experimental framework, a so called observational data.

Prior Rubin’s work, analysis of causal relationships outside of experimental data was ‘strictly forbidden’. This means, that it has been less than half of a century since major contributions in modern statistical science.

It takes time for researchers to catch-up on these developments and to change curriculum of applied statistics courses. Most statistical curricula are built on a long tradition that is rooted in classical methods like hypothesis testing, linear regression, and analysis of variance (the most abused statistical method). These topics became standard more than 70 years ago and are still seen as an essential foundation. Updating curricula is often slow, especially in large institutions.

Another reason is that analysis of causal relationships is conceptually harder — it requires students to think beyond formulas, develop understanding about causal assumptions, causal designs, and causal interpretations. This is a bigger cognitive leap than teaching well-defined procedures like computing a p-value.

Modern statistical approaches are more abstract and less tidy from a theoretical standpoint while classical methods often rely on clean, linear assumptions that are easier to teach and test.

How the curriculum of applied statistics education should change?

The most important thing is that the curricula shifts from technical and theoretical details to modern statistical thinking and the use of methods and techniques to derive data-insights in a scientifically objective way.

Students should learn basics about causality in statistics, e.g., how statistical science defines the cause, what is a causal design and how to design an objective causal design.

Is there a course that consists of such applied statistics curriculum?

Yes, our founder Dr. Ana Kolar has been developing such curricula for the past 10 years. Some of her in-person and online courses can be attended at University of Helsinki. For those interested in online learning, you can find available courses here.

Dr. Kolar uses experiential learning approach to teaching which enables deep learning. She believes that students should have the opportunity to deepen their knowledge during the learning process, because this is the only path to knowledge. She is rated as an excellent teacher and she is fun too!

Can I learn about how to analyse causal relationships by studying books and articles?

Yes. There are plenty of books, articles and online lectures that one can learn from, but without a proper guidance, it will take years or even a decade to grasp foundations of causal inference in its completeness.

The causal inference is one of the most complex topics in statistics. It is a highly demanding field of study that requires a heavy use of ‘thinking’ and a holistic approach when developing understanding about how to develop a causal design and satisfy causal assumptions. The use of modern statistical thinking plays a significant role in this process. 

The post Do you wonder how statistics lies? appeared first on Tarastats Statistical Consultancy.

]]>
Data science without causal inference is like a fish without water https://www.tarastats.com/data-science-needs-causal-inference/ Sun, 15 May 2016 12:18:30 +0000 http://new.tarastats.com/?p=383 Learning about foundational concepts of causal inference is crucial for data analytics and AI because most questions of interest are causal in their nature. Whether we perform impact evaluations, A/B testing, quality control or clinical trials, causal inference is the method of choice. Causal inference is one of the most complex data inference methods, but […]

The post Data science without causal inference is like a fish without water appeared first on Tarastats Statistical Consultancy.

]]>

Learning about foundational concepts of causal inference is crucial for data analytics and AI because most questions of interest are causal in their nature.

Whether we perform impact evaluations, A/B testing, quality control or clinical trials, causal inference is the method of choice.

Causal inference is one of the most complex data inference methods, but with high rewards in terms of provided insights. In its essence, the theory and methods behind causal inference enable us to analyse causal relationships and thus fully unlock the value that data holds. 

Being familiar with causal inference methods and techniques also equips us with problem solving skills that are crucial for analysing data in a scientifically objective way.

The main problem of causal inference is missing data and how to handle it. Because incomplete data is a typical problem in data analytics (often associated with biased results), it is important to become familiar with causal inference methods and techniques — the knowledge that enhances our skills for dealing effectively with incomplete data.

 

Understanding incomplete data is the path to handle missing data.

For a long time, causal inference was allowed to be performed only within a randomised  experimental framework. Due to recent developments in statistical-methodological science, we are now able to perform causal inference also with observational data, i.e., a non-randomised experimental data.

In recent years, causal inference has become one of the most popular methods for analysing data. However, many still struggle with complexities of causal inference conceptual framework.

The conceptual framework of causal inference provides foundational knowledge about the required causal reasoning as also the use of modern statistical thinking when designing studies and analysing data in a causal effect fashion. Such foundational knowledge is critical to be able to analyse causal relationships in a scientifically objective way.

Causal inference also provides us with understanding of the impact that study designs have on trustworthiness of obtained data insights as also on capacity to unlock the value from data.

Some examples of questions that causal inference can answer:

Not all causal questions can be answered

For example, do black students perform better in education attainment than white or Hispanic students?

In this example, the race is considered to be the cause. However, because we cannot manipulate such a cause, meaning that there is no simple intervention with which we could transform a white person into a black person, results of such study cannot be called causal effects, but rather associations which are conditional on a set of covariates used in comparative analysis.

Another example of a cause that cannot be manipulated is sex. We cannot give a magic pill to an individual and transform him/her into an opposite sex.

Within the randomised experimental framework we use intervention to manipulate units of one group in comparison. For example, we apply intervention to units of one group (usually called a treated group), while not applying it to units of another group, i.e., control group. The intervention is well-formulated, known cause.

The known cause is the cause that can be manipulated. 

When we can define the known cause, we are able to use causal inference methods and techniques to perform causal effect studies also with observational data. 

However, we must make sure that when using observational data, we make all the effort to come up with comparable groups, meaning, to have two or more approximately identical groups of units with respect to important covariates, which can differ only with respect to applied intervention, i.e., the known cause.

Selection of covariates

A careful selection is of utmost importance, in order to be able to reconstruct observational data structure to mimic a data structure of a randomised experiment. Such reconstruction is a complex task, but in its essence it requires a reconstruction of an assignment mechanism of observational data in a way to mimic an assignment mechanism of randomised experimental data.

What is Assignment Mechanism?

In a two group experimental randomised design, units are assigned to either Group 1 or Group 2, popularly called a treated and a control group. The mechanism which assigns units randomly is called an assignment mechanism. 

Because with observational data such assignment mechanism either does not exist or it is broken, it is important to reconstruct it in a way to mimic an assignment mechanism of the randomised experiment.

The process of reconstructing the assignment mechanism can be in many ways considered an art work. Yes, science requires art! However, this ‘art’ requires from us to be well-familiar with the necessary causal inference assumptions and ways to satisfy them.

Causal Inference without assumptions is mission impossible

There is a set of causal assumptions that are required to be satisfied if we want to obtain trustworthy conclusions of causal effect estimates. Justifying causal assumptions is challenging. It requires creative thinking, modern statistical thinking and understanding about the science of causal thinking.

Understanding assumptions and how to justify them is of great importance when designing causal inference studies because effectiveness of causal designs depends on capacity to justify the required assumptions. 

The more effective the causal design is, the better we can justify required assumptions. The importance of a good design is of such that “Sometimes the design effort can be so extensive that a description of it, with no analyses of any outcome data, can be itself publishable” – Donald B. Rubin (2008) For Objective Causal Inference, design trumps analysis. The Annals of Applied Statistics.

To be able to design causal inference studies effectively, it is important to get familiar with conceptual foundations of causal inference. Causal inference is not an algorithm and neither an equation, but a methodological and analytical approach for analysing causal relationships that requires heavy use of ‘human-mind’ software. Learn more about causal inference’s foundations here.

The post Data science without causal inference is like a fish without water appeared first on Tarastats Statistical Consultancy.

]]>