probability theory is not just a problem of rolling dice in mathematical homework, but something that can actually predict the future and guide us to deal with real problems such as disaster warning and disease detection - it is the crystal ball in the hands of mathematicians. This article is translated from the article Mathematics and Crystal Balls6, written by Joseph Malkevitch, Honorary Professor of Mathematics and Computer Science, City University of New York, and will be divided into two push articles.
Written by | Joseph Malkevitch (Honor Professor of Mathematics and Computer Science, York College, City University of New York)
Compiled by | Shi Hao
People want to foresee the future, such as wanting to know what the weather will be tomorrow; wanting to know whether we have enough savings for our future retirement life; wanting to know how our friendship with friends will develop; or knowing what courses to learn in college can bring us happiness. In fact, using mathematical tools may be more effective in predicting the future than using crystal balls. We all like the time to be safe and avoid unpleasant times as much as possible. However, everything people do can have negative results, so we are always looking for ways to avoid these risks. And mathematics often helps us.
Crystal Ball丨Image Source: iStock
Trouble caused by disaster prediction
For those who want to live a "carefree" life, there are few places that will not be attacked by the "outside" dangerous. For example, in the United States, some areas are prone to tornadoes and heavy snow; other places may cause disasters such as floods, hurricanes, and earthquakes. In these areas that are affected by the "action of nature", it would be great if we can make disaster predictions so that there are no casualties and property losses are minimized. Today, some mathematical models have been developed to assist in the handling of various natural disasters. The most common prediction model we rely on is the increasingly accurate weekly weather forecast. These forecasts are given by scientists based on data from satellites, land-based monitoring and sensor systems. On the other hand, these predictions also rely on atmospheric models based on partial differential equation theory and numerical methods for solving these equations—the rapid advancement of computational power and theory makes these reports more reliable.
An interesting example happened in Italy in 2009. A group of Italian geologists and a government official came to trial for allegedly failing to give appropriate warnings to the 2009 Italian earthquake that killed 309 people. Geologists around the world are worried. They understand that although we have made great progress in the technology of trying to warn, we predict that only has the probability of occurrence rather than a certainty . After trial, seven people were convicted of manslaughter and sentenced to six years in prison. To the relief of scientists around the world, in 2014, the Court of Appeal released these geologists and commuted sentences to government officials. But the relatives of the deceased still condemn the government's act of eliminating themselves in court.
Lacquela earthquake, local government offices were also destroyed | Image source: wiki
When a big storm comes, if the weather forecaster fails to issue a serious enough warning of the potential danger, should they be blamed? Sometimes, due to forecast reasons, transportation systems that may be damaged by major storms are shut down first, which can cause huge logistical and economic problems for many people. If the storm does not come as expected, such things happen from time to time, it will be a little disappointing for many people. But on the other hand, when people may be rescued, some people do not respond strongly enough, which involves the issue in the "Italy Earthquake". Of course, compared with weather forecasts, earthquake prediction is too much behind.
Another example is infectious diseases.Some infectious diseases (such as influenza ) will become popular every year, and there are also periodic times that occur once every few years. For children, having whooping cough or measles may cause death; for the elderly, they do not know whether some of the vaccines they received when they were young are still effective. In addition, if the flu appears at this time, the elderly may be more severely affected. Because it may not be a big deal for young people to get influenza, but the elderly to get influenza will cause pneumonia and or suffer from other life-threatening diseases. So, should parents vaccinate their children? Should the elderly get the flu vaccine?
Although some people have allergic reactions, judging from the long-term vaccination history, the vaccine has greatly extended people's life span and improved the quality of life.
Risk behavior
Recently, the American Powerball (Power Ball) lottery launched an amazing prize of $1.6 billion! A major reason practitioners are confident in lottery and casino gambling is that mathematics tells them that as long as there are many consumers, these "industries" will flourish. If a person has a crystal ball that can help him choose the right lottery, he will make a fortune.
Whether as an individual or a group, when thinking about the future, people always have expectations for the future—sometimes it is beautiful, sometimes it seems less attractive. One of the contributions of mathematics to understanding what the future can bring to us is to bring the concept of "expected value (expected value)". When mathematicians used this term, they had an extremely precise definition in their minds, but there were many subtleties in this definition. One reason people are nervous about the future is that they are not sure what will happen in the future. The future involves randomness and probability (route, danomness, stochasticness, probability) . In order to understand the meaning of expected values, we must first talk about probability theory.
We often hear expressions about the future similar to the following:
The probability of raining is 70%;
The probability of another earthquake occurs at this location is one in one million;
throw a pair of uniform dice, and the probability that the sum of the points of the two dice is equal to 7 is 1/6;
What does this statement mean? To answer this question, we have to return to the two pillars of mathematics - basic mathematics and applied mathematics. Basic Mathematics establishes an ideological and conceptual system based on definitions and axioms (rules system) , and then derives mathematical facts and theorems from these constructs. Applied mathematics takes these maths and tries to use them to see the world. In the following text, I will try to tell it in a relatively informal way, trying to avoid "cumbersome" mathematical symbols and "formal" definitions.
First of all, probability applies to the domain of finite results, and can also be used on the domain of infinite results. To help understand this sentence, here are different examples to explain. For example, the definition domain of limited results:
Ms. Susan wants twins, and there are four possibilities for her child to be born in total - two boys were born one after another; two girls were born one after another; first a girl and then a boy; first a boy and then a girl. We only consider these four possible orders.
For infinite results domains, for example, we might catch a salmon that is spawning in a river in the western United States and weigh the salmon. The possible result of weighing is a real number between the weight range that the salmon can achieve—maybe one of infinite weights.
From a mathematical perspective, we can simulate these possibilities and form a set M by imagining the results of a certain "experiment" or actual observation. The set M can be finite or infinite.For each result m in the finite result set M, we will assign a real number to it, called the probability of the result m, which is called P(m). These real numbers cannot be allocated in a completely arbitrary way, they must follow specific attributes or axioms:
A: The values of P(m) range from 0 to 1, including 0 and 1.
B: The sum of all P(m) corresponding to m in M must be equal to 1.
Please note that if the probability of m occurring is P(m), then the probability of complementary event m’, that is, the probability that m will not occur is 1-P(m). That is to say, if the coin can be front or reverse, and the probability of the reverse is 2/5, then the probability of the front is 3/5.
Also, since we only have limited results, we don't need any ideas related to the limit (calculus) to do the calculations.
For set M with infinite results, we require that the two conditions listed above remain true. If there is a minimum value in P(m), for all results, their probability sum cannot be 1. Because no matter how small a finite number is, there will always be a sum greater than 1 if added in infinite times. Therefore, there are other subtle ways to deal with probability on infinite sets.
But it is worth noting that we are able to find an infinite set on which the probability of each individual event result can be non-zero. This is possible because there are infinite positive numbers, and the sum of the sequences is 1.
When we are using mathematics, we must extract mathematical concepts from a chaotic, undefined world of terms and axioms and explain their meanings.
If you give a data set, for example, your weight within 30 days (maybe measured at the same time every morning) . People tend to observe the fluctuations of these numbers—these numbers probably won’t be the same. If you want to understand the "rules" of these numbers, one way is to calculate some typical values, such as "average", which is a very attractive number. The average is usually obtained by adding up all the values and dividing them by the number of trials. The problem with
Using a single number to represent a huge data set is that sets of numbers expressing different things tend to have the same single number as their representative. For example, the average value of 5, 5, 5, 5, 5 is 5, and the average value of -3, -3, -3, 13, 13, 13, 13 is also 5. One of the early developments in the use of numbers in science and statistics was the recognition that measuring the same number “independently” multiple times may be more reliable than measuring only one number at a time. Due to the measuring device and the "artificial" process, measurements inevitably produce some errors in any case, but one can make the measurement as reliable as possible.
is similar to the mean of the random value, which is a quantity called "expected value". Suppose in a certain game, you have a 3/10 chance to win $3 and a 7/10 chance to win $4. By weighting the result with the probability of the result, you can see how much money you will win if you play such a game. In the above case, the expected value is that you will get $3 in the time of 3/10 and you will get $4 in the time of 7/10, so
expectation = 3(3/10)+ 4(7/10)= (9/10)+ (28/10)= 37/10 = 3.70.
If you need to pay $3.75 to play this game, you will lose 5 cents for every single time you play on average. When you win $3, you actually lose 75 cents; when you win $4, you earn 25 cents. But because the frequency of winning or losing is different, the probability of the result is different, and you will lose an average of 5 cents. Note that 3.70 is not the result of the game, nor is it the probability.
conditional probability
Sometimes, the implementation of the "experiment" will affect the probability of events occurring. There are two black balls and two white balls in the box of
. Consider the following two different plans::
Solution A: Stir the box and disrupt the ball.Select a ball from the box, then return the first ball, continue to stir, and remove the second ball.
Solution B: Stir the box and disrupt the ball. Select the first ball from the box and then the second ball.
There is no doubt that the probability of you getting two black balls depends on which solution you use. In Plan B, if the first one you draw is a white ball, it is impossible to get two black balls.
For Scheme A, you can only get two black balls when you draw the first black ball and the second time you draw it is also a black ball. Therefore, the probability of drawing BB (B means you draw the black ball) can be calculated by calculating P(BB)=(1/2)(1/2)=1/4. In the calculation of the probability of drawing two black balls in Scheme B, we need to analyze the situation:
The first ball drawn is black, and the second ball is also black. Therefore, the possibility that the first ball is black is 2/4=1/2. Now there are 3 balls left: two white balls and one black ball. The probability of the black ball being drawn is 1/3. With this in mind, we can see that the probability of two black balls being drawn is (1/2)(1/3)= 1/6.
This simple problem is related to the most basic but exquisite problem in probability theory, namely conditional probability. This concept can be traced back to the earliest days of studying randomness. If we use modern symbols to represent it, P(A|B) represents the probability of A occurring in the case where B occurs. For example, when we take two balls out of the box, given that the first ball is white, the probability that the second ball is black is 2/3. We can also regard P(A|B) as P(A∩B)/P(B) = P (the first ball is black, the second ball is black)/(P (the second ball is black)=(1/3)/(1/2)=2/3
How do you "define" or consider the value of P(X|Y)? In other words, what we are looking for is the probability that X occurs after Y occurs, P(X|Y) is the probability that X and Y occur simultaneously divided by the probability that Y occurs. Note that in this calculation, P(Y) is used as the denominator. For calculating P(Y|X), we calculate Y and X Probability of occurrence (same probability as X and Y occurs) , but we divide by P(X). We are looking for the effect of "the part of X occurs" on "the occurrence of Y and X occurs simultaneously".
Bayesian theorem
Many people will confuse P(A|B) and P(B|A) These two conditional probabilities are usually different. For example, if Event A means a drug test is positive, Event B means the patient has this disease. Then the probability that a patient with this disease is positive and the probability that a person who has this disease is positive is completely different.
medical tests may be very accurate, but when a disease is relatively rare, just because the test result is positive, this does not mean that the person must have this disease. An example will help reveal the relevant problems.
Assuming that a disease (D) is very rare, the incidence rate of the general population is 0.005, indicating that 5 out of 1,000 people suffer from this disease. Assuming the diagnosis of disease D The test is a blood test. When a person really has disease D, the probability of returning a positive index of disease D is 0.99. But the bad thing is that when a person does not have disease D, the test may also have a positive result (i.e., diseased) , with a probability of 0.05, which is relatively low. Note that 0.99 and 0.05 cannot be added because these two are not complementary events.
gives three different numbers here, and we will use these numbers to deduce some other numbers through some "laws" of probability. Let us introduce some symbols to clarify our ideas. Symbols have both advantages and disadvantages. These symbols can make the concept clearer because there are many similar but different meanings. In order to distinguish them, a large number of symbols must be used.
T represents an event where the test result is positive, regardless of whether the person is sick or not;
P(D) represents the probability of a person getting sick;
P(T|D) represents the probability of a person being positive when he is sick;
P(T|D') represents the probability of a person being positive even if he is not sick;
Based on the above information, we can write down the values of these three different probabilities:
P (D) = 0.005
P (T | D) = 0.99
P(T | D ')= 0.05
When the test result is positive, the patient wants to know the chance of getting sick, but please note that the answer is not one of the numbers given above! However, we can infer this number through probability theory.
In addition to other probabilistic tools, we will also use a "fact" called Bayesian theorem or Bayesian formula . This result was proposed by Thomas Bayes (Thomas Bayes, 1702-1761) , but it was not published during his lifetime. Today, Bayesian is famous for its statistical applications because of the statistical applications of terms such as " Bayesian inference) " and "Bayesian statistics (Bayesian statistics) ".
Thomas Bayes
Bayes's results are shown in the neon light below:
Bayes Theorem丨Image source: wiki
Although we may only know P(B|A), this result allows us to calculate other conditional probabilities related to the problem P(B|A).
Return to the above diagnosis situation, let's see what we can infer.
First use the concept of complementary events, and the sum of the probability of events and complementary events is 1. We have:
P(D') = 1-P(D) = 1-.0005 = 0.995 (the probability that someone will not get sick)
P(T'|D) = 1-P(T|D) = 1-0.99 = 0.01 (the probability that someone is sick but has not detected a positive)
P(T'|D') = 1-P(T|D') = 1-P(T|D') = 1-0.05 = 0.95 (The probability that someone is not sick or tested positive)
Now, let's look at some other probabilities worth paying attention to. For example, the probability of positive feedback whether you have a disease or not, and the probability of negative feedback whether you have a disease or not. There are two ways to get positive feedback. One is to get a positive test result when you get a disease; the other is to get a positive test result when you don’t get a disease. We can use symbols to represent:
P(T) = P(T|D)P(D) + P(T|D')P(D') = (0.99)(0.005) + (0.05)(0.995) = 0.00495 +0.04975 = 0.0547
P(T ')= P(T ' | D)P(D)+ P(T ' | D ')P(D ')=(0. 01)(0.005)+(0.95)(0.995)=0.9453 (Here we add the probability that a person is sick but has not tested positive, and the probability that he has not been diagnosed or has not been diagnosed.)
We need to check the correctness of the calculation. Logically speaking, 0.0547+0.9453 should add up to 1, and it does add up to 1! Maybe these numbers look a bit surprising - the probability of getting a positive test result is quite small, but this just reflects that few people suffer from this disease.
However, so far, we have not gotten the number we are really interested in - if a person tests positive, what is the probability of his illness? Do you need to feel scared if a person is tested positive? This is where we need to use Bayesian results.
P(D|T) = (P(T|D))(P(D))/P(T) = (0.99)(0.005)/(0.0547) = 0.0904936 ≈0.0905
Therefore, even if the probability of detecting this disease in the test is high, only a small number of people who test positive do suffer from the disease. This result is because this disease is very rare.Usually, suspected patients will do another independent test to see if they are really sick so as to avoid unnecessary treatment. The result of
Bayesian can also be used to obtain three other conditional probabilities, two of which can also be obtained by using the fact that "the sum of the probability of an event and its complementary event is 1".
P(D ' | T)= 0.9095 (probability of positive test but not ill)
P(D ' | T ')= 0.99995 (probability of negative test but not ill)
P(D | T ')= 0.00005 (Probability of negative test and disease)
The last number can be calculated using Bayesian results, as shown below:
P(D | T ')=(P(T ' | D))(P(D))/ P(T ')= 0.01(0.005)/ 0.9453 = 0.00005
Yes, although the symbols and calculations here are complicated, these can help the patient and his doctor correctly view what it means to get a positive result in rare disease tests.
(To be continued)
References
[1] Beniston, M, From Turbulence to Climate: Numerical Investigations of the Atmosphere with a Hierarchy of Models, Springer, Berlin, 1998.
[2] Daston, L., Classical Probability During the Enlightenment, Princeton U. Press, Princeton, 1988.
[3] Falk, R., and M. Bar-Hillel, Probabilistic dependence between events. The Two-Year College Mathematics Journal. 14 (1983) 240-7.
[4] Falk, R., Conditional probabilities: insights and difficulties. In Proceedings of the Second International Conference on Teaching Statistics 1986, pp 292-297.
[5] Falk, R., Misconceptions of statistical significance. Journal of structural learning. March, 1986.
[6] Gelman, A. and J. Carlin, H. Stern, D. Rubin, Bayesian Data Analysis (2nd edition), Chapman & Hall/CRC, Philadelphia, 2003
[7] Hacking, I., The Emergence of Probability, Cambridge U. Press, New York, 2006.
[8] Hald, A., A History of Mathematical Statistics from 1750 to 1930, Wiley, New York, 1998.
[9] Hald, A., A History of Probability and Statistics and Their Applications Before 1750., Wiley, New York, 2003.
[10] Mayo, D., Experimental Knowledge, University of Chicago Press, Chicago, 1996.
[11] Mayo, D., Error and Inference: Recent Exchanges on Experimental Reasoning, Reliability, and the Objectivity and Rationality of Science, Cambridge University Press, New York, 2010.
[12] Roulstone, I. and J. Norbury, Invisible in the Storm: the role of mathematics in understanding weather, Princeton U. Press, Princeton, 2013.
[13] Stigler, S., The History of Statistics: The Measurement of Uncertainty Before 1900, Harvard U. Press, Cambridge, 1990.
[14] van Plato, J., Creating Modern Probability: Its Mathematics, Physics and Philosophy in Historical Perspective, Cambridge U. Press, New York, 1994.