◇ NodrizaDesclasificados

AAWSAP DIRD An Introduction to the Statistical Drake Equation March 11 2010

Departamento de Guerra (EE.UU.) · 2010 · Documento · Release 06
⚠ Texto extraído por OCR de la fuente oficial — puede contener errores de reconocimiento. El documento original es la autoridad.
                    UNCLASSIFIED/ /POil OFFl@IAI:: WSli 0Nk¥


                         Defense
                         Intelligence
                         Reference
                         Document
                         Acquisition Threat Support
11 March 20 10

!COD : 1 December 2009




                         An Introduction to the
                         Statistical Drake Equation




                    UNCLASSIFIED/ /FOA OFFICIO L: 1PiF ON! Y
                   UNCLASSIFIED//P81t 8FFIEIAL Y&E &ttbl/


An Introduction to the Statistical Drake Equation




Prepared by:

Acquisition Support Division (DW0-3)
Defense Warning Office
Directorate for Analysis
Defense Intelligence Agency

Author:



AAP Person 80

Administrative Note

COPYRIGHT WARNING: Further dissemination of the photographs in this publication is not authorized .




This product is one in a series of advanced technology reports produced in FY 2009
under the Defense Intelligence Agency, Defense Warning Office's Advanced Aerospace
Weapon System Applications (AAWSA) Program. Comments or questions pertaining to
this document should be addressed t o !AAP Person 1              ~ AAWSA Prog ram
Manager, Defense Intelligence Agency, ATTN: CLAR/DWO-3, Bldg 6000, Washington,
DC 20340-5100.



ii
                   UNCLASSIFIED//F8R 8FFl61tlib Wlilli Qtlb¥
                          UNCLASSIFIED/ /POil OFFl@IAL WSli 0Nk¥



Contents
1. Introduction .......................................................................................................iv

2. The Key Question: How Far are They ? .............................................................. 4

3. Computing N By Virtue of the Drake Equation (1961) ........................................ 7

4. The Drake Equation is Over-Simplified ............................................................. 10

5. The Statistical Drake Equation ......................................................................... 11

6. Solving the Statistical Drake Equation By Virtue of the Central Limit Theorem
      (CLT) of Statistics .......................................................... ,.................................... 13

7. An Example Explaining the Statistical Drake Equation ..................................... 15

8. Finding the Probability Distribution of the Et-Distance By Virtue of the Statistical
      Drake Equation ................................................................................................. 18

9. The "Data Enrichment Principle" as the Best CLT Consequence Upon the
      Statistical Drake Equation (Any Number of Factors Allowed) ........................... 23

10. Conclusions .................................................................................................... 23

Appendix A: Proof of Shannon's 1948 Theorem Stating That the Uniform
                     Distribution is the "Most Uncertain" One Over a Finite Range of
                     Values ............................................................................................. 25

Appendix B: Original Text of the Author's Paper #IAC-08-A4.1.4 Entitled the
                     Statistical Drake Equation ............................................................... 28

References ........................................................................................................... 55




iii
                          UNCLASSIFIED/ fEOR OFFICIO! 11SF ON! X
                 UNCLASSIFIED/ /POil OFFl@IAI:: WSli 0Nk¥


An Introduction to the Statistical Drake Equation
1. Introduction
SETI (an acronym for "Search for Extraterrestrial Intelligence") is a relatively
new branch of scientific research, having begun only in 1959. Its goal is to
ascertain whether alien civilizations exist in the universe, how far from us
they exist, and possibly how much more advanced than us they may be.

As of 2009, the only physical tools we know that could help us get in touch
with aliens are the electromagnetic waves an alien civilization could emit and
we could detect. This forces us to use the largest radiotelescopes on Earth for
SETI research, because the higher our collecting area of electromagnetic
radiation is, the higher our sensitivity is (that is, the farther in space we can
probe). Yet, even by using the largest radiotelescopes on Earth (the 310-meter
dish at Arecibo, for instance), we cannot search for aliens beyond, say, a few
hundred light years away. This is a very, very small amount of space around us
within our galaxy, the Milky Way, that is about 100,000 light years in diameter.
Thus, current SETI can cover only a very tiny fraction of the galaxy, and it is
not surprising that in the past 50 years of SETI searches, NO extraterrestrial
civilization was discovered. Quite simply, we did not get far enough!

This demands the construction of much more powerful and radically new
radiotelescopes. Rather than big and heavy metal dishes, whose mechanical
problems hamper SETI research too much, we are now turning to "software
radiotelescopes," where a large number of small dishes (ATA = Allen
Telescope Array, and ALMA= Atacama Large Millimeter/submillimeter Array)
or even just of simple dipoles (LOFAR = Low Frequency Array) using state-of­
the- art electronics and very- high-speed computing can outperform the
classical radiotelescopes in many regards. The final dream in this field is the
SKA ( = Square Kilometer Array), currently being designed and expected to be
completed around 2020.

2. The Key Question: How Far are They?
But still, the key question remains: how far are they?

Or, more correctly, how far do we expect the NEAREST extraterrestrial civilization to be
from t he Solar System in the galaxy?

This question was first faced in a scientific manner back in 1961 by the same scientist
who also was t he first experimental SETI rad io astronomer ever: t he American, Fra nk
Donald Drake (born 1930). He first considered the shape and size of the galaxy where
we are living: the Milky Way. This is a spiral ga laxy measuring some 100,000 light
years in diameter and some 16,000 light years in thickness of the Ga lactic Disk at half­
way from its center. That is:

The diameter of the galaxy is (about) 100,000 light years, (abbreviated ly) i.e., its
radius, R Gala.1y , is about 50,000 ly.



iv
                 UNCLA SSI FIED/ /FOR OFFI&IJ.k Wili Ql'II.¥
                   UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

The thickness of the Galactic Disk at half-way from its center, h a " 1t'-'Y ' is about 16,000 ly.

The volume of the galaxy may then be approximated as the volume of the
corresponding cylinder, i.e.

                                                                                                 (1)

Now consider the sphere around us having a radius r. The volume of such a sphere is

                                                       4 ( Ef Distance ) 3                      (2)
                                  V o ur _ Sphere =
                                                       3n          -    2


In the last equation, we had to divide the distance "ET_Distance" between ourselves
and the nearest ET civilization by 2 because we are now going to make the
unwarranted assumption that all ET civilizations are equally spaced from each
other in the galaxy! This is a crazy assumption, clearly, and should be replaced by
more scientifically-grounded assumptions as soon as we know more about our Galactic
Neighborhood. At the moment, however, this is the best guess that we can make, and
so we shall take it for granted, although we are aware that this is a weak point in the
reasoning.

Furthermore, let us denote by N the total number of civilizations now living in the
galaxy, including ourselves. Of course, this number N is unknown. We only know that
N ~ 1 since one civilization does at least exist!

Having thus assumed that ET civilizations are UNIFORMLY SPACED IN THE GALAXY, we
can then write down the proportion:

                                          V a a/a.,y      V o ,, r _ Spher e
                                          -- = -~-                                              (3)
                                              N

That is, upon replacing both (1) and (2) into (3):

                                                                                      3
                                         2               -4 i'l" ( _Ef- _
                                                                        Dis __
                                                                            lance )
                                    i'l" R Ga /axy h =   3                  2
                                                                                                (4)
                                          N                             1


The last equation contains two unknowns: N and ET_Distance, and so we don't know
which one it is better to solve for.

However, we may suppose that, by resorting to the (rather uncertain) knowledge that
we have about the Evolution of the galaxy through the last 10 billion years or so, we
might somehow compute an approximate value for N.

Then, we may solve (4) for ET_ Distance thus obtaining the (AVERAGE) DISTANCE
BETWEEN ANY PAIR OF NEIGHBORING CIVILIZATIONS IN THE GALAXY (DISTANCE
LAW)




5
                   UNCLASSIFIED/ ;'FQA QFFICiIPL. !Iii 011! Y
                  UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥


                                                       3 6 R2      h     C
                                Ef_ ff,stance (N) =        VN
                                                            Ca!a,y                         (5)
                                                                        VN
where the posit ive constant C is defined by

                              C = V6 R ~ala.,y h ca /axy ,., 28845 light years .           (6)

Equations (5) and (6) are the starting point to understand t he orig in of the Drake
equation that we discuss in detail in Section 3 of th is paper.

Let us just complete this section by pointing out three different numerical cases of the
distance law (5):

•   We know that we exist, so N may not be smaller tha n 1, i.e., N ~ 1. Suppose then
    that we are alone in the galaxy, i.e., that N=l. Then the distance law (5) yields as
    distance to the nearest civilization from us just the constant C, i.e., 28,845 light
    years. Th is is about the distance in between ourselves and the center of the galaxy
    (i. e. the Galactic Bulge) . Thus, this result seems to suggest that, if we do not find
    any extraterrestrial civilization around us in these outskirts of the galaxy where we
    live, we should look around the Galactic Center first. And this is indeed what is
    happening, i.e., many SETI searches are actually point ing the antennas towards the
    Galactic Center, looking for beacons (see, for instance ref. [1]).

•   Suppose next that N=l000, i.e. there are about a thousand extraterrestrial
    communicating civilizations in the whole galaxy right now. Then the distance law (5)
    yields an average distance of 2,885 light yea rs. This is a distance that most
    radiotelescopes in Earth may not reach for SETI searches right now: hence the need
    to build larger radiotelescopes, like ALMA, LOFAR and the SKA.

•   Suppose finally that N=l00000O, i.e., there are a million communicating civilizations
    now in the galaxy. Then the dist ance law (5) yields an average dista nce of 288 light
    years. Th is is with in the (upper) range of distances that our current rad iotelescopes
    may reach for SETI searches, and that justifies all SETI searches that have been
    done so far in t he first fifty years of SETI (1960-2010).

In conclusion, interpolating the above three special cases of N, we may say that the
distance law (5) yields t he following key diagram of the average ET distance vs. the
assumed number of communicating civilizations, N, in the galaxy right now (Figure 1):




6
                  UNCLASSIFIED/ J'FOA. OFlilCI0 I. P!SE ON! X
                              UNCLASSIFIED/ /FOR OFFIEiIAk Wlii QPU,¥


                      Av era ge DIST A CE o f the nea res t ET c iviliza tion vs . th e ASSUM ED NUMBffi of ET c ivilizations in th e Gah
            200
    Cl)
    0:::
    <(
    UJ
    >­      175
    !­
    :I:
    0
    :i      150
    .!::
     !'.l
    g
            125
     ~
     !'l
     =
            100
                      l


            750
                       \
            500



            250
                           ' "--             ~




              0
                  0          I 00000      200000     300000    400000     500000     600000     700000     800000    900000    I000000
                                       ASS MED NUMB ER of civiliz ations in the Galaxy (that is, Nin the Drake equation)

Figure 1. DISTANCE LAW; i.e., the Average Distance (plot along the ver-1:ical axis in light years) Versus
the NUMBER of Communicating Civilizations ASSUMED to Exist in the Galaxy Right Now


3. Computing N By Virtue of the Drake Equation (1961)
In the previous section, the problem of finding how close the nearest ET civilization may
be was "solved" by reducing it to the computation of N, the total number of
extraterrestrial civilizations now existing in this galaxy. In this section the famous
Drake equation is described, that was proposed back in 1961 by Frank Dona ld Drake
(born 1930) to estimate the numerical value of N. We believe that no better
introductory description of the Drake equations exists other than the one given by Carl
Sagan in his 1983 book "Cosmos" (ref. [2]), in its turn based on the famous TV series
"Cosmos." So, in this paragraph we report Carl Sagan's description of the Drake
equation unabridged.

"But is there anyone out there to talk to? With a third or a half a trillion stars in our
Milky Way galaxy alone, could ours be the only one accompanied by an inhabited
planet? How much more likely it is that technical civilizations are a cosm ic
commonplace, that the galaxy is pulsing and humming with advanced societies, and,
therefore, that the nearest such culture is not so very far away - perhaps transmitting
from antennas established on a planet of a naked-eye star just next door. Perhaps
when we look up at the sky at nig ht, near one of those faint pinpoints of light is a world
on which someone quite different from us is then glancing id ly at a star we call the Sun
and entertaining, for just a moment, an outrageous speculation.



7
                              UNCLASSIFIED//FOR OFFIEiIAk Wlii QPllsV
                  UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

It is very hard to be sure. There may be several impediments to the evolution of a
technical civilization. Planets may be rarer than we think. Perhaps the origin of life is
not so easy as our laboratory experiments suggest. Perhaps the evolution of advanced
life forms is improbable. Or it may be that complex life forms evolve more readily, but
intelligence and technical societies require an unlikely set of coincidences - just as the
evolution of the human species depended on the demise of the dinosaurs and the ice­
age recession of the forests in whose trees our ancestors screeched and dimly
wondered. Or perhaps civilizations arise repeatedly, inexorably, on innumerable planets
in the Milky Way, but are generally unstable; so all but a tiny fraction are unable to
survive their technology and succumb to greed and ignorance, pollution and nuclear
war.

It is possible to explore this great issue further and make a crude estimate of N, the
number of advanced civilizations in the galaxy. We define an advanced civilization as
one capable of radio astronomy. Th is is, of course, a parochial if essential definition.
There may be countless worlds on wh ich the inhabitants are accomplished linguists or
superb poets but indifferent radio astronomers. We will not hear from them. N can be
written as the product or multiplication of a number of factors, each a kind of filter,
every one of which must be sizable for there to be a large number of civilizations:

•   Ns, the number of stars in the Milky Way galaxy.

•   fp, the fraction of stars that have planetary systems.

• ne, the number of planets in a given system that are ecologically suitable for life.
•   fl, the fraction of otherwise suitable planets on which life actually arises.

•   fi, the fraction of inhabited planets on which an intelligent form of life evolves.

•   fc, the fraction of planets inhabited by intelligent beings on which a communicative
    techn ical civilization develops .

•   fl, the fraction of planetary lifetime graced by a technical civilization.

Written out, the equation reads

                                   N = Ns • jjJ • ne • fl ·ft · Jc· fL                          (7)


All of the f's are fractions, having values between O and 1; they will pare down the
large value of Ns.

To derive N we must estimate each of these quantities. We know a fa ir amount about
the early factors in the equation, the number of stars and planetary systems. We know
very little about the later factors, concerning the evolution of intelligence or the lifetime
of technical societies. In these cases our estimates will be little better than guesses. I
invite you, if you disagree with my estimates below, make your own choices and see
what implications your alternative suggestions have for the number of advanced
civilizations in the galaxy. One of the great virtues of this equation, due to Frank Drake
of Cornell, is that it involves subjects ranging from stellar and planetary astronomy to
organic chemistry, evolutionary biology, history, politics and abnormal psychology.
Much of the Cosmos is in the span of the Drake equation.

8
                  UNCLASSIFIED//FQA QFFICiIPL. !Iii 011! Y
                      UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

We know Ns, the number of stars in the Milky Way galaxy, fairly well, by careful counts
of stars in a small but representative region of the sky. It is a few hundred billion; some
recent estimates place it at 4 x 10 11 . Very few of these stars are of the massive short­
lived variety that squander their reserves of thermonuclear fuel. The great majority
have lifetimes of billions or more years in which they are shining stably, providing a
suitable energy source for the energy and evolution of life on nearby planets.

There is evidence that planets are a frequent accompaniment of star formation: in the
satellite systems of Jupiter, Saturn and Uranus, which are like miniature solar systems;
in theories of the origin of the planets; in studies of double stars; in observations of
accretion disks around stars; and is some preliminary investigations of gravitational
perturbations of nearby stars. 1 Many, perhaps even most, stars may have planets. We
take the fraction of stars that have planets, fp, as roughly equa l to 1/3. Then the total
number of planetary systems in the galaxy would be Ns fp                    ~
                                                                 1.3 x 10 11 (the symbol                     ~
means "approximately equal to"). If each system were to have about ten planets, as
ours does, the tota l number of worlds in the galaxy would be more than a trillion, a vast
arena for the cosmic drama.

In our own solar system there are several bodies t hat may be suitab le for life of some
sort: the Earth certainly, and perhaps Mars, Titan and Jupiter. Once life originates, it
tends to be very adaptable and tenacious. There must be many different environments
suitable for life in a given planetary system. But conservatively we choose ne=2. Then
the number of planets in the galaxy su itable for life becomes Ns fp ne 3 x 10 11 .        ~
Experiments show that under the most common cosmic conditions the molecular basis
of life is read ily made, the building blocks of molecules able to make copies of
themselves. We are now on less certain grounds; there may, for example, be
impediments in the evolution of the genetic code, although I think this is unlikely over
billions of years of primeval chemistry . We choose fl~ 1/3, implying a total number of
planets in the Milky Way on which life has arisen at least once as Ns fp ne fl~ 1 x 10 11 ,
a hundred billion inhabited worlds . That in itself is a remarkable conclusion. But we are
not yet fin ished.

The choices of fi and fc are more difficult. On the one hand, many individually unlikely
steps had to occur in biologica l evolution and human history for our present intelligence
and technology to develop. On the other hand, there must be quite different pathways
to an advanced civilization of specified capab ilities. Considering the apparent difficulty
in the evolution of large organisms, represented by the Cambrian explosion, let us
choose fix fc = 1/100, meaning that only 1 per cent of planets on wh ich life arises
actually produce a technical civilization. This estimate represents some middle ground
among the varying scientific options. Some think that the equivalent of the step from
the emergence of trilobites to the domestication of fire goes like a shot in all planetary
systems; others th ink t hat, even given ten or fifteen billion years, the evolution of a
technical civi lization is unlikely. This is not a subject on wh ich we can do much
experimentation as long as our investigations are limited to a single planet. Multiplying

1 Carl Sagan was writ ings these lines back in the 1970's, when no extrasolar planets had been discovered yet. The

first such discovery occurred in 1995, wh en Michel Mayor and Didier Queloz, working at the "Observatoire de Haute
Provence" in France, discovered the first extrasolar planet orbiting the nearby star 51 Peg. This first extrasolar
planet was hence named 51 Peg B. Many more extrasolar planets were discovered around nearby stars ever since .
As of Apri l 2009, 347 extrasolar planets (exoplanets) are listed in the Extrasolar Planets Encyclopaed ia.



9
                      UNCLASSIFIED/ /POil Offl@IAL W&i 9Nls¥
                  UNCLASSIFIED/ /FOR OFFIEiIAk Wlii QPU,¥


these factors together, we find Ns fp ne fl fi fc ~ 1 x 109 , a billion planets on which
technical civilizations have arisen at least once. But that is very different from saying
that there are a billion planets on which technical civilizations now exist. For this we
must also estimate fl.

What percentage of the lifetime of a planet is marked by a technical civil ization? The
Earth has harbored a technical civilization characterized by radio astronomy for only a
few decades out of a lifetime of a few billion years. So far, then, for our planet fl is less
than 1/108 , a millionth of a percent. And it is hardly out of the question that we might
destroy ourselves tomorrow. Suppose this were a typical case, and the destruction so
complete that no other technical civilization - of the human or any other species - were
able to emerge in the five or so billion years remaining before the Sun dies. Then Ns fp
            ~
ne fl fi fc fl  10, and, at a given time there would be only a tiny smattering, a handful,
a pitiful few technical civilizations in the galaxy, the steady state number maintained as
emerging societies replace those recently self-immolated. The number N might be even
as small as 1 if civilizations tend to destroy themselves soon after reaching a
technological phase; there might be no one for us to talk with but ourselves. And that
we do but poorly. Civilizations would take billions of years of tortuous evolution, and
then snuff themselves out in an instant of unforgivable neglect.

But consider the alternative, the prospect that at least some civilizations learn to live
with technology; that the contradictions posed by the vagaries of past brain evolution
are consciously resolved and do not lead to self destruction; or that, even if major
disturbances occur, they are reveres in the subsequent billions of years of biological
evolution. Such societies might live to a prosperous old age, their lifetimes measured
perhaps on geological or stellar evolutionary time scales. If 1 percent of civilizations can
survive technological adolescence, take the proper fork at this critical historical branch
point and achieve maturity, then fl ~ 1/100, N ~ 107 , and the number of extant
civilizations in t he galaxy is in the millions. Thus, for all our concern about t he possible
unrel iabi lity of our estimates of the early factors in the Drake equation, which involve
astronomy, organic chemistry and evolutionary biology, the principal uncertainty comes
to economics and politics and what, on Earth, we call human nature. It seems fairly
clear that if self-destruction is not the overwhelmingly preponderant fate of galactic
civilizations, then the sky is softly humming with messages from the stars.

These estimates are stirring . They suggest that the receipt of a message from space is,
even before we decode it, a profoundly hopeful sign. It means that someone has
learned to live with high technology; that it is possible to survive technological
adolescence. This alone, quite apart from the contents of the message, provides a
powerful justification for the search for other civilizations.

4. The Drake Equation is Over-Simplified
In the nearly fifty years (1961-2009) elapsed since Frank Drake proposed his equation,
a number of scientists and writers tried to find out which numerical values of its seven
independent variables are more real istic in agreement with our present-day knowledge.
Thus there is a considerable amount of literature about the Drake equation nowadays,
and, as one can easily imagine, the results obtained by the various authors largely
differ from one another. In other words, the value of N, that various authors obtained
by different assumptions about the astronomy, the biology and the sociology implied by
the Drake equation, may range from a few tens (in the pessimist's view) to some

10
                  UNCLASSIFIED/ /FOR OFFIEiIAk Wlii QPII.¥
                   UNCLASSIFIED/ /FOR OFFIEil.t.k W~i Q~lls¥

million or even billions in the optimist's opinion. A lot of uncertainty is thus affecting our
knowledge of N as of 2010. In all cases, however, the final result about N has always
been a sheer number, i.e., a positive integer number ranging from 1 to millions or
billions. This is precisely the aspect of the Drake equation that th is author regarded as
" too simpl istic" and improved mathematically in his paper #IAC-08-A4.1.4, entitled
"The Statistical Drake Equation" and presented on October 1st , 2008, at the 59 th
International Astronautical Congress (IAC) held in Glasgow, Scotland, UK, September
29 th thru October 3rd , 2008. That paper is attached herewith as Appendix B. Newcomers
to SETI and to the Drake equation, however, may find that paper too difficult to be
understood mathematically at a first reading. Thus, I shall now expla in the content of
that paper "by speaking easily." I thank the reader for his or her attention .

5. The Statistical Drake Equation
We start by an examp le.

Consider the first independent variable in the Drake equation (7), i. e., Ns, the number
of stars in the Milky Way galaxy. Astronomers tell us that approximately there should
be about 350 millions stars in the galaxy. Of course, nobody has counted (or even seen
in the photographic plates) all the stars in the galaxy! There are too many practical
difficulties preventing us from doing so: just to name one, the dust clouds that don't
allow us to see even the Galactic Bulge (i.e. the central region of the galaxy) in the
visible light (although we may "see it" at radio frequencies like the famous neutral
hydrogen line at 1420 MHz). So, it doesn't make any sense to say that Ns = 350 x 106 ,
or, say (even worse) that the number of stars in the galaxy is (say) 354,233,321, or
similar fanciful exact integer numbers. That is just silly and non-scientific. Much more
scientific, on the contrary, is to say that the number of stars in the galaxy is 350 million
plus or minus, say, 50 millions (or whatever values the astronomers may regard as
more appropriate, since this is just an example to let the reader understand the
difficulty).

Thus, it makes sense to REPLACE each of the seven independent variables in the Drake
equation (7) by a MEAN VALUE (350 millions, in the above example) PLUS OR MINUS A
CERTAIN STANDARD DEVIATION (SO millions, in the above example) .

By doing so, we have made a great step ahead: we have abandoned the too-simplistic
equation (7) and replaced it by something more sophisticated and scientifically more
serious: the STATISTICAL Drake equation. In other words, we have transformed the
classical and simplistic Drake equation (7) into an advanced statistica l tool for the
investigation of a host of facts hardly known to us in detail. In other words still:

•    We replace each independent variable in (7) by a RANDOM VARIABLE, labeled
     D, (from Drake).

•    We assume that the MEAN VALUE of each D, is the same numerical value previously
     attributed to the corresponding independent variable in (7).

•    But now we also ADD A STANDARD DEVIATION cr 0 ; on each side of the mean value,
     that is provided by the knowledge gathered by scientists in each discipline
     encompassed by each D,.


11
                   UNCLASSIFIED//FQA QFFICilOL. !!ii ONI X
                  UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

Having so done, the next question is:

How can we find out the PROBABILITY DISTRIBUTION for each D; ?

For instance, shall that be a Gaussian, or what?

This is a difficult question, for nobody knows, for instance, the probability distribution of
the number of stars in the galaxy, not to mention the probability distribution of the
other six variables in the Drake equation (7).

There is a brilliant way to get around this difficulty, though.

We start by excluding the Gaussian because each variable in the Drake equation is a
POSITIVE (or, more precisely, a non-negative) random variable, while the Gaussian
applies to REAL random variables only. So, the Gaussian is out. Then, one might
consider the large class of well-studied and positive probability densities called "the
gamma distributions," but it is then unclear why one should adopt the gamma
distributions and not any other. The solution to th is apparent conundrum comes from
Shannon's Information Theory and a theorem that he proved in 1948: "The probability
distribution having maximum entropy ( = uncertainty) over any FINITE range of real
values is the UNIFORM distribution over that range," This is proven in Appendix A of the
present document.

So, at this point, we assume that each of the seven D; in (7) is a UNIFORM random
variable, whose mean value and standard deviation is known by the scientists working
in the respective field (let it be astronomy, or biology, or sociology). Notice that, for
such a uniform distribution, the knowledge of the mean value Po; and of the standard
deviation u 0 , automatically determines the RANGE of that random variable in between
its lower (called a;) and upper (called b;) limits: in fact these limits are given by the
equations


                                                                                            (8)


(the "surprising" factor ✓3 in the above equations comes from the definitions of mean
value and standard deviation: please see equations (12), (15) and (17) in Appendix B
for the relevant proof). So the uniform distribution of each random variable D; is
perfectly determined by its mean value and standard deviation, and so are all its other
properties.

The next problem is the following:

OK, since we now know everything about each uniformly distributed D;, what is the
probability distribution of N, given that N is the product (7) of all the D1 ?

In other words, not only do we want to find the analytical expression of the probability
density function of N, but we also want to relate its mean value f-lN to all mean values
µ0
    of the D;, and its standard deviation
     ,                                      u to all standard deviations
                                                N                          u  of the D;.
                                                                               0
                                                                                   ,




12
                  UNCLASSIFIED/ fFOR: QFFICiIO L. lalii ODIL.¥
                   UNCLASSIFIED/ /FOR OFFI&I.t.k W~i Q~lk¥

This is a difficult problem.

It occupied the author's mind for no less than about ten years (1997 -2007).

It is actually an ANALYTICALLY UNSOLVABLE problem, in that, to the best of this
author's knowledge, it is IMPOSSIBLE to find an analytic expression for any FINITE
PRODUCT of uniform random variablesD; . This result is proven in Sections 2 thru 3.3 of
Appendix B (unfortunately!) .

6. Solving the Statistical Drake Equation By Virtue of the
Central Limit Theorem (CLT) of Statistics
The solution to the problem of finding the analytical expression for the probability
density function of N in the statistica l Drake equation was found by this author in
September 2007. The key steps are the following:

•    Take the natural logs of both sides of the statistical Drake equation (7). This
     changes the product into a sum.

•    The mean values and standard deviations of the logs of the random variables D;
     may all be expressed analytically in terms of the mean values and standard
     deviations of the D; .

•    Recall the Central Limit Theorem (CLT) of statistics, stating that (loosely speaking) if
     you have a SUM of independent random variables, each of which is ARBITRARILY
     DISTRIBUTED (hence, also including uniformly distributed), then, when the number
     of terms in the sum increases indefinitely (i.e. for a sum of random variables
     infinitely long) .. . the SUM RANDOM VARIABLE TENDS TO A GAUSSIAN.

•    Thus, the natural log of N tends to a Gaussian.

•    Thus, N tends to the LOGNORMAL DISTRIBUTION.

•    The mean value and standard deviations of this lognormal distribution of N may all
     be expressed analytically in terms of the mean values and standard deviations of
     the logs of the D, already found previously.

This result is fundamental.

All the relevant equations are summarized in the following Table 1. This table is actually
the same as Table 2 of the author's original paper IAC-08-A4.1.4, entitled "The
Statistical Drake Equation" and presented by him at the International Astronautical
Congress (IAC) held in Glasgow, UK, on October l5t, 2008. This orig inal paper is
reproduced in Appendix B.

To sum up, not only is it found that N approaches the completely known lognormal
distribution for an INFINITY of factors in the statistical Drake equation (7), but the way
is paved to further applications by removing the cond ition that the number of terms in
the product (7) must be FINITE.



13
                   UNCLASSIFIED/ /FOR 8FFI&l.t.k Wii QNk¥
                 UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

This possibility of ADDING ANY NUMBER OF FACTORS IN THE DRAKE EQUATION (7)
was not envisaged, of course, by Frank Drake back in 1961, when "summarizing" the
evolution of life in the galaxy in SEVEN simple STEPS. But today, the number of factors
in the Drake equation should already be increased: for instance, there is no mention in
the original Drake equation of the possibility that asteroidal impacts might destroy the
life on Earth at any time, and this is because the demise of the dinosaurs at the K/T
impact had not been yet understood by scientists in 1961, and was so only in 1980!

In practice, the number of factors should INCREASE as much as necessary in order to
get better and better estimates of N as long as our scientific knowledge increases. This
is called the "Data Enrichment Principle" and believe should be the next important goal
in the study of the statistical Drake equation.

Finally, a numerical example explaining how the statistical Drake equation works in the
practice will be given in the next section.




14
                 UNCLASSIFIED/ /rOR orr1e1At HSE OrtLY
                   UNCLASSIFIED/ /FOR OFFIEiIAk Wlii QPU,¥

Table 1. Summary of the Properties of the Lognormal Distribution That Applies
 to the Random Variable N = Number of ET Communicating Civilizations in the
                                   Galaxy

Random variable                                  N = number of communicating ET
                                                 civilizations in galaxy
Probab ility distribution                        Loq norma l
Probabi lity density function                                                               (1n(11h,J'
                                                                     1        1             -~
                                                 JN (n)= - · &                          e                    (n ;::-: 0)
                                                                     n    27!<:T

Mean va lue                                               a'
                                                          -
                                                 (N )= eµe 2
Variance                                         a 1 = e2µ ea' (ea' - I)
Standard deviation                                         a'
                                                 a N = eµ e2 .J ea 2 - 1
All the moments, i. e. k-th moment                                       k ·-
                                                                              2   a'
                                                 (N k)= ekµ e                     2

Mode ( = abscissa of the log normal peak)                     _
                                                 n nnde = npeak -
                                                                          _ ~I - a- 2
                                                                           e e
Value of the Mode Peak                                                        a'
                                                 f N (n lTl)dc ) = &1 ·e -p ·e2
                                                                    2,r (l
Median (= fifty-fifty probability value for      rred ian = m = e"
N)
Skewness                                           K    ( ,    )                                                 e-6pe-3a'
                                                 _ 3_ = ea + 2
                                                 (K4)¾         .,                       k'        - lne3a2 + 3e2a' +6ea' +61
Kurtosis                                             K                   ,,                  .,          ., "}
                                                 _       4_     = e4a- + 2 e3a- + 3 e-u- - 6
                                                 (K2)2
Expression of 1-1 in terms of the lower (a;)
                                                 µ = ± (Y;) = ± b;[in(b;) - 1]- a;[in (a;) -1]
and upper (b ;) limits of the Drake                  i=I      i= I          b; - a;
uniform input random va riables 0 ;
Expression of <:T 2 in terms of the lower (a;)                 7                   7
                                                                                            a;b; [ln (b; )- in(a;)f
and upper (b;) limits of the Drake               a 2 =La~ = L l-
                                                              i= I                i=I               (b; -a}
uniform input random variables 0 ;


7. An Example Explaining the Statistical Drake Equation
To understand how things work in practice for the statistical Drake equation, please
consider the following table 2. It is made up of three columns:

•    The first column on the left lists the seven input sheer numbers that also become

•    The mean values (middle column).

•    Finally the last column on the right lists the seven input standard deviations.



15
                   UNCLASSIFIED/ 1re1t OFFl@IAL WSE 9Nk¥
                       UNCLASSIFIED/ /FOR OFFIEil.t.k W&i Q~lls¥

The bottom line is the classical Drake equation (7). We see that, for this particular set
of seven inputs, the classical Drake equation (i.e. the product of the seven numbers)
yields a total of 3500 communicating extraterrestrial civilizations existing in the galaxy
right now.


                        rs := 350 -109                              s :=     s


                              0                                                                     10
                       fp := 100                                µfp := fp                 ofp := 100

                                                                                                      l
                       ne := 1                                  µne := ne                 crne := -
                                                                                                  /3
                                                                                          crfl := ~
                                0
                       fl := -                                  µfl := fl
                              100                                                                100

                       fi := ~                                  µfi := fi                 on := ~
                            100                                                                  100

                                                                                          crfc := ~
                               0
                       fc:= -                                   µfc := fc
                             100                                                                  100

                       fL := 10000                              µfl. := tL                crtL := 1000
                               1010                                                               1010


                                    _1   := Ks -fp -ne-fl -fi .fc .ff,           = 3500

Table 2. Input Values (i.e. mean values and standard deviations) for the Seven Drake Uniform Random
Variables Di . The first column on the left lists the seven input sheer numbers that also become the mean values
(middle column). Finally the last column on the right lists the seven input standard deviations . The bottom line is
the classical Drake equation (7).

The statistical Drake equation, however, provides a much more articulated answer than
just the above sheer number N = 3500. In fact, a MathCad code written by this author
and capable of performing all t he numerical calculations required by the statistical
Drake equation for a given set of seven input mean va lues plus seven input standard
deviations, yields for N the lognormal distribution (thin curve) plotted in Figure 2. We
see immediately that the peak of this thin curve (i.e. the mode) falls at about
 n rmde;;;; npeak = eµ e-a' ""250 (this is equation (99) of Appendix B), while the median (fifty­

fifty value spl itting the lognormal density in two parts with equal undergoing areas) falls
at about nm,d ian;;;; eµ ""1740 . These seem to be smaller values than N = 3500 provided by
the classical Drake equations, but it's a wrong impression due to a poor "intuitive"
understanding of what statistics is! In fact, neither the mode nor the median are the
" really important" values: the really important value for N is the MEAN VALUE! Now if
you look at t he thin curve in Figure 2 below (i.e. the lognorma l distribution arisin g from
the Central Li mit Theorem), you see that this curve has a LONG TAIL ON THE RIGHT! In
other words, it does NOT immediately go down to nearly zero beyond the peak of the
mode. Thus, when you actually compute the mean value, you should not be too



16
                       UNCLASSIFIED// POR OPPICll<L ti.!! er~LV
                     UNCLASSIFIED/ /FOR OFFI&IAk Wlii QPU,¥

                                                    (]''


surprised to find out that it equals (N) = e 11 e2 : : : 4589 .559          ~ 4590 communicating
civilizations now in the galaxy. This is the important number, and it is HIGHER than the
3500 provided by the classical Drake equation. Thus, in conclusion, THE STATISTICAL
EXTENSION of the classical Drake equation INCREASES OUR HOPES to find an
extraterrestria I civi Iization !


                             PROBABILITY DENSITY FUNCTION OF N




                                  1000           2000               3000                           4000
                                  N = Number of ET Civilizations in Galaxy

Figure 2. Comparing the Two Probability Density Functions of the Random Variable N Found (1)
Without Resorting to the CLT at All (thick curve) and (2) Using the CLT and the Relevant Lognormal
Approximation (thin curve).

Even more so our hopes are increased when we go on to consider the standard
deviation associated with the mean value 4590. In fact, the standard deviation is given
                                                                     ,,.,
by equation (97) of Appendix B. This yields            2 ✓e" 2 -1 = 11195 and so the
                                                      er"' = e11 e
expected number of N may actually be even much higher than the 4590 provided by
the mean value alone! The "upper limit of the one-sigma confidence interval" (as
statisticians call it), i.e. the sum 4590+11195 = 15,785, yields a higher number still!
(Note: the "lower limit of the one-sigma confidence interval is ZERO because the
lognormal distribution is POSITIVE (or, more correctly, non-negative)). Finally, the
reader should note that the thick curve depicted in Figure 2 is just the NUMERICAL
solution of the statistical Drake equation for a FINITE number of 7 input factors. Figure
2 actually shows that this curve "is well interpolated" by the lognormal distribution (thin
curve), i.e., by the neat analytical expression provided by the Central Limit Theorem for
an INFINITE number of factors in the Drake equation. That is, in conclusion, Figure 2
visually shows that taking 7 factors or an infinity of factors "is almost the same thing"
already for a value as small as 7.



17
                     UNCLASSIFIED/ /FOil OFFI@IAb W&i QNI.¥
                  UNCLASSIFIED/ /FOR OFFIEil.t.k Wii Q~lls¥


8. Finding the Probability Distribution of the Et-Distance By
Virtue of the Statistical Drake Equation
Having solved the statistical Drake equation by finding the lognormal distribution, we
are now in a position to solve the ET-DISTANCE problem by resorting to statistics again,
rather than just to the purely deterministic Distance Law (5), as we did in Section 2.
This is "scientifically more serious" than just the purely deterministic Distance Law (5)
inasmuch as the new statistical Distance Law will yield a PROBABILITY DENSITY for the
Distance, with the relevant mean value and standard deviation . In other words, the
Distance Law (5) itself becomes a random variable whose probability distribution, mean
value and standard deviation must be computed by "replacing" into (5) the fact that N
is now known to follow the lognormal distribution . This is mathematically described in
detail in Section 7 of Appendix A.

The important new result is the PROBABILITY DENSITY FOR THE DISTANCE, the
equation of which is




                                                                                           (9)


holding for r ;,: O. This is equation (114) of Appendix B.

Starting from this equation, the MEAN VALUE OF THE random variable ET_ DISTANCE is
computed as
                                                                         J.I   cr2

                                            (Er_Distance) =Ce 3 e""jg                    (10)


which is equation (119) of Appendix B, and finally the ET_DISTANCE STANDARD
DEVIATION

                                                                   _!!. a2Kr2
                                      CJ ET_Distanu: -
                                                         -
                                                               C e 3 e 18 e 9 - l        (11)


which is equation ( 123) of Appendix B. Of course, all other descriptive statistical
quantities, such as moments, cumulants etc. can be computed upon starting from the
probability density (9), and the resu lt is Table two hereafter, that is Table 3 of Appendix
B.

Finally, to complete this section, as well as this "introduction to the statistical Drake
equation," the numerical values that equations (10) and (11) yield for the Input Table 1
are determined. They are, respectively:

                                                        _}!_     CTl

                               r,11ea11 _ m l ue   = Ce 3 e 18 ;::: 2,670 light years    (12)



18
                  UNCLASSIFIED/ /FOR OFFIEil.t.k Wii QNls¥
                UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

which is equation (153) of Appendix B, and

                                            I'   (12   ~


                         a ET_Di stono, =C e 3 e •8 ~ e9 - 1 a:: 1,309 light years   (13)

which is equation (154) of Appendix B.




19
                UNCLASSIFIED/ /FOR &FFIGIAk lalii ODIL:¥
                    UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

 Table 2. Summary of the Properties of the Probability Distribution That Applies
 to the Random Variable ET_Distance Yielding the (average} Distance Between
        Any Two Neighboring Communicating Civilizations in the Galaxy

Random variable                                   ET_ Distance between any two neighboring
                                                  ET civilizations in galaxy assuming they are
                                                  UNIFORMLY distributed throughout the
                                                  whole qalaxy volume.
Probability distribution                          Unnamed
Probability density function
                                                                                                                    _(•{ 6 R[ata1'.I l,Gola')]-µ          J
                                                                                 3               I                                        2u 2
                                                  fET_Distana,(r)           =-;. · Ji; CT ·e
Numerical constant C related to the Milky
                                                  C = 3 6 R 8alruy h Gala.,y "" 28,845 light years
Way size
Mean value                                                                                   µ           u2

                                                  (Er_Distance) = C e- 3 e18
Variance
                                                    2           -
                                                  CTET_Distance -
                                                                             2       -3p 9
                                                                          C e 2 ea [ e a
                                                                                                         2
                                                                                                                    9 -
                                                                                                                         2
                                                                                                                                  l
                                                                                                                                  I

Standard deviation                                                               _ji_       er ✓ a                  2

                                                                      -
                                                  CTET_Distance -         C e 3 e 18 e 9 _ l
All the moments, i.e. k-th moment                                                                    - k!!.             k2,u2

                                                  (Er_Dis tan eek) = c k e                                   3 e             18


Mode ( = abscissa of the log normal peak)                                            _ji_            er
                                                  rnnde   =       rpeak = Ce 3 e
                                                                                                     9


Value of the Mode Peak                            Peak Value of fET_ Distanu:(r) =
                                                                                      !!. -a'                 3
                                                  = fET_Distana,Cr,rnde) = cfi; CT
                                                                                   ·e  3 . e 18


Median ( = fifty-fifty probability value for N)                                    _ji_
                                                  ~dian = m = Ce 3
Skewness                                                                                                     ,,.,                 5,,.2          ,,., ]

                                                                                        e-µ [ e 2 - 3e 18 +2e 6
                                                  __!!i_ =
                                                          3                                                                                                    3
                                                  (K4 )2                     8a 2          5a 2    4a 2      a2   2a2                                         ]2
                                                                    C 3 [ e9            -4e- 9 -3e      +12 e - 6e- 9         9                  3

Ku rtosis                                                            4 a2                   a'                      2 a2
                                                   K4 _ 9 +2e - 3 +3e - 9 -6
                                                  (K2)2 - e
Expression of µin terms of the lower (ai)         µ= I            (r;) = I b; [ln (b;)- 1]- a;[ln(a;)- 1]
and upper (bi) limits of the Drake uniform
                                                          i=I               ;~1                               b; - a;
input random variables Di
                                                              7              7
Expression of CT 2 in terms of the lower (ai)                                                a;b; [ln (b;)- ln(a;)f
                                                  CT 2 = Io} =I i
and upper (bi) limits of t he Drake uniform                i= I              i=I                             (b; - a; )2
input random variables Di

 20
                    UNCLASSIFIED/ J'FQA QFFICiIPL. !Iii 011! Y
     UNCLASSIFIED/ /FOR OFFI&IAk Wlii QPU,¥




21
     UNCLASSIFIED// POil OFFI@IAb W&i 91'11.¥
                       UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

It is clarifying to draw the graph of the ET_Distance probability density (9):


                                          DISTANCE OF NEAREST Ef_ CIVILIZA TION
     5.63 ·10-20




                           500      1000                1500 2000        2500 3000        3500        4000   4500   5000
                                                         ET _Distance from Earth (light years)
Figure 3. The Probability of Finding the Nearest Extraterrestrial Civilization at the distance r From Earth
(in light years) if the Values Assumed in the Drake Equation are Those Shown in Input Table 1. The
relevant probability density function fET_Disian.,( r) is given by equation (9). Its mode (peak abscissa) equals 1933
light years, but its mean value is higher since the curve has a long tail on the right : the mean value equals in fact
2670 light years. Finally, the standard deviation equals 1309 light years: THIS IS GOOD NEWS FOR SETI,
inasmuch as the nearest ET galaxy civilization might lie at just 1 sigma = 2670-1309 = 1361 light years from us.

From Figure 3 we see that the probability of finding extraterrestrials is practically zero
up to a distance of about 500 light years from Earth. Then it starts increasing with the
increasing distance from Earth, and reaches its maximum at

                                                                    _l!_   ~
                                          r=de         = r peak = C e 3 e 9 :::: 1,933 light years.                        (14)

This is the MOST LIKELY VALUE of the distance at which we can expect to find the
nearest extraterrestrial civilization.

It is not the mean va lue of the probabil ity distribution (9) for fET_Di stan.,( r). In fact, the
probability density (9) has an infinite tail on the right, as clearly shown in Figure 3, and
hence its mean value must be higher than its peak value. As given by (10) and (12), its
                                     ....!:!..   0'2

mean value is r,,,en,,_mlue = Ce c::,2670 light years. This is the MEAN (va lue of the)
                                          3 e 18

DISTANCE at which we can expect to find extraterrestria ls.



22
                       UNCLASSIFIED/ ;'FQA QFFICiIPL. !Iii 011! Y
                  UNCLASSIFIED/ /FOR OFFIEiIAk Wlii QPU,¥

After having found the above two distances (1933 and 2670 light years, respectively),
the next natural question that arises is: "what is the range, back and forth around the
mean value of the distance, within which we can expect to find extraterrestria ls with
"the highest hopes?" The answer to this question is given by the notion of standard
deviation that we already found to be given by (11) and (13),

                                                I'    0'2    ~

                         CT ET_Distance = Ce     3 e •8 V e 9    -1 ""1309 light years .

More precisely, this is the so-ca lled 1-sigma (distance) level. Probability theory then
shows that the nearest extraterrestrial civilization is expected to be located within this
ra nge, i.e. within the two distances of (2670-1309) = 1361 light years and
(2670+1309) = 3979 light years, with probability given by the integral of fET_oistance(r)
taken in between these two lower and upper limits, that is:

                                 i3979 1igh1 years

                                  l36 1lig btycars
                                                     fET Di siance (r) dr:::: 0.75 = 75 %
                                                         -
                                                                                            (15)

In plain words: with 75 percent probability, the nearest extraterrestrial civilization is
located in between the distances of 1361 and 3979 light years from us, having assumed
the input values to the Drake Equation given by table 1. If we change those input
values, then all the numbers change again, of course.

9. The "Data Enrichment Principle" as the Best CLT
Consequence Upon the Statistical Drake Equation (Any
Number of Factors Allowed)
As a fitting climax to all the statistical equations developed so far, let us now state our
"DATA ENRICHMENT PRINCIPLE." It simply states that "The Higher the Number of
Factors in the Statistical Drake equation, The Better."

Put in this simple way, it simply looks like a new way of saying that the CLT lets the
random variable Y approach the normal distribution when the number of terms in the
sum (4) approaches infinity. And this is the case, indeed.

10. Conclusions
We have sought to extend the classical Drake equation to let it encompass Statistics
and Probability.

This approach appears to pave the way to future, more profound investigations
intended not only to associate "error bars" to each factor in the Drake equation, but
especially to increase the number of factors themselves. In fact, this seems to be the
only way to incorporate into the Drake equation more and more new scientific
information as soon as it becomes available. In the long run, the Statistical Drake
equation might just become a huge computer code, growing in size and especially in
the depth of the scientific information it contains. It would thus be Humanity's first
"Encyclopaedia Galactica."




23
                  UNCLASSIFIED/ /FOR OFFIEiIAk Wlii QPlls¥
                 UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

Unfortunately, to extend the Drake equation to Statistics, it was necessary to use a
mathematical apparatus that is more sophi sticated than just the simple product of
seven numbers.




24
                 UNCLASSIFIED/ /FOR 8FFIGl.t.k W&i QPII.¥
                                   UNCLASSIFIED/ /FOR OFFIEil.t.k W~i Q~lls¥


Appendix A: Proof of Shannon's 1948 Theorem Stating
That the Uniform Distribution is the "Most Uncertain" One
Over a Finite Range of Values

Information Theory was initiated by Claude Shannon (1916-2001) in his well-known
1948 two papers:
      Rqrutted   mm commi:ms from              B • 1-S,'l •st Ii         •ral Jour:!li,
      \ GI. 1 . pp.   9--!13 , 62J..,iS , July, Occober. I




                                                  A Mathematical Theory of Communication
                                                                                   By C. E SHANNO


In this Appendix, we wish to draw attention to a couple of theorems that Shannon
proves on pages 36 and 37 of his work, and read, respectively (note that Shannon
omits the upper and lower limits of all integrals in the first theorem: they are minus
infinity and plus infinity, respectively):

      5. Letp(x) bea one-dimensionaldis :i ti The onno p(x) ginng ammcimumentropy 1:.ubjec-tto the
         c dition tha e     dan:l devia on ofx be fixed at a i- G ian. To show this we must maximize

                                                                               H(x) = -      fp    x) logp(x) d


                                                                   <T-    =j p            x- dx   and 1 =/ p x)dx

          as cons               •. This requi~ . by                           calculu~ of variations maxunizing

                                                               / [- p x)logp(x)                   >.p(x     11p(x)] dx.

               e condition or lus is
                                                                               - 1- logp(x
                                                       •ng the constllllf5 to sat'                 the




and




25
                                   UNCLASSIFIED/ /FOR: QFFIGI0ls Uii 0111 X
                       UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

      7. If x is limited to a half line (p(x) = 0 for x < 0) and the fin.t moment of x is ihted at a:

                                                              a=   f     p (x dx

         then the maximwn en opy OCCW1! when

                                                              p(x)     =,! - f.r/al
                                                                          a
         and 1s equal to logea.

Now, we wish to point out that there is a third possible case, other than the two given
by Shannon. This is the case when the probabi lity density function p(x) is limited to a
FINITE INTERVAL a::; x::; b. This is obviously the case with any physical POSITIVE
ra ndom variable, such as a distance, or the number N of extraterrestrial communicating
civilizations in the,". And it is easy to prove that for any such finite random variable the
maximum entropy distribution is the UNIFORM distribution over a -5, x -5, b. Shannon did
not bother to prove this simple theorem in his 1948 papers since he probably regarded
it as too trivial. But we prefer to point out this theorem since, in the language of the
statistica l Drake equation, it sounds like:

"Since we don't know what the probability distribution of any one of the Drake random
variables D; is, it is safer to assume that each of them has the maximum possible
entropy over a; -5,x-5, h; , i. e., that D; is UNIFORM LY distributed there.

The proof of th is theorem is along the same lines as for the previous two cases
discussed by Shannon:

We start by assuming t hat a; -5, x -5, b; .

We then form the linear combination of the entropy integral plus the normalization
condition for D;




where i is a Lagrange multipli er.

Performing the variation, one finds

                                             - Iogp(x)- 1+1 = 0 that is: p(x)= e ,1,- i .

App lying the normalization condition (constraint) to the last expression for p(x) yields


                                  I = f. b, p ( x ) dx = f. b; e,l,-1 dx = e2-1 f.b; dx = e,1,-1 ( b; - a,. )
                                        a1               a1                       a1


that yields

                                                               ,l,-1          1
                                                              e        = --
                                                                        h; - a;



26
                       UNCLASSIFIED/ /POlt Offl@IAL WSi 9Nk¥
                 UNCLASSIFIED/ /FOR OFFI&IAk Wlii QPU,¥

and finally

                               p(x) = -   1-    with a; s x s b;
                                      b; - a;

showi ng that the maximum-entropy probability distribution over any FINITE int erval
a; sxsb; is the UNIFORM distrib ution .




27
                 UNCLASSIFIED/ /FOR OFFI&IAk Wlii QNk¥
                      UNCLASSIFIED/ /FOR OFFIEil.t.k W&i Q~lls¥


Appendix B: Original Text of the Author's Paper #IAC-08-
A4.1.4 Titled the Statistical Drake Equation

                                                IA C-08-A4.1.4


           THE STATISTICAL DRAKE EQUATION
                                             Claudio Maccone
          Co-Vice Chair, SETI Permanent Study Group, International Academy ofAstronautics

                       Address: Via Martorelli, 43 - Torino (Turin) 10155 - Italy
                      URL: http://www.maccone.com/ - E-mail: [email protected]

ABSTRACT. We provide th statistical generalization of the Drake equation.

From a simple product of seven positive numbers, the Drake eq uation is now turned into the product of seven
positive random variables. We call this "the Statistical Drake Equation ," The mathematical consequences of
this transformation are then derived. The proof of our results is based on the Central Limit Theorem (CLT) of
Statistics. In loose terms, the CLT states that the sum of any number of independent random variables, each of
which may be ARBITRARILY distributed, approaches a Gaussian (i.e. u01mal) random variable. This is called
the Lyapunov Form of the CLT, or the Lindeberg Form of th CLT, depending on the mathematical constraints
assumed on the third moments of the various probability distributions. In conclusion , we show that:
 l) The new random variable N, yielding the number of communicating civilizations in the Galaxy, follows the
     LOGNORMAL distribution . Then, as a consequence, the mean value of this lognormal distribution is the
     ordinary Nin the Drake equation. The standard deviation , mode, and all the moments of this lognormal N
     are found also.
2) The seven factors in the ordinary Drake equation now become seven positive random variables. The
     probability distribution of each random variable may be ARBITRARY. The CLT in the so-called
     Lyapunov or Lindeberg forms (that both do not assume the factors to be identically distributed) allows for
     that. In other words the CLT "translates" into our statistical Drake equation by allowing an arbitrary
     probability distribution fo r each factor. This is both physically realistic and practicalJ y very useful, of
     course.
3) An application of our statistical Drake equation then follows. The (average) DISTANCE between any two
     neighboring and communicating civilizations in the Galaxy may be shown to be inversely proportional to
     the cubic root of N. Then, in our approach, this distance becomes a new random variable. We derive the
     relevant probability density function , apparently previously unknown and dubbed "Maccone distribution"
     by Paul Davies.
4) DATA ENRICHMENT PRINCIPLE. It should be noticed that ANY positive number of random vai'iables
     in the Statistical Drake Equation is compatible with the CLT. So, our generalization allows for many more
     factors to be added in the future as long as more refined scientific knowledge about each factor will be
     known to the scientists. This capability to make room for more future factors in the statistical Drake
     equation we call the "Data Enrichment Principle", and we regard it as the key to more profound future
     results in the fields of Astrobiology and SETI.

Finally, a practical example is given of how our statistical Drake equation works numerically. We work out in
detail the case where each of the seven random variables is uniformly distributed around its own mean value
and has a given standard deviation. For instance, the number of stars in the Galaxy is assumed to be uniformly
distributed around (say) 350 billions with a standard deviation of (say) I billion. Then, the resulting lognormal
distribution of N is computed numerically by virtue of a MathCad file that the author has written. This shows

28
                      UNCLASSIFIED/; P"OR OP"P"lelltt li!I! er•tv
                      UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

that the mean value of the lognormal random variable N is actually of the same order as the classical N given
by the ordinary Drake equation, as one might expect from a good statistical generalization.

1. INTRODUCTION                                                number of civilizations now transmitting and
                                                               receiving, and this implies an estimate of"how
     The Drake equation is a now famous result                 long will a technological civilization live?"
(see ref. [1] for the Wikipedia summary) in the                that nobody can make at the moment. Also,
fields of SETI (the Search for ExtraTetTestial                 are they going to destroy themselves in a
Intelligence, see ref. [2]) and Astrobiology (see ref.         nuclear war, and thus live only a few decades
[3]). Devised in l 960, the Drake equation was the             of technological civi lization? Or are they
first scientific attempt to estimate the number N of           slowly becoming wiser, reject war, speak a
ExtraTerrestrial civilizations in the Galaxy with              single language (like English today), and
which we might come in contact. Frank D. Drake                 merge into a single "nation", thus living in
(see ref. [4]) proposed it as the product of seven             peace for ages? Or will robots take over one
factors:                                                       day making "flesh animals" disappear forever
                                                               (the so-called "post-biological universe")?
            N=Ns-fp-ne·fl·fi·fc·fL.                 (1)             No one knows ...

Where:                                                         But let us go back to the Drake equation (1).
I)   Ns is the estimated number of stars in our           In the fifty years of its existence, a number of
   Galaxy.                                                suggestions have been put forward about the
2) fp is the fraction (= percentage) of such stars        different numeric values of its seven factors. Of
   that have planets.                                     course, every different set of these seven input
3) ne is the number "Earth-type" such planets             numbers yields a different value for N, and we can
   around the given star; in other words, ne is           endlessly play that way. But we claim that these
   number of planets, in a given stellar system,          are like ... children plays!
   on which the chemical conditions exist for life
   to begin its course: they are "ready for life,"             We claim the classical Drake equation (1), as
4) fl is fraction(= percentage) of such "ready for        we shall call it from now on to distinguish it from
   life" planets on which life actually starts and        our statistical Drake equation to be introduced in
   grows up (but not yet to the "intelligence"            the coming sections, well, the classical Drake
   level).                                                equation is scientifically inadequate in one regard
5) fl is the fraction (= percentage) of such              at least: it just handles sheer numbers and does not
   "planets with life fom1S" that actually evolve         associate an error bar to each of its seven factors.
   until some form of "intelligent civilization"          At the very least, we want to associate an error
   emerges (like the first, historic human                bar to each D;.
   civilizations on Earth).
6) Jc is the fraction (= percentage) of such                   Well, we have thus reached STEP ONE in our
   "planets with civilizations" where the                 improvement of the classical Drake equation:
   civilizations evolve to the point of being able        replace each sheer number by a probability
   to communicate across the interstellar                 distribution!
   distances with other (at least) similarly
   evolved civilizations. As far as we know in                The reader is now asked to look at the flow
   2008, this means that they must be aware of            chart in the next page as a guide to this paper,
   the Maxwell equations governing radio waves,           please.
   as well as of computers and radioastronomy
   (at least).                                            2.   STEP 1: LETTING EACH FACTOR
7) fl is the fraction of galactic civilizations alive          BECOME A RANDOM VARIABLE
   at the time when we, poor humans, attempt to
   pick up their radio signals (that they throw out           In this paper we adopt the notations of the
   into space just as we have done since l 900,           great book "Probability, Random Variables and
   when Marconi started the transatlantic                 Stochastic Processes" by Athanasios Papoulis
   transmissions). In other words, fl is the              (1921-2002), now re-published as Papoulis-Pillai,


29
                      UNCLASSIFIED/ ,'FOR: QFFIGI CL. Wii ODIL:¥
                     UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

ref. [5]. The advantage of this notation is that it
makes a neat distinction between probabilistic (or                                                          (3)
statistical: it's the same thing here) variables,
always denoted by capitals, from non-probabilistic
(or "detenninistic") variables, always denoted by      Of course, N now becomes a (positive) random
lower-case letters. Adopting the Papoulis notation     variable too, having its own (positive) mean value
also is a tribute to him by this author, who was a     and standard deviation. Just as each of the D; has its
Fulbright Grantee in the United States with him at     own (positive) mean value and standard deviation ...
the Polytechnic [nstitute (now Polytechnic             ... the natural question then arises: how are the seven
University) of New York in the years 1977-78-79.       mean values on the right related to the mean value on
                                                       the left?
    We thus introduce seven new (positive)             ... and how are the seven standard deviations on the
random variables D; ("D" from "Drake") defined         right related to the standard deviation on the left?
                                                             Just take the next step ...
as
                                                       3.   STEP 2: INTRODUCING LOGS TO
                      D 1 = Ns                              CHANGE THE PRODUCT INTO A SUM
                      D2 =fp
                      D3 =ne                                Products of random variables are not easy to
                                                       handle in probability theory. It is actually much
                      D4 = fl                    (2)
                                                       easier to handle sums of random variables, rather
                      Ds = fl                          than products, because:
                      D6 =Jc                                1) The probability density of the sum of two or
                                                                more independent random variables is the
                      D7   =fL                                  convolution of the relevant probabifay
                                                                densities (worry not about the equations,
so that our STATISTICAL Drake equation may be                   right now) .
simply rewritten as                                         2) The Fourier transform of the convolution
                                                                simply is the product of the Fourier
                                                                transforms (again, worry not about the
                                                                equations,       at        this       point)




30
                     UNCLASSIFIED//EOR OFFICIO! !!SF ON! X
             UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥



                                                I
                                                    1. Introduction
                                                                        I
                                                            I
                              2. Step 1: Letting each factor become a random
                                                            I
                              2.1.   Step 2: Introducing logs to change the product into a
                                                                                                      I
                                                            I
                              2.2.   Step 3: The transformation law of random variables.
                                                                                                      I
                                                            I
                              3. Step 4:   Assuming the easiest input distribution for
                                           each D1: the uniform distribution.


                                                            I
                              3 .1. Step 5: A numerical example of the Statistical Drake equation
                                             with uniform distributions for the Drake random
                                             variables 0 ;.


                                                            I


     3.2. Step 6 : Computing the logs of the
                    7 uniformly distributed
                    Drake random variables
                   O; .

                          I
     3.3. Step 7: Finding the probability
                   density function of N, but
                   only numerically not
                   analytically.


                          I

             I   DEAD END!
                                     I                    4. The Central Limit Theorem (CLT) of Statistics.
                                                                                     I
                                                          s. LOGNORMAL distribution as the probability
                                                                distribution of the number N of
                                                                communicating ExtraTerrestrial Civilizations
                                                                in the Galaxy.
                                                                                     I
                                                          6. Compari ng the CLT r esults with the Non- CLT
                                                             results, and discarding the Non-CLT approach.



                                                          7. DISTANCE to the nearest ExtraTerrestriai
                                                             Civilization as a probability distribution ( Paul
                                                             Davies dubbed that the Maccone distribution).


                                                                                     I
                                                          7. 1 Classical, non-probabilistic derivation of the
                                                               Distance to the nearest ET Civilization.



                                                          7.2    Probabilistic derivation of probability density
                                                                 function for nearest ET Civilization Distance.



                                                          7.3    Statistical properties of the distribution.


                                                          7.4    Numerical example of the distribution.
                                                                                     I
                                                          8.     DATA ENRICHMENT PRINCIPLE as the best
                                                                 CLT consequence upon the Drake equation :
                                                                 ;mx number of factors allowed for.



31
             UNCLASSIFIED/ ;'FQA QFFICilPL. PP&i ODI! Y
                        UNCLASSIFIED/ /FOR OFFI&IAk Wlii QPU,¥

So, let us take the natural logs of both sides of the       type of probability density function (pdt) for the last
Statistical Drake equation (3) and change it into a         seven of equations (5), then we must compute the
sum:                                                        (new and different) pdf of the logs of such random
                                                            variables. And the pdf of these logs certainly is not
                                                            gamma-type any more.

                                                                  It is high time now to remind the reader of a
                                                            certain theorem that is proved in probability courses,
It is now convenient to introduce eight new (positive)      but, unfortunately, does not seem to have a specific
random variables defined as follows:                        name. It is the transformation law (so we shall call
                                                            it, see, for instance, ref. [5]) allowing us to compute
                                                            the pdf of a certain new random variable Y that is a
                    Y=ln(N)
               {                                      (5)   known function Y = g(X) of another random
                   Y;=ln(D;) i=l,...,7.
                                                            variable X having a known pdf. In other words, if the
                                                            pdf fx (x) of a certain random variable X is known,
Upon inversion, the first equation of (5) yields the
important equation, that will be used in the sequel         then the pdf fr(Y) of the new random variable Y,
                                                            related to X by the functional relationship
                                                      (6)
                                                                                  y = g(X)                       (8)
We are now ready to take STEP THREE.
                                                            can be calculated according to this rule:
   STEP 3: THE TRANSFORMATION LAW                           1) First invert the corresponding non-probabilistic
OF RANDOM VARIABLES                                             equation y = g(x) and denote by X; (y) the
                                                                various real roots resulting from the this
    So far we did not mention at all the problem:               inversion.
"which probability disttibution shall we attach to          2) Second, take notice whether these real roots may
each of the seven (positive) random variables D; ?"             be either finitely- or infinitely-many, according
                                                                to the nature of the function y = g(x).
    It is not easy to answer this question because we       3) Third, the probability density function of Y is
do not have the least scientific clue to what                   then given by the (finite or infinite) sum
probability distributions fit at best to each of the
seven points listed in Section 1.
                                                                                                                 (9)
     Yet, at least one trivial error must be avoided:
claiming that each of those seven random variables
must have a Gaussian (i.e. normal) distribution. In
                                                                where the summation extends to all roots x;(Y) and
fact, the Gaussian distribution, having the well­
known bell-shaped probability density function                   lg' (x; (y)~ is the absolute value of the first

                                                                derivative of g(x) where the i-th root x;(Y) has
                                                                been replaced instead of x.
                                          (a ;,: o)   (7)
                                                            Since we must use this transformation law to transfer
                                                            from the D; to the Y; = ln(D;), it is clear that we
has its independent variable y ranging between --oo
and oo and so it can apply to a real random variable        need to start from a D; pdf that is as simple as
Y only, and never to positive random variables like         possible. The gamma pdf is not responding to this
those in the statistical Drake equation (3). Period.        need because the analytic expression of the
                                                            transformed pdf is very complicated (or, at least, it
     Searching again for probability density functions      looked so to this author in the first instance). Also,
that represent positive random variables, an obvious        the gamma distribution has two free parameters in it,
choice would be the gamma distributions (see, for           and this "complicates" its application to the various
instance, ref. [6]). However, we discarded this choice      meanings of the Drake equation. Tn conclusion, we
too because of a different reason: please keep in mind      discarded the gamma distributions and confined
that, according to (5), once we selected a particular

32
                        UNCLASSIFIED/ /FOil OFFI@IAb W&i! QNI.¥
                                 UNCLASSIFIED/ /fOR OFFI&IAk Wlii QPU,¥

ourselves to the simpler uniform distribution instead,                                 _ (b; -a;)(a; +a;b; +b; ) _ a; +a;b; +b;
as shown in the nest section.
                                                                                       -        3 (b; - a;)      -       3
4. STEP 4: ASSUMING THE EASIEST
   I PUT DISTRIBUTION FOR EACH D; :                                               The second moment of the uniform distribution is
   THE U IFORM DISTRIBUTION                                                       thus


     Let us now suppose that each of the seven D; is                                                                a; +a;b; + b;
                                                                                                ( un itiorm D .2) = ---'---'-'-----'-         (13)
distributed UNIFORMLY in the interval ranging                                                              -  I             3
from the lower limit a; ~ 0 to the upper limit
b; ~a;·                                                                           From (12 and (13) we may now derive the variance
                                                                                  of the uniform distribution
     This is the same as saying that the probability
density function of each of the seven Drake random
variables D; has the equation                                                     c:r~nifi>rnLD; = ( uniform_D ;2)-( uniform_ D ;) 2


     funi rom, o. (x) = - 1-            with O::; a; s; x s; b;            (10)
                                                                                       =
                                                                                           a; + a;b; + b;2
                                                                                            ?

                                                                                                                                              ( 14)
              - '       h;-a;
                                                                                                    3                        4           12
as it follows at once from the normalization condition
                                                                                  Upon taking the square root of both sides of ( 14 ), we
                  fti;
                       b
                       , funi rol'TlLD- (x )dx = 1 .
                                    I
Fuente: archivo UAP oficial del gobierno de EE.UU. (dominio público) · war.gov/ufo ↗ · ver en el archivo de Nodriza

← Todos los desclasificados

Nodriza · Cosprax — archivo oficial, sin curaduría: la fuente del gobierno es la autoridad