For probability distributions and on the same finite set,with equality exactly when on the support of .
A probability space is a measure space whose total measure is one. Its measurable subsets are events, and its measure is their probability.
An event is a measurable subset of a probability space.
For a point uniformly distributed in the unit disk, its polar coordinates are independent with densities
For three independent points from any rotationally symmetric planar distribution with nonzero radii, the origin lies in their triangle exactly when their three angles do not all lie in one semicircle. Three independent uniform angles have this property with probability .
If two pi-systems satisfy the probability factorization pairwise, their generated sigma-algebras are independent. For one event in the first family, the events in the second sigma-algebra satisfying the factorization form a Dynkin system; apply Dynkin lemma, then repeat with the roles reversed.
For a sequence of random variables , the tail sigma-algebra isIts events are unaffected by changing any finite initial segment.
The tail sigma-algebra of a sequence of independent random variables is trivial: every tail event has probability zero or one.
The product measure is characterized on measurable rectangles depending on finitely many coordinates by the product of their component measures.
A random variable is a measurable function from a probability space to a measurable space.
Random variables are independent when their generated sigma-algebras are independent sigma-algebras. Equivalently, their joint distribution is the product of their marginal distributions.
A family of independent random variables is independent and identically distributed when every variable has the same probability distribution.
The Poisson arrivals see time averages property says that an external Poisson process independent of a stationary system sees the system-state distribution at its arrival times. Thus the fraction of arrivals that accept a state-dependent admission rule is .
An queue has Poisson arrivals, independent identically distributed service times, and one server. If service has mean and , its mean busy period isThis follows by decomposing the busy period into the first service and the busy-period descendants of arrivals during that service.
In a stable regenerative queue,where is mean population, is throughput, and is mean sojourn time. Renewal reward proves it by identifying the area under the population path with the sum of customer sojourn times.
The main modes include almost-sure convergence, convergence in probability, convergence in , and convergence in distribution.
Random variables converge in distribution, or converge weakly, to whenfor every bounded continuous function . For real random variables this is equivalent to convergence of the distribution functions at every continuity point of the limiting distribution function.
A sequence of random variables converges almost surely when the set of outcomes on which pointwise convergence fails has probability zero.
Random vectors converge weakly to exactly when every linear projection converges weakly to . The reverse implication follows because these one-dimensional limits give pointwise convergence of the multivariate characteristic functions.
Convergence in probability is equivalent toThe forward implication splits at a fixed error threshold; the reverse implication is Markov's inequality.
From in probability, choose with
. The first Borel--Cantelli lemma then gives almost surely.
. The first Borel--Cantelli lemma then gives almost surely.
Independent indicators with success probabilities converge to zero in probability, but successes occur infinitely often almost surely by the second Borel--Cantelli lemma.
If in probability and , then in for every . The bound gives uniform integrability of the th powers, while convergence in probability controls their bounded part.
For independent identically distributed with mean zero and variance one,The numerator obeys the central limit theorem, the denominator converges almost surely to one by the strong law, and Slutsky's theorem combines them.
Binomial distributions with trial count tending to infinity and success probability tending to zero converge to a Poisson law when their product converges.
A simple symmetric random walk starting from haswhere the increments are independent and satisfy
An integrable adapted process is a martingale when .
If every increment of a discrete martingale takes values in , thenIterated conditioning shows that the increments are independent symmetric signs, so is a simple symmetric random walk.
An integrable adapted process is a supermartingale when .
If a supermartingale dominates an adapted reward process , optional sampling gives for every bounded stopping time .
A random time is a stopping time when for every .
For a discrete martingale and stopping time ,The indicator is measurable at time , so the stopped process is a martingale. If it is uniformly bounded and almost surely, bounded convergence gives directly.
For simple symmetric random walk stopped on hitting or and , choose byThen are martingales. Solving the two boundary equations gives
For bounded stopping times , a supermartingale satisfies ; equality holds for a martingale.
In discrete time, a process is predictable or previsible when is -measurable for every .
For a square-integrable martingale transform and positive conditional increment variance,
For a stopping time ,The indicator is predictable, so a stopped integrable martingale is a martingale.
Let be a simple symmetric random walk, let be its first hitting time of the positive integer , and reverse every increment after . The resulting path is again a simple symmetric random walk. Consequently,
If is generated by independent symmetric signs , every martingale has predictable coefficients such thatOne may take .
For a bounded stopping time and a square-integrable representation ,The cross terms vanish because predictable multiples of distinct independent signs are orthogonal martingale differences.
The increments of a supermartingale are predictable and nonnegative.
Every integrable discrete-time supermartingale has the form , where is a martingale and is predictable, integrable, nondecreasing, and starts at zero.
The finite-horizon Snell envelope of rewards is defined backward by and . It is the least supermartingale dominating .
For the Doob compensator of a Snell envelope, : compensation grows only on the stopping region .
The first time the Snell-envelope compensator is about to become positive is an optimal stopping time; before it the martingale part equals both value and reward at stopping.
If , only finitely many occur almost surely. For independent events, divergence implies infinitely many occur. Use a union bound for the first and for the second.
For a nonnegative random variable , Tonelli theorem givesConsequently, for each , exactly when ; compare the integral on successive intervals of length .
For independent rate-one exponential variables and ,almost surely. The first follows by applying both Borel--Cantelli lemmas to thresholds ; the second uses
.
.
A stochastic process is a family of random variables indexed by time or another ordered parameter.
A Gaussian process is a stochastic process whose every finite vector of values has a multivariate normal distribution. Its law is determined by its mean and covariance functions.
A standard Brownian motion starts at zero, has almost surely continuous paths, and has independent Gaussian increments
A standard Brownian bridge is the centered Gaussian process on . Its covariance is and it satisfies .
For every real ,is a martingale. Conditional expectation factors at time because the Gaussian increment is independent of the past and has exponential moment .
For , shifting Brownian drift bymakes under an equivalent measure . The discounted stock is then a martingale.
The time-zero price of the payoff is
For a Brownian motion with drift , the process on has the same law as . Thereforehave the same distribution, which equates the corresponding geometric-Brownian lookback payoffs.
A centered Gaussian process with almost surely continuous paths and covariance is a standard Brownian motion.
If is Brownian motion, thenis also Brownian motion. Its covariance is , and continuity at zero follows from almost surely.
For ,This follows from the Brownian reflection principle weighted by the Cameron-Martin theorem for a linear drift.
For ,
Reflecting a Brownian path after its first hit of a level preserves Wiener measure and maps an endpoint to .
On paths through time , translating Brownian motion by the linear drift changes Wiener measure by the density
For a twice differentiable integrable test function and Brownian motion ,is a martingale. This is the one-dimensional generator form of Dynkin's formula.
Let be continuous with . Ifis a martingale for every real , then conditional expectations solveThus is independent of and distributed as , so is Brownian motion.
Pointwise convergence of characteristic functions to a function continuous at zero implies convergence in distribution.
A rate- Poisson process has independent stationary increments. Conditional on , arrival times are the order statistics of independent uniform points on .
Independently retaining a point at time with probability gives a Poisson count of mean .
Conditioned on a total of Poisson points, the number falling in a region of relative intensity is .
With empty initial state, Poisson arrivals of rate and exponential lifetimes of rate produce a Poisson occupancy of mean at time .
The state probabilities satisfy and , yielding the Poisson law of mean .
Conditioned on , the arrival times are distributed as the order statistics of independent uniform points on . Their conditional joint density onis , obtained by dividing the Poisson-process density by .
For independent Poisson processes of rates and , the number of type-two arrivals before the first type-one arrival satisfiesThis follows by marking the superposed process by its arrival type.
Let be the points of a rate- Poisson process and let independent identically distributed marks be independent of the process. ThenCondition on , use uniform order statistics, and sum the resulting exponential series.
With , the age of a rate- Poisson process has meanGiven , the last arrival is the maximum of uniform points and , including .
A Fokker-Planck equation evolves a density by drift and diffusion in conservation-law form.
For an integrable stopping time and i.i.d. integrable increments, the expected stopped sum equals expected stopping time times mean increment.
A Poisson point process has independent counts on disjoint sets and Poisson count in each set with mean given by its intensity measure.
The union of independent Poisson point processes with diffuse intensity measures is a Poisson point process with intensity measureOn every finite collection of disjoint measurable sets, this follows by adding independent Poisson counts and checking that the resulting counts remain independent.
For a Poisson point process with intensity measure and a square-integrable measurable function ,The second identity follows because distinct points contribute the square of the mean, while the diagonal contribution is .
Let be a nonnegative measurable intensity function on a space with reference measure . A Non-homogeneous Poisson point process has independent counts on disjoint measurable sets andwhenever the integral is finite.
The intensity function is the density of the expected counting measure: .
Let be a Poisson point process with intensity measure and let be measurable. If the pushforwardis locally finite, then the image counting measure is Poisson with intensity . If one requires a simple point process, the pushforward must also be diffuse, so distinct source points do not produce coincident image points with positive probability.
For a homogeneous Poisson point process of intensity in , mapping to , where is the unit-ball volume, gives a rate- Poisson process on . Mapping to instead gives rate .
For intensity in , let be the distance to the th closest point and put . Then is the th arrival time of a rate- Poisson process, so
Independently assigning a mark to each point of a Poisson point process, with mark probabilities allowed to depend measurably on the point, produces independent Poisson point processes for the different marks. If the original intensity is and mark has probability , its intensity is .
Independently retaining a point at with probability thins a Poisson point process of intensity to one of intensity ; retained and rejected processes are independent.
Independent and identically distributed random variables with a continuous probability distribution realize each of their possible strict orderings with equal probability.
A probability distribution assigns probabilities to the possible values of a random variable.
The beta distribution on has density proportional to and mean .
The cumulative distribution function of a real random variable isIt is nondecreasing, right-continuous, tends to zero at , and tends to one at .
If distribution functions converge pointwise to a continuous distribution function , then the convergence is uniform. Choose finitely many points whose successive -increments and two tails are small, then use monotonicity to squeeze between its values at adjacent points.
The Cauchy distribution with location and scale has probability density functionThe standard Cauchy distribution has , , characteristic function , and infinite first absolute moment.
If are independent standard Cauchy random variables, thenis again standard Cauchy. Indeed, its characteristic function is .
A probability distribution is stable when every linear combination of independent copies has the same distribution up to a location and scale change. The Cauchy distribution is stable with stability parameter one.
For a nonnegative integer-valued random variable , its probability generating function isThe generating function of a sum of independent random variables is the product of their generating functions.
For a discrete random variable , its probability mass function isIt is nonnegative and sums to one over the possible values of .
For a distribution function , its generalized inverse or quantile function is
If is uniform on , then has distribution function . Indeed, right continuity givesapart from irrelevant endpoint cases.
If at every continuity point of , then the generalized inverses satisfyat every continuity point of . Since a monotone function has only countably many discontinuities, evaluating all quantiles at one uniform random variable gives an almost-sure coupling of the corresponding convergence in distribution.
A discrete probability distribution is supported on a finite or countable set.
On the convention supported by , a geometric variable of parameter has mass , mean , and variance .
When independent uniform draws are made from types, the time until every type has appeared decomposes as a sum of independent geometric waiting times with parametersConsequentlyand Chebyshev's inequality gives in probability.
On the failures-before-the-th-success convention, a negative binomial variable with success probability has probability mass functionIt has expected value and variance .
For the th success to occur after exactly failures, the final trial must be a success, while the preceding trials contain failures and successes. Choosing the failure locations gives the factor .
The binomial distribution counts successes in independent Bernoulli trials with common success probability.
The multinomial distribution gives the category counts from independent trials with category probabilities .
The Poisson distribution with mean lambda assigns mass .
Independent Poisson counts conditioned on their sum have a multinomial distribution.
For each group , let independent cell counts have Poisson means . Conditional on their total , the cells have a multinomial distribution with probabilitiesIncluding a free group main effect in the Poisson log-linear model makes its profiled likelihood for the remaining parameters proportional to the corresponding multinomial likelihood. The two fits therefore give the same fitted proportions and likelihood-ratio comparisons.
For large , a Poisson random variable of mean is approximately normal with mean and variance .
A continuous probability distribution is described by a density with respect to Lebesgue measure.
A probability density function is a nonnegative measurable function satisfyingIts integral over the whole sample space is one.
The sum of independent exponential variables of rate has the Erlang distribution with densityConvolving with one more exponential proves the formula inductively.
If is a positive integer and are independent uniform variables, thenhas the gamma distribution with shape and rate , because each summand is exponential with rate .
The exponential distribution with rate has densityIt is the gamma distribution with shape one.
If , thenThus an exponential remaining lifetime has the same law regardless of the elapsed lifetime.
If and are independent exponential variables with the same rate, then is uniform on and is independent of .
For independent exponential variables of rates and , let and let be the minimizing index. ThenConsequently is exponential with rate , , and and are independent.
If is geometric on with parameter , independently of unit-rate exponential variables , thenis exponential with rate . Its moment-generating function is for .
A continuous uniform distribution has constant density on its interval of support.
Write . The two pieces in a uniform split are independent exactly whenSetting shows that for some . Thus the pieces are independent exactly when they are independent exponential variables of the same rate; equivalently, the total has gamma density .
The centred Laplace distribution with rate has density .
In the modelmaximum likelihood minimizes the sum of absolute residualsThe scale estimate is .
A distribution symmetric about is invariant under ; when its mean exists, it equals .
A median is a point with at least half the probability on either side.
The moment-generating function is where finite and encodes moments through derivatives at zero.
The joint distribution function is ; mixed differentiation gives a joint density when sufficiently smooth.
The normal distribution has density proportional to .
The absolute value of a standard normal random variable has density
A positive random variable is log-normal when its logarithm has a normal distribution. If and , then .
The standard normal distribution is and has cumulative distribution function .
The standard normal distribution function is
A variable is the sum of squares of independent standard normal variables.
If , coupling by extra independent squared normals gives with , so every fixed upper quantile increases with the degrees of freedom.
The Weibull distribution has density on .
For uniform on , .
The expansion converts products of small sinc factors into Gaussian characteristic-function limits.
A uniform variable on has mean zero and variance .
A lifetime exponential with rate survives for at least duration with probability .
The expected value is the probability-weighted average of a random variable, defined by a sum or an integral when it exists.
The variance of a square-integrable random variable is
For independent square-integrable random variables,
The standard deviation is the nonnegative square root of the variance.
The covariance of square-integrable random variables isIt determines the variance of a difference through
For a random vector with finite second moments, its covariance matrix isIt is positive semidefinite, and the variance of is .
For random variables with finite positive variances,
For nonnegative , ; for nonnegative integer , .
Expectation is linear without requiring independence: whenever the expectations exist.
For independent identically distributed variables of variance , the mean of samples has variance .
Probability inequalities bound event probabilities using moments or other tractable quantities.
For a nonnegative integer-valued random variable , Markov inequality givesConsequently, proves that some outcome has , while shows that with probability tending to one.
For a nonnegative random variable , Chebyshev's inequality givesHence implies that with probability tending to one.
Applying Markov's inequality to givesOptimizing this expression is the basic Chernoff-bound method.
The union bound says .
Conditional probability is when .
For a finite partition with positive probabilities,
A conditional distribution is the probability law of one random quantity given information about another.
A random partition is a probability distribution on partitions of a finite or countable set.
A restricted-growth string encodes a set partition by numbering blocks in order of first appearance.
A singleton block is a block containing exactly one element.
An exchangeable random partition has a law invariant under finite permutations of the underlying labels.
In a fair random root-to-leaf path, a specified leaf at depth is reached with probability .
A Markov chain has a future conditional distribution depending on the present state alone.
A birth-death chain is a discrete-time Markov chain on nonnegative integers that can move only to the same state or a neighbouring state.
Independent Markov chains with transition matrices and form a product chain with transition matrix . If their stationary distributions are and , the product distribution is stationary.
The transition matrix of a finite Markov chain has entries
A Markov chain is irreducible when every state can reach every other state along a path of positive-probability transitions.
A state is aperiodic when the greatest common divisor of its positive-probability return times is one. A Markov chain is aperiodic when all its states are aperiodic.
A state is recurrent when a Markov chain started there returns to it with probability one.
A state is transient when its probability of ever returning is less than one.
The first return time to a state is
The mean recurrence time of is . In a finite irreducible Markov chain with stationary distribution , it equals .
If two finite Markov chains satisfythen every positive-probability path and return cycle for also exists for . Consequently irreducibility and aperiodicity pass from to . Recurrence and mean return time need not do so, because the added transitions and changed probabilities can create escape routes or make returns less frequent.
For independent random variables with a common law, satisfies . The next-state law therefore depends on the past only through .
For i.i.d. Bernoulli variables, hides which summand is the newest one. Given , the previous value can reveal and thereby change the conditional law of , so need not be a Markov chain.
A Markov chain is reversible when its stationary flow satisfies detailed balance.
Detailed balance is the identity pi_i P_ij = pi_j P_ji for every pair of states.
A hitting probability is the probability of reaching one set before another and solves a discrete harmonic boundary problem.
Suppose an urn starts with green balls and red balls. Drawing green removes one green ball, while drawing red adds one green and one red ball. The difference remains two, and is harmonic for the green-ball chain. Stopping on hitting or and then letting tend to infinity gives termination probability
For a target set , the expected hitting time satisfies on andoutside , whenever the expectation is finite.
The expected number of visits to a state before absorption satisfies the reward equationwith zero boundary data at absorbing states where counting stops.
For the chain with absorbing, equal-probability nearest-neighbour moves between , and a probability- self-loop at , starting from the expected absorption time is , the probability of visiting before is , and the expected number of visits to before absorption is .
A conditional hitting time measures time to a target under a specified successful hitting event and can be computed by weighted first-step equations.
A periodic chain can return to a state only at times sharing a common divisor greater than one.
A lazy chain stays put with positive probability, removing periodicity without changing invariant distributions.
A finite irreducible aperiodic Markov chain converges to its unique invariant distribution.
A state is recurrent when the chain returns to it almost surely after starting there.
A recurrent state is positive recurrent when its expected return time is finite.
A recurrent state is null recurrent when its expected return time is infinite.
A state is recurrent exactly when the sum over time of its return probabilities diverges.
For nearest-neighbour random walk on with upward probability , the probability of ever hitting zero from is .
A random walk on a graph moves along an incident edge at each step.
A simple random walk on a locally finite graph chooses each neighbouring vertex with equal probability.
For the symmetric walk and ,Since the central binomial coefficient satisfies , this expectation is of order .
For independent symmetric walks on , the probability that all are at the origin at time isThe Borel-Cantelli lemmas therefore show that simultaneous returns occur only finitely often almost surely when .
On a finite undirected graph, stationary mass is proportional to vertex degree.
A biased nearest-neighbor walk steps right and left with unequal probabilities.
Let a discrete-time chain move up with probability and down with probability away from zero, while at zero it moves up with probability and stays with probability . It is transient for , null recurrent for , and positive recurrent for . In the positive-recurrent case its invariant distribution isand aperiodicity from the self-loop at zero implies convergence to this distribution.
Gambler's ruin studies a nearest-neighbour random walk stopped on reaching either endpoint of a finite interval.
For a symmetric random walk started at and stopped on first reaching or , the mean duration isIn particular, . This follows either from the first-step recurrence or by stopping the martingale .
At an almost surely finite stopping time, a Markov process restarts from its stopped state with the same transition law and independently of the history conditional on that state.
Let and . Before returning to , the chain visits zero times with probability and, for , exactly times with probabilityWhen , the expected number of visits is .
For an irreducible positive-recurrent Markov chain with invariant distribution , the expected number of visits to during one return cycle from to isCombining this with the two-state excursion visit law gives
The Markov property says that, conditional on the present state, the future evolution is independent of the past.
A continuous-time Markov chain waits an exponential time of rate in state and then jumps according to probabilities .
A finite-state process has generator exactly when, for every function on the state space,is a martingale with respect to the process filtration.
A Q-matrix has nonnegative off-diagonal entries and row sums zero. Its diagonal entry is the negative total rate out of the state.
For a continuous-time Markov chain, and the matrices satisfy . On a finite state space with Q-matrix , .
For a finite-state continuous-time Markov chain with Q-matrix , the backward and forward equations are and , respectively.
For transition rates , the generator acts on a function byWhenever the expectations are finite,Choosing polynomial functions gives differential equations for the moments.
Conditioned on a finite jump path , the holding times are independent exponentials with the corresponding rates. Multiplying the probability of occupying at time by turns its simplex density into a product symmetric under reversal of the time portions and the state sequence.
For , the fixed-time skeleton has matrix . If is the holding rate, then for ,Thus diverges exactly when does, proving recurrence equivalence for an irreducible chain.
The jump chain records the successive states visited by a continuous-time Markov chain while discarding its holding times.
On , let the chain jump from every to with probability and to with probability , and jump from zero to one. It is recurrent exactly when ; at it is null recurrent.
For holding rates , a measure satisfies exactly when is invariant for the jump chain. Thus when normalization is possible.
Even when the jump chain has invariant probability , the continuous-time invariant measure may have infinite mass if holding rates become too small, producing infinite mean return time.
A recurrent jump chain with infinite invariant measure can yield a positive-recurrent continuous-time chain when .
If the jump chain visits a state of finite holding rate infinitely often, the sum of its independent positive holding times at that state diverges almost surely, preventing explosion.
Explosion occurs when infinitely many jumps take place in finite time.
Conditional on a jump path , if , then the sum of exponential holding times has finite expectation and is finite almost surely.
Suppose a transient nearest-neighbour jump chain on has uniformly bounded expected visits to each state and the holding rate at grows geometrically. Thenso the continuous-time chain explodes almost surely.
Under the generator convention, an invariant distribution satisfies .
An explosive chain can have a summable solution of despite being transient. Nonexplosion is needed for the usual equivalence between invariant distributions and positive recurrence.
A batch-birth linear-death process has transitionsWriting and taking for , its master equation is
A birth-death process moves only between neighbouring nonnegative integer states, with state-dependent birth and death rates.
The Yule process is the pure birth process with rate in population state . Starting from one individual, it is nonexplosive because the independent holding times have divergent total almost surely.
Given , the birth times in a Yule process have the law of the order statistics of independent variables with densityThis follows by dividing the joint holding-time density by
.
.
For individual birth and death rates and one initial individual, the finite-time extinction probability satisfiesWhen ,
For birth rates and death rates , a reversible measure satisfiesThus , and it is an invariant distribution when the resulting measure is summable and the chain is nonexplosive.
The th falling factorial moment is , where .
If is Poisson with mean , then ; in particular, and .
Approximating the population law by a Poisson law replaces its first two factorial moments by and , closing a quadratic first-moment equation as a Riccati equation for .
For unit jumps with birth and death rates , the Fokker--Planck drift and diffusion in the conventionare and .
Near a stable deterministic equilibrium, replace the drift by its linearization and the diffusion coefficient by its equilibrium value. The resulting Ornstein-Uhlenbeck approximation has a stationary normal distribution.
If boundary fluxes decay sufficiently fast, two integrations by parts givewhile the diffusion term contributes zero to the first moment.
For constant birth rate and death rate away from zero, the chain is transient when , null recurrent when , and positive recurrent when .
The queue length in an queue is a birth-death process with constant arrival rate and service rate while nonempty.
If potential arrivals form a Poisson process of rate and an arrival seeing customers joins with probability , the queue length is a birth-death process with birth rate and death rate for . Its stationary ratios, when normalizable, are
An queue is positive recurrent exactly when its traffic intensity is below one.
When , the equilibrium queue length has distribution and mean .
At equilibrium, the departure process of a stable queue is Poisson with the same rate as the arrival process.
If a served customer returns with probability , thinning service completions gives effective external departure rate for the queue-length process.
In equilibrium, a Bernoulli-feedback queue has the same queue-length process as an ordinary queue with service rate , so its external departures form a Poisson process of rate .
An queue has Poisson arrivals of rate , independent exponential service times of rate , and infinitely many servers. Its occupancy is a birth-death process with birth rate and death rate in state .
For every , an queue is positive recurrent and has the Poisson distribution with mean as its unique stationary distribution:
In stationarity, the departure process of an queue is a Poisson process of rate . By time reversibility, the departures before any fixed time are independent of the occupancy at that time.
If and , thenwith independent summands. The binomial term counts surviving initial customers; Poisson thinning gives the surviving later arrivals.
Writing , the stationary empty probability is . Alternating idle periods of mean with busy periods and applying the renewal-reward theorem gives
Two states of a Markov chain communicate when each can be reached from the other with positive probability in finitely many steps.
A communicating class is closed when no transition from it can reach a state outside it.
In a finite chain, a communicating class that is not closed is transient: after it is left, it need not be visited again and is eventually left almost surely.
A stationary distribution satisfies .
A renewal process with independent identically distributed nonnegative interarrival times has renewal epochsand counting process .
For independent identically distributed regenerative cycles with finite mean length, the long-run reward rate equals expected reward per cycle divided by expected cycle length.
If , then the strong law and the inequalities imply almost surely.
The renewal interval containing has lengthSelection by a fixed observation time favours longer interarrival times.
For a nonnegative random variable with , its size-biased version is defined byfor bounded measurable . If has density , then has density .
For exponential interarrival times of rate , the age and residual life at time converge jointly to independent rate- exponential variables. Their sum therefore converges to the gamma distribution with density , the size-biased exponential law.
At every deterministic , the containing interval stochastically dominates an unselected interarrival time:Conditioning on the first interarrival gives a renewal equation; subtracting the constant tail probability leaves a renewal equation with a nonnegative forcing term.
The residual lifetime is the time remaining until the next renewal.
For non-arithmetic inter-renewal times with finite positive mean , the excess converges in distribution to with
If an independent superposition of a Poisson process and a non-arithmetic finite-mean renewal process is itself a renewal process, then its first interarrival is exponential. Equating the limiting excess law computed as a renewal excess with the minimum of the Poisson and second-process excesses yields an integral equation whose survival-function solutions are exponential.
With bounded integer lifetimes, residual lifetime decreases deterministically and resets randomly at zero.
The invariant residual-life mass at i is proportional to the probability that a lifetime exceeds i.
Variance equals the variance of a conditional mean plus the mean conditional variance.
The characteristic function of a real random variable isIt determines the probability distribution of .
Independence makes the characteristic function of a sum equal the product of the individual characteristic functions.
Dividing a centered sum by the square root of its total variance gives unit variance and is the natural central-limit scaling.
An alphabet is a finite nonempty set of symbols from which strings or source outputs are formed.
For discrete random variables, conditional entropy is
If independent random variables take values in a finite additive group, thenIndeed, , and symmetrically for . Independence is essential: taking makes the sum constant.
The mutual information between discrete random variables is
The information supplied by two observations decomposes as
Send a uniform bit twice through independent binary symmetric channels of crossover probability . The first received bit supplies bits. Since the received bits disagree with probability , the additional information in the second bit is
The capacity of a discrete memoryless channel is
If every row of an -output channel matrix is a permutation of every other row and all column sums agree, the uniform input achieves a uniform output and
A binary symmetric channel with crossover probability has capacitybits per channel use.
For a discrete random variable with probabilities , its Shannon information entropy is
The binary entropy function is
Huffman's algorithm repeatedly combines the least weights for a -ary prefix code and yields an optimal expected length.
Some optimal prefix tree places the two least probable symbols as sibling leaves of maximum depth, which is the inductive basis of Huffman's algorithm.
An optimal prefix code minimizes probability-weighted codeword length among all prefix codes for the source.
If symbols of probabilities have codeword lengths and in an optimal prefix code, then . Otherwise, exchanging their codewords preserves the prefix property and changes the expected length bycontradicting optimality.
Every optimal binary prefix code has two codewords of maximal length that differ only in their final bit. Indeed, choose a deepest leaf in the prefix tree. If its sibling were not a codeword, maximality of the depth would make that sibling subtree empty, so the chosen word could be shortened by deleting its final bit. This would preserve prefix-freeness and reduce expected length.
Distinct tree shapes can have the same minimum expected length, particularly when Huffman merges have ties.
For a largest binary-symbol probability , prevents a length-one word, while forces one in a Huffman tree.
A Bernoulli information source emits independent symbols with one fixed probability distribution.
The information rate is the limiting entropy per emitted symbol, when that limit exists.
A source is reliably encodable at rate if there are block encoders with at most outputs and corresponding decoders whose probability of reconstructing the first source symbols incorrectly tends to zero as .
For a source and fixed positive integer , form the block source . Whenever the original information rate exists, the blocked source has information rate , because
For a stationary source, typical long blocks have probability approximately ; equivalently normalized self-information converges to the entropy rate.
For binary prefix codes, the optimal expected length obeys .
Unicity distance is the ciphertext length at which key equivocation reaches zero and the key becomes uniquely determined under the source model.
Key equivocation is the conditional entropy remaining about a key after observing ciphertext symbols.
For ciphertext alphabet and source entropy , each symbol has redundancy , leading in Shannon's ideal model to .
If distinct keys transform every ciphertext into equally probable plaintexts through a probability-preserving source symmetry, key equivocation never vanishes and unicity distance is infinite.
An estimator of a parameter is unbiased when for every parameter value in the model.
Observations or errors are heteroscedastic when their variances are not all equal.
For a model with fitted parameters and maximized likelihood , the Akaike information criterion is
For a normal linear model with known error variance , the scaled Akaike criterion differs by a model-independent constant from
If are independent vectors and is a rank- orthogonal projection, thenThe first term is squared approximation bias, while is fitted-model variance and is irreducible new-response noise.
If is the rank- orthogonal projection onto a normal linear model's column space, thenAdding gives , so Mallows' is unbiased for independent-copy prediction error even when the projection model is misspecified.
A shape parameter changes the form of a probability distribution without merely translating it or rescaling its random variable.
A density has sufficient statistic and cumulant function .
The cumulant function, or log-partition function, is the normalizing termwith a sum in the discrete case.
Differentiating the normalization identity gives
The mean parameter corresponding to the natural parameter is
For an expected value parameter and a shape parameter , the inverse Gaussian distribution has probability density functionIts mean is and its variance is .
For fixed positive integer , the failures-before-the-th-success distribution is a one-parameter exponential family withand carrier .
With dispersion and known weight , an exponential dispersion density has exponentIts mean is and its variance is .
A generalized linear model takes independent responses from an exponential dispersion family and relates their means to linear predictors by . The response distribution, systematic component, link, dispersion, and known observation weights together specify the model.
The linear predictor of a generalized linear model is the linear combination of the explanatory variables.
A link function connects the response mean to the linear predictor through .
Binomial regression models independent and relates to a linear predictor. Logistic and probit regression use the logit and inverse-normal links respectively.
For a generalized linear model with variance function , the Pearson estimator is
The canonical link identifies the linear predictor with the exponential family's natural parameter: , where .
For a Poisson mean , the natural parameter is , so the canonical link is the logarithmic link .
A Poisson regression models independent count responses byIts likelihood is .
For fixed exposures , a Poisson exposure model takes independent observationsIt has mean and variance and is therefore heteroscedastic when the exposures differ.
For , unweighted least squares through the origin and maximum likelihood giveBoth estimators are unbiased.
The two variances areSince by Cauchy-Schwarz, maximum likelihood has no larger variance, with equality when all positive exposures are equal.
For testing against a larger alternative, approximate size- upper-tail tests reject whenor, respectively,
Logistic regression models the conditional log odds as an affine function of predictors.
For binary independent responses, logistic regression sets
At an interior maximum-likelihood fit of a logistic-regression model containing an intercept, the intercept score function isConsequently the sum of the fitted probabilities equals the number of observed successes exactly.
For a positive semidefinite trace ball , the quadratic-form hypothesis class isIts empirical logistic risk is convex as a function of .
For independent grouped counts , grouped-binomial logistic regression specifiesFitting the proportions with binomial weights gives the same likelihood.
With treatment coding, the intercept describes the reference factor level and each displayed factor coefficient is a contrast against that level. Changing the reference level reparametrizes the same fitted model: fitted probabilities and likelihood do not change, but coefficients representing different pairwise contrasts can have different standard errors and p-values.
Binary probit regression sets , where is the standard normal distribution function. Its intercept score is a weighted sum of residuals, so an intercept does not generally force .
IRLS fits a generalized linear model by repeatedly solving a weighted least-squares approximation to its score equations.
Independent Poisson variables conditioned on their sum are multinomial, with cell probabilities proportional to their rates.
, , and full-rank gives .
Orthogonal projections of an isotropic Gaussian vector are independent. If is an orthogonal projection of rank and , then
Linear regression models a response mean as a linear combination of predictors. Ordinary least squares chooses to minimize the sum of squared residuals.
The design matrix has one row for each observation and one column for each regression coefficient, so the linear predictor is .
A regression coefficient is a component of measuring the change in the linear predictor associated with its design-matrix column.
An interaction term lets the effect of one predictor depend on another. Inthe slope with respect to is .
If but the predictor is observed as , with independent centred errors, the population regression slope isMeasurement error in the predictor therefore shrinks the slope toward zero, while independent response noise changes its uncertainty but not its population value.
In R,
lm(y ~ 1, data=d) fits an intercept-only linear model, lm(y ~ x1 + x2, data=d) adds quantitative predictors, and a factor predictor is expanded into indicator columns.With an intercept present, a categorical predictor having represented levels adds linearly independent indicator columns and therefore model degrees of freedom.
The hat matrixis the orthogonal projection onto the column space of a full-rank design matrix, and the fitted response is .
If is the least-squares estimate after deleting observation , Cook's distance isIt measures the deleted estimate in the same quadratic metric as the coefficient confidence ellipsoid.
Multicollinearity means that predictor columns are strongly linearly related. It inflates coefficient variances and can make individual effects imprecise even when a joint test strongly rejects that all corresponding coefficients vanish.
In simple linear regression with a nonconstant predictor,
For a nonconstant predictor, putThe minimized residual sum of squares isEach centered sum can be recovered in constant time from the five raw sums , , , , and .
Among linear unbiased estimators in a homoscedastic linear model, least squares has minimum variance.
Weighted least squares minimizes ; inverse-variance weights give the efficient estimator under known heteroscedasticity.
Generalized least squares minimizes for a known covariance shape .
Whitening multiplies a model by a square root of its precision matrix so that the transformed errors have identity covariance.
Least-squares differentiation gives , or its precision-weighted analogue.
Under finite moments and a nonsingular limiting design moment, laws of large numbers make least-squares estimators converge to the true coefficient.
A one-way normal model assigns each factor level its own mean while assuming independent Gaussian errors with a common variance.
A cell-means model uses one coefficient per factor level and no intercept, so each coefficient directly equals a group mean.
A linear contrast is a coefficient-weighted sum of cell means whose coefficients total zero, used to test interpretable differences among groups.
For genotypes , full dominance of allele imposes .
With allele counts , no dominance makes the genotype mean affine in count and imposes .
For , leverage is and residual variance is .
The risk ratio compares event probabilities in two groups:
When worsening is the adverse event, drug efficacy relative to a control group is
The odds ratio is the ratio of two event odds. For rare events it is close to the corresponding risk ratio, because when is small.
Additive main effects encode mutual independence in contingency-table Poisson models. An interaction permits association of its factors.
A saturated log-linear model has enough parameters to reproduce every observed contingency-table cell count and therefore has zero residual deviance.
For treatment groups Control, LD, and SD and binary outcomes, equal LD and SD efficacy can be imposed while retaining separate group totals by using separate treatment main effects but one shared treated-versus-control interaction with outcome. Comparing this five-parameter model with the saturated six-parameter model gives a one-degree-of-freedom analysis of deviance for nested generalized linear models.
An MLE maximizes sample likelihood. Regularly, ; parameter-dependent support can change both rate and limit.
For observed data , the likelihood function is the joint probability mass or density regarded as a function of the model parameter .
The log-likelihood is the logarithm of the likelihood as a function of the model parameter. Its maximizers are the same because the logarithm is strictly increasing.
For independent exponential observations with rate , the MLE is and has asymptotic variance .
The logistic location family has density and distribution function
LDA assumes class-conditional Gaussian distributions with a shared covariance matrix, producing affine log odds.
Quadratic discriminant analysis uses class-conditional Gaussian distributions with class-dependent covariance matrices. Their log density ratio is a quadratic polynomial in the observation.
The bootstrap approximates a sampling distribution by repeatedly sampling with replacement from the empirical distribution.
The risk of an estimator is its expected loss as a function of the unknown parameter. Under quadratic loss,
The mean squared error of an estimator is
Every estimator with finite second moment satisfies
An estimator is admissible when no other estimator has risk no larger at every parameter value and strictly smaller at at least one value.
A minimax estimator minimizes the supremum of its risk over the parameter space.
The Cramer-Rao inequality lower-bounds estimator variance by inverse Fisher information, with a derivative correction for bias.
The score is the gradient of the log-likelihood, .
Under regularity permitting differentiation under the integral, .
The Fisher information is and equals for an independent identically distributed sample.
Fisher information tensorizes when the information in an -observation model is the sum of the information contributions from its observations. For identically distributed observations this meansIndependence and the mean-zero score identity make the cross terms vanish.
Fisher scoring replaces the observed negative Hessian in Newton iteration by the Fisher information:
For , the score and information areConsequently one Fisher-scoring step sends every interior starting value directly to the MLE .
Under standard regularity conditions, .
For , the one-observation score is and the information matrix is .
Statistical hypothesis testing compares data against a null hypothesis using a controlled rejection probability.
The null hypothesis is the set of parameter values or probability distributions against which a statistical test controls its probability of rejection.
The significance level is the chosen upper bound on a test's probability of rejecting the null hypothesis when it is true.
A p-value is the probability, computed under a null hypothesis, of obtaining a test statistic at least as unfavorable to that hypothesis as the observed value.
A joint hypothesis test imposes several parameter restrictions simultaneously and accounts for dependence among their estimators; separate one-parameter confidence intervals do not generally determine its result.
For a test with rejection region , its power function isIts supremum over the null parameter space is the size of the test.
A level- test is uniformly most powerful when its power is at least that of every other level- test at every parameter value in the alternative.
If two tests have the same size and each has strictly greater power than the other at some alternative, neither test is uniformly most powerful.
For observed cell counts and null expected counts , the statistic isWith specified null cell probabilities and sufficiently large expected counts, its null limit is .
For an by contingency table, the statistichas asymptotically a chi-squared distribution with degrees of freedom under independence.
For an approximately normal estimator, the one-parameter Wald statisticis approximately standard normal under . A two-sided test reports .
For a regular -parameter model, the Wald statistic for a candidate isUnder the true parameter it converges in distribution to .
For a full-row-rank matrix and the null hypothesis , putUnder the null, asymptotic normality and Slutsky theorem give .
In a regular scalar model under the null, Taylor expansion about the MLE givesA uniform law of large numbers for the observed information makes
and the Fisher information used in the Wald statistic converge to the same positive limit. Their ratio therefore converges in probability to one.
and the Fisher information used in the Wald statistic converge to the same positive limit. Their ratio therefore converges in probability to one.
A likelihood-ratio test rejects where the alternative likelihood is sufficiently large relative to the null likelihood.
For two parameter values and , the likelihood ratio is .
A generalized likelihood-ratio test compares the likelihood maximized under the null with the likelihood maximized over the full parameter space.
The deviance of a fitted generalized linear model isFor regular nested models, the deviance reduction is asymptotically chi-squared with degrees of freedom equal to the difference in fitted dimensions.
For independent unit-variance samples, testing equal means uses and rejects for large .
Likelihood-ratio tests of different null subspaces can have crossing rejection regions; comparing their quadratic-form geometry exhibits observations accepted by one and rejected by another.
The score test evaluates the likelihood gradient at the restricted estimator and scales its quadratic form by inverse Fisher information.
A restricted maximum-likelihood estimator maximizes likelihood over the null parameter space rather than the full model.
Under a regular simple -parameter null, the normalized score converges to , so its information-standardized squared norm converges to .
If , then .
For independent observations, has the distribution.
If , then converges in distribution to .
For testing one simple hypothesis against another, a test that rejects for sufficiently large likelihood ratio has greatest power among all tests of no greater size.
A one-parameter family has a monotone likelihood ratio in a statistic when is nondecreasing in whenever .
For a family with a monotone likelihood ratio in , a level- upper-tail test in is a uniformly most powerful test for the corresponding one-sided hypothesis.
Flooring an exponential variable of rate produces a geometric variable on the nonnegative integers with success probability .
Coarsening continuous observations changes the delta-method variance; for floored exponential data the rate estimator has limiting variance .
For independent and prior , the posterior has variance and meanAs , and , so fixed-level posterior credible intervals agree asymptotically with the corresponding known-variance confidence intervals.
For nested Gaussian linear models differing by parameters, the statistichas an distribution under the reduced model.
A Galton--Watson process forms each generation by giving every current individual an independent number of offspring with a common distribution.
If is the offspring probability-generating function, the eventual extinction probability is the smallest fixed point of in .
For offspring mean and variance , Taylor expansion at one givesfor the extinction probability .
If nonnegative quantities satisfythen iteration givesFor , the growth factor is bounded by .
For an irreducible positive recurrent Markov chain with stationary distribution , the expected return time to state is . More generally, expected occupation rewards in a return cycle equal their stationary rate times the expected cycle length.
A Markov additive process couples a Markov chain to an additive functional whose increments depend on the chain.
Two subvectors of a jointly multivariate normal distribution are independent exactly when their cross-covariance matrix is zero.
For means , standard deviations , and correlation , the standardized quadratic form in the density is
For independent with known zero mean,
Conditionally on ,
Any jointly normal random variables with zero covariance are independent.
Orthogonal projections of an isotropic Gaussian vector onto orthogonal subspaces are independent.
If is a sufficient statistic and is an estimator with finite variance, then has the same expectation and no greater variance. Indeed, the tower property of conditional expectation preserves the mean and the law of total variance gives
For independent , the unbiased estimator of has mean square error . Conditioning on the sufficient statistic giveswhose mean square error is .
A statistic is sufficient for a parameter when the conditional distribution of the full data given is independent of that parameter.
For independent observations with known variance, is minimal sufficient for . Factorization proves sufficiency, while the likelihood ratio for two samples is independent of exactly when their sums agree.
A sufficient statistic is minimal sufficient when it is a function of every other sufficient statistic, up to null sets. Its level sets form the coarsest sufficient partition of the sample space.
A statistic is sufficient exactly when the likelihood factors into a parameter-dependent function of the statistic times a parameter-free function of the sample.
A Monte Carlo method approximates a deterministic quantity using averages of simulated random variables and controls the approximation with probabilistic limit theorems.
Suppose is a target probability density and is a proposal density with . Draw and an independent , and accept whenThe accepted value has density , the acceptance probability is , and the expected number of proposals is .
To estimate , sample independently from a reference density whose support covers that of , and useThe summands have expectation , and the strong law of large numbers gives almost-sure convergence under integrability.
For an unnormalized target density, self-normalized importance sampling uses weights and estimates expectations by . Resampling a point with probability proportional to produces the associated weighted empirical distribution.
Statistical inference estimates unknown parameters and quantifies uncertainty from observed data.
For an empirical distribution function based on an independent sample from a continuous distribution ,where is a Brownian bridge.
For a symmetric bandwidth- averaging kernel and a density with bounded first derivative, the interior pointwise bias is at most a constant times . Near a support boundary, an uncorrected symmetric kernel can instead have nonvanishing bias.
The standard error of an estimator is the standard deviation of its sampling distribution, or an estimate of that standard deviation.
Under regularity conditions, minus twice the logarithm of a generalized likelihood ratio converges under the null hypothesis to a chi-squared distribution whose degrees of freedom equal the difference in parameter dimensions.
For an by contingency table, Pearson's sum of squared observed-minus-fitted counts divided by fitted counts is asymptotically under independence.
Statistical decision theory compares decision rules through the expected loss they incur under each parameter value.
Under zero-one loss, a Bayes classifier assigns an observation to a class of greatest posterior probability. For two classes with prior probabilities and densities , it chooses class one exactly when .
For two multivariate normal densities, the Bayes log posterior odds are a quadratic polynomial. Equal covariance matrices cancel the quadratic term and give linear discriminant analysis; unequal covariance matrices give quadratic discriminant analysis.
A binary Bayes classifier under zero-one loss is unique up to null sets when the posterior class probabilities tie only on a null set. For two distinct nonsingular Gaussian distributions, the tie set is the zero set of a nonzero quadratic polynomial and has Lebesgue measure zero.
For a prior , the integrated risk is . Its infimum over decision rules is the Bayes risk, attained by a Bayes rule when one exists.
A least favorable prior maximizes the Bayes risk over the allowed priors. If a Bayes rule has constant risk equal to the minimax value, its prior is least favorable.
A minimax rule minimizes the worst-case risk .
An equalizer rule has the same risk at every parameter value. A Bayes equalizer rule is minimax when its constant risk bounds from below the worst-case risk of every competing rule.
For two distinct nonsingular Gaussian class distributions, vary the class-zero prior continuously. The class-zero and class-one error probabilities of the corresponding Bayes classifier cross. At a crossing they are equal, making the classifier an equalizer rule, a minimax decision rule, and the associated prior a least favorable prior.
If a rule has constant risk and a sequence of priors has Bayes risks tending to , then the rule is minimax. Every competing rule has worst-case risk at least every one of those Bayes risks.
A Student t confidence interval replaces unknown Gaussian scale by an independent residual estimate.
In a two-parameter Gaussian simple linear regression centered at , the intercept estimator is . With residual standard deviation ,which gives the usual two-sided interval from the corresponding quantiles.
Bayesian statistics updates a prior distribution to a posterior distribution using likelihood.
A prior probability describes uncertainty about an event or parameter before the current observation is incorporated.
The posterior density is proportional to likelihood times prior density.
A point-null mixture prior assigns positive mass to one parameter value and distributes the remaining mass continuously over alternatives. Posterior point mass is obtained by dividing the null component of the marginal density by the full marginal density.
For the equal point-null mixture,
With a fixed diffuse alternative prior, a frequentist p-value and the posterior probability of a point null can order evidence differently. For and , the p-value exceeds the null posterior near zero, while sufficiently far in the tails the null posterior exceeds the p-value because it has a slightly slower Gaussian decay.
An improper prior is a nonnegative prior kernel with infinite total mass. It can still produce a proper posterior after multiplication by the likelihood and normalization.
For a Poisson observation of mean and a gamma prior with shape and rate , the posterior is gamma with shape and rate .
A gamma prior combined with exponential observations gives a gamma posterior.
A beta prior combined with Bernoulli successes and failures gives posterior .
A Bayes estimator minimizes posterior expected loss.
Under squared error loss, the Bayes estimator is the posterior mean.
Under , the posterior expected loss is minimized atprovided the conditional inverse moment is finite and positive.
Posterior expected loss averages the loss over the posterior and is minimized pointwise in the data.
Markov chain Monte Carlo constructs an ergodic Markov chain whose stationary distribution is the target law and uses its post-burn-in states as approximate dependent samples.
For a joint density , a two-coordinate Gibbs update draws
The joint target density is stationary for a full Gibbs sweep, since
For independent , a flat prior on , and an exponential prior of rate on , the full conditionals areandwhere the gamma distribution is parametrized by shape and rate.
For an estimator and leave-one-out versions , the jackknife bias estimate and corrected estimator areIf , the correction leaves bias .
For observations , the unbiased sample variance isFor independent identically distributed observations with finite variance , it converges in probability to .
An estimator is asymptotically normal when a rescaled estimation error converges in distribution to a normal law.
Slutsky's theorem combines convergence in distribution with convergence in probability through continuous algebraic operations.
For estimating equations, asymptotic covariance often has the form , called a sandwich covariance.
If is asymptotically normal, differentiability gives the variance multiplied by .
For a centered bivariate normal sample,This is the multivariate delta method for .
For an independent sample from , the maximum likelihood estimator is andThe parameter-dependent support makes the convergence rate rather than , so the regular maximum-likelihood central limit theorem does not apply.
Asymptotic relative efficiency compares the leading asymptotic variances of two consistent estimators.
A Wald interval centers at an asymptotically normal estimator and uses a consistent estimated standard error times a normal quantile.
The best affine predictor has slope covariance divided by predictor variance and intercept chosen to match means.
The residual variance after projection onto one factor is , a covariance Schur complement.
Codex Wiki