As a preliminary, consider the 'football field' model of an unbiased random walk. Let's say there is a 30% chance of an event E. Either it will occur, or it won't.
Now, imagine we are on the 30 yard line of a football field, and (continuously) randomly walking; reaching the opponent's goal line means E occured, retreating to our own goal line means E failed. It can be shown via probability theory, that we will 'score' with 30% probability (thank goodness). The formula is: (30 - 0) / (100 - 0)
That is, the distance to our own goal line, divided by the field length.
In the problem above, we receive the message "E = .5". This message is an event; call it F. F has occurred, and we are so informed. What was the a priori probability of F?
To get that, we return to the football field model. We are on the 30 yard line (E = .3). What is the chance that we will cross the 50 yard line (E =.5)? Shorten the football field model, to a total 'field length' of 50 yards. Then, similarly to the above, we compute "the probability we will cross the 50 yard line before our own goal line":
30/50 = .6
So P(F) = .6 and I(F) = log(1/.6) = .74 Hence the message "E occurs with probability .5" contains .74 bits of information.
This part I'm not so sure about.
We could figure the updated probability of E as .46, i.e. the weighted average between .3 and .5, then proceed as before, to get I(F) = log (1/(30/46)) = .62 bits
Or, use .74 bits, as previously, then 80% of that = .59 bits
Mark
Didn't find your answer? Ask the community — no account required.
C
Curt Welch
I assume you were trying to say, log2(1/0.3) instead of log(1/3).
Of course, because the information of a certain event (log2(1/1)) is zero. So log2(1/.3) - 0 is of course the same thing as log2(1/.3).
Yes, it is tricky. The point of the message seems to have gone over your head. Nothing is assumed, or known about the prior probability of the message showing up so we can't calculate the information content when it shows up in this problem. The problem is more complex than that.
What I was trying to address (and which I have no idea if information theory in any of its forms has formally addressed because I'm no expert in the field), is how we would measure the information content in an event which created an update to our prior knowledge about some future event which had not yet happened.
Simple information theory that you get in an introduction doesn't address this case at all. You are talking about nothing but the simple information theory that measures the information content of an event when it happened, based on our prior knowledge of it's probability of happen. That just doesn't apply to this case.
Right, but we don't know the prior probability of F in this problem, so we can't calculate that information value.
That's what they used to say about numbers as well! :)
I don't pretend to know much about information theory, but I do know the general concept that information is never negative. And in the simple form of information, where we have a prior probability (
C
Curt Welch
Well, this is an interesting approach, but it makes an assumption not given by the problem. The assumption is that the underlying system controlling event F is the same as your football field model.
I could for example pick another assumption, and come up with a different answer using your approach. Lets say that event E is is a green light flashing and NOT E is a red light flashing. It's controlled by a computer connected to some highly random external source. But the computer is programed to produce event E with probability .3, and event NOT E, with a probability of .7.
But, the program is actually a bit more complex. Before the program produces the answer, a human flips a coin, and enters the result into the computer. For heads, the computer lights a yellow light, and changes it's probability of producing event E (the green light) to .5. For tails, the yellow light does not light, and the probability of E (the green light) is .1. So half the time, the probability is .5, and half the time the probability is .5.
We don't know how the coin toss comes out, and we don't know when it will happen. We only know that E (green light) or not E (red light) will happen, and that the yellow light might or might not, come on some extended time before one of the other two happens.
So we call the yellow light, your event F.
At the start, the probability of E happening is .3 (.5*.1 + .5 * .5). We don't know how the coin toss will come out, so the probability is the long term average of the two possible events.
But, when the green light comes on, we know the probability of E happening changes to .5.
So how much information does the green light tell us using your way of calculating it? Well, the a priori probably of the green light coming on was .5, so the information must be log2(1/.5) or 1 bit.
So, you made up your story about the cause of message F, and came up with a probability of .74, and I made up my story about the cause of message F, and came up with 1 bit. Both your story, and my story, fit the facts of the problem, but we came up with different answers.
In short, we know nothing about the probability of receiving message F, and as such, we can't calculate the amount of information in F, about F, unless we make assumptions not given in the problem. All we can calculate, is the amount of information in F, about E. How does F change the amount of information we will receive when E happens? That is a number we can calculate (as I did in my previous post). That's the only number we can calculate given the nature of the problem without making assumptions not specified in the problem.
Is it reasonable, or valid, to talk about how much information message F contains about future event E? In my limited exposure to information theory I've not seen this idea discussed. (As the person who posted the question pointed out the idea was not covered in their text book as well). But the way I outlined it seems to be perfectly logical and reasonable to me.
Informally, we talk all the time about message F containing information about E. Formally, the normal use of information theory you get in an intro text book doesn't cover this. So, is this an idea that is covered by more advanced information theory? Or is it something that has just not been looked at? Or is it something that just can't be looked at like this with formal information theory for some reason?
It's true, I saw it on Fernwood Tonight. Fernwood's best scientist (and only barber) analyzed the polarization of light by the circular weave of polyester fabric, and found it selectively permeable to harmful radiation. To test his theory he dressed up lab rats in custom made tiny leisure suits. You didn't see that one? It was right after the show with the woman who was violated by a beam of light from outer space, and just before the one with the irate guy who could not sleep because some woman down the street was on her roof all night yelling "Come back! Come back!"
Now that's information.
-- Michael
J
jon
Z
zzbunker
None. All probabilities are dynamic. It's just a matter of how stable the distribution is Since events that occur with probability. 0.5 is how the word "bit" was even first evented.
M
Mark-T
That may indeed be a fatal weakness.
huh?
OK
Where does .1 come from?
? The green light IS event E.
.3
I think you have a point here, despite the errors.
Not quite. More precisely, we try to calculate the information conveyed (contained) by F.
log(1/.3) - log(1/.5) bits
It may simply be ill posed, after all. We just don't know the a priori probability of F, hence can't make any statements about it.
My random walk model may only apply to some subset of cases.
Nah, it's not some hole in information theory. I'd say it's an example which no one thought to discuss in a text book, being ill posed, or too trivial.
Mark
M
Mark-T
Good proofread.
snip
A message which tells you something you already know. e.g. "the probability of E is .3" Information is the amount of 'surprise', loosely speaking.
As trivial as it may seem, some people don't get that. I once had a communications project, with some error coding. The algorithm required zero insertion, at various points in the block, at both receiver and sender. No problem, easy enough to code.
Of course, there was no need to transmit those zeros, they carried no information. But preposterously, the client insisted we actually use bandwidth to send them, saying the block length was part of the spec! This person was a project manager, a position of responsibility. I thought: "Are you allowed to go out alone? Do you cross the street without an adult to hold your hand?
It was a lesson in the limits of human intelligence.
Mark
C
Curt Welch
Typo. I meant to write half is .5 and half is .1
It's just a number I made up so that the average of the two worked out to be .3 so it fit the original problem description (a probability that changed from .3 to .5 when an update message was received).
Again, a stupid mistake on my part. It's not easy to follow my logic when I make so many errors is it? That should have read:
But, when the YELLOW light comes on, we know the probability of E happening changes to .5.
YELLOW
YELLOW
Right, I was trying to talk about the yellow light, not the green light. The information in the yellow light (event F) is 1 bit in my example because it had a 50/50 chance of happening.
Yeah, the story makes a lot more sense when you don't fill it with errors.
I guess when we talk about language, we make the distinction of what the message tells us about other things. But in information theory, the norm seems to be limited to what the arrival of the message tells us about itself (when it arrives, the probability of it arriving changes from what it was before, to 100%).
Language however is typically used to make reference to other events - so it updates our information (probability distributions) about potential future events. It seems valid to me to talk about the information content of a message related to all the probabilities it updates.
And, as such, the message "carries" with it information based on what it tells us. But what's more interesting about this is that the message doesn't really contain the information as much as we measure it's "information content" based on how much it changes us (or changes the receiver in general).
It's not trivial unless you just consider it an invalid question. Then it's trivial only in the sense that it's nonsense - it has no answer.
But it's only ill posed, if you choose to limit your definition of information in a way that makes it ill posed.
And since the normal formal definition of information seems to be based on changes to expected probability of events, it doesn't seem possible to me to exclude this example as an invalid question. In the way the question was asked, where event F was an English language statement containing information about figure event E, it drags in all the complexity of language and meaning.
But when you transform it to something like my machine example with blinking lights, and how our expected probabilities of future lights (red and green) change based on if the yellow light flashes, it's hard to ignore the question of how much information the yellow light contains.
The yellow light in my example has a probability of 50% so when we see the yellow light, it gives us 1 bit of information about the yellow light. But it also adjusts our expectations of the soon to follow red or green lights. Should we ignore that adjustment to our probability and pretend it is meaningless to ask how much information it contained based on how it adjusted our expectation?
This seems to me to be a question that information theory can't write off as ill formed.
For example, lets look at the system I talked about above from a different perspective. Each run of the system produces one of 4 results with the following probabilities:
Event Probability Bits
Green .5 * .1 = .05 4.322 Red .5 * .9 = .45 1.152 Yellow Green .5 * .5 = .25 2 Yellow Red .5 * .5 = .25 2
But if we ignore the yellow light we see these possible results:
Event Probability Bits
Green .3 1.737 Red .7 .515
Entropy of channel .882 bits/symbol
If we ignore the Red and Green lights:
Yellow .5 1.0 No Yellow .5 1.0
Entropy of channel 1.0 bits/symbol
But if we combine them together ignoring their correlations, we should be able to add the Entropy to end up with a total of 1.882 bits per combined symbol. But since the Yellow light has a conditional probability effect on the Red and Green, the result turns out to be less information and we only end up with a total of 1.734 bits per symbol. .148 bits per symbol is lost.
The fact that the yellow and the red/green lights are correlated, means there is less total information in the channel. So, when we see the yellow light, it tells us something about the red/green that we would not have known if we just waited for the red green light.
The yellow light has 1 bit per symbol on its own.
But it also tells us that Green changes from .3 to .5, and the red changes from .7 to .5. So the net change in total entropy of the red/green event is:
.882 (aka .3 vs .5) to 1 (.5 vs .5) which is a gain of .118 bits.
But we also have to consider the not-yellow pseudo event. Which changes the probability of the red/green event from .3/.7 to .1/.9 which changes the entropy from:
.882 to .469 or a reduction of .413 bits.
So with the addition of the yellow light, the entropy of the red/green light had a net change of:
.5 * .118 + .5 * -.413 = -.147 bits
And the new entropy of the red/green event when it happens after yellow/not-yellow, is:
.5 * .469 + .5 * 1 = .735 or .882 - .147 = .735
So the yellow/not-yellow event had a entropy of 1 bit on it's own, plus
-.147 bits of information about the up coming red/green event. So the yellow/not-yellow event had a total entropy effect of 1-.147 or .853 on the combined events.
So, you can say the yellow/not-yellow had .853 bits of entropy and the red-green has .882 giving a total of 1.735 which matches what we started with (other than a .001 rounding error).
The point of all this, is that information theory can't produce consistent results unless it acknowledges the effects of conditional probabilities. In the above, if you ignored the correlation of the two events and assumed they were uncorrelated with a 50/50 probability of the first and a .3/.7 for the second, the information channel had an entropy of 1 + .882 or
1.882. But if you look at the actually probability of all combinations, you see the information channel has a real entropy of 1.734 bits. To explain this difference, you have to acknowledge that the first event (the yellow/net-yellow) carries with it some information about the soon to happen second event, which reduces the entropy of the both events combined.
And if you acknowledge all this, it does seem to me, that it becomes meaningful to ask, how much information the yellow light alone, tells us about the green light. And since the yellow light tells us the probability of the green light has changed from .3 to .5, we can say the green light has told us .737 bits of information about the green light (even if we don't know the probability of this event (which was .5 in my example)). In effect, we are just getting some of the the information about what state the environment is in (in regards to how the state will effect the production of the green light event) early.