Harmonic entropy: Difference between revisions
add way more zeta HE stuff |
|||
| Line 38: | Line 38: | ||
So for example, if we want to express the probability that the incoming dyad "400 cents" is perceived as the JI basis interval "5/4," we would write that as the conditional probability | So for example, if we want to express the probability that the incoming dyad "400 cents" is perceived as the JI basis interval "5/4," we would write that as the conditional probability | ||
<math>\newcommand{\cent}{\text{¢}}</math> | <math>\displaystyle \newcommand{\cent}{\text{¢}}</math> | ||
<math>P(J=5/4|C=400\cent)</math> | <math>\displaystyle P(J=5/4|C=400\cent)</math> | ||
Or, in general, if we want to write the conditional probability that some incoming dyad of <math>c</math> cents is perceived as the JI basis interval <math>j</math>, we would write that as | Or, in general, if we want to write the conditional probability that some incoming dyad of <math>c</math> cents is perceived as the JI basis interval <math>j</math>, we would write that as | ||
<math>P(J=j|C=c)</math> | <math>\displaystyle P(J=j|C=c)</math> | ||
which notationally, we will often abbreviate as | |||
<math>\displaystyle P(j|c)</math> | |||
Note that at this point, we haven't yet specified what the particular probability distribution is. There are different ways to do this, which are described in more detail below. Generally, most approaches involve each JI interval's probability being assigned based on how close it is to <math>c</math> (closer dyads are given a larger probability), and how simple it is (simple dyads are given a higher probability, if distance is the same). | Note that at this point, we haven't yet specified what the particular probability distribution is. There are different ways to do this, which are described in more detail below. Generally, most approaches involve each JI interval's probability being assigned based on how close it is to <math>c</math> (closer dyads are given a larger probability), and how simple it is (simple dyads are given a higher probability, if distance is the same). | ||
| Line 51: | Line 55: | ||
Once we have decided on a probability distribution, we can finally evaluate the Shannon entropy. For a random variable <math>X</math>, the Shannon entropy is defined as: | Once we have decided on a probability distribution, we can finally evaluate the Shannon entropy. For a random variable <math>X</math>, the Shannon entropy is defined as: | ||
<math>H(X) = -\sum_{x \in X} P( | <math>\displaystyle H(X) = -\sum_{x \in X} P(x) \log_b P(x)</math> | ||
where the different <math>x</math> are taken from the sample space of <math>X</math>, and <math>b</math> is the base of the log. Different choices of <math>b</math> simply change the units in which entropy is given, the most common values being 2 and e, denoting "bits" and "nats". We will omit the base going forward, for simplicity. | where the different <math>x</math> are taken from the sample space of <math>X</math>, and <math>b</math> is the base of the log. Different choices of <math>b</math> simply change the units in which entropy is given, the most common values being 2 and e, denoting "bits" and "nats". We will omit the base going forward, for simplicity. | ||
| Line 57: | Line 61: | ||
In our case, we want to find the entropy of the random variable <math>J</math> of JI intervals, given a particular choice of incoming dyad in cents. The corresponding quantity that we want is: | In our case, we want to find the entropy of the random variable <math>J</math> of JI intervals, given a particular choice of incoming dyad in cents. The corresponding quantity that we want is: | ||
<math>H(J| | <math>\displaystyle H(J|c) = -\sum_{j \in J} P(j|c) \log P(j|c)</math> | ||
Note that above, the summation is only taken on the <math>j</math> from the sample space of <math> | Note that above, the summation is only taken on the <math>j</math> from the sample space of <math>J</math> (i.e. the set of JI basis intervals), whereas the parameter <math>c</math> is treated as constant within the summation (and is taken as the free parameter to the function). | ||
Since the parameter <math>c</math> is the free parameter, sometimes the above is notated as | Since the parameter <math>c</math> is the free parameter, sometimes the above is notated as | ||
<math>\text{HE}(c) = H(J| | <math>\displaystyle \text{HE}(c) = H(J|c)</math> | ||
which makes more explicit that <math>c</math> is the argument to the harmonic entropy function, which is equal to the entropy of <math>J</math>, conditioned on <math> | which makes more explicit that <math>c</math> is the argument to the harmonic entropy function, which is equal to the entropy of <math>J</math>, conditioned on the incoming dyad of <math>c</math> cents. | ||
==Probability Distributions== | ==Probability Distributions== | ||
| Line 78: | Line 82: | ||
Other spreading functions have also been explored, such as the use of the heavy-tailed [https://en.wikipedia.org/wiki/Laplace_distribution Laplace distribution], sometimes described as the "Vos function" in Paul's writings. These two functions are part of the [https://en.wikipedia.org/wiki/Generalized_normal_distribution Generalized normal distribution] family, which has a parameter not only for the variance but for the kurtosis. However, for simplicity, we will assume the Gaussian distribution as the spreading function for the remainder of this article, so that the spreading function for an incoming dyad <math>c</math> can be written as follows: | Other spreading functions have also been explored, such as the use of the heavy-tailed [https://en.wikipedia.org/wiki/Laplace_distribution Laplace distribution], sometimes described as the "Vos function" in Paul's writings. These two functions are part of the [https://en.wikipedia.org/wiki/Generalized_normal_distribution Generalized normal distribution] family, which has a parameter not only for the variance but for the kurtosis. However, for simplicity, we will assume the Gaussian distribution as the spreading function for the remainder of this article, so that the spreading function for an incoming dyad <math>c</math> can be written as follows: | ||
<math>S(x-c) = \frac{1}{s\sqrt{2\pi}} e^{-\frac{(x-c)^2}{2s^2}}</math> | <math>\displaystyle S(x-c) = \frac{1}{s\sqrt{2\pi}} e^{-\frac{(x-c)^2}{2s^2}}</math> | ||
where the notation <math>S(x-c)</math> is chosen to make clear that we are translating <math>S(x)</math> to be centered around the incoming dyad <math>c</math>, which is now the mean of the Gaussian. | where the notation <math>S(x-c)</math> is chosen to make clear that we are translating <math>S(x)</math> to be centered around the incoming dyad <math>c</math>, which is now the mean of the Gaussian. | ||
| Line 95: | Line 99: | ||
For discrete sets of JI basis ratios, the log-frequency spectrum can be divided up into '''domains''' assigned to each ratio. Each ratio is assigned a domain with lower bound equal to the mediant of itself and its nearest lower neighbor, and likewise with upper bound equal to the mediant of itself and its nearest upper neighbor. If no such neighbor exists, <math>\pm \infty</math> is used instead. Mathematically, this can be represented via the following expression: | For discrete sets of JI basis ratios, the log-frequency spectrum can be divided up into '''domains''' assigned to each ratio. Each ratio is assigned a domain with lower bound equal to the mediant of itself and its nearest lower neighbor, and likewise with upper bound equal to the mediant of itself and its nearest upper neighbor. If no such neighbor exists, <math>\pm \infty</math> is used instead. Mathematically, this can be represented via the following expression: | ||
<math>\displaystyle P(j|c) = \int_{\cent(j_l)}^{\cent(j_u)} S(x-c) dx</math> | |||
<math>P( | |||
where <math>S(x-c)</math> is the spreading function associated with c, <math>j_l</math> and <math>j_u</math> are the domain lower and upper bounds associated with JI basis ratio <math>j</math>, and <math>\cent(f) = 1200\log_2(f)</math>, or the "cents" function converting frequency ratios to cents. Typically, <math>j_l</math> is set equal to the mediant of <math>j</math> and its nearest lower neighbor (if it exists), or <math>-\infty</math> if not; likewise with <math>j_u</math> and its nearest upper neighbor. | where <math>S(x-c)</math> is the spreading function associated with c, <math>j_l</math> and <math>j_u</math> are the domain lower and upper bounds associated with JI basis ratio <math>j</math>, and <math>\cent(f) = 1200\log_2(f)</math>, or the "cents" function converting frequency ratios to cents. Typically, <math>j_l</math> is set equal to the mediant of <math>j</math> and its nearest lower neighbor (if it exists), or <math>-\infty</math> if not; likewise with <math>j_u</math> and its nearest upper neighbor. | ||
| Line 113: | Line 116: | ||
While it's still an open conjecture that this pattern holds for arbitrarily large <math>N</math>, the assumption is sometimes made that this is the case, and hence that for these basis ratio sets, <math>\frac{1}{\sqrt{nd}}</math> "approximations" to the width are sufficient to estimate domain-integral Harmonic Entropy. | While it's still an open conjecture that this pattern holds for arbitrarily large <math>N</math>, the assumption is sometimes made that this is the case, and hence that for these basis ratio sets, <math>\frac{1}{\sqrt{nd}}</math> "approximations" to the width are sufficient to estimate domain-integral Harmonic Entropy. | ||
This modifies the expression for the probabilities <math>P( | This modifies the expression for the probabilities <math>P(j|c)</math> as follows, noting that for now the "probabilities" won't sum to 1: | ||
<math>\ | <math>\displaystyle Q(j|c) = \frac{S(\cent(j)-c)}{\sqrt{j_n \cdot j_d}}</math> | ||
where the <math> | where the <math>Q</math> notation now represents that these "probabilities" are unnormalized, and <math>j_n</math> and <math>j_d</math> are the numerator and denominator, respectively, of JI basis ratio <math>j</math>. Again, the set of basis rationals here is assumed to be all of those rationals of Tenney Height ≤ <math>N</math> for some <math>N</math>. | ||
A similar observation for the use of Weil-bounded subsets of the rationals suggests domain widths of <math>\frac{1}{\max(n,d)}</math>, yielding instead the following formula: | A similar observation for the use of Weil-bounded subsets of the rationals suggests domain widths of <math>\frac{1}{\max(n,d)}</math>, yielding instead the following formula: | ||
<math>\ | <math>\displaystyle Q(j|c) = \frac{S(\cent(j)-c)}{\max(j_n, j_d)}</math> | ||
where this time the set of basis rationals is assumed to be all of those of Weil Height ≤ <math>N</math> for some <math>N</math>. | where this time the set of basis rationals is assumed to be all of those of Weil Height ≤ <math>N</math> for some <math>N</math>. | ||
| Line 127: | Line 130: | ||
In both cases, the general approach is the same: the value of the spreading function, taken at the value of <math>\cent(j)</math>, is divided by some sort of "complexity" function representing how much weight is given to that rational number. While the two complexity functions considered thus far were derived empirically by observing the asymptotic behavior of various height-bounded subsets of the rationals, we can generalize this for arbitrary basis sets of rationals and arbitrary complexities as follows: | In both cases, the general approach is the same: the value of the spreading function, taken at the value of <math>\cent(j)</math>, is divided by some sort of "complexity" function representing how much weight is given to that rational number. While the two complexity functions considered thus far were derived empirically by observing the asymptotic behavior of various height-bounded subsets of the rationals, we can generalize this for arbitrary basis sets of rationals and arbitrary complexities as follows: | ||
<math>\ | <math>\displaystyle Q(j|c) = \frac{S(\cent(j)-c)}{\|j\|}</math> | ||
where <math>\|j\|</math> denotes a complexity function mapping from rational numbers to non-negative reals. | where <math>\|j\|</math> denotes a complexity function mapping from rational numbers to non-negative reals. | ||
| Line 133: | Line 136: | ||
As these "probabilities" don't sum to 1, the result is not a probability distribution at all, invalidating the use of the Shannon Entropy. To rectify this, the distribution is normalized so that the probabilities do sum to 1: | As these "probabilities" don't sum to 1, the result is not a probability distribution at all, invalidating the use of the Shannon Entropy. To rectify this, the distribution is normalized so that the probabilities do sum to 1: | ||
<math>P( | <math>\displaystyle P(j|c) = \frac{Q(j|c)}{\sum_{j \in J} Q(j|c)}</math> | ||
which is equal to the unnormalized probability, divided by the sum of all unnormalized probabilities. This definition of <math>P( | which is equal to the unnormalized probability, divided by the sum of all unnormalized probabilities. This definition of <math>P(j|c)</math> is then used directly to compute the entropy. | ||
This approach to assigning probabilities to basis rationals is useful because it hypothetically makes it possible to consider the HE of sets of rationals which are dense in the reals, or even the entire set of positive rationals, although the best way to do this is a subject of ongoing research. | This approach to assigning probabilities to basis rationals is useful because it hypothetically makes it possible to consider the HE of sets of rationals which are dense in the reals, or even the entire set of positive rationals, although the best way to do this is a subject of ongoing research. | ||
| Line 166: | Line 169: | ||
The '''Harmonic Rényi Entropy of order a''' of an incoming dyad can be defined as follows: | The '''Harmonic Rényi Entropy of order a''' of an incoming dyad can be defined as follows: | ||
<math>\text{HE}_a(c) = H_a(J | <math>\displaystyle \text{HE}_a(c) = H_a(J|c) = \frac{1}{1-a} \log \sum_{j \in J} P(j|c)^a</math> | ||
Being a q-analog, it is noteworthy that Rényi entropy converges to Shannon entropy in the limit as <math>a \to 1</math>, a fact which can be verified using L'Hôpital's rule as found [http://www.sonycsl.co.jp/person/nielsen/Note-HopitalRuleShannonRenyiTsallis.pdf here]. | Being a q-analog, it is noteworthy that Rényi entropy converges to Shannon entropy in the limit as <math>a \to 1</math>, a fact which can be verified using L'Hôpital's rule as found [http://www.sonycsl.co.jp/person/nielsen/Note-HopitalRuleShannonRenyiTsallis.pdf here]. | ||
| Line 174: | Line 177: | ||
In a musical context, by considering the incoming dyad as analogous to a cryptographic code which is attempting to be "cracked" by an intelligent auditory system, we can consider that the analogous "worst-case attacker" would be a "best-case auditory system" which has complete awareness of the probability distribution for any incoming dyad. This analogy would view such an auditory system as actively attempting to choose the most probable rational, rather than drawing a rational at random weighted by the distribution. | In a musical context, by considering the incoming dyad as analogous to a cryptographic code which is attempting to be "cracked" by an intelligent auditory system, we can consider that the analogous "worst-case attacker" would be a "best-case auditory system" which has complete awareness of the probability distribution for any incoming dyad. This analogy would view such an auditory system as actively attempting to choose the most probable rational, rather than drawing a rational at random weighted by the distribution. | ||
The use of | The use of <math>a=∞</math> min-entropy would reflect this view. In contrast, the use of <math>a=1</math> Shannon entropy reflects a much "dumber" process which performs no such analysis and perhaps doesn't even seek to "choose" any sort of "victor" rational at all. As the parameter a interpolates between these two options, it can be interpreted as the extent to which the rational-matching process for incoming dyads is considered to be "intelligent" and "active" in this way. | ||
Some psychoacoustic effects naturally fit into this paradigm, such as the virtual pitch integration process, which actually does attempt to find a single victor when matching incoming chords with chunks of the harmonic series. Other psychoacoustic effects, such as that of beatlessness, may instead be better viewed as "dumb" processes whereby nothing in particular is being "chosen," but where a more uniform distribution of matching rational numbers for a dyad simply generates a more discordant sonic effect. Different values of a can differentiate between the predominance given to these two types of effect in the overall construct of psychoacoustic concordance. | Some psychoacoustic effects naturally fit into this paradigm, such as the virtual pitch integration process, which actually does attempt to find a single victor when matching incoming chords with chunks of the harmonic series. Other psychoacoustic effects, such as that of beatlessness, may instead be better viewed as "dumb" processes whereby nothing in particular is being "chosen," but where a more uniform distribution of matching rational numbers for a dyad simply generates a more discordant sonic effect. Different values of a can differentiate between the predominance given to these two types of effect in the overall construct of psychoacoustic concordance. | ||
| Line 182: | Line 185: | ||
==Examples== | ==Examples== | ||
===a=0: Harmonic Hartley Entropy=== | ===a=0: Harmonic Hartley Entropy=== | ||
<math>H_0(J| | <math>\displaystyle H_0(J|c) = \log |J|</math> | ||
where <math>|J|</math> is the cardinality of the set of basis rationals. This assumes, in essence, an "infinitely dumb" auditory system which can do no better than picking a rational number from a uniform distribution completely at random. All dyads have the same Harmonic Hartley Entropy. The Hartley Entropy is sometimes called the "max-entropy," and is useful mainly as an upper bound on the other forms of entropy: all Rényi Entropies are always guaranteed to be less than the Hartley Entropy. | where <math>|J|</math> is the cardinality of the set of basis rationals. This assumes, in essence, an "infinitely dumb" auditory system which can do no better than picking a rational number from a uniform distribution completely at random. All dyads have the same Harmonic Hartley Entropy. The Hartley Entropy is sometimes called the "max-entropy," and is useful mainly as an upper bound on the other forms of entropy: all Rényi Entropies are always guaranteed to be less than the Hartley Entropy. | ||
| Line 191: | Line 194: | ||
===a=1: Harmonic Shannon Entropy (Harmonic Entropy)=== | ===a=1: Harmonic Shannon Entropy (Harmonic Entropy)=== | ||
<math>H_1(J| | <math>\displaystyle H_1(J|c) = -\sum_{j \in J} P(j|c) \log P(j|c)</math> | ||
This is Paul's original Harmonic Entropy. Within the cryptographic analogy, this can be thought of as an auditory system which simply selects a rational at random from the incoming distribution, weighted via the distribution itself. | This is Paul's original Harmonic Entropy. Within the cryptographic analogy, this can be thought of as an auditory system which simply selects a rational at random from the incoming distribution, weighted via the distribution itself. | ||
| Line 200: | Line 203: | ||
===a=2: Harmonic Collision Entropy=== | ===a=2: Harmonic Collision Entropy=== | ||
<math>H_2(J | <math>\displaystyle H_2(J|c) = -\log \sum_{j \in J} P(j|c)^2 = -\log (J_1 = J_2|c)</math> | ||
where <math>J_1</math> and <math>J_2</math> are two independent and identically distributed random variables of JI basis ratios, conditioned on the same incoming dyad <math> | where <math>J_1</math> and <math>J_2</math> are two independent and identically distributed random variables of JI basis ratios, conditioned on the same incoming dyad <math>c</math>, and the collision entropy is the same as the negative log of the probability that the two JI variables produce the same outcome. | ||
[[File:HE_Tenney_N_10000_s_17cents_a=2.png]] | [[File:HE_Tenney_N_10000_s_17cents_a=2.png]] | ||
| Line 209: | Line 212: | ||
===a=∞: Harmonic Min-Entropy=== | ===a=∞: Harmonic Min-Entropy=== | ||
<math>H_\infty(J | <math>\displaystyle H_\infty(J|c) = -\log \max_{j \in J} P(j|c)</math> | ||
This is the min-entropy, which simply takes the negative log of the largest probability in the distribution. This can be thought of as representing the "strength" of the incoming dyad from being "deciphered" by a "best-case" auditory system. The name "min-entropy" reflects that the <math>a=\infty</math> case is guaranteed to be a lower bound among all Rényi entropies. | This is the min-entropy, which simply takes the negative log of the largest probability in the distribution. This can be thought of as representing the "strength" of the incoming dyad from being "deciphered" by a "best-case" auditory system. The name "min-entropy" reflects that the <math>a=\infty</math> case is guaranteed to be a lower bound among all Rényi entropies. | ||
| Line 225: | Line 228: | ||
The Harmonic Rényi Entropy is defined as | The Harmonic Rényi Entropy is defined as | ||
<math>\text{HE}_a(c) = H_a(J | <math>\displaystyle \text{HE}_a(c) = H_a(J|c) = \frac{1}{1-a} \log \sum_{j \in J} P(j|c)^a</math> | ||
As before, we can write <math>P( | As before, we can write <math>P(j|c)</math> as follows: | ||
<math>P( | <math>\displaystyle P(j|c) = \frac{Q(j|c)}{\sum_{j \in J} Q(j|c)}</math> | ||
where <math> | where <math>Q(j|c)</math> is the "unnormalized" probability, and the denominator above is the sum of these unnormalized probabilities, so that all of the <math>P(j|c)</math> sum to 1. | ||
To simplify notation, we first rewrite the denominator as a "normalization" function: | To simplify notation, we first rewrite the denominator as a "normalization" function: | ||
<math>\psi(c) = \sum_{j \in J} | <math>\displaystyle \psi(c) = \sum_{j \in J} Q(j|c)</math> | ||
and putting back into the original equation, we get | and putting back into the original equation, we get | ||
<math>H_a(J | <math>\displaystyle H_a(J|c) = \frac{1}{1-a} \log \left( \sum_{j \in J} \left( \frac{Q(j|c)}{\psi(c)} \right)^a \right)</math> | ||
Since <math>\psi(c)</math> is the same for each basis ratio, we can pull it out of the summation to obtain: | Since <math>\psi(c)</math> is the same for each basis ratio, we can pull it out of the summation to obtain: | ||
<math>H_a(J | <math>\displaystyle H_a(J|c) = \frac{1}{1-a} \log \left( \frac{\sum_{j \in J} Q(j|c)^a}{\psi(c)^a} \right)</math> | ||
To simplify notation further, we can also rewrite the numerator, the sum of "raw" (unnormalized) pseudo-probabilities, as a function: | To simplify notation further, we can also rewrite the numerator, the sum of "raw" (unnormalized) pseudo-probabilities, as a function: | ||
<math>\rho_a(c) = \sum_{j \in J} | <math>\displaystyle \rho_a(c) = \sum_{j \in J} Q(j|c)^a</math> | ||
Finally, we put this all together to obtain a simplified version of the Harmonic Rényi Entropy equation: | Finally, we put this all together to obtain a simplified version of the Harmonic Rényi Entropy equation: | ||
<math>\text{HE}_a(c) = H_a(J | <math>\displaystyle \text{HE}_a(c) = H_a(J|c) = \frac{1}{1-a} \log \left( \frac{\rho_a(c)}{\psi(c)^a} \right)</math> | ||
| Line 262: | Line 265: | ||
===Convolution product for <math>\psi(c)</math>=== | ===Convolution product for <math>\psi(c)</math>=== | ||
<math>\psi(c)</math>, the normalization function, is written as follows: | <math>\displaystyle \psi(c)</math>, the normalization function, is written as follows: | ||
<math>\psi(c) = \sum_{j \in J} | <math>\displaystyle \psi(c) = \sum_{j \in J} Q(j|c)</math> | ||
Again, <math> | Again, <math>Q(j|c)</math> is defined as follows: | ||
<math>\ | <math>\displaystyle Q(j|c) = \frac{S(\cent(j)-c)}{\|j\|}</math> | ||
We can rewrite the above equation as a convolution with a delta distribution: | We can rewrite the above equation as a convolution with a delta distribution: | ||
<math>\ | <math>\displaystyle Q(j|c) = \left(S \ast \frac{\delta_{-\cent(j)}}{\|j\|}\right)(-c)</math> | ||
Putting this back into the original summation, we obtain | Putting this back into the original summation, we obtain | ||
<math>\psi(c) = \sum_{j \in J} \left(S \ast \frac{\delta_{-\cent(j)}}{\|j\|}\right)(-c)</math> | <math>\displaystyle \psi(c) = \sum_{j \in J} \left(S \ast \frac{\delta_{-\cent(j)}}{\|j\|}\right)(-c)</math> | ||
We note that the left factor in the convolution product is always the same <math>S(-c)</math>, which is not dependent on <math>j</math> in any way. Since convolution distributes over addition, we can factor the <math>S</math> out of the summation to obtain | We note that the left factor in the convolution product is always the same <math>S(-c)</math>, which is not dependent on <math>j</math> in any way. Since convolution distributes over addition, we can factor the <math>S</math> out of the summation to obtain | ||
<math>\psi(c) = \left[S \ast \left(\sum_{j \in J} \frac{\delta_{-\cent(j)}}{\|j\|}\right)\right](-c)</math> | <math>\displaystyle \psi(c) = \left[S \ast \left(\sum_{j \in J} \frac{\delta_{-\cent(j)}}{\|j\|}\right)\right](-c)</math> | ||
We can clean up this notation by defining the auxiliary distribution K: | We can clean up this notation by defining the auxiliary distribution K: | ||
<math>K(c) = \sum_{j \in J} \frac{\delta_{-\cent(j)}}{\|j\|}</math> | <math>\displaystyle K(c) = \sum_{j \in J} \frac{\delta_{-\cent(j)}}{\|j\|}</math> | ||
Which leaves us with the final expression: | Which leaves us with the final expression: | ||
<math>\psi(c) = \left[S \ast K\right](-c)</math> | <math>\displaystyle \psi(c) = \left[S \ast K\right](-c)</math> | ||
===Convolution product for <math>\rho_a(c)</math>=== | ===Convolution product for <math>\rho_a(c)</math>=== | ||
The derivation for <math>\rho_a(c)</math> proceeds similarly. Recall the function is written as follows: | The derivation for <math>\rho_a(c)</math> proceeds similarly. Recall the function is written as follows: | ||
<math>\rho_a(c) = \sum_{j \in J} | <math>\displaystyle \rho_a(c) = \sum_{j \in J} Q(j|c)^a</math> | ||
The expression for each <math> | The expression for each <math>Q(j|c)^a</math> is: | ||
<math>\ | <math>\displaystyle Q(j|c)^a = \frac{S(\cent(j)-c)^a}{\|j\|^a}</math> | ||
We can again express this as a convolution, this time of the function <math>S^a(-c)</math>, meaning the spreading function S taken to the a'th power, and a delta distribution: | We can again express this as a convolution, this time of the function <math>S^a(-c)</math>, meaning the spreading function S taken to the a'th power, and a delta distribution: | ||
<math>\ | <math>\displaystyle Q(j|c)^a = \left(S^a \ast \frac{\delta_{-\cent(j)}}{\|j\|^a}\right)(-c)</math> | ||
Putting this back into the original summation and factoring as before, we obtain | Putting this back into the original summation and factoring as before, we obtain | ||
<math>\rho_a(c) = \left[S^a \ast \left(\sum_{j \in J} \frac{\delta_{-\cent(j)}}{\|j\|^a}\right)\right](-c)</math> | <math>\displaystyle \rho_a(c) = \left[S^a \ast \left(\sum_{j \in J} \frac{\delta_{-\cent(j)}}{\|j\|^a}\right)\right](-c)</math> | ||
And again we clean up notation by defining the auxiliary distribution | And again we clean up notation by defining the auxiliary distribution | ||
<math>K^a(c) = \sum_{j \in J} \frac{\delta_{-\cent(j)}}{\|j\|^a}</math> | <math>\displaystyle K^a(c) = \sum_{j \in J} \frac{\delta_{-\cent(j)}}{\|j\|^a}</math> | ||
so that | so that | ||
<math>\rho_a(c) = \left[S^a \ast K^a\right](-c)</math> | <math>\displaystyle \rho_a(c) = \left[S^a \ast K^a\right](-c)</math> | ||
We have now succeeded in representing <math>\rho_a(c)</math> as a convolution. | We have now succeeded in representing <math>\rho_a(c)</math> as a convolution. | ||
| Line 333: | Line 336: | ||
Taking all of this, we can rewrite the original expression for Harmonic Rényi Entropy as follows: | Taking all of this, we can rewrite the original expression for Harmonic Rényi Entropy as follows: | ||
<math>\text{HE}_a(c) = H_a(J | <math>\displaystyle \text{HE}_a(c) = H_a(J|c) = \frac{1}{1-a} \log \left( \frac{\left[S^a \ast K^a\right](-c)}{\left[S \ast K\right]^a(-c)} \right)</math> | ||
where the expression | where the expression | ||
<math>\left[S \ast K\right]^a(-c)</math> | <math>\displaystyle \left[S \ast K\right]^a(-c)</math> | ||
represents the convolution of <math>S</math> and </math>K, taken to the <math>a</math>'th power, and flipped backwards. Note that if <math>S(x)</math> is a symmetrical (even) spreading function, and if for each ratio <math>n/d</math> in <math>J</math>, if the inverse <math>d/n</math> is also in <math>J</math>, then the above convolution will also be symmetrical, and we also have | represents the convolution of <math>S</math> and </math>K, taken to the <math>a</math>'th power, and flipped backwards. Note that if <math>S(x)</math> is a symmetrical (even) spreading function, and if for each ratio <math>n/d</math> in <math>J</math>, if the inverse <math>d/n</math> is also in <math>J</math>, then the above convolution will also be symmetrical, and we also have | ||
<math>\left[S \ast K\right]^a(-c) = \left[S \ast K\right]^a(c)</math> | <math>\displaystyle \left[S \ast K\right]^a(-c) = \left[S \ast K\right]^a(c)</math> | ||
We have succeeded in representing Harmonic Rényi Entropy in simple terms of two convolution products, each of which can be computed in <math>O(N log N)</math> time. | We have succeeded in representing Harmonic Rényi Entropy in simple terms of two convolution products, each of which can be computed in <math>O(N log N)</math> time. | ||
=Extending HE to N= | =Extending HE to <math>N=\infty</math>= | ||
All of the models described above involve a finite set of rational numbers, bounded by some complexity function, and where the complexity is less than some max value <math>N</math>. | All of the models described above involve a finite set of rational numbers, bounded by some complexity function, and where the complexity is less than some max value <math>N</math>. | ||
It so happens that | It so happens that we are more or less able to analytically continue this definition to the situation where <math>N=\infty</math>. More precisely, we are able to analytically continue the exponential of HE, which yields the same relative interval rankings as standard HE. | ||
The only technical caveat is that we use the HE of the "unnormalized" probability distribution. However, in the large limit of <math>N</math>, this appears to agree closely with the usual HE. We go into more detail below about this. | |||
Our basic approach is: rather than weighting intervals by <math>(nd)^{0.5}</math>, we choose a different exponent, such as <math>(nd)^2</math>. For an exponent which is large enough (we will show that it must be greater than 1), HE does indeed converge as <math>N \to \infty</math>, and we show that this yields an expression related to the [[The_Riemann_Zeta_Function_and_Tuning|Riemann Zeta function]]. We can then use the analytic continuation of the zeta function to obtain an analytically continued curve for the <math>(nd)^{0.5}</math> weighting, which we then show empirically does indeed appear to be what HE converges on for large values of <math>N</math>. | |||
This enables us to speak cognizantly of the harmonic entropy of an interval as measured against ''all'' rational numbers. | This enables us to speak cognizantly of the harmonic entropy of an interval as measured against ''all'' rational numbers. | ||
==Background: Unnormalized Entropy== | ==Background: Unnormalized Entropy== | ||
Our derivation only analytically continues the entropy function for the "unnormalized" set of probabilities, which we previously wrote as <math>Q(j|c)</math>. For this definition to be philosophically perfect, we would want to analytically continue the entropy function for the normalized sense of probabilities, previously written as <math>P(j|c)</math>. | |||
However, in practice, the "unnormalized entropy" appears to be an extremely good approximation to the normalized entropy for large values of <math>N</math>. The resulting curve has the same minima and maxima as HE, the same general shape, and for all intents and purposes looks exactly like HE, just shifted on the y-axis. | However, in practice, the "unnormalized entropy" appears to be an extremely good approximation to the normalized entropy for large values of <math>N</math>. The resulting curve has the same minima and maxima as HE, the same general shape, and for all intents and purposes looks exactly like HE, just shifted on the y-axis. | ||
| Line 376: | Line 383: | ||
For now, our derivation is limited to the case of <math>\sqrt{nd}</math> Tenney-weighted rationals, although it may be possible to derive a similar result for <math>\max(n,d)</math> weighting as well. | For now, our derivation is limited to the case of <math>\sqrt{nd}</math> Tenney-weighted rationals, although it may be possible to derive a similar result for <math>\max(n,d)</math> weighting as well. | ||
Additionally, because it simplifies the derivation, we will use ''unreduced rationals'' in our basis set, meaning that we will even allow unreduced fractions such as <math>4/2</math> so long as <math>nd<N</math> for our bound N. The use of HE with unreduced rationals | Additionally, because it simplifies the derivation, we will use ''unreduced rationals'' in our basis set, meaning that we will even allow unreduced fractions such as <math>4/2</math> so long as <math>nd<N</math> for our bound N. The use of HE with unreduced rationals has previously been studied by Paul and shown to be not that much different than HE with reduced ones. However, we will easily show later unreduced and reduced rationals converge to the same thing in the limit as <math>N \to \infty</math>, up to a constant multiplicative scaling. | ||
===Definition of the Unnormalized Harmonic Rényi Entropy=== | |||
Let's start by recalling the original definition for Harmonic Rényi Entropy, using complexity-normalization probabilities: | |||
<math>\displaystyle \text{HE}_a(c) = \frac{1}{1-a} \log \sum_{j \in J} P(j|c)^a</math> | |||
Remember also that the definition of <math>P(j|c)</math> is as follows: | |||
<math>\displaystyle P(j|c) = \frac{Q(j|c)}{\sum_{j \in J} Q(j|c)}</math> | |||
where the <math>Q(j|c)</math> is the "unnormalized probability" - the raw value of the spreading function, evaluated at the ratio in question, divided by the ratio's complexity. The above equation tells us that the normalized probability is equal to the unnormalized probability, divided by the sum of all unnormalized probabilities. | |||
<math>\text{HE}_a(c) = | Putting the two together, we get | ||
<math>\displaystyle \text{HE}_a(c) = \frac{1}{1-a} \log \sum_{j \in J} \left( \frac{Q(j|c)}{\sum_{j \in J} Q(j|c)} \right)^a</math> | |||
Now, for us to define the unnormalized HE, we simply take the standard Rényi entropy equation, and replace the normalized probabilities with unnormalized ones, yielding | |||
<math>\displaystyle \text{UHE}_a(c) = \frac{1}{1-a} \log \sum_{j \in J} Q(j|c)^a</math> | |||
Using our convolution theorem from before, we can express the above as | |||
<math>\displaystyle \text{UHE}_a(c) = \frac{1}{1-a} \log \left( S^a \ast K^a \right)(-c)</math> | |||
where, as before, <math>S^a</math> is our spreading function, taken to the <math>a</math>'th power, and <math>K^a</math> is our convolution kernel, with the weights on the delta functions taken to the <math>a</math>'th power as described previously. | |||
Note that if <math>S</math> is symmetric, as in the case of the Gaussian or Laplace distributions, then the inverted argument of <math>(-c)</math> on the end is redundant, and can be replaced by <math>(c)</math>. | |||
Lastly, it so happens that it will be much easier to understand our analytic continuation if we look at the exponential of the UHE, times <math>(1-a)</math>, rather than the UHE itself. The reasons for this will become clear later. If we do so, we get | |||
<math>\displaystyle \exp((1-a) \text{UHE}_a(c)) = \left( S^a \ast K^a \right)(-c)</math> | |||
Note that this function is simply a monotonic transformation of the original, and so preserves the exact same concordance ranking on all intervals. | |||
===Analytic Continuation of the Convolution Kernel=== | |||
The definition for <math>K</math> is: | The definition for <math>K</math> is: | ||
<math>\displaystyle K(c) = \sum_{j \in J} \ | <math>\displaystyle K(c) = \sum_{j \in J} \frac{\delta_{-\cent(j)}}{\|j\|}</math> | ||
where <math>\|j\|</math> represents the "complexity" of the JI basis ratio <math>j</math>. In the particular case of Tenney weighting, we get: | where <math>\|j\|</math> represents the "complexity" of the JI basis ratio <math>j</math>. In the particular case of Tenney weighting, we get: | ||
<math>\displaystyle K(c) = \sum_{j \in J} \ | <math>\displaystyle K(c) = \sum_{j \in J} \frac{\delta_{-\cent(j)}}{(j_n \cdot j_d)^{0.5}}</math> | ||
where <math>j_n</math> and <math>j_d</math> are the numerator and denominator of <math>j</math>, respectively. | where <math>j_n</math> and <math>j_d</math> are the numerator and denominator of <math>j</math>, respectively. | ||
| Line 399: | Line 437: | ||
Although it may seem odd, we can take the Fourier transform of the above to obtain the following expression: | Although it may seem odd, we can take the Fourier transform of the above to obtain the following expression: | ||
<math>\displaystyle \mathcal{F}\left\{K(c)\right\}(\omega) = \sum_{j \in J} \ | <math>\displaystyle \mathcal{F}\left\{K(c)\right\}(\omega) = \sum_{j \in J} \frac{e^{i \omega \cent(j)}}{(j_n \cdot j_d)^{0.5}}</math> | ||
Furthermore, for simplicity, we can change the units, so that rather than the argument being given in cents, it is given in "natural" units of "[https://en.wikipedia.org/wiki/Neper nepers]", a technique often used by Martin Gough in his work on [[Logarithmic_approximants|Logarithmic approximants]]. The representation of any interval in nepers is given by simply taking is natural logarithm. Doing so, by defining the change of variables <math>c = \frac{1200}{\log(2)}n</math>, we obtain | Furthermore, for simplicity, we can change the units, so that rather than the argument being given in cents, it is given in "natural" units of "[https://en.wikipedia.org/wiki/Neper nepers]", a technique often used by Martin Gough in his work on [[Logarithmic_approximants|Logarithmic approximants]]. The representation of any interval in nepers is given by simply taking is natural logarithm. Doing so, by defining the change of variables <math>c = \frac{1200}{\log(2)}n</math>, we obtain | ||
<math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \sum_{j \in J} \ | <math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \sum_{j \in J} \frac{e^{i \omega \log (j_n/j_d)}}{(j_n \cdot j_d)^{0.5}}</math> | ||
We can treat the presence of the logarithm within the exponential function as changing the base of the exponential, so that we get | We can treat the presence of the logarithm within the exponential function as changing the base of the exponential, so that we get | ||
<math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \sum_{j \in J} \ | <math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \sum_{j \in J} \frac{(j_n/j_d)^{i \omega}}{(j_n \cdot j_d)^{0.5}}</math> | ||
We can also factor each term in the summation to obtain | We can also factor each term in the summation to obtain | ||
<math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \sum_{j \in J} \left[ \ | <math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \sum_{j \in J} \left[ \frac{{j_n}^{i \omega}}{j_n^{0.5}} \cdot \frac{{j_d}^{-i \omega}}{j_d^{0.5}} \right]</math> | ||
which we can rewrite as | which we can rewrite as | ||
<math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \sum_{j \in J} \left[ \ | <math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \sum_{j \in J} \left[ \frac{1}{{j_n}^{0.5 -i \omega}} \cdot \frac{1}{{j_d}^{0.5 + i \omega}} \right]</math> | ||
| Line 424: | Line 462: | ||
Bounding by <math>\max(n,d) < N</math> is the same as specifying that <math>j_n < N</math> and <math>j_d < N</math>. Doing so, we get | Bounding by <math>\max(n,d) < N</math> is the same as specifying that <math>j_n < N</math> and <math>j_d < N</math>. Doing so, we get | ||
<math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \sum_{1<j_n, j_d<N} \left[ \ | <math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \sum_{1<j_n, j_d<N} \left[ \frac{1}{{j_n}^{0.5 -i \omega}} \cdot \frac{1}{{j_d}^{0.5 + i \omega}} \right]</math> | ||
We can now factor the above product to obtain: | We can now factor the above product to obtain: | ||
<math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \left[ \sum_{j_n=1}^N \ | <math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \left[ \sum_{j_n=1}^N \frac{1}{{j_n}^{0.5 -i \omega}} \right] \cdot \left[ \sum_{j_d=1}^N\frac{1}{{j_d}^{0.5 + i \omega}} \right]</math> | ||
Now, we can see that as <math>N \to \infty</math> above, the summations do not converge. However, incredibly enough, each of the above expressions has a very well-known analytic continuation, which is the Riemann zeta function. | Now, we can see that as <math>N \to \infty</math> above, the summations do not converge. However, incredibly enough, each of the above expressions has a very well-known analytic continuation, which is the Riemann zeta function. | ||
| Line 436: | Line 474: | ||
To perform the analytic continuation, we temporarily change the <math>0.5</math> in the denominator to some value <math>\sigma> 1</math>. This is equivalent to changing our original <math>\sqrt{nd}</math> weighting to some other exponent, such as <math>(nd)^2</math> or <math>(nd)^{1.5}</math>. Doing this causes both of the summations above to converge, so that we obtain | To perform the analytic continuation, we temporarily change the <math>0.5</math> in the denominator to some value <math>\sigma> 1</math>. This is equivalent to changing our original <math>\sqrt{nd}</math> weighting to some other exponent, such as <math>(nd)^2</math> or <math>(nd)^{1.5}</math>. Doing this causes both of the summations above to converge, so that we obtain | ||
<math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \left[ \sum_{j_n=1}^\infty \ | <math>\displaystyle \mathcal{F}\left\{K(n)\right\}(\omega) = \left[ \sum_{j_n=1}^\infty \frac{1}{{j_n}^{\sigma -i \omega}} \right] \cdot \left[ \sum_{j_d=1}^\infty\frac{1}{{j_d}^{\sigma + i \omega}} \right]</math> | ||
It has been well known for more than a century that both of these summations converge to the Riemann zeta function, so that we get | It has been well known for more than a century that both of these summations converge to the Riemann zeta function, so that we get | ||
| Line 452: | Line 490: | ||
Furthermore, although the above series doesn't converge for <math>\sigma = 0.5</math>, we can simply use the analytic continuation of the Riemann zeta function to obtain a meaningful function at that point, so that our original convolution kernel can be written as | Furthermore, although the above series doesn't converge for <math>\sigma = 0.5</math>, we can simply use the analytic continuation of the Riemann zeta function to obtain a meaningful function at that point, so that our original convolution kernel can be written as | ||
<math>K(n) = \mathcal{F}^{-1}\left\{| \zeta(0.5+\omega) |^2\right\}(n)</math> | <math>\displaystyle K(n) = \mathcal{F}^{-1}\left\{| \zeta(0.5+\omega) |^2\right\}(n)</math> | ||
which is the inverse Fourier transform of the squared absolute value of the Riemann zeta function, taken at the critical line. | which is the inverse Fourier transform of the squared absolute value of the Riemann zeta function, taken at the critical line. | ||
Lastly, to do some cleanup, we previously went with <math>\max(n,d)<N</math> bounds, rather than <math>\sqrt{nd}<N</math> bounds, despite using <math>\sqrt{nd}</math> weighting on ratios. However, it is easy to show above that regardless of which bounds you use, both | Lastly, to do some cleanup, we previously went with <math>\max(n,d)<N</math> bounds, rather than <math>\sqrt{nd}<N</math> bounds, despite using <math>\sqrt{nd}</math> weighting on ratios. However, it is easy to show above that regardless of which bounds you use, both choices converge to the same function when <math>\sigma > 1</math> in the limit as <math>N > \infty</math>. Since these series agree on this right half-plane of the Riemann zeta function, they share the same analytic continuation, so that we get the same result, despite using our technique to simplify the derivation. | ||
It is likewise easy to show that the function <math>K^a(n)</math>, taken from the numerator of our original Harmonic Rényi Entropy convolution expression, can be expressed as | It is likewise easy to show that the function <math>K^a(n)</math>, taken from the numerator of our original Harmonic Rényi Entropy convolution expression, can be expressed as | ||
<math>K^a(n) = \mathcal{F}^{-1}\left\{|\zeta(0.5a+\omega) |^2\right\}(n)</math> | <math>\displaystyle K^a(n) = \mathcal{F}^{-1}\left\{|\zeta(0.5a+\omega) |^2\right\}(n)</math> | ||
so that the choice of <math>a</math> simply changes our choice of vertical slice of the Riemann zeta function. | so that the choice of <math>a</math> simply changes our choice of vertical slice of the Riemann zeta function. | ||
We can put this back into | ===Analytic Continuation of Unnormalized Harmonic Rényi Entropy=== | ||
We can put this back into our equation for the Unnormalized Harmonic Rényi Entropy. To do so, we will continue with our change of units from cents to nepers, corresponding to a change of our variable from <math>c</math> to <math>n</math>. We will likewise assume the spreading probability distribution <math>S</math> has been scaled to reflect the new choice of units. | |||
Our original equation was | |||
<math>\displaystyle \exp((1-a) \text{UHE}_a(n)) = \left( S^a \ast K^a \right)(-n)</math> | |||
Using our expression for <math>K^a</math> as <math>N \to \infty</math>, we get | |||
<math>\displaystyle \exp((1-a) \text{UHE}_a(n)) = \left( S^a \ast \mathcal{F}^{-1}\left\{|\zeta(0.5a+\omega)|^2\right\} \right)(-n)</math> | |||
To simplify this, we will define an auxiliary notation for the zeta function as follows: | |||
<math>\zeta_\sigma(\omega) = \zeta(\sigma + i\omega)</math> | |||
yielding the simplified expression: | |||
<math>\displaystyle \exp((1-a) \text{UHE}_a(n)) = \left( S^a \ast \mathcal{F}^{-1}\left\{|\zeta_{0.5a}|^2\right\} \right)(-n)</math> | |||
We can simplify the expression of the above if we likewise take the Fourier transform of <math>S</math>. If we do, we obtain the [https://en.wikipedia.org/wiki/Characteristic_function_(probability_theory) characteristic function] of the distribution, which is typically denoted by <math>\phi(\omega)</math>. We will use the following definitions: | We can simplify the expression of the above if we likewise take the Fourier transform of <math>S</math>. If we do, we obtain the [https://en.wikipedia.org/wiki/Characteristic_function_(probability_theory) characteristic function] of the distribution, which is typically denoted by <math>\phi(\omega)</math>. We will use the following definitions: | ||
<math>\phi(\omega) = \mathcal{F}\left\{S(n)\right\}(\omega)</math> | <math>\displaystyle \phi(\omega) = \mathcal{F}\left\{S(n)\right\}(\omega)</math> | ||
<math>\displaystyle \phi_a(\omega) = \mathcal{F}\left\{S(n)^a\right\}(\omega)</math> | |||
Doing so, and noting that convolution becomes multiplication in the Fourier domain, we get | Doing so, and noting that convolution becomes multiplication in the Fourier domain, we get | ||
<math>\displaystyle \text{ | <math>\displaystyle \exp((1-a) \text{UHE}_a(n)) = \mathcal{F}^{-1}\left\{\phi_a \cdot |\zeta_{0.5a}|^2\right\}(-n)</math> | ||
Lastly, we note that for any function <math>f(x)</math>, we have <math>\mathcal{F}\left\{f(-x)\right\} = \mathcal{F}\left\{\overline {f(x)} \right\}</math>, where <math>\overline x</math> is complex conjugation. For simplicity's sake, we can this write as <math>\mathcal{F}\left\{\overline f\right\}</math>. Putting that all together, we get | |||
<math>\displaystyle \exp((1-a) \text{UHE}_a(n)) = \mathcal{F}^{-1}\left\{\overline \phi_a \cdot |\zeta_{0.5a}|^2\right\}</math> | |||
where we can drop the overline on <math>|\zeta_{0.5a}|^2</math> because it is purely real, and its complex conjugate is itself. | |||
==Examples== | |||
It is very easy to see empirically that our expression does seem to converge be the thing that UHE converges on in the limit of large N. Here are some examples for different values of s and a, showing that as N increases it converges on the our analytically continued "zeta HE." | |||
In all these examples, the left plot is a series of plots of UHE at different N, with the zeta HE being a slightly thicker line in the background. The right plot is the largest N plotted against zeta HE, being almost a perfect line. | |||
<math>\ | Note also that we don't plot the UHE directly, but rather <math>\exp((1-a)UHE)</math>, as described previously. The units have also been converted back to cents. Each HE function has been scaled so that the minimum entropy is 0 and the maximum entropy is 1. | ||
Lastly, note that, due to our multiplication of UHE by <math>(1-a)</math> above, the UHE would be flipped upside down relative to what we're used to, with higher values corresponding to more concordant intervals. In the pictures above we flipped it back upside down, for consistency with the earlier pictures. | |||
===s=0.5%, a=1.00001=== | |||
[[File:ExpUHE vs zeta s=0.5%.png|800px]] | |||
===s=1%, a=1.00001=== | |||
[[File:ExpUHE vs zeta s=1%.png|800px]] | |||
===s=1.5%, a=1.00001=== | |||
[[File:ExpUHE vs zeta s=1.5%.png|800px]] | |||
Note that in all these plots, the value of <math>a</math> is chosen to be <math>1.00001</math> rather than exactly <math>1</math>, so as to avoid that <math>(1-a)</math> term becoming 0. Similar results are seen for other choices of <math>a</math>: | |||
''s=1%, a=2.2''<br> | |||
[[File:ExpUHE vs zeta s=1% a=2.2.png|800px]] | |||
Note that you have to be careful if you choose to compute the analytically continued <math>a=2</math> numerically: this corresponds to the vertical slice of zeta function at <math>\Re(s) = 1</math>, where there is a pole. | |||
' | ==Why not Normalized HE?== | ||
On the surface, everything we did with the convolution theorem, and subsequent analytic continuation, should appear to work for normalized HE as well. For example, let's review our result for the exp of unnormalized HE: | |||
<math>\displaystyle \exp((1-a) \text{UHE}_a(n)) = \mathcal{F}^{-1}\left\{\overline \phi_a \cdot |\zeta_{0.5a}|^2\right\}</math> | |||
Our original expression for normalized HE was | |||
<math>\displaystyle \exp((1-a) \text{HE}_a(n)) = \frac{\left[S^a \ast K^a\right](-c)}{\left[S \ast K\right]^a(-c)}</math> | |||
Naively, you might expect we could simply apply the same analytic continuation technique to the numerator and denominator. If we do so, we would get | |||
<math>\displaystyle \exp((1-a) \text{HE}_a(n)) = \frac{\mathcal{F}^{-1}\left\{\overline \phi_a \cdot |\zeta_{0.5a}|^2\right\}}{\mathcal{F}^{-1}\left\{\overline \phi \cdot |\zeta_{0.5}|^2\right\}^a}</math> | |||
However, to see why this doesn't work, let's compare the analytically continued version of the denominator (i.e. the normalization term) with the finite versions. Look at the following picture, which is a plot of <math>\mathcal{F}^{-1}\left\{\overline \phi \cdot |\zeta_{0.5}|^2\right\}^a</math>: | |||
[[File:HE_normalization_terms.png|800px]] | |||
This picture shows how the denominator changes as <math>N</math> increases: you can see that in general, the function is shifted upward, increasing without bound. The thin plots reflect this for N=1000, 5000, 10000, 50000, and 100000, where you can see them increasing. | |||
You will note that the denominator also looks exactly like unnormalized HE, just upside down. Normalized HE is the quotient of two functions that both look like this, which are slightly different. This quotient produces the usual HE curve, which is flipped upside down relative to the denominator, and which also increases without bound. That all these functions increase without bound is just another way to state that these things generally don't converge as <math>N \to \infty</math>. | |||
However, look at what happens with our analytic continuation, which is given by the thicker blue line at the bottom. Despite our sequence of finite-<math>N</math> denominator terms increasing on the y-axis, the analytically continued version suddenly "snaps" back to zero. Although the curve shape is roughly the same, the vertical offset is almost completely eliminated when the analytic continuation is done. | |||
The problem here is that the original HE function was the quotient of two very large, strictly positive functions - the numerator and denominator. However, performing the analytic continuation on each separately has caused both to "snap" back to zero, so that the denominator, while retaining the same shape, now has points where it touches the x-axis. As a result, the quotient of the two will have poles where the denominator is zero. | |||
The resulting quotient of analytically continued functions looks like this, and does not remotely resemble HE: | |||
[[File:NaiveHEanalyticcontinuation.png|800px]] | |||
Those "spikes" are poles where the denominator is zero. | |||
The problem is that we're really stretching the boundaries of complex analysis with this. With unnormalized HE, we were able to analytically continue the Fourier transform of exp-UHE to obtain a concrete expression in terms of the Riemann zeta function. While complex analysis makes no guarantees on the behavior of the Fourier transform of a holomorphic function, we did see the result seemed to converge on exp-UHE in the limit of large <math>N</math> when transforming back from the Fourier domain, confirming empirically that our analytically continued expression seemed to make sense. | |||
But in the case of "normalized HE," we analytically continued the Fourier transforms of the numerator and denominator, separately, transformed both out of the Fourier domain, and then took the quotient. Complex analysis makes no guarantee on the behavior of the quotient of two Fourier transforms of holomorphic functions, and in this case the behavior is very strange. A different approach to analytically continuing the expression would be required. | |||
This | This same principle explains why we plotted the exp of UHE, rather than UHE itself. Were we to take the log of finite UHE, we would be taking the log of a strictly positive function. However, the analytically continued exp-UHE snaps back to the x-axis, so that there are points where the function is zero or even negative. Taking the log of the analytically continued exp-UHE would yield a complex-valued function where it is negative, due to this snapping effect. However, looking at exp-UHE directly has no such problem. | ||
Finally, it is noteworthy that for <math>a>2</math>, we end up looking at slices of the zeta function for which <math>\Re(s)>1</math>. This is where our original unnormalized HE function should converge as <math>N \to \infty</math>, corresponding to the region where the Riemann zeta function Dirichlet series converges. For these values of <math>a</math>, the exp-UHE ''is'' positive. So, we can take the log again and look at the usual UHE. This can be useful for plotting, since exp-UHE tends to "flatten" out the curve for high values of <math>a</math>, whereas taking the log accentuates the minima and maxima (and more closely resembles the usual HRE). | |||
''' | |||
== | =To Do= | ||
There are a number of things that need to be added to this article. Below are listed some for reference: | |||
* 3HE, both for finite HE and for <math>N -> \infty</math> | |||
* infinite zeta-UHE with reduced rationals only | |||
* write-up of new free parameter, the <math>(nd)^\sigma</math> weighting exponent | |||
* write-up of how this new free parameter corresponds to a for generalized normal distributions | |||
* write-up of fast computation for infinite zeta-UHE, perhaps with a zeta table | |||
* addition of many more pictures | |||
=References= | =References= | ||