N2D3P9: Difference between revisions

Cmloegcmluin (talk | contribs)
Created page with "'''<math>\text{N2D3P9}</math>''' or Entoo-Deethree-Peenine, is a fictional character in the Star Wars franchise. In an alternative timeline, the young Anakin Skywalker assembl..."
 
Cmloegcmluin (talk | contribs)
No edit summary
Line 54: Line 54:
<math>\text{N2D3P9}</math> was developed or discovered rather late in the development of Sagittal notation. So what did we use previously, to decide which ratios should get the simple symbols? We used actual data on ratio usage from [http://www.huygens-fokker.org/microtonality/scales.html the Huygens-Fokker Foundation's scale archive], kindly provided by [[Manuel Op de Coul]].
<math>\text{N2D3P9}</math> was developed or discovered rather late in the development of Sagittal notation. So what did we use previously, to decide which ratios should get the simple symbols? We used actual data on ratio usage from [http://www.huygens-fokker.org/microtonality/scales.html the Huygens-Fokker Foundation's scale archive], kindly provided by [[Manuel Op de Coul]].


All scales in the archive were treated equally, as we didn't have any information about their relative importance. Each occurrence of a pitch ratio in a scale was counted as one vote for that ratio. Then the ratios were grouped into 5-rough pitch classes and a single figure obtained for each 5-rough ratio (representing the class). There were 29,403 votes, allocated to 820 5-rough ratios.
All scales in the archive were treated equally, as we didn't have any information about their relative importance. Each occurrence of a pitch ratio in a scale was counted as one vote for that ratio. Then the ratios were grouped into 5-rough pitch classes and a single figure obtained for each 5-rough superunison ratio (representing the class). There were 29,403 votes, allocated to 820 5-rough ratios.


Like the frequency of use of letters in an alphabet, when sorted in order of decreasing popularity, the ratios obeyed an approximate [https://en.wikipedia.org/wiki/Zipf%27s_law Zipf's law] distribution, with the Nth most popular ratio having votes proportional to approximately <math>\frac{1}{N^{1.37}}</math>. This meant that about half the ratios had only one vote each, and three quarters of them had 3 votes or less. Such low numbers of votes meant that the data on the less popular ratios was vulnerable to "historical noise". In other words, the position of such a ratio in the list might not be a good predictor of its relative frequency of use in the future.
Like the frequency of use of letters in an alphabet, when sorted in order of decreasing popularity, the ratios obeyed an approximate [https://en.wikipedia.org/wiki/Zipf%27s_law Zipf's law] distribution, with the Nth most popular ratio having votes proportional to approximately <math>\frac{1}{N^{1.37}}</math>. This meant that about half the ratios had only one vote each, and three quarters of them had 3 votes or less. Such low numbers of votes meant that the data on the less popular ratios was vulnerable to "historical noise". In other words, the position of such a ratio in the list might not be a good predictor of its relative frequency of use in the future.
Line 77: Line 77:
Rather than attempt to fit functions to the exact counts of votes for each ratio, the functions were fit to the rank indices of each ratio; in other words, a function only needed to sort ratios the same as the actual data, and within each rank position it was unimportant how close its estimate of votes was. In technical parlance, the goal was to minimize the [https://en.wikipedia.org/wiki/Spearman%27s_rank_correlation_coefficient Spearman’s rank coefficient] between the estimated ranks and the actual ranks. For purposes of comparing competing functions, minimizing Spearman’s rank coefficient could be simplified to minimizing the sum of squared differences between the ranks. But because fitting to the simpler ratios which had more votes is more important, a Zipf's-law weighting was applied to the ranks by taking their reciprocals before calculating their squared differences. A [https://en.wikipedia.org/wiki/Ranking#Fractional_ranking_(%221_2.5_2.5_4%22_ranking) fractional ranking] strategy was used to ensure that stretches of the data with tied vote counts did not distort the measurement.
Rather than attempt to fit functions to the exact counts of votes for each ratio, the functions were fit to the rank indices of each ratio; in other words, a function only needed to sort ratios the same as the actual data, and within each rank position it was unimportant how close its estimate of votes was. In technical parlance, the goal was to minimize the [https://en.wikipedia.org/wiki/Spearman%27s_rank_correlation_coefficient Spearman’s rank coefficient] between the estimated ranks and the actual ranks. For purposes of comparing competing functions, minimizing Spearman’s rank coefficient could be simplified to minimizing the sum of squared differences between the ranks. But because fitting to the simpler ratios which had more votes is more important, a Zipf's-law weighting was applied to the ranks by taking their reciprocals before calculating their squared differences. A [https://en.wikipedia.org/wiki/Ranking#Fractional_ranking_(%221_2.5_2.5_4%22_ranking) fractional ranking] strategy was used to ensure that stretches of the data with tied vote counts did not distort the measurement.


The overall strategy, then, was to minimize this weighted rank correlation, while also minimizing the complexity of the function, to avoid overfitting. The 5-rough-ratio notational popularity ranking function that had been used by the creators of Sagittal was <math>\text{sopfr}</math> ([https://mathworld.wolfram.com/SumofPrimeFactors.html sum of prime factors with repetition]), and as simple as this function is, it does a remarkably good job of estimating the rank of pitch ratios. For comparison, the weighted sum of squares <math>\text{sopfr}</math> gives for the Scala stats is about 0.026, while the weighted sum of squares <math>\text{N2D3P9}</math> gives is about 0.010. Functions giving sums of squares as low as 0.008 were found, however, these functions were so complex that they probably were fitting to noise in the Scala stats instead of to the true nature of musical pitch. An informal “chunk” metric was devised to compare function complexity in terms of fit to the data, with considered functions ranging from one chunk (<math>\text{sopfr}</math>) to eight chunks; the winning function <math>\text{N2D3P9}</math> has five chunks.
The overall strategy, then, was to minimize this weighted rank correlation, while also minimizing the complexity of the function, to avoid overfitting. An earlier 5-rough-ratio notational popularity ranking function that had been used by the creators of Sagittal was <math>\text{sopfr}</math> ([https://mathworld.wolfram.com/SumofPrimeFactors.html sum of prime factors with repetition]), and as simple as this function is, it does a remarkably good job of estimating the rank of pitch ratios. For comparison, the weighted sum of squares <math>\text{sopfr}</math> gives for the Scala stats is about 0.026, while the weighted sum of squares <math>\text{N2D3P9}</math> gives is about 0.010. Functions giving sums of squares as low as 0.008 were found, however, these functions were so complex that they probably were fitting to noise in the Scala stats instead of to the true nature of musical pitch. An informal “chunk” metric was devised to compare function complexity in terms of fit to the data, with considered functions ranging from one chunk (<math>\text{sopfr}</math>) to eight chunks; the winning function <math>\text{N2D3P9}</math> has five chunks.


Several techniques were used to find and decide on <math>\text{N2D3P9}</math> as the best 5-rough ratio notational popularity rank estimation function. Initial observations about shortcomings of <math>\text{sopfr}</math>, such as its failure to differentiate balanced ratios from their imbalanced equivalents — such as <math>\frac{11}{5}</math> versus <math>\frac{55}{1}</math> — or those with different prime limits such as <math>\frac{13}{5}</math> and <math>\frac{11}{7}</math>, despite those pairs of ratios exhibiting remarkably different actual ranks in the Scala stats, formed the basis of the investigation. Psychoacoustic plausibility of functions was used as a top-down guide for experimentation. [https://en.wikipedia.org/wiki/Mathematical_optimization Optimization] tools such as [https://www.microsoft.com/en-us/microsoft-365/blog/2009/09/21/new-and-improved-solver/ Excel's Evolutionary Solver] were used to navigate toward ideal values for each parameter. A brute-force technique was also utilized whereby nearly 2 billion functions combined out of constituent "submetrics" were checked automatically. In the end, one of the functions generated from the brute-force checker was recognized as being re-writable in a much simpler form with parameter values rounded to whole numbers without doing much damage to its sum-of-squares, and thus <math>\text{N2D3P9}</math> was born.
Several techniques were used to find and decide on <math>\text{N2D3P9}</math> as the best 5-rough ratio notational popularity rank estimation function. Initial observations about shortcomings of <math>\text{sopfr}</math>, such as its failure to differentiate balanced ratios from their imbalanced equivalents — such as <math>\frac{11}{5}</math> versus <math>\frac{55}{1}</math> — or those with different prime limits such as <math>\frac{13}{5}</math> and <math>\frac{11}{7}</math>, despite those pairs of ratios exhibiting remarkably different actual ranks in the Scala stats, formed the basis of the investigation. Psychoacoustic plausibility of functions was used as a top-down guide for experimentation. [https://en.wikipedia.org/wiki/Mathematical_optimization Optimization] tools such as [https://www.microsoft.com/en-us/microsoft-365/blog/2009/09/21/new-and-improved-solver/ Excel's Evolutionary Solver] were used to navigate toward ideal values for each parameter. A brute-force technique was also utilized whereby nearly 2 billion functions combined out of constituent "submetrics" were checked automatically. In the end, one of the functions generated from the brute-force checker was recognized as being re-writable in a much simpler form with parameter values rounded to whole numbers without doing much damage to its sum-of-squares, and thus <math>\text{N2D3P9}</math> was born.