networks
Is it really a power law?
It is 2007, and "scale-free" has become the most successful brand in network science: hundreds of papers report power-law degree distributions in the internet, metabolic networks, citation graphs — usually on the evidence of a roughly straight line on a log-log plot, fit by least squares, which is statistically indefensible. Test power-law claims rigorously: maximum-likelihood fits with principled estimates of where the tail begins, goodness-of-fit against the hypothesis itself, and likelihood-ratio comparisons against the boring alternatives — lognormal, stretched exponential — that mimic straight lines. Then audit a famous published claim and report what survives. Theories of network growth, robustness, and epidemic threshold all lean on these tails. If the power laws are artifacts, a decade of theory has been explaining a plotting choice.
Who this problem belongs to
The two figures whose methods fit it best, out of 48 in contention.
Clauset co-authored, with Shalizi and Newman, the 2009 SIAM Review paper that is essentially this problem: maximum-likelihood estimation of the power-law exponent, a principled estimator for where the tail begins, Kolmogorov-Smirnov goodness-of-fit via bootstrap, and likelihood-ratio tests against lognormal and exponential alternatives, applied to re-audit dozens of published 'scale-free' claims and find most unsupported. His subsequent network-science career (with Newman and Shalizi) built the toolkit this problem asks for from scratch, including the software (powerlaw) practitioners still use. There is no better-matched carrier in this pool: the problem is not analogous to his work, it is a description of it, down to the audit-a-famous-claim structure.
Shalizi is the second author of the Clauset-Shalizi-Newman 2009 paper that this problem restates almost verbatim: rigorous maximum-likelihood fitting, principled tail-cutoff selection, and likelihood-ratio comparison against lognormal and exponential alternatives, deployed specifically to re-examine celebrated 'scale-free network' claims and show most fail. As a statistician working at the boundary of complex systems and rigorous inference, his broader body of work is a running critique of exactly the kind of hype-driven, methodologically loose empirical claim this problem describes — power laws as the paradigm case. He is essentially a co-author of the solution this problem is asking for, making him one of the two strongest possible carriers alongside Clauset.
Fought here
In the mind map
The same ideas, as concepts rather than history — in John's ML knowledge map.
48 figures are scored on this problem. Draw it in a battle to see where you land.