2012 · 678 citations · 40 references
Privacy ProtectionEngineeringUsable SecurityInformation SecurityInformation ForensicsAnonymized CorpusCorpus LinguisticsPseudonymizationInformation RetrievalData ScienceData MiningComputational LinguisticsData AnonymizationUser-chosen PasswordsShannon EntropyIdentity-based SecurityData PrivacyComputer ScienceMillion PasswordsPrivacyPrivacy LeakageData SecurityCryptographySecurity MeasurementPhishingPartial Guessing Metrics
The study aims to develop a statistical framework for estimating password guessing difficulty from a massive anonymized corpus. The authors use anonymized password histograms of ~70 million Yahoo! users to construct partial guessing metrics, including a new guesswork variant parameterized by an attacker's success rate, replacing Shannon entropy.
We report on the largest corpus of user-chosen passwords ever studied, consisting of anonymized password histograms representing almost 70 million Yahoo! users, mitigating privacy concerns while enabling analysis of dozens of subpopulations based on demographic factors and site usage characteristics. This large data set motivates a thorough statistical treatment of estimating guessing difficulty by sampling from a secret distribution. In place of previously used metrics such as Shannon entropy and guessing entropy, which cannot be estimated with any realistically sized sample, we develop partial guessing metrics including a new variant of guesswork parameterized by an attacker's desired success rate. Our new metric is comparatively easy to approximate and directly relevant for security engineering. By comparing password distributions with a uniform distribution which would provide equivalent security against different forms of guessing attack, we estimate that passwords provide fewer than 10 bits of security against an online, trawling attack, and only about 20 bits of security against an optimal offline dictionary attack. We find surprisingly little variation in guessing difficulty; every identifiable group of users generated a comparably weak password distribution. Security motivations such as the registration of a payment card have no greater impact than demographic factors such as age and nationality. Even proactive efforts to nudge users towards better password choices with graphical feedback make little difference. More surprisingly, even seemingly distant language communities choose the same weak passwords and an attacker never gains more than a factor of 2 efficiency gain by switching from the globally optimal dictionary to a population-specific lists.
40
A Mathematical Theory of Communication
Claude E. Shannon · Bell System Technical Journal · 1948 · 78.4K citations
Power-Law Distributions in Empirical Data
SIAM Review · 2009 · 7.4K citations · Full text
Power-law distributions in empirical data
Aaron Clauset, Cosma Rohilla Shalizi, M. E. J. Newman · Figshare · 2018 · 6.7K citations · Full text
Transmission of Information<sup>1</sup>
R. V. L. Hartley · Bell System Technical Journal · 1928 · 1.8K citations
A large-scale study of web password habits
Dinei Florêncio, Cormac Herley · 2007 · 1K citations