2008 · 12 citations · 24 references
CartographyGeographic Information SystemsGeospatial SemanticsInformation RetrievalData ScienceGeographic Information RetrievalVolunteered Geographic InformationGeographyData IntegrationGeographic CoverageGreat BritainDigital GeographySemantic WebPublic HealthGeospatial DataSocial SciencesWeb Coverage
In this paper, we describe a methodology to estimate the geographic coverage of the web without the need for secondary knowledge or complex geo-tagging. This is achieved by randomly selecting toponyms from the Ordnance Survey 50K gazetteer to create search queries and thus gather document counts from various web sources for Great Britain. The same gazetteer is then used to geo-code the results and enable mapping. To validate our approach, and demonstrate the effects of geo/non-geo and geo/geo ambiguity, we mapped the selected toponyms to Geograph, a community project that contains user generated geo-tagged photographs of the UK. Although success varies with resolution, the proposed approach is likely sufficient to be reliably used by applications exploring the geographic coverage of the web for cases where references to settlements are likely to be common. In our case, we applied the method to produce maps of web coverage for a range of sources at a resolution of 30km.
24
Introduction to the Special Issue on the Web as Corpus
Adam Kilgarriff, Gregory Grefenstette · Computational Linguistics · 2003 · 920 citations · Full text
Engineering, Semantic Web, Semantics +22
Einat Amitay, Nadav Har’El, Ron Sivan et al. · 2004 · 528 citations
Philip Resnik, Noah A. Smith · Computational Linguistics · 2003 · 525 citations · Full text