Publication | Closed Access
Judging a site by its content
28
Citations
20
References
2011
Year
Unknown Venue
Abuse DetectionEngineeringMachine LearningInformation ForensicsContent CreationCommunicationText MiningNatural Language ProcessingSpam FilteringSocial MediaInformation RetrievalData ScienceData MiningPattern RecognitionContent AnalysisPhysical WorldThreat DetectionUser-generated ContentKnowledge DiscoveryWebometricsComputer ScienceLegitimate Web PageContent AuditingTextual ContentArtsPhishing
Physical‑world cues reliably distinguish safe from unsafe situations, but the ambiguous nature of the Internet makes it difficult for users to differentiate scams from legitimate web pages. The study investigates training classifiers that automatically detect malicious web pages by leveraging textual content, structural tags, page links, visual appearance, and URLs. Using a large web‑mail provider’s labeled data, the authors extract these features, train classifiers, and assess the predictive strength of each feature set. The resulting classifiers cut the error rate by more than half compared to URL‑only models and outperform previous, more constrained approaches.
The physical world is rife with cues that allow us to distinguish between safe and unsafe situations. By contrast, the Internet offers a much more ambiguous environment; hence many users are unable to distinguish a scam from a legitimate Web page. To help address this problem, we explore how to train classifiers that can automatically identify malicious Web pages based on clues from their textual content, structural tags, page links, visual appearance, and URLs. Using a contemporary labeled data feed from a large Web mail provider, we extract such features and demonstrate how they can be used to improve classification accuracy over previous, more constrained approaches. In particular, by analyzing the full content of individual Web pages, we more than halve the error rate obtained by a comparably trained classifier that only extracts features from URLs. By training classifiers on different sets of features, we are further able to assess the strength of clues provided by these different sources of information.
| Year | Citations | |
|---|---|---|
Page 1
Page 1