A web page classifier library based on random image content analysis using deep learning

L. Espinosa-Leal, A. Lendasse, K.-M. Björk, A. Akusok

Research output: Chapter in Book/Report/Conference proceedingConference article in proceedingsScientificpeer-review

5 Citations (Scopus)

Abstract

In this paper we present a methodology and the corresponding Python library 1 for the classification of webpages. The method retrieves a fixed number of images from a given webpage, and based on them classifies the webpage into a set of established classes with a given probability. The library trains a random forest model built upon the features extracted from images by a pre-trained neural network. The implementation is tested by recognizing weapon class webpages in a curated list of 3859 websites. The results show that the best method of classifying a webpage among the classes of interest is to assign the class according to the maximum probability of any image belonging to this (weapon) class being above the threshold, across all the retrieved images. Our finding can have an important impact in the treatment of internet addictions.

Original languageEnglish
Title of host publicationPETRA '18: Proceedings of the 11th PErvasive Technologies Related to Assistive Environments Conference
PublisherAssociation for Computing Machinery ACM
Pages13-16
Number of pages4
ISBN (Electronic)978-1-4503-6390-7
DOIs
Publication statusPublished - 2018
MoE publication typeA4 Article in a conference publication

Keywords

  • Computer vision
  • Deep learning
  • Webpage classification

Fingerprint

Dive into the research topics of 'A web page classifier library based on random image content analysis using deep learning'. Together they form a unique fingerprint.

Cite this