Feature Extraction and Feature Selection: Reducing Data Complexity with Apache Spark

Sisiaridis, Dimitrios; Markowitch, Olivier

Computer Science > Databases

arXiv:1712.08618 (cs)

[Submitted on 11 Dec 2017]

Title:Feature Extraction and Feature Selection: Reducing Data Complexity with Apache Spark

Authors:Dimitrios Sisiaridis, Olivier Markowitch

View PDF

Abstract:Feature extraction and feature selection are the first tasks in pre-processing of input logs in order to detect cyber security threats and attacks while utilizing machine learning. When it comes to the analysis of heterogeneous data derived from different sources, these tasks are found to be time-consuming and difficult to be managed efficiently. In this paper, we present an approach for handling feature extraction and feature selection for security analytics of heterogeneous data derived from different network sensors. The approach is implemented in Apache Spark, using its python API, named pyspark.

Subjects:	Databases (cs.DB); Cryptography and Security (cs.CR)
Cite as:	arXiv:1712.08618 [cs.DB]
	(or arXiv:1712.08618v1 [cs.DB] for this version)
	https://doi.org/10.48550/arXiv.1712.08618
Journal reference:	International Journal of Network Security & Its Applications (IJNSA), Vol.9, No.6, November 2017

Submission history

From: Dimitrios Sisiaridis [view email]
[v1] Mon, 11 Dec 2017 21:02:21 UTC (2,140 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.DB

< prev | next >

new | recent | 2017-12

Change to browse by:

cs
cs.CR

References & Citations

DBLP - CS Bibliography

listing | bibtex

Dimitrios Sisiaridis
Olivier Markowitch

export BibTeX citation

Computer Science > Databases

Title:Feature Extraction and Feature Selection: Reducing Data Complexity with Apache Spark

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Databases

Title:Feature Extraction and Feature Selection: Reducing Data Complexity with Apache Spark

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators