Recognizing and Extracting Cybersecurity Entities from Text
dc.contributor.author | Hanks, Casey | |
dc.contributor.author | Maiden, Michael | |
dc.contributor.author | Ranade, Priyanka | |
dc.contributor.author | Finin, Tim | |
dc.contributor.author | Joshi, Anupam | |
dc.date.accessioned | 2022-07-19T20:42:17Z | |
dc.date.available | 2022-07-19T20:42:17Z | |
dc.date.issued | 2022-08-02 | |
dc.description | Proceedings of the 39th International Conference on Machine Learning, Baltimore, Maryland, USA | en_US |
dc.description.abstract | Cyber Threat Intelligence (CTI) is information describing threat vectors, vulnerabilities, and attacks and is often used as training data for AI-based cyber defense systems such as Cybersecurity Knowledge Graphs (CKG). There is a strong need to develop community-accessible datasets to train existing AI-based cybersecurity pipelines to efficiently and accurately extract meaningful insights from CTI. We have created an initial unstructured CTI corpus from a variety of open sources that we are using to train and test cybersecurity entity models using the spaCy framework and exploring self-learning methods to automatically recognize cybersecurity entities. We also describe methods to apply cybersecurity domain entity linking with existing world knowledge from Wikidata. Our future work will survey and test spaCy NLP tools, and create methods for continuous integration of new information extracted from text. | en_US |
dc.description.sponsorship | This research was supported by grants from NSA and the National Science Foundation (No. 2114892). | en_US |
dc.description.uri | https://par.nsf.gov/biblio/10416967-recognizing-extracting-cybersecurity-entities-from-text | en_US |
dc.format.extent | 7 pages | en_US |
dc.genre | conference papers and proceedings | en_US |
dc.genre | postprints | en_US |
dc.identifier | doi:10.13016/m2i1n3-2log | |
dc.identifier.uri | http://hdl.handle.net/11603/25196 | |
dc.language.iso | en_US | en_US |
dc.relation.isAvailableAt | The University of Maryland, Baltimore County (UMBC) | |
dc.relation.ispartof | UMBC Computer Science and Electrical Engineering Department Collection | |
dc.relation.ispartof | UMBC Faculty Collection | |
dc.relation.ispartof | UMBC Student Collection | |
dc.rights | This item is likely protected under Title 17 of the U.S. Copyright Law. Unless on a Creative Commons license, for uses protected by Copyright Law, contact the copyright holder or the author. | en_US |
dc.subject | UMBC Ebiquity Research Group | |
dc.title | Recognizing and Extracting Cybersecurity Entities from Text | en_US |
dc.type | Text | en_US |
dcterms.creator | https://orcid.org/0000-0002-6593-1792 | en_US |
dcterms.creator | https://orcid.org/0000-0002-8641-3193 | en_US |