Extraction of entities from cyber threat intelligence reports using large language models

Гусейнли, А.

Рақамли технологияларнинг назарий ва амалий масалалари · 2026-yil

Annotatsiya

Cyber Threat Intelligence reports combine analytical prose with dense technical indicators, making structured entity extraction a challenging but operationally valuable task. This paper presents a comparative evaluation of three large language models - Claude Sonnet 4.6, GPT-5.4, and LLaMA 4 Scout - on a manually annotated corpus of 21 real-world CTI reports across 15 entity types and 1284 ground truth instances. The paper evaluates zero-shot and few-shot prompting conditions and studies the effect of iterative prompt refinement, focusing on explicit format constraints for cryptographic hash entities. Results show that Claude Sonnet 4.6 and GPT-5.4 achieve comparable performance under zero-shot conditions, with LLaMA 4 Scout trailing by a substantial margin. Few-shot prompting consistently reduces hallucination rates but yields mixed F1 results, with exemplar cardinality emerging as a critical design factor. Entity extraction difficulty varies substantially across types, with technical indicator categories showing near-perfect performance and semantic categories such as tool and target sector posing the greatest challenges across all evaluated models.

Maqola ma’lumotlari
MualliflarГусейнли, А.
JurnalРақамли технологияларнинг назарий ва амалий масалалари
Nashr sanasi2026-08-03
Jild9
Son3
Betlar7-15
TilRus
DOI10.62132/ijdt.v9i3.393

Kalit so‘zlar

киберразведка, распознавание именованных сущностей, большие языковые модели, инженерия промптов, cyber threat intelligence, named entity recognition, large language models, prompt engineering

Ilmiy soha

Рақамли технологияларнинг назарий ва амалий масалалари jurnalidan boshqa maqolalar

Рақамли технологияларнинг назарий ва амалий масалалари — barcha maqolalar