Nowadays, Large Language Models (LLMs) are widely included in many daily applications and services, where they play a crucial role in supporting users for automated decision-making tasks. As their adoption grows, so does their vulnerability to prompt injection attacks, which use natural language instructions to bypass safety constraints and also manipulate model behavior. To address these issues, the paper proposes a hybrid framework for prompt injection detection that combines domain-adaptive pretraining (DAPT) of RoBERTa encoder with a set of rule-based text patterns to detect suspicious adversarial interaction modes within prompt inputs. The proposed approach is evaluated through an empirical analysis across multiple prompt injection datasets. Experimental results show that the hybrid model outperforms text-only baselines and maintains reasonable performance under distributional shift. An ablation study is conducted to evaluate the individual component contribution.
Linguistic Pattern-Enriched Transformers for Prompt Injection Detection in LLM Systems
De Maio Carmen.
;Senatore Sabrina.;
2026
Abstract
Nowadays, Large Language Models (LLMs) are widely included in many daily applications and services, where they play a crucial role in supporting users for automated decision-making tasks. As their adoption grows, so does their vulnerability to prompt injection attacks, which use natural language instructions to bypass safety constraints and also manipulate model behavior. To address these issues, the paper proposes a hybrid framework for prompt injection detection that combines domain-adaptive pretraining (DAPT) of RoBERTa encoder with a set of rule-based text patterns to detect suspicious adversarial interaction modes within prompt inputs. The proposed approach is evaluated through an empirical analysis across multiple prompt injection datasets. Experimental results show that the hybrid model outperforms text-only baselines and maintains reasonable performance under distributional shift. An ablation study is conducted to evaluate the individual component contribution.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


