Independent research · AI security

Context changes whether a message is an attack.

I built a detector that scores untrusted text alongside the instructions governing how an agent should use it. The same sentence can be a harmless quotation in one setting and an attempt to redirect the agent in another.

Two modelsModernBERT and BGE reranker
Two matched testsRogue Security 5k and AgentHarm
One input pairTrusted context + untrusted content

The question

Can the model use context?

A string-only detector sees the suspicious sentence but not what the application asked the agent to do. I framed detection as a pairwise classification problem: the trusted task and policy form one input; retrieved pages, tool output, email, or files form the other.

I built separate training and benchmark data paths, source-aware normalization, hard-negative generation, threshold selection on validation data, and benchmark-specific evaluation. This made it possible to compare stored model runs on the same prepared test records without mixing training and test material into one score.

Selected evaluations

Results by dataset

ModernBERT and BGE were evaluated on the same prepared files at a fixed 0.70 decision threshold. The figures below show named stored runs, not a pooled benchmark.

Matched stored comparison at threshold 0.70
DatasetModelROC-AUCF1False-positive rate
Rogue Security 5k 5,000 recordsModernBERT0.9740.9006.1%
Rogue Security 5k 5,000 recordsBGE reranker0.9410.8288.1%
AgentHarm 352 recordsModernBERT0.8790.79433.0%
AgentHarm 352 recordsBGE reranker0.8100.77635.2%

The two datasets test different behavior. The AgentHarm false-positive rate shows that a threshold selected for one setting should not be carried into another without calibration.

Separate model variant

Canonical test split

A different ModernBERT checkpoint, trained on collected real datasets, reached 0.9955 ROC-AUC and 0.9710 F1 on the 2,101-record S-Labs test split at its stored 0.25 threshold.

The prepared training set contained 440,338 rows. An audit found zero exact matches with the test text after case folding and whitespace normalization. That check does not establish that there were no semantic near-duplicates. This result belongs to its own model variant and is not part of the comparison above.

What the work established

A useful screening component needs a deployment policy.

Context-aware scoring improved the two matched stored comparisons. The wider benchmark suite showed variable transfer across domains, so I treated the model as one screening signal. The application still needs rules for what to do with a flagged document, a calibrated threshold for its traffic, and monitoring for false alerts.

This case study uses aggregate results only. Raw research datasets and model weights are not published here.

Back to selected work

Click outside the figure or press Escape to close.