Why it matters
Adversa's IICL study evaluates a few-shot jailbreak that interleaves benign and harmful demonstrations and uses short output-field labels to shift model behavior. Across more than 3,500 probes, ten models, and seven ablations, results vary materially with example order and field names; the work is vendor-authored and its model-specific attack rates should be independently reproduced.
My takeaway: Add structural few-shot attacks to safety regression suites, vary demonstration order, labels, templates, and model versions, and report both per-query and per-attempt success with uncertainty. Keep a held-out attack set and rerun it after every system-prompt, scaffold, or model change.