https://github.com/anakintano/compactation-poisoning
We measure whether an indirect prompt injection buried earlier in a conversation survives an LLM-driven context compaction step (the mechanism production agent platforms use to summarize long conversation histories and discard the original turns), and whether it retains behavioral force afterwards. Across 994 trials over three open-weight summarize
https://github.com/anakintano/compactation-poisoning
alignment-research context-compaction research responsible-ai
Last synced: 18 days ago
JSON representation
We measure whether an indirect prompt injection buried earlier in a conversation survives an LLM-driven context compaction step (the mechanism production agent platforms use to summarize long conversation histories and discard the original turns), and whether it retains behavioral force afterwards. Across 994 trials over three open-weight summarize
- Host: GitHub
- URL: https://github.com/anakintano/compactation-poisoning
- Owner: Anakintano
- Created: 2026-07-01T05:30:01.000Z (26 days ago)
- Default Branch: master
- Last Pushed: 2026-07-01T05:48:09.000Z (26 days ago)
- Last Synced: 2026-07-01T07:23:13.324Z (26 days ago)
- Topics: alignment-research, context-compaction, research, responsible-ai
- Language: Python
- Homepage: https://github.com/Anakintano/compactation-poisoning
- Size: 1.3 MB
- Stars: 1
- Watchers: 0
- Forks: 0
- Open Issues: 0