Channels¶
pikit.channels ¶
Indirect prompt-injection channels.
A channel hides an (optionally attack-worded) payload inside an external
data artifact the model later reads — a web page, a retrieved document, an
email — and returns the full prompt the target would receive. Channels are
orthogonal to attacks and compose freely (attack × channel).
Each channel subclasses :class:pikit.base.Channel and registers itself
under a short key. Import this package to populate the registry, then use
channels.get(key) / channels.list().
References¶
Indirect prompt injection was introduced by Greshake et al., "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection" (AISec 2023).
Channel ¶
Bases: ABC
An indirect injection carrier.
Where an :class:Attack controls how a payload is worded, a Channel
controls where and how the payload is hidden — inside an external data
artifact the model later reads (a web page, a retrieved document, an
email). The two are orthogonal and compose freely: word a payload with an
attack, then embed it with a channel.
Subclasses implement :meth:taint, which returns the tainted data
artifact itself (the web page / document / email). This is what an
attacker actually controls and what an agent's compromised tool would
return. The concrete :meth:embed is a convenience that prepends an
instruction to the tainted artifact to form a full prompt.
Two delivery modes are supported:
- text mode (:meth:
taint) — operates on a plain-text representation of the artifact. This is the simulation default and requires no real files. - file mode (:meth:
taint_file) — operates on a real file (.html,.eml,.pdf,.ics, …). The tainted output is a file whose format matches what a real agent would encounter. For binary formats (PDF, XLSX) subclasses may use helper libraries; the default implementation reads the file as text and delegates to :meth:taint.
taint
abstractmethod
¶
Hide payload inside data, returning the tainted artifact.
Parameters¶
data:
The clean external data (page HTML, document body, email text).
payload:
The injected instruction to hide. May be the raw attacker task
or the output of an :class:Attack (e.g. attack.inject("", task))
to combine wording with carrier.
Returns¶
str
The tainted artifact — the data with the payload hidden inside,
not a full prompt. Feed this to an agent's compromised tool,
or use :meth:embed to turn it into a prompt.
taint_file ¶
Hide payload inside a real file, returning the output path.
Reads the carrier file at path, injects the payload, and writes
the tainted file. For text-based formats (HTML, Markdown, YAML,
CSV, Python, etc.) the default implementation reads the file as
text, calls :meth:taint, and writes the result. Subclasses that
target binary formats (PDF, XLSX) override this to use format-
specific libraries.
Parameters¶
path:
Path to the clean carrier file.
payload:
The injected instruction to hide.
output_path:
Where to write the tainted file. When None, writes to
<path>.tainted.<ext>.
Returns¶
str The path to the tainted file.
extract ¶
Extract the visible text a model would see from a tainted artifact.
In text mode this is typically the identity (the artifact is text). Subclasses for structured formats may parse the artifact to produce the model-visible text. Useful for defenders and for verifying that a payload is present.
extract_file ¶
Read a tainted file and return the text a model would see.
Default implementation reads the file as text and calls
:meth:extract. Binary-format subclasses override to parse.
embed ¶
Taint data and prepend instruction to form a full prompt.
Convenience for the non-agent case: returns instruction followed
by the tainted artifact — the full prompt a target would receive.
pikit.channels.webpage ¶
Webpage channel: hide the payload inside HTML a model scrapes/reads.
Models that summarize or answer questions over fetched web pages ingest the raw HTML, including parts a human browser never renders. This channel hides the payload in such a region so it is invisible on screen yet present in the text the model processes.
pikit.channels.document ¶
Document channel: hide the payload in a retrieved document or email body.
Models in RAG pipelines and email assistants read document/message text verbatim. This channel plants the payload in a plausible location within that text — a footnote, mid-body, or appended trailer.
DocumentChannel ¶
pikit.channels.markdown ¶
Markdown channel: hide the payload in Markdown a model reads.
Assistants that summarize or answer over Markdown (READMEs, wiki pages, issues, notes) ingest source that renders differently than it reads. This channel hides the payload where it is easy to miss on a rendered page.
MarkdownChannel ¶
pikit.channels.code_comment ¶
Code-comment channel: hide the payload inside source-code comments.
Coding assistants and code-review/analysis agents read source files whole, comments included. An instruction planted in a comment is easy for a human reviewer to skim past but is fully present in what the model processes.
CodeCommentChannel ¶
Bases: Channel
Embed the payload in a code comment.
Parameters¶
style:
* "hash" — # payload (Python, shell, Ruby, YAML).
* "slashes" — // payload (C, Java, JS, Go, Rust).
* "block" — /* payload */ (C-family block comment).
position:
"end" (default) appends the comment after the code; "start"
prepends it (e.g. a fake file-header directive).
pikit.channels.skills ¶
Skills channel: hide the payload inside an Agent Skill (SKILL.md).
Agent Skills are Markdown files with a YAML frontmatter (name /
description) plus a body of instructions. An agent reads the
description to decide whether to load a skill, then follows the body.
An attacker who can publish or edit a skill hides instructions in the
description (executed the moment the agent considers the skill) or in
the body (executed on load) — a topical indirect-injection vector as agents
increasingly auto-discover and load third-party skills.
SkillsChannel ¶
Bases: Channel
Embed the payload in a SKILL.md skill definition.
Parameters¶
method:
* "description" — append the payload to the frontmatter
description (read during skill selection, before load).
* "body" — append the payload to the skill body.
* "instructions" — insert the payload disguised as a numbered
step within the body's instructions.
pikit.channels.unicode_hidden ¶
Unicode-hidden channel: encode the payload as invisible characters.
The payload is rendered as characters that occupy no visible space, so the tainted data looks identical to the clean data to a human, yet the model still receives the hidden instruction in its token stream.
Two schemes are provided, both losslessly decodable (see :func:decode):
zero_width— each payload byte becomes 8 zero-width characters (ZWSP for 0, ZWNJ for 1).unicode_tags— each ASCII char is mapped into the Unicode Tags block (U+E0000+), which renders as nothing in most contexts.
pikit.channels.structured_data ¶
Structured-data channel: hide the payload inside JSON / CSV / TSV.
Agents that query APIs, databases, or spreadsheets ingest structured responses. A payload planted in a field value, a column name, or a comment line blends in naturally and is processed verbatim by the model.
StructuredDataChannel ¶
Bases: Channel
Embed the payload in a JSON / CSV / TSV data artifact.
Parameters¶
fmt:
"json" (default), "csv", or "tsv".
method:
* "field_value" — inside a data field value.
* "field_name" — as a new key / column header.
* "comment" — a comment-style line (CSV/TSV # payload;
JSON "_comment" key).
taint_file ¶
Inject the payload into a real structured-data file.
Detects the format from the file extension (.json, .csv,
.tsv) and injects accordingly, writing a valid tainted file.
pikit.channels.pdf_metadata ¶
PDF-metadata channel: hide the payload inside PDF metadata fields.
When a model reads a PDF — via text extraction, a retrieval pipeline, or a summarisation tool — the document's metadata (Title, Author, Subject, Keywords) is often included in the text stream. A payload planted in a metadata field is invisible in the rendered page content yet present in what the model processes.
File mode: :meth:taint_file uses :mod:pypdf to inject the payload
into a real PDF's /Info dictionary, producing a valid tainted .pdf.
PDFMetadataChannel ¶
Bases: Channel
Embed the payload in a PDF metadata field.
The channel works on a plain-text representation of PDF metadata
(key: value lines, one per field). When the raw PDF bytes are not
available — which is the common case in simulation — this text form
is what the extraction pipeline would hand to the model.
Parameters¶
field:
The metadata field to inject into:
* "title" — the Title field.
* "author" — the Author field.
* "subject" — the Subject field.
* "keywords" — the Keywords field.
* "custom" — a custom X-Comment field.
taint_file ¶
Inject the payload into a real PDF's metadata dictionary.
Uses :mod:pypdf to read the PDF, overwrite the target metadata
field with the payload, and write the tainted PDF.
extract_file ¶
Return page text plus metadata as a model-facing PDF reader might.
PDF files are binary, so the base text-file implementation cannot
reliably expose an injected /Info value to the agent testbed.
pikit.channels.log_file ¶
Log-file channel: hide the payload inside system / application logs.
Debug and monitoring agents read log files to diagnose issues. A payload planted in a log entry — disguised as a warning, an error trace, or a system message — blends in with the noise and is processed verbatim by the model.
LogFileChannel ¶
pikit.channels.email_headers ¶
Email-header channel: hide the payload inside email header fields.
The existing document channel hides payloads in the email body.
This channel targets the headers — X-headers, Reply-To, Subject, and
custom fields. Email assistants and triage agents that parse headers
for routing, filtering, or display ingest them verbatim, making headers
an effective indirect-injection surface distinct from the body.
EmailHeadersChannel ¶
Bases: Channel
Embed the payload in an email header field.
Works on a plain-text email representation (headers followed by a blank line and the body), which is what an agent's email-parsing tool would produce after MIME decoding.
Parameters¶
field:
* "x_header" — a custom X-Note header (default).
* "reply_to" — the Reply-To header.
* "subject" — appended to the Subject header.
* "custom" — a fully custom X-Instructions header.
pikit.channels.calendar_event ¶
Calendar-event channel: hide the payload inside a calendar invite.
Scheduling agents that read .ics files or calendar APIs process event metadata — title, description, location, attendee notes. A payload planted in one of these fields is ingested verbatim when the agent summarises, triages, or acts on the event.
File mode: :meth:taint_file operates on real .ics files using
iCalendar standard field names (SUMMARY, DESCRIPTION, LOCATION,
NOTE).
CalendarEventChannel ¶
Bases: Channel
Embed the payload in a calendar event field.
Works on a plain-text calendar event representation (field key:
value lines), which is what an agent's calendar tool would return
after parsing an .ics file or calling a calendar API.
Parameters¶
field:
* "title" — the event Title (default).
* "description" — the event Description.
* "location" — the event Location.
* "attendee_note" — an attendee Note field.
taint_file ¶
Inject the payload into a real .ics calendar file.
Uses iCalendar standard property names (SUMMARY, DESCRIPTION,
LOCATION, NOTE).
pikit.channels.config_file ¶
Config-file channel: hide the payload inside configuration files.
Agents that audit, review, or deploy applications read YAML, TOML, and .env configuration files. A payload planted in a config value or comment is ingested verbatim — and because config files are trusted by convention, the model may be especially susceptible.
ConfigFileChannel ¶
Bases: Channel
Embed the payload in a configuration file.
Parameters¶
fmt:
* "yaml" (default) — key: value lines.
* "toml" — key = "value" lines.
* "env" — KEY=value lines.
method:
* "value" — appended to an existing config value.
* "comment" — a comment line (# payload).
* "new_key" — a new config key whose value is the payload.
pikit.channels.translation ¶
Translation channel: hide the payload inside a translation tool's output.
Translation agents read source text and produce translated output. When the source text is attacker-controlled (e.g. user-submitted content, scraped web text), a payload planted in the source survives into the translation output and is then processed by downstream agents that consume the translated text.
TranslationChannel ¶
Bases: Channel
Embed the payload in a translation output.
Works on a plain-text translation result representation (Source:
... / Translation: ...), which is what a translation tool would
return to the agent.
Parameters¶
method:
* "source" — the payload is in the source text (default).
* "translation" — the payload appears in the translated
output (as if the translator was tricked into passing it
through verbatim).
* "note" — a translator's note appended to the output.
pikit.channels.spreadsheet ¶
Spreadsheet channel: hide the payload inside spreadsheet cell data.
Agents that analyse spreadsheets (Excel, Google Sheets, CSV-as-sheet) read cell values, formulas, and comments. A payload planted in a cell value, a cell comment, or a sheet name blends in with legitimate data and is processed verbatim by the model.
File mode: :meth:taint_file operates on real .csv files.
For .xlsx files, :mod:openpyxl is used to inject cell values and
comments.
SpreadsheetChannel ¶
Bases: Channel
Embed the payload in a spreadsheet cell representation.
Works on a plain-text spreadsheet representation (A1: value lines,
one per cell), which is what a spreadsheet-reading tool would return
to the agent after converting a .xlsx / .gsheet to text.
Parameters¶
method:
* "cell_value" — appended to an existing cell value (default).
* "cell_comment" — a cell comment (A1 [comment]: payload).
* "sheet_name" — the payload becomes a sheet tab name.
taint_file ¶
Inject the payload into a real spreadsheet file.
Supports .csv files (text-based) and .xlsx files (via
:mod:openpyxl).