Skip to content

Channels

pikit.channels

Indirect prompt-injection channels.

A channel hides an (optionally attack-worded) payload inside an external data artifact the model later reads — a web page, a retrieved document, an email — and returns the full prompt the target would receive. Channels are orthogonal to attacks and compose freely (attack × channel).

Each channel subclasses :class:pikit.base.Channel and registers itself under a short key. Import this package to populate the registry, then use channels.get(key) / channels.list().

References

Indirect prompt injection was introduced by Greshake et al., "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection" (AISec 2023).

Channel

Bases: ABC

An indirect injection carrier.

Where an :class:Attack controls how a payload is worded, a Channel controls where and how the payload is hidden — inside an external data artifact the model later reads (a web page, a retrieved document, an email). The two are orthogonal and compose freely: word a payload with an attack, then embed it with a channel.

Subclasses implement :meth:taint, which returns the tainted data artifact itself (the web page / document / email). This is what an attacker actually controls and what an agent's compromised tool would return. The concrete :meth:embed is a convenience that prepends an instruction to the tainted artifact to form a full prompt.

Two delivery modes are supported:

  • text mode (:meth:taint) — operates on a plain-text representation of the artifact. This is the simulation default and requires no real files.
  • file mode (:meth:taint_file) — operates on a real file (.html, .eml, .pdf, .ics, …). The tainted output is a file whose format matches what a real agent would encounter. For binary formats (PDF, XLSX) subclasses may use helper libraries; the default implementation reads the file as text and delegates to :meth:taint.

taint abstractmethod

taint(data: str, payload: str) -> str

Hide payload inside data, returning the tainted artifact.

Parameters

data: The clean external data (page HTML, document body, email text). payload: The injected instruction to hide. May be the raw attacker task or the output of an :class:Attack (e.g. attack.inject("", task)) to combine wording with carrier.

Returns

str The tainted artifact — the data with the payload hidden inside, not a full prompt. Feed this to an agent's compromised tool, or use :meth:embed to turn it into a prompt.

taint_file

taint_file(path: str, payload: str, output_path: Optional[str] = None) -> str

Hide payload inside a real file, returning the output path.

Reads the carrier file at path, injects the payload, and writes the tainted file. For text-based formats (HTML, Markdown, YAML, CSV, Python, etc.) the default implementation reads the file as text, calls :meth:taint, and writes the result. Subclasses that target binary formats (PDF, XLSX) override this to use format- specific libraries.

Parameters

path: Path to the clean carrier file. payload: The injected instruction to hide. output_path: Where to write the tainted file. When None, writes to <path>.tainted.<ext>.

Returns

str The path to the tainted file.

extract

extract(tainted_data: str) -> str

Extract the visible text a model would see from a tainted artifact.

In text mode this is typically the identity (the artifact is text). Subclasses for structured formats may parse the artifact to produce the model-visible text. Useful for defenders and for verifying that a payload is present.

extract_file

extract_file(path: str) -> str

Read a tainted file and return the text a model would see.

Default implementation reads the file as text and calls :meth:extract. Binary-format subclasses override to parse.

embed

embed(instruction: str, data: str, payload: str) -> str

Taint data and prepend instruction to form a full prompt.

Convenience for the non-agent case: returns instruction followed by the tainted artifact — the full prompt a target would receive.

pikit.channels.webpage

Webpage channel: hide the payload inside HTML a model scrapes/reads.

Models that summarize or answer questions over fetched web pages ingest the raw HTML, including parts a human browser never renders. This channel hides the payload in such a region so it is invisible on screen yet present in the text the model processes.

WebpageChannel

WebpageChannel(method: str = 'comment')

Bases: Channel

Embed the payload in a non-rendered region of an HTML page.

Parameters

method: * "comment" — inside an HTML comment <!-- payload -->. * "hidden_div" — a display:none div. * "alt_attr" — the alt text of an <img>.

pikit.channels.document

Document channel: hide the payload in a retrieved document or email body.

Models in RAG pipelines and email assistants read document/message text verbatim. This channel plants the payload in a plausible location within that text — a footnote, mid-body, or appended trailer.

DocumentChannel

DocumentChannel(method: str = 'footnote')

Bases: Channel

Embed the payload in the body of a document or email.

Parameters

method: * "footnote" — appended as a footnote-style note at the end. * "inline" — inserted near the middle of the body. * "appended" — plainly appended after the body (most naive).

pikit.channels.markdown

Markdown channel: hide the payload in Markdown a model reads.

Assistants that summarize or answer over Markdown (READMEs, wiki pages, issues, notes) ingest source that renders differently than it reads. This channel hides the payload where it is easy to miss on a rendered page.

MarkdownChannel

MarkdownChannel(method: str = 'comment')

Bases: Channel

Embed the payload in a low-visibility part of a Markdown document.

Parameters

method: * "comment" — an HTML comment <!-- payload --> (not rendered). * "link_title" — the title attribute of an inline link. * "reference" — a trailing reference-style link definition.

pikit.channels.code_comment

Code-comment channel: hide the payload inside source-code comments.

Coding assistants and code-review/analysis agents read source files whole, comments included. An instruction planted in a comment is easy for a human reviewer to skim past but is fully present in what the model processes.

CodeCommentChannel

CodeCommentChannel(style: str = 'hash', position: str = 'end')

Bases: Channel

Embed the payload in a code comment.

Parameters

style: * "hash" — # payload (Python, shell, Ruby, YAML). * "slashes" — // payload (C, Java, JS, Go, Rust). * "block" — /* payload */ (C-family block comment). position: "end" (default) appends the comment after the code; "start" prepends it (e.g. a fake file-header directive).

pikit.channels.skills

Skills channel: hide the payload inside an Agent Skill (SKILL.md).

Agent Skills are Markdown files with a YAML frontmatter (name / description) plus a body of instructions. An agent reads the description to decide whether to load a skill, then follows the body. An attacker who can publish or edit a skill hides instructions in the description (executed the moment the agent considers the skill) or in the body (executed on load) — a topical indirect-injection vector as agents increasingly auto-discover and load third-party skills.

SkillsChannel

SkillsChannel(method: str = 'description')

Bases: Channel

Embed the payload in a SKILL.md skill definition.

Parameters

method: * "description" — append the payload to the frontmatter description (read during skill selection, before load). * "body" — append the payload to the skill body. * "instructions" — insert the payload disguised as a numbered step within the body's instructions.

pikit.channels.unicode_hidden

Unicode-hidden channel: encode the payload as invisible characters.

The payload is rendered as characters that occupy no visible space, so the tainted data looks identical to the clean data to a human, yet the model still receives the hidden instruction in its token stream.

Two schemes are provided, both losslessly decodable (see :func:decode):

  • zero_width — each payload byte becomes 8 zero-width characters (ZWSP for 0, ZWNJ for 1).
  • unicode_tags — each ASCII char is mapped into the Unicode Tags block (U+E0000+), which renders as nothing in most contexts.

UnicodeHiddenChannel

UnicodeHiddenChannel(scheme: str = 'zero_width', position: str = 'end')

Bases: Channel

Embed the payload as invisible Unicode characters within the data.

Parameters

scheme: "zero_width" (default) or "unicode_tags". position: Where to splice the hidden run: "end" (default), "start", or "middle".

decode

decode(text: str) -> str

Recover a payload hidden by either scheme from text.

Useful in tests and for defenders building detectors. Characters that are not part of a hidden payload are ignored.

pikit.channels.structured_data

Structured-data channel: hide the payload inside JSON / CSV / TSV.

Agents that query APIs, databases, or spreadsheets ingest structured responses. A payload planted in a field value, a column name, or a comment line blends in naturally and is processed verbatim by the model.

StructuredDataChannel

StructuredDataChannel(fmt: str = 'json', method: str = 'field_value')

Bases: Channel

Embed the payload in a JSON / CSV / TSV data artifact.

Parameters

fmt: "json" (default), "csv", or "tsv". method: * "field_value" — inside a data field value. * "field_name" — as a new key / column header. * "comment" — a comment-style line (CSV/TSV # payload; JSON "_comment" key).

taint_file

taint_file(path: str, payload: str, output_path=None) -> str

Inject the payload into a real structured-data file.

Detects the format from the file extension (.json, .csv, .tsv) and injects accordingly, writing a valid tainted file.

pikit.channels.pdf_metadata

PDF-metadata channel: hide the payload inside PDF metadata fields.

When a model reads a PDF — via text extraction, a retrieval pipeline, or a summarisation tool — the document's metadata (Title, Author, Subject, Keywords) is often included in the text stream. A payload planted in a metadata field is invisible in the rendered page content yet present in what the model processes.

File mode: :meth:taint_file uses :mod:pypdf to inject the payload into a real PDF's /Info dictionary, producing a valid tainted .pdf.

PDFMetadataChannel

PDFMetadataChannel(field: str = 'title')

Bases: Channel

Embed the payload in a PDF metadata field.

The channel works on a plain-text representation of PDF metadata (key: value lines, one per field). When the raw PDF bytes are not available — which is the common case in simulation — this text form is what the extraction pipeline would hand to the model.

Parameters

field: The metadata field to inject into: * "title" — the Title field. * "author" — the Author field. * "subject" — the Subject field. * "keywords" — the Keywords field. * "custom" — a custom X-Comment field.

taint_file

taint_file(path: str, payload: str, output_path=None) -> str

Inject the payload into a real PDF's metadata dictionary.

Uses :mod:pypdf to read the PDF, overwrite the target metadata field with the payload, and write the tainted PDF.

extract_file

extract_file(path: str) -> str

Return page text plus metadata as a model-facing PDF reader might.

PDF files are binary, so the base text-file implementation cannot reliably expose an injected /Info value to the agent testbed.

pikit.channels.log_file

Log-file channel: hide the payload inside system / application logs.

Debug and monitoring agents read log files to diagnose issues. A payload planted in a log entry — disguised as a warning, an error trace, or a system message — blends in with the noise and is processed verbatim by the model.

LogFileChannel

LogFileChannel(level: str = 'warn', position: str = 'end')

Bases: Channel

Embed the payload as a log entry.

Parameters

level: The log level prefix for the fake entry: * "info" — [INFO] payload * "warn" — [WARN] payload * "error" — [ERROR] payload * "debug" — [DEBUG] payload position: Where to splice the fake entry: "end" (default) or "middle".

pikit.channels.email_headers

Email-header channel: hide the payload inside email header fields.

The existing document channel hides payloads in the email body. This channel targets the headers — X-headers, Reply-To, Subject, and custom fields. Email assistants and triage agents that parse headers for routing, filtering, or display ingest them verbatim, making headers an effective indirect-injection surface distinct from the body.

EmailHeadersChannel

EmailHeadersChannel(field: str = 'x_header')

Bases: Channel

Embed the payload in an email header field.

Works on a plain-text email representation (headers followed by a blank line and the body), which is what an agent's email-parsing tool would produce after MIME decoding.

Parameters

field: * "x_header" — a custom X-Note header (default). * "reply_to" — the Reply-To header. * "subject" — appended to the Subject header. * "custom" — a fully custom X-Instructions header.

pikit.channels.calendar_event

Calendar-event channel: hide the payload inside a calendar invite.

Scheduling agents that read .ics files or calendar APIs process event metadata — title, description, location, attendee notes. A payload planted in one of these fields is ingested verbatim when the agent summarises, triages, or acts on the event.

File mode: :meth:taint_file operates on real .ics files using iCalendar standard field names (SUMMARY, DESCRIPTION, LOCATION, NOTE).

CalendarEventChannel

CalendarEventChannel(field: str = 'title')

Bases: Channel

Embed the payload in a calendar event field.

Works on a plain-text calendar event representation (field key: value lines), which is what an agent's calendar tool would return after parsing an .ics file or calling a calendar API.

Parameters

field: * "title" — the event Title (default). * "description" — the event Description. * "location" — the event Location. * "attendee_note" — an attendee Note field.

taint_file

taint_file(path: str, payload: str, output_path=None) -> str

Inject the payload into a real .ics calendar file.

Uses iCalendar standard property names (SUMMARY, DESCRIPTION, LOCATION, NOTE).

pikit.channels.config_file

Config-file channel: hide the payload inside configuration files.

Agents that audit, review, or deploy applications read YAML, TOML, and .env configuration files. A payload planted in a config value or comment is ingested verbatim — and because config files are trusted by convention, the model may be especially susceptible.

ConfigFileChannel

ConfigFileChannel(fmt: str = 'yaml', method: str = 'value')

Bases: Channel

Embed the payload in a configuration file.

Parameters

fmt: * "yaml" (default) — key: value lines. * "toml" — key = "value" lines. * "env" — KEY=value lines. method: * "value" — appended to an existing config value. * "comment" — a comment line (# payload). * "new_key" — a new config key whose value is the payload.

pikit.channels.translation

Translation channel: hide the payload inside a translation tool's output.

Translation agents read source text and produce translated output. When the source text is attacker-controlled (e.g. user-submitted content, scraped web text), a payload planted in the source survives into the translation output and is then processed by downstream agents that consume the translated text.

TranslationChannel

TranslationChannel(method: str = 'source')

Bases: Channel

Embed the payload in a translation output.

Works on a plain-text translation result representation (Source: ... / Translation: ...), which is what a translation tool would return to the agent.

Parameters

method: * "source" — the payload is in the source text (default). * "translation" — the payload appears in the translated output (as if the translator was tricked into passing it through verbatim). * "note" — a translator's note appended to the output.

pikit.channels.spreadsheet

Spreadsheet channel: hide the payload inside spreadsheet cell data.

Agents that analyse spreadsheets (Excel, Google Sheets, CSV-as-sheet) read cell values, formulas, and comments. A payload planted in a cell value, a cell comment, or a sheet name blends in with legitimate data and is processed verbatim by the model.

File mode: :meth:taint_file operates on real .csv files. For .xlsx files, :mod:openpyxl is used to inject cell values and comments.

SpreadsheetChannel

SpreadsheetChannel(method: str = 'cell_value')

Bases: Channel

Embed the payload in a spreadsheet cell representation.

Works on a plain-text spreadsheet representation (A1: value lines, one per cell), which is what a spreadsheet-reading tool would return to the agent after converting a .xlsx / .gsheet to text.

Parameters

method: * "cell_value" — appended to an existing cell value (default). * "cell_comment" — a cell comment (A1 [comment]: payload). * "sheet_name" — the payload becomes a sheet tab name.

taint_file

taint_file(path: str, payload: str, output_path=None) -> str

Inject the payload into a real spreadsheet file.

Supports .csv files (text-based) and .xlsx files (via :mod:openpyxl).