Skip to content

Agent

pikit.agent

Agent environments for testing prompt-injection attacks.

An agent wraps a :class:~pikit.targets.Target and exposes a run() that returns a human-readable :class:Trace. Three kinds ship:

  • chat — a plain assistant (no tools); direct injection via the user message.
  • tool — a general tool-calling agent; indirect injection via a compromised tool's return value (the taint map).
  • scenario agents — email / rag / browser — preconfigured with a realistic toolset and an observable sink.

Like the rest of pikit, agents register under short keys; use get_agent(key) to fetch the class, then construct it.

Agent

Agent(target: Target, *, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None)

Bases: ABC

Base class for agents under test.

Parameters

target: The model backend (:class:~pikit.targets.Target). system: Optional system prompt. defenses: Optional :class:DefenseHooks applied at the three insertion points.

run abstractmethod

run(user_message: str, **kwargs) -> Trace

Run the agent on user_message and return the :class:Trace.

Trace dataclass

Trace(steps: List[TraceStep] = list(), final_text: str = '')

An ordered record of an agent run, for human inspection.

sink_calls property

sink_calls: List[TraceStep]

Tool-call steps that hit a sink (observable action).

tainted_steps property

tainted_steps: List[TraceStep]

Steps whose data was the injected artifact.

to_dict

to_dict() -> Dict[str, Any]

Return the structured trace used in JSON/JSONL experiment output.

TraceStep dataclass

TraceStep(kind: str, text: str = '', tool_name: Optional[str] = None, args: Optional[dict] = None, content: Optional[str] = None, tainted: bool = False, is_sink: bool = False, decision: Optional[str] = None)

One step in an agent run.

to_dict

to_dict() -> Dict[str, Any]

Return a JSON-serializable representation of this trace step.

DefenseHooks dataclass

DefenseHooks(system: Optional[Defense] = None, tool_result: Optional[Defense] = None, user: Optional[Defense] = None)

Optional defenses applied at three points of the agent loop.

Parameters

system: Hardens the system prompt (defends against the model being talked out of its instructions). tool_result: Hardens untrusted tool output before it re-enters the model — the key defense position for indirect injection. user: Hardens the incoming user message (defends against direct injection).

CallableAgentAdapter

CallableAgentAdapter(runner: AgentCallable, *, system: Optional[str] = None, taint: Optional[Dict[str, str]] = None, defenses: Optional[DefenseHooks] = None)

Bases: Agent

Wrap a user-supplied agent callable as a pikit agent.

The callable receives the hardened user message. It may return either a final text string or a fully populated :class:Trace. Optional keyword arguments expose the adapter context to framework integrations: system, taint, and defenses.

Returning a Trace is recommended whenever the external framework can capture tool calls. Returning text still provides a useful minimal adapter for direct-injection testing.

Tool dataclass

Tool(name: str, description: str, func: Callable[..., Any], parameters: dict = (lambda: {'type': 'object', 'properties': {}})(), is_sink: bool = False, category: str = 'general')

A callable tool exposed to an agent's model.

Parameters

name, description, func: Identity, model-facing description, and the underlying callable. parameters: JSON-schema for the arguments. Auto-derived if not given. is_sink: Marks an externally-observable action (e.g. send_email). The trace highlights when a sink fires — the key signal for judging whether an injection succeeded. category: Tool category for pool-based selection. One of: "web", "email", "file", "code", "knowledge", "communication", "general".

to_schema

to_schema() -> dict

Return the provider-agnostic {name, description, parameters}.

tool

tool(name: Optional[str] = None, *, description: Optional[str] = None, is_sink: bool = False, parameters: Optional[dict] = None, category: str = 'general') -> Callable[[Callable], Tool]

Decorator turning a plain function into a :class:Tool.

The tool name defaults to the function name; description to its docstring. parameters is auto-derived from type hints unless given. category tags the tool for pool-based selection by scenario agents.

Examples

@tool(description="Fetch a URL and return its body.", category="web") ... def fetch_url(url: str) -> str: ... return "..." isinstance(fetch_url, Tool) True

build_system_prompt

build_system_prompt(role: str, tools: List[Tool], *, instructions: str = '', safety: Optional[str] = None) -> str

Build a structured system prompt from the agent's tool set.

Parameters

role: One-line description of the agent's role (e.g. "a web browsing assistant"). tools: The tools the agent can call. instructions: How the agent should use its tools to accomplish tasks. safety: Optional safety constraint text. Defaults to a standard indirect- injection warning.

all_tools

all_tools() -> List[Tool]

Return all tools in the pool.

get_tool

get_tool(name: str) -> Optional[Tool]

Return a single tool by name, or None if not found.

get_tools

get_tools(names: List[str]) -> List[Tool]

Return a list of tools by name (skips unknown names).

tools_by_category

tools_by_category(category: str) -> List[Tool]

Return all tools in a given category.

data_source_tools

data_source_tools() -> List[Tool]

Return all data-source tools (potential taint points).

sink_tools

sink_tools() -> List[Tool]

Return all sink tools (externally-observable actions).

tool_names

tool_names() -> List[str]

Return all tool names in the pool.

categories

categories() -> List[str]

Return all distinct categories in the pool.

pikit.agent.base

Agent base class and the execution :class:Trace.

run() returns a :class:Trace rather than a bare string. The trace is the artifact a human reads to judge — manually — whether an injection succeeded: it shows every model turn, tool call, and tool result, and highlights when a sink fired or a step carried tainted data. The library deliberately renders no verdict (no evaluator/scoring); it makes the signals easy to see and offers structured accessors so you can write your own one-line assertion.

TraceStep dataclass

TraceStep(kind: str, text: str = '', tool_name: Optional[str] = None, args: Optional[dict] = None, content: Optional[str] = None, tainted: bool = False, is_sink: bool = False, decision: Optional[str] = None)

One step in an agent run.

to_dict

to_dict() -> Dict[str, Any]

Return a JSON-serializable representation of this trace step.

Trace dataclass

Trace(steps: List[TraceStep] = list(), final_text: str = '')

An ordered record of an agent run, for human inspection.

sink_calls property

sink_calls: List[TraceStep]

Tool-call steps that hit a sink (observable action).

tainted_steps property

tainted_steps: List[TraceStep]

Steps whose data was the injected artifact.

to_dict

to_dict() -> Dict[str, Any]

Return the structured trace used in JSON/JSONL experiment output.

Agent

Agent(target: Target, *, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None)

Bases: ABC

Base class for agents under test.

Parameters

target: The model backend (:class:~pikit.targets.Target). system: Optional system prompt. defenses: Optional :class:DefenseHooks applied at the three insertion points.

run abstractmethod

run(user_message: str, **kwargs) -> Trace

Run the agent on user_message and return the :class:Trace.

pikit.agent.loop

The provider-agnostic function-calling loop shared by tool agents.

run_tool_loop

run_tool_loop(target: Target, user_message: str, tools: List[Tool], *, system: Optional[str] = None, hooks: Optional[DefenseHooks] = None, taint: Optional[Dict[str, str]] = None, max_steps: int = 8, **target_kwargs) -> Trace

Drive target.chat over tools until it stops calling tools.

Parameters

taint: Map of tool_name -> artifact. When the model calls a tainted tool, the loop returns the artifact as that tool's result instead of invoking the real function (the indirect-injection delivery point, and it avoids real side effects during a test). **target_kwargs: Forwarded to target.chat() on every call (e.g. temperature=0.7).

pikit.agent.hooks

Defense insertion points for agents.

The same prevention-style :class:~pikit.base.Defense objects used for direct injection can be slotted into three points of an agent's data flow. The most valuable for indirect injection is tool_result — the layer through which an attacker's tainted artifact re-enters the model.

DefenseHooks dataclass

DefenseHooks(system: Optional[Defense] = None, tool_result: Optional[Defense] = None, user: Optional[Defense] = None)

Optional defenses applied at three points of the agent loop.

Parameters

system: Hardens the system prompt (defends against the model being talked out of its instructions). tool_result: Hardens untrusted tool output before it re-enters the model — the key defense position for indirect injection. user: Hardens the incoming user message (defends against direct injection).

pikit.agent.adapters

Adapters for exercising externally implemented agents with pikit.

The built-in scenarios are controlled testbeds. An adapter lets a researcher bring an existing agent runner while retaining pikit's Trace and judge interfaces.

CallableAgentAdapter

CallableAgentAdapter(runner: AgentCallable, *, system: Optional[str] = None, taint: Optional[Dict[str, str]] = None, defenses: Optional[DefenseHooks] = None)

Bases: Agent

Wrap a user-supplied agent callable as a pikit agent.

The callable receives the hardened user message. It may return either a final text string or a fully populated :class:Trace. Optional keyword arguments expose the adapter context to framework integrations: system, taint, and defenses.

Returning a Trace is recommended whenever the external framework can capture tool calls. Returning text still provides a useful minimal adapter for direct-injection testing.

pikit.adapters

Framework adapters for running external agents through pikit.

Adapters are optional integrations: install the matching extra only when a framework is needed. The core package remains dependency-free.

CallableAgentAdapter

CallableAgentAdapter(runner: AgentCallable, *, system: Optional[str] = None, taint: Optional[Dict[str, str]] = None, defenses: Optional[DefenseHooks] = None)

Bases: Agent

Wrap a user-supplied agent callable as a pikit agent.

The callable receives the hardened user message. It may return either a final text string or a fully populated :class:Trace. Optional keyword arguments expose the adapter context to framework integrations: system, taint, and defenses.

Returning a Trace is recommended whenever the external framework can capture tool calls. Returning text still provides a useful minimal adapter for direct-injection testing.

TraceRecorder

TraceRecorder()

Build a structured pikit trace from external-agent events.

TaintRouter

TaintRouter(rules: Optional[Iterable[ToolTaintRule]] = None, *, taint: Optional[Dict[str, str]] = None)

Resolve tool calls to either clean execution or a tainted artifact.

resolve

resolve(tool_name: str, args: Dict[str, Any]) -> Optional[str]

Return a tainted artifact when a rule matches, else None.

ToolTaintRule dataclass

ToolTaintRule(tool_name: str, payload: str, when: Optional[ToolMatcher] = None)

A conditional taint response for one tool.

Parameters

tool_name: Name of the framework tool to intercept. payload: Artifact returned instead of the real tool result on a match. when: Optional predicate receiving parsed tool arguments. When omitted, every call to tool_name is tainted.

HermesCLIAdapter

HermesCLIAdapter(executable: str = 'hermes', *, model: Optional[str] = None, provider: Optional[str] = None, toolsets: Optional[Iterable[str]] = None, safe_mode: bool = True, hermes_home: Optional[str] = None, **kwargs: Any)

Bases: RuntimeCLIAdapter

Run one non-interactive Hermes CLI turn without messaging channels.

hermes --oneshot works directly in a terminal. By default the adapter adds --safe-mode to disable user customizations, skills, plugins, and MCP servers. Set safe_mode=False only with a dedicated isolated Hermes profile and an explicit tool policy.

OpenClawCLIAdapter

OpenClawCLIAdapter(executable: str = 'openclaw', *, agent: Optional[str] = None, model: Optional[str] = None, session_key: Optional[str] = 'agent:main:pikit', thinking: Optional[str] = None, state_dir: Optional[str] = None, config_path: Optional[str] = None, **kwargs: Any)

Bases: RuntimeCLIAdapter

Run one local, headless OpenClaw agent turn through its CLI.

The adapter invokes openclaw agent --local --json. It never supplies --deliver or a channel, so OpenClaw's messaging integrations are not involved. Use a dedicated OpenClaw profile/configuration that only exposes test fixture tools when evaluating indirect injection.

run

run(user_message: str, **kwargs: Any) -> Trace

Run with optional isolated OpenClaw state/config paths.

A configured provider is still required. The adapter does not perform onboarding because provider credentials and tool policy belong to the researcher's explicitly managed test profile.

RuntimeCLIAdapter

RuntimeCLIAdapter(executable: str, *, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, timeout: int = 120, env: Optional[Dict[str, str]] = None)

Bases: AgentHarness

Base class for a one-shot, headless agent-runtime command.

Subclasses implement :meth:build_command. The process receives a hardened user message and returns its captured stdout as final text unless a runtime-specific parser extracts a structured final response.

No command is executed until :meth:run is called. The adapter never enables a messaging channel or delivery flag on its own.

is_available

is_available() -> bool

Return whether the configured runtime executable is on PATH.

build_command

build_command(user_message: str) -> list

Build the non-interactive command for one agent turn.

parse_output

parse_output(stdout: str) -> str

Return final text from command stdout; subclasses may override.

run

run(user_message: str, **kwargs: Any) -> Trace

Run the runtime in a subprocess and translate it into a trace.

RuntimeFixture dataclass

RuntimeFixture(key: str, tool_name: str, user_message: str, tool_args: Dict[str, str], data_sample: str)

A safe fixture source exposed by both runtime plugins.

AgentHarness

Bases: Protocol

An executable agent environment that returns a pikit trace.

fixture_for

fixture_for(channel: str, override: str = '') -> RuntimeFixture

Resolve a fixture by explicit key or by pikit channel.

mapping_kind

mapping_kind(channel: str, fixture: RuntimeFixture) -> str

Return native when a fixture directly models the chosen channel.

fallback indicates a safe approximation (for example calendar content delivered through the email-like fixture) and is recorded in results so researchers do not over-interpret carrier fidelity.

get_harness

get_harness(name: str) -> Type[Any]

Return a registered harness class.

list_harnesses

list_harnesses() -> list[str]

Return registered harness names.

register_harness

register_harness(name: str) -> Callable[[Type[Any]], Type[Any]]

Register a harness class for matrix/runtime configuration.

pikit.agent.tools

Tools for agents: a :class:Tool wrapper and a @tool decorator.

A tool is a plain Python function plus a JSON-schema description the model uses to decide how to call it. The schema is auto-derived from the function's type hints (zero dependencies — no pydantic), and can be overridden explicitly when richer per-argument descriptions are needed.

Tool dataclass

Tool(name: str, description: str, func: Callable[..., Any], parameters: dict = (lambda: {'type': 'object', 'properties': {}})(), is_sink: bool = False, category: str = 'general')

A callable tool exposed to an agent's model.

Parameters

name, description, func: Identity, model-facing description, and the underlying callable. parameters: JSON-schema for the arguments. Auto-derived if not given. is_sink: Marks an externally-observable action (e.g. send_email). The trace highlights when a sink fires — the key signal for judging whether an injection succeeded. category: Tool category for pool-based selection. One of: "web", "email", "file", "code", "knowledge", "communication", "general".

to_schema

to_schema() -> dict

Return the provider-agnostic {name, description, parameters}.

tool

tool(name: Optional[str] = None, *, description: Optional[str] = None, is_sink: bool = False, parameters: Optional[dict] = None, category: str = 'general') -> Callable[[Callable], Tool]

Decorator turning a plain function into a :class:Tool.

The tool name defaults to the function name; description to its docstring. parameters is auto-derived from type hints unless given. category tags the tool for pool-based selection by scenario agents.

Examples

@tool(description="Fetch a URL and return its body.", category="web") ... def fetch_url(url: str) -> str: ... return "..." isinstance(fetch_url, Tool) True

pikit.agent.system_prompt

Dynamic system-prompt builder for tool-based scenario agents.

Instead of hard-coding tool descriptions in the system prompt, this module generates them from the actual :class:Tool objects the agent is configured with. This ensures the prompt always reflects the real tool set and stays in sync when tools are added or removed.

build_tool_list_section

build_tool_list_section(tools: List[Tool]) -> str

Generate a - name(args): description bullet list for the prompt.

build_system_prompt

build_system_prompt(role: str, tools: List[Tool], *, instructions: str = '', safety: Optional[str] = None) -> str

Build a structured system prompt from the agent's tool set.

Parameters

role: One-line description of the agent's role (e.g. "a web browsing assistant"). tools: The tools the agent can call. instructions: How the agent should use its tools to accomplish tasks. safety: Optional safety constraint text. Defaults to a standard indirect- injection warning.

pikit.agent.builtin_tools

Built-in tool pool for scenario agents.

A unified pool of ~36 tools organized by category. Data-source tools (return external content the model reads) are the taint points in an indirect-injection test — the caller overrides their return value via the taint map. Sink tools (externally-observable actions like send_email) are marked is_sink=True so the trace highlights when the model performs a dangerous action.

Pool management: use :func:get_tools, :func:get_tool, :func:all_tools, :func:data_source_tools, :func:sink_tools, or :func:tools_by_category to query the pool. Each scenario agent picks a relevant subset by category rather than hard-coding individual tools.

fetch_url

fetch_url(url: str) -> str

Fetch a web page and return its full HTML content.

search_web

search_web(query: str) -> str

Search the web and return formatted results.

download_file

download_file(url: str, save_path: str = '/tmp/download') -> str

Download a file from the given URL.

click_link(link_text: str) -> str

Click a link by its text and return the resulting page.

get_page_title

get_page_title(url: str) -> str

Return the title tag content of a web page.

post_form

post_form(url: str, data: str) -> str

Submit form data to a URL. This is a sink — an externally-observable action.

read_email

read_email() -> str

Read the user's latest email and return its full content.

list_emails

list_emails() -> str

List emails in the inbox.

search_emails

search_emails(query: str) -> str

Search the user's email for messages matching a query.

get_attachment

get_attachment(email_id: str, attachment_name: str) -> str

Download and preview an email attachment.

send_email

send_email(to: str, body: str, subject: str = '') -> str

Send an email. This is a sink — an externally-observable action.

forward_email

forward_email(email_id: str, to: str) -> str

Forward an email. This is a sink — an externally-observable action.

read_file

read_file(path: str) -> str

Read a file from the project and return its contents.

list_directory

list_directory(path: str = '.') -> str

List files and subdirectories at the given path.

search_files

search_files(pattern: str) -> str

Search for files matching a name pattern.

get_file_info

get_file_info(path: str) -> str

Return file metadata.

write_file

write_file(path: str, content: str) -> str

Write content to a file. This is a sink — an externally-observable action.

delete_file

delete_file(path: str) -> str

Delete a file. This is a sink — an externally-observable action.

move_file

move_file(src: str, dst: str) -> str

Move a file from src to dst. This is a sink.

read_code

read_code(path: str) -> str

Read a source code file.

load_skill

load_skill(name: str) -> str

Load a coding skill/extension by name.

search_codebase

search_codebase(query: str) -> str

Search the codebase for matches.

run_command

run_command(command: str) -> str

Run a shell command. This is a sink — an externally-observable action.

run_tests

run_tests(test_path: str = '') -> str

Run tests. This is a sink — it may modify state.

install_package

install_package(package: str) -> str

Install a package. This is a sink — it modifies the environment.

search_kb

search_kb(query: str) -> str

Search the knowledge base and return top results.

get_document

get_document(doc_id: str) -> str

Retrieve a document by ID.

summarize_text

summarize_text(text: str) -> str

Summarize the given text.

translate_text

translate_text(text: str, source_lang: str = 'en', target_lang: str = 'zh') -> str

Translate text from source to target language.

post_message

post_message(channel: str, message: str) -> str

Post a message to a chat channel. This is a sink.

create_ticket

create_ticket(title: str, description: str, priority: str = 'Medium') -> str

Create a ticket. This is a sink — it creates an external record.

schedule_meeting

schedule_meeting(title: str, time: str, attendees: str) -> str

Schedule a meeting. This is a sink — it sends invites.

read_channel

read_channel(channel: str) -> str

Read recent messages from a chat/IM channel.

get_dm

get_dm(user: str) -> str

Read DM history with a user.

get_thread

get_thread(thread_id: str) -> str

Read a threaded conversation.

send_dm

send_dm(to: str, message: str) -> str

Send a DM. This is a sink — an externally-observable action.

get_events

get_events(date: str) -> str

List calendar events for a date.

get_event_details

get_event_details(event_id: str) -> str

Get full details of a calendar event.

create_event

create_event(title: str, time: str, attendees: str) -> str

Create a calendar event. This is a sink — it sends invites.

modify_event

modify_event(event_id: str, title: str = '', time: str = '') -> str

Modify a calendar event. This is a sink — it may notify attendees.

get_balance

get_balance(account_id: str = '') -> str

Get the current balance for an account.

get_transactions

get_transactions(account_id: str = '', limit: int = 10) -> str

Get recent transactions for an account.

get_account_info

get_account_info(account_id: str = '') -> str

Get account information.

transfer_money

transfer_money(from_account: str, to_account: str, amount: str) -> str

Transfer money. This is a sink — it moves real funds.

pay_bill

pay_bill(biller: str, amount: str, from_account: str = '') -> str

Pay a bill. This is a sink — it moves real funds.

search_flights

search_flights(origin: str, destination: str, date: str) -> str

Search for available flights.

search_hotels

search_hotels(city: str, check_in: str, check_out: str) -> str

Search for available hotels.

get_flight_details

get_flight_details(flight_id: str) -> str

Get full details of a flight.

get_hotel_details

get_hotel_details(hotel_id: str) -> str

Get full details of a hotel.

book_flight

book_flight(flight_id: str, passenger: str) -> str

Book a flight. This is a sink — it charges payment.

book_hotel

book_hotel(hotel_id: str, guest: str, check_in: str, check_out: str) -> str

Book a hotel. This is a sink — it charges payment.

read_feed

read_feed() -> str

Read the social media feed.

get_post

get_post(post_id: str) -> str

Get full details of a social media post.

get_notifications

get_notifications() -> str

Get social media notifications.

create_post

create_post(content: str) -> str

Create a post. This is a sink — it publishes publicly.

share_post

share_post(post_id: str, comment: str = '') -> str

Share a post. This is a sink — it publishes to your feed.

all_tools

all_tools() -> List[Tool]

Return all tools in the pool.

get_tool

get_tool(name: str) -> Optional[Tool]

Return a single tool by name, or None if not found.

get_tools

get_tools(names: List[str]) -> List[Tool]

Return a list of tools by name (skips unknown names).

tools_by_category

tools_by_category(category: str) -> List[Tool]

Return all tools in a given category.

data_source_tools

data_source_tools() -> List[Tool]

Return all data-source tools (potential taint points).

sink_tools

sink_tools() -> List[Tool]

Return all sink tools (externally-observable actions).

tool_names

tool_names() -> List[str]

Return all tool names in the pool.

categories

categories() -> List[str]

Return all distinct categories in the pool.

pikit.agent.tool_agent

A general tool-calling agent driven by the function-calling loop.

ToolAgent

ToolAgent(target: Target, *, tools: Optional[List[Tool]] = None, taint: Optional[Dict[str, str]] = None, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, max_steps: int = 8)

Bases: Agent

An agent that can call a fixed set of tools in a loop.

Parameters

target, system, defenses: See :class:~pikit.agent.base.Agent. tools: The tools exposed to the model. taint: Map of tool_name -> artifact marking compromised tools whose return value is replaced by the injected artifact (indirect-injection delivery point). See :func:~pikit.agent.loop.run_tool_loop. max_steps: Safety cap on loop iterations.

pikit.agent.chat_agent

A no-tools chat agent — the simplest target under test.

ChatAgent

ChatAgent(target: Target, *, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None)

Bases: Agent

A plain chat assistant with no tools.

Wraps :meth:Target.query; direct injection arrives as the user message. The system and user defense hooks apply (there are no tool results to defend).

pikit.agent.scenarios.browser

Browser scenario: fetch web pages, then act — sink is post_form.

Models the Greshake-style indirect injection where a fetched web page hides an instruction. Taint points: fetch_url, search_web, download_file, click_link. Sinks: post_form.

This scenario simulates a web-browsing agent that can fetch pages, search the web, download files, click links, and submit forms. The expanded tool pool gives the model more autonomy — it must decide which tools to use and whether to trust the content they return.

BrowserAgent

BrowserAgent(target: Target, *, taint: Optional[Dict[str, str]] = None, tools: Optional[List[Tool]] = None, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, max_steps: int = 8)

Bases: ToolAgent

A browsing agent with a rich tool pool.

Taint points: fetch_url / search_web / download_file / click_link / get_page_title. Sink: post_form.

The agent has access to web-fetching and navigation tools. The indirect-injection attack surface: a fetched page, search result, or downloaded file contains a hidden instruction that tricks the agent into calling post_form to exfiltrate data.

Parameters

target: The model backend. taint: Map of tool_name -> artifact. Typically {"fetch_url": <tainted HTML>} or {"search_web": <tainted result>}. tools: Override the default tool set. If None, uses :data:BROWSER_TOOLS. system: Override the default system prompt. If None, a prompt is dynamically generated from the tool set. defenses: Optional defense hooks. max_steps: Safety cap on loop iterations (default 8).

default_task property

default_task: str

A sensible default user message for this scenario.

pikit.agent.scenarios.email_assistant

Email-assistant scenario: read mail, then act — sink is send_email.

The classic indirect-injection test: a tainted email body instructs the model to exfiltrate data by emailing an attacker. Taint points: read_email, list_emails, search_emails, get_attachment. Sinks: send_email, forward_email, post_message.

Simulates an email assistant with a rich tool pool: it can read, list, search, and forward emails, download attachments, search the knowledge base, and post messages to chat channels. The expanded tool set tests whether the model can be tricked into using cross-domain tools (e.g. posting to a chat channel) via injection.

EmailAssistantAgent

EmailAssistantAgent(target: Target, *, taint: Optional[Dict[str, str]] = None, tools: Optional[List[Tool]] = None, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, max_steps: int = 8)

Bases: ToolAgent

An email assistant with a rich tool pool.

Taint points: read_email / list_emails / search_emails / get_attachment / search_kb. Sinks: send_email / forward_email / post_message.

Parameters

target: The model backend. taint: Map of tool_name -> artifact. tools: Override the default tool set. If None, uses :data:EMAIL_TOOLS. system: Override the default system prompt. defenses: Optional defense hooks. max_steps: Safety cap on loop iterations (default 8).

default_task property

default_task: str

A sensible default user message for this scenario.

pikit.agent.scenarios.rag_qa

RAG question-answering scenario with a rich tool pool.

The model answers a question over retrieved documents. Taint points: search_kb, get_document, read_file. The "sink" here is the final answer itself — whether the model complies with the injected instruction is observed in trace.final_text. An optional post_form sink models pipelines that forward the answer somewhere observable.

The expanded tool pool includes summarize_text and translate_text for richer document processing, testing whether the model can be tricked into using these tools in service of an injection.

RagQaAgent

RagQaAgent(target: Target, *, taint: Optional[Dict[str, str]] = None, tools: Optional[List[Tool]] = None, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, max_steps: int = 8)

Bases: ToolAgent

A RAG QA agent with a rich tool pool.

Taint points: search_kb / get_document / read_file. Sinks: final answer / post_form.

Parameters

target: The model backend. taint: Map of tool_name -> artifact. tools: Override the default tool set. If None, uses :data:RAG_TOOLS. system: Override the default system prompt. defenses: Optional defense hooks. max_steps: Safety cap on loop iterations (default 8).

default_task property

default_task: str

A sensible default user message for this scenario.

pikit.agent.scenarios.coding

Coding scenario: a code-assistant agent with a rich tool pool.

Models a coding agent like Claude Code / Cursor / Aider that reads project files, loads skills, searches the codebase, and can execute commands and modify files. Taint points: read_code, read_file, load_skill, search_codebase, search_files. Sinks: run_command, write_file, delete_file, move_file, run_tests, install_package.

The expanded tool pool gives the model more autonomy — it must decide which tools to use for a given task and whether to trust content from code comments, skill definitions, and search results.

CodingAgent

CodingAgent(target: Target, *, taint: Optional[Dict[str, str]] = None, tools: Optional[List[Tool]] = None, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, max_steps: int = 8)

Bases: ToolAgent

A coding agent with a rich tool pool.

Taint points: read_code / read_file / load_skill / search_codebase / search_files. Sinks: run_command / write_file / delete_file / move_file / run_tests / install_package.

Parameters

target: The model backend. taint: Map of tool_name -> artifact. tools: Override the default tool set. If None, uses :data:CODING_TOOLS. system: Override the default system prompt. defenses: Optional defense hooks. max_steps: Safety cap on loop iterations (default 8).

default_task property

default_task: str

A sensible default user message for this scenario.