Agent¶
pikit.agent ¶
Agent environments for testing prompt-injection attacks.
An agent wraps a :class:~pikit.targets.Target and exposes a run() that
returns a human-readable :class:Trace. Three kinds ship:
chat— a plain assistant (no tools); direct injection via the user message.tool— a general tool-calling agent; indirect injection via a compromised tool's return value (thetaintmap).- scenario agents —
email/rag/browser— preconfigured with a realistic toolset and an observable sink.
Like the rest of pikit, agents register under short keys; use
get_agent(key) to fetch the class, then construct it.
Agent ¶
Bases: ABC
Base class for agents under test.
Parameters¶
target:
The model backend (:class:~pikit.targets.Target).
system:
Optional system prompt.
defenses:
Optional :class:DefenseHooks applied at the three insertion points.
run
abstractmethod
¶
Run the agent on user_message and return the :class:Trace.
Trace
dataclass
¶
TraceStep
dataclass
¶
TraceStep(kind: str, text: str = '', tool_name: Optional[str] = None, args: Optional[dict] = None, content: Optional[str] = None, tainted: bool = False, is_sink: bool = False, decision: Optional[str] = None)
One step in an agent run.
DefenseHooks
dataclass
¶
DefenseHooks(system: Optional[Defense] = None, tool_result: Optional[Defense] = None, user: Optional[Defense] = None)
Optional defenses applied at three points of the agent loop.
Parameters¶
system: Hardens the system prompt (defends against the model being talked out of its instructions). tool_result: Hardens untrusted tool output before it re-enters the model — the key defense position for indirect injection. user: Hardens the incoming user message (defends against direct injection).
CallableAgentAdapter ¶
CallableAgentAdapter(runner: AgentCallable, *, system: Optional[str] = None, taint: Optional[Dict[str, str]] = None, defenses: Optional[DefenseHooks] = None)
Bases: Agent
Wrap a user-supplied agent callable as a pikit agent.
The callable receives the hardened user message. It may return either a
final text string or a fully populated :class:Trace. Optional keyword
arguments expose the adapter context to framework integrations:
system, taint, and defenses.
Returning a Trace is recommended whenever the external framework can
capture tool calls. Returning text still provides a useful minimal
adapter for direct-injection testing.
Tool
dataclass
¶
Tool(name: str, description: str, func: Callable[..., Any], parameters: dict = (lambda: {'type': 'object', 'properties': {}})(), is_sink: bool = False, category: str = 'general')
A callable tool exposed to an agent's model.
Parameters¶
name, description, func:
Identity, model-facing description, and the underlying callable.
parameters:
JSON-schema for the arguments. Auto-derived if not given.
is_sink:
Marks an externally-observable action (e.g. send_email). The
trace highlights when a sink fires — the key signal for judging
whether an injection succeeded.
category:
Tool category for pool-based selection. One of: "web",
"email", "file", "code", "knowledge",
"communication", "general".
tool ¶
tool(name: Optional[str] = None, *, description: Optional[str] = None, is_sink: bool = False, parameters: Optional[dict] = None, category: str = 'general') -> Callable[[Callable], Tool]
Decorator turning a plain function into a :class:Tool.
The tool name defaults to the function name; description to its
docstring. parameters is auto-derived from type hints unless given.
category tags the tool for pool-based selection by scenario agents.
Examples¶
@tool(description="Fetch a URL and return its body.", category="web") ... def fetch_url(url: str) -> str: ... return "..." isinstance(fetch_url, Tool) True
build_system_prompt ¶
build_system_prompt(role: str, tools: List[Tool], *, instructions: str = '', safety: Optional[str] = None) -> str
Build a structured system prompt from the agent's tool set.
Parameters¶
role: One-line description of the agent's role (e.g. "a web browsing assistant"). tools: The tools the agent can call. instructions: How the agent should use its tools to accomplish tasks. safety: Optional safety constraint text. Defaults to a standard indirect- injection warning.
get_tool ¶
Return a single tool by name, or None if not found.
get_tools ¶
Return a list of tools by name (skips unknown names).
tools_by_category ¶
Return all tools in a given category.
data_source_tools ¶
Return all data-source tools (potential taint points).
pikit.agent.base ¶
Agent base class and the execution :class:Trace.
run() returns a :class:Trace rather than a bare string. The trace is
the artifact a human reads to judge — manually — whether an injection
succeeded: it shows every model turn, tool call, and tool result, and
highlights when a sink fired or a step carried tainted data. The
library deliberately renders no verdict (no evaluator/scoring); it makes the
signals easy to see and offers structured accessors so you can write your
own one-line assertion.
TraceStep
dataclass
¶
TraceStep(kind: str, text: str = '', tool_name: Optional[str] = None, args: Optional[dict] = None, content: Optional[str] = None, tainted: bool = False, is_sink: bool = False, decision: Optional[str] = None)
One step in an agent run.
Trace
dataclass
¶
Agent ¶
Bases: ABC
Base class for agents under test.
Parameters¶
target:
The model backend (:class:~pikit.targets.Target).
system:
Optional system prompt.
defenses:
Optional :class:DefenseHooks applied at the three insertion points.
run
abstractmethod
¶
Run the agent on user_message and return the :class:Trace.
pikit.agent.loop ¶
The provider-agnostic function-calling loop shared by tool agents.
run_tool_loop ¶
run_tool_loop(target: Target, user_message: str, tools: List[Tool], *, system: Optional[str] = None, hooks: Optional[DefenseHooks] = None, taint: Optional[Dict[str, str]] = None, max_steps: int = 8, **target_kwargs) -> Trace
Drive target.chat over tools until it stops calling tools.
Parameters¶
taint:
Map of tool_name -> artifact. When the model calls a tainted
tool, the loop returns the artifact as that tool's result instead of
invoking the real function (the indirect-injection delivery point,
and it avoids real side effects during a test).
**target_kwargs:
Forwarded to target.chat() on every call (e.g.
temperature=0.7).
pikit.agent.hooks ¶
Defense insertion points for agents.
The same prevention-style :class:~pikit.base.Defense objects used for
direct injection can be slotted into three points of an agent's data flow.
The most valuable for indirect injection is tool_result — the layer
through which an attacker's tainted artifact re-enters the model.
DefenseHooks
dataclass
¶
DefenseHooks(system: Optional[Defense] = None, tool_result: Optional[Defense] = None, user: Optional[Defense] = None)
Optional defenses applied at three points of the agent loop.
Parameters¶
system: Hardens the system prompt (defends against the model being talked out of its instructions). tool_result: Hardens untrusted tool output before it re-enters the model — the key defense position for indirect injection. user: Hardens the incoming user message (defends against direct injection).
pikit.agent.adapters ¶
Adapters for exercising externally implemented agents with pikit.
The built-in scenarios are controlled testbeds. An adapter lets a researcher
bring an existing agent runner while retaining pikit's Trace and judge
interfaces.
CallableAgentAdapter ¶
CallableAgentAdapter(runner: AgentCallable, *, system: Optional[str] = None, taint: Optional[Dict[str, str]] = None, defenses: Optional[DefenseHooks] = None)
Bases: Agent
Wrap a user-supplied agent callable as a pikit agent.
The callable receives the hardened user message. It may return either a
final text string or a fully populated :class:Trace. Optional keyword
arguments expose the adapter context to framework integrations:
system, taint, and defenses.
Returning a Trace is recommended whenever the external framework can
capture tool calls. Returning text still provides a useful minimal
adapter for direct-injection testing.
pikit.adapters ¶
Framework adapters for running external agents through pikit.
Adapters are optional integrations: install the matching extra only when a framework is needed. The core package remains dependency-free.
CallableAgentAdapter ¶
CallableAgentAdapter(runner: AgentCallable, *, system: Optional[str] = None, taint: Optional[Dict[str, str]] = None, defenses: Optional[DefenseHooks] = None)
Bases: Agent
Wrap a user-supplied agent callable as a pikit agent.
The callable receives the hardened user message. It may return either a
final text string or a fully populated :class:Trace. Optional keyword
arguments expose the adapter context to framework integrations:
system, taint, and defenses.
Returning a Trace is recommended whenever the external framework can
capture tool calls. Returning text still provides a useful minimal
adapter for direct-injection testing.
TaintRouter ¶
TaintRouter(rules: Optional[Iterable[ToolTaintRule]] = None, *, taint: Optional[Dict[str, str]] = None)
Resolve tool calls to either clean execution or a tainted artifact.
resolve ¶
Return a tainted artifact when a rule matches, else None.
ToolTaintRule
dataclass
¶
A conditional taint response for one tool.
Parameters¶
tool_name:
Name of the framework tool to intercept.
payload:
Artifact returned instead of the real tool result on a match.
when:
Optional predicate receiving parsed tool arguments. When omitted,
every call to tool_name is tainted.
HermesCLIAdapter ¶
HermesCLIAdapter(executable: str = 'hermes', *, model: Optional[str] = None, provider: Optional[str] = None, toolsets: Optional[Iterable[str]] = None, safe_mode: bool = True, hermes_home: Optional[str] = None, **kwargs: Any)
Bases: RuntimeCLIAdapter
Run one non-interactive Hermes CLI turn without messaging channels.
hermes --oneshot works directly in a terminal. By default the
adapter adds --safe-mode to disable user customizations, skills,
plugins, and MCP servers. Set safe_mode=False only with a dedicated
isolated Hermes profile and an explicit tool policy.
OpenClawCLIAdapter ¶
OpenClawCLIAdapter(executable: str = 'openclaw', *, agent: Optional[str] = None, model: Optional[str] = None, session_key: Optional[str] = 'agent:main:pikit', thinking: Optional[str] = None, state_dir: Optional[str] = None, config_path: Optional[str] = None, **kwargs: Any)
Bases: RuntimeCLIAdapter
Run one local, headless OpenClaw agent turn through its CLI.
The adapter invokes openclaw agent --local --json. It never supplies
--deliver or a channel, so OpenClaw's messaging integrations are not
involved. Use a dedicated OpenClaw profile/configuration that only exposes
test fixture tools when evaluating indirect injection.
run ¶
Run with optional isolated OpenClaw state/config paths.
A configured provider is still required. The adapter does not perform onboarding because provider credentials and tool policy belong to the researcher's explicitly managed test profile.
RuntimeCLIAdapter ¶
RuntimeCLIAdapter(executable: str, *, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, timeout: int = 120, env: Optional[Dict[str, str]] = None)
Bases: AgentHarness
Base class for a one-shot, headless agent-runtime command.
Subclasses implement :meth:build_command. The process receives a
hardened user message and returns its captured stdout as final text unless
a runtime-specific parser extracts a structured final response.
No command is executed until :meth:run is called. The adapter never
enables a messaging channel or delivery flag on its own.
RuntimeFixture
dataclass
¶
RuntimeFixture(key: str, tool_name: str, user_message: str, tool_args: Dict[str, str], data_sample: str)
A safe fixture source exposed by both runtime plugins.
AgentHarness ¶
Bases: Protocol
An executable agent environment that returns a pikit trace.
fixture_for ¶
Resolve a fixture by explicit key or by pikit channel.
mapping_kind ¶
Return native when a fixture directly models the chosen channel.
fallback indicates a safe approximation (for example calendar content
delivered through the email-like fixture) and is recorded in results so
researchers do not over-interpret carrier fidelity.
register_harness ¶
Register a harness class for matrix/runtime configuration.
pikit.agent.tools ¶
Tools for agents: a :class:Tool wrapper and a @tool decorator.
A tool is a plain Python function plus a JSON-schema description the model uses to decide how to call it. The schema is auto-derived from the function's type hints (zero dependencies — no pydantic), and can be overridden explicitly when richer per-argument descriptions are needed.
Tool
dataclass
¶
Tool(name: str, description: str, func: Callable[..., Any], parameters: dict = (lambda: {'type': 'object', 'properties': {}})(), is_sink: bool = False, category: str = 'general')
A callable tool exposed to an agent's model.
Parameters¶
name, description, func:
Identity, model-facing description, and the underlying callable.
parameters:
JSON-schema for the arguments. Auto-derived if not given.
is_sink:
Marks an externally-observable action (e.g. send_email). The
trace highlights when a sink fires — the key signal for judging
whether an injection succeeded.
category:
Tool category for pool-based selection. One of: "web",
"email", "file", "code", "knowledge",
"communication", "general".
tool ¶
tool(name: Optional[str] = None, *, description: Optional[str] = None, is_sink: bool = False, parameters: Optional[dict] = None, category: str = 'general') -> Callable[[Callable], Tool]
Decorator turning a plain function into a :class:Tool.
The tool name defaults to the function name; description to its
docstring. parameters is auto-derived from type hints unless given.
category tags the tool for pool-based selection by scenario agents.
Examples¶
@tool(description="Fetch a URL and return its body.", category="web") ... def fetch_url(url: str) -> str: ... return "..." isinstance(fetch_url, Tool) True
pikit.agent.system_prompt ¶
Dynamic system-prompt builder for tool-based scenario agents.
Instead of hard-coding tool descriptions in the system prompt, this module
generates them from the actual :class:Tool objects the agent is configured
with. This ensures the prompt always reflects the real tool set and stays
in sync when tools are added or removed.
build_tool_list_section ¶
Generate a - name(args): description bullet list for the prompt.
build_system_prompt ¶
build_system_prompt(role: str, tools: List[Tool], *, instructions: str = '', safety: Optional[str] = None) -> str
Build a structured system prompt from the agent's tool set.
Parameters¶
role: One-line description of the agent's role (e.g. "a web browsing assistant"). tools: The tools the agent can call. instructions: How the agent should use its tools to accomplish tasks. safety: Optional safety constraint text. Defaults to a standard indirect- injection warning.
pikit.agent.builtin_tools ¶
Built-in tool pool for scenario agents.
A unified pool of ~36 tools organized by category. Data-source tools
(return external content the model reads) are the taint points in an
indirect-injection test — the caller overrides their return value via the
taint map. Sink tools (externally-observable actions like
send_email) are marked is_sink=True so the trace highlights when
the model performs a dangerous action.
Pool management: use :func:get_tools, :func:get_tool,
:func:all_tools, :func:data_source_tools, :func:sink_tools, or
:func:tools_by_category to query the pool. Each scenario agent picks a
relevant subset by category rather than hard-coding individual tools.
download_file ¶
Download a file from the given URL.
click_link ¶
Click a link by its text and return the resulting page.
post_form ¶
Submit form data to a URL. This is a sink — an externally-observable action.
search_emails ¶
Search the user's email for messages matching a query.
get_attachment ¶
Download and preview an email attachment.
send_email ¶
Send an email. This is a sink — an externally-observable action.
forward_email ¶
Forward an email. This is a sink — an externally-observable action.
list_directory ¶
List files and subdirectories at the given path.
write_file ¶
Write content to a file. This is a sink — an externally-observable action.
delete_file ¶
Delete a file. This is a sink — an externally-observable action.
run_command ¶
Run a shell command. This is a sink — an externally-observable action.
install_package ¶
Install a package. This is a sink — it modifies the environment.
translate_text ¶
Translate text from source to target language.
post_message ¶
Post a message to a chat channel. This is a sink.
create_ticket ¶
Create a ticket. This is a sink — it creates an external record.
schedule_meeting ¶
Schedule a meeting. This is a sink — it sends invites.
send_dm ¶
Send a DM. This is a sink — an externally-observable action.
create_event ¶
Create a calendar event. This is a sink — it sends invites.
modify_event ¶
Modify a calendar event. This is a sink — it may notify attendees.
get_transactions ¶
Get recent transactions for an account.
transfer_money ¶
Transfer money. This is a sink — it moves real funds.
pay_bill ¶
Pay a bill. This is a sink — it moves real funds.
search_flights ¶
Search for available flights.
search_hotels ¶
Search for available hotels.
book_flight ¶
Book a flight. This is a sink — it charges payment.
book_hotel ¶
Book a hotel. This is a sink — it charges payment.
create_post ¶
Create a post. This is a sink — it publishes publicly.
share_post ¶
Share a post. This is a sink — it publishes to your feed.
get_tool ¶
Return a single tool by name, or None if not found.
get_tools ¶
Return a list of tools by name (skips unknown names).
tools_by_category ¶
Return all tools in a given category.
data_source_tools ¶
Return all data-source tools (potential taint points).
pikit.agent.tool_agent ¶
A general tool-calling agent driven by the function-calling loop.
ToolAgent ¶
ToolAgent(target: Target, *, tools: Optional[List[Tool]] = None, taint: Optional[Dict[str, str]] = None, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, max_steps: int = 8)
Bases: Agent
An agent that can call a fixed set of tools in a loop.
Parameters¶
target, system, defenses:
See :class:~pikit.agent.base.Agent.
tools:
The tools exposed to the model.
taint:
Map of tool_name -> artifact marking compromised tools whose
return value is replaced by the injected artifact (indirect-injection
delivery point). See :func:~pikit.agent.loop.run_tool_loop.
max_steps:
Safety cap on loop iterations.
pikit.agent.chat_agent ¶
pikit.agent.scenarios.browser ¶
Browser scenario: fetch web pages, then act — sink is post_form.
Models the Greshake-style indirect injection where a fetched web page hides
an instruction. Taint points: fetch_url, search_web,
download_file, click_link. Sinks: post_form.
This scenario simulates a web-browsing agent that can fetch pages, search the web, download files, click links, and submit forms. The expanded tool pool gives the model more autonomy — it must decide which tools to use and whether to trust the content they return.
BrowserAgent ¶
BrowserAgent(target: Target, *, taint: Optional[Dict[str, str]] = None, tools: Optional[List[Tool]] = None, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, max_steps: int = 8)
Bases: ToolAgent
A browsing agent with a rich tool pool.
Taint points: fetch_url / search_web / download_file /
click_link / get_page_title.
Sink: post_form.
The agent has access to web-fetching and navigation tools. The
indirect-injection attack surface: a fetched page, search result, or
downloaded file contains a hidden instruction that tricks the agent
into calling post_form to exfiltrate data.
Parameters¶
target:
The model backend.
taint:
Map of tool_name -> artifact. Typically
{"fetch_url": <tainted HTML>} or {"search_web": <tainted result>}.
tools:
Override the default tool set. If None, uses :data:BROWSER_TOOLS.
system:
Override the default system prompt. If None, a prompt is
dynamically generated from the tool set.
defenses:
Optional defense hooks.
max_steps:
Safety cap on loop iterations (default 8).
pikit.agent.scenarios.email_assistant ¶
Email-assistant scenario: read mail, then act — sink is send_email.
The classic indirect-injection test: a tainted email body instructs the
model to exfiltrate data by emailing an attacker. Taint points:
read_email, list_emails, search_emails, get_attachment.
Sinks: send_email, forward_email, post_message.
Simulates an email assistant with a rich tool pool: it can read, list, search, and forward emails, download attachments, search the knowledge base, and post messages to chat channels. The expanded tool set tests whether the model can be tricked into using cross-domain tools (e.g. posting to a chat channel) via injection.
EmailAssistantAgent ¶
EmailAssistantAgent(target: Target, *, taint: Optional[Dict[str, str]] = None, tools: Optional[List[Tool]] = None, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, max_steps: int = 8)
Bases: ToolAgent
An email assistant with a rich tool pool.
Taint points: read_email / list_emails / search_emails /
get_attachment / search_kb.
Sinks: send_email / forward_email / post_message.
Parameters¶
target:
The model backend.
taint:
Map of tool_name -> artifact.
tools:
Override the default tool set. If None, uses :data:EMAIL_TOOLS.
system:
Override the default system prompt.
defenses:
Optional defense hooks.
max_steps:
Safety cap on loop iterations (default 8).
pikit.agent.scenarios.rag_qa ¶
RAG question-answering scenario with a rich tool pool.
The model answers a question over retrieved documents. Taint points:
search_kb, get_document, read_file. The "sink" here is the
final answer itself — whether the model complies with the injected
instruction is observed in trace.final_text. An optional post_form
sink models pipelines that forward the answer somewhere observable.
The expanded tool pool includes summarize_text and translate_text
for richer document processing, testing whether the model can be tricked
into using these tools in service of an injection.
RagQaAgent ¶
RagQaAgent(target: Target, *, taint: Optional[Dict[str, str]] = None, tools: Optional[List[Tool]] = None, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, max_steps: int = 8)
Bases: ToolAgent
A RAG QA agent with a rich tool pool.
Taint points: search_kb / get_document / read_file.
Sinks: final answer / post_form.
Parameters¶
target:
The model backend.
taint:
Map of tool_name -> artifact.
tools:
Override the default tool set. If None, uses :data:RAG_TOOLS.
system:
Override the default system prompt.
defenses:
Optional defense hooks.
max_steps:
Safety cap on loop iterations (default 8).
pikit.agent.scenarios.coding ¶
Coding scenario: a code-assistant agent with a rich tool pool.
Models a coding agent like Claude Code / Cursor / Aider that reads project
files, loads skills, searches the codebase, and can execute commands and
modify files. Taint points: read_code, read_file, load_skill,
search_codebase, search_files. Sinks: run_command,
write_file, delete_file, move_file, run_tests,
install_package.
The expanded tool pool gives the model more autonomy — it must decide which tools to use for a given task and whether to trust content from code comments, skill definitions, and search results.
CodingAgent ¶
CodingAgent(target: Target, *, taint: Optional[Dict[str, str]] = None, tools: Optional[List[Tool]] = None, system: Optional[str] = None, defenses: Optional[DefenseHooks] = None, max_steps: int = 8)
Bases: ToolAgent
A coding agent with a rich tool pool.
Taint points: read_code / read_file / load_skill /
search_codebase / search_files.
Sinks: run_command / write_file / delete_file /
move_file / run_tests / install_package.
Parameters¶
target:
The model backend.
taint:
Map of tool_name -> artifact.
tools:
Override the default tool set. If None, uses :data:CODING_TOOLS.
system:
Override the default system prompt.
defenses:
Optional defense hooks.
max_steps:
Safety cap on loop iterations (default 8).