09. Lifecycle Hooks & Enterprise Observability
⭐️ Love Smoke Monkey Harness? Please support the project with a star on GitHub! 🌐 Interactive Simulator: Test the agent loop and stream live at smoke-monkey-harness.vercel.app
Enterprise applications require rigorous security boundaries, audit logging, telemetry, and resilient error recovery.
Smoke Monkey provides Lifecycle Hooks to inspect and mutate every model call and tool execution, paired with a Structured Error Hierarchy (AgentErrorInfo) that eliminates opaque string failures.
Lifecycle Hooks
Hooks allow embedding applications to observe and guard the two core activities of an agent: calling language models and executing tools.
flowchart LR
TurnStart([Agent Turn]) --> BMC["beforeModelCall\n(Prompt inspection / Redaction)"]
BMC --> LLMCall[Model Provider Execution]
LLMCall --> AMC["afterModelCall\n(Token usage & cost telemetry)"]
AMC --> BTC["beforeToolCall\n(Path allowlisting & Security guard)"]
BTC --> ToolExec[Tool Execution]
ToolExec --> ATC["afterToolCall\n(Audit logs & Performance timing)"]
ATC --> TurnEnd([Next Turn / Complete])
Hook Definition Example
import { createAgent } from '@smoke-monkey/harness';
const agent = createAgent({
workspacePath: process.cwd(),
hooks: {
// 1. Guard tool execution before permissions or calls run
async beforeToolCall({ toolName, input, userId }) {
// Path Allowlist Security Policy
if (toolName === 'write_file' && input.path.includes('.env')) {
return { block: true, reason: 'Modifying environment secrets is strictly forbidden.' };
}
// Dynamic Argument Redaction (strip credentials)
if (input.content) {
return { input: { ...input, content: redactPii(input.content) } };
}
},
// 2. Audit log after tool execution settles
async afterToolCall({ toolName, result, durationMs, blocked, error }) {
auditLogger.record({
tool: toolName,
durationMs,
success: !blocked && !error,
error: error?.message,
});
},
// 3. Inspect or block outgoing prompts
async beforeModelCall({ messages, tools, step }) {
metrics.gauge('agent.step', step);
metrics.gauge('agent.prompt_messages', messages.length);
},
// 4. Track LLM token usage and dollar cost
async afterModelCall({ usage, durationMs, error }) {
costTracker.record({
promptTokens: usage.promptTokens,
completionTokens: usage.completionTokens,
durationMs,
failed: Boolean(error),
});
},
},
});
Hook Semantics & Execution Rules
| Hook | Return Options | Effect | Error Behavior |
|---|---|---|---|
beforeModelCall |
{ messages?, tools? }{ block: true, reason } |
Mutates outgoing messages or aborts the entire run. | Fail-Closed: If the hook throws, the run immediately aborts with hook_blocked. |
beforeToolCall |
{ input }{ block: true, reason } |
Replaces tool arguments or cancels tool execution. | Fail-Closed: If the hook throws, the tool is blocked and the error is returned to the LLM. |
afterModelCall |
void | Observability and cost telemetry. | Fail-Open: If the hook throws, it is logged and ignored. |
afterToolCall |
void | Observability and audit logging. | Fail-Open: If the hook throws, it is logged and ignored. |
Placement Insight:
beforeToolCallexecutes before the interactive permission prompt. If a security policy hook rejects a tool call, the user is never unnecessarily prompted.
Structured Error Hierarchy (AgentErrorInfo)
Instead of throwing generic string errors (where a provider rate limit, a permission denial, and an out-of-bounds error look identical), Smoke Monkey classifies every failure into a typed AgentErrorInfo structure:
export interface AgentErrorInfo {
code: string; // Stable machine code (e.g. 'provider_rate_limited')
layer: AgentErrorLayer; // 'provider' | 'tool' | 'run' | 'hook' | 'permission' | 'transport'
severity: AgentErrorSeverity; // 'info' | 'warning' | 'error' | 'fatal'
message: string; // User-facing message (never a raw stack trace)
retryable: boolean; // True if immediate or delayed retry could succeed
hint?: string; // Actionable advice for the user or UI
details?: unknown; // Optional debug details
}
Error Layers & Severity Codes
| Layer | Common Codes | Severity | Meaning |
|---|---|---|---|
provider |
provider_rate_limitedprovider_authprovider_timeout |
warning / fatal |
Upstream LLM issues. Rate limits are flagged retryable: true. Bad API keys are fatal. |
tool |
tool_failedtool_blockedtool_not_found |
warning |
Scoped to a single tool turn. The loop continues and allows the agent to self-heal. |
run |
run_repeated_errorrun_hard_stoprun_empty_responses |
fatal |
A loop guard halted the run to prevent infinite costs or thrashing. |
permission |
permission_denied |
warning |
Human rejected a mutating command. Scoped to the tool call. |
hook |
hook_blockedhook_failed |
fatal / warning |
A security hook stopped the execution. |
Non-Terminal Warnings vs Terminal Failures
The harness distinguishes between temporary snags and terminal failures:
1. run.warning (Non-Terminal)
Fires when an operation encounters a transient issue (such as a 429 rate limit or retryable network drop). The agent is still actively retrying, allowing the UI to show an informational banner:
agent.on('run.warning', (e) => {
const { code, message, hint } = e.data.errorInfo;
toast.warning(`${message} (${hint})`);
});
2. run.failed (Terminal)
Fires only when retries have exhausted or an unrecoverable failure has occurred:
agent.on('run.failed', (e) => {
const { code, retryable, hint } = e.data.errorInfo;
console.error(`Task aborted: ${code}. ${hint}`);
if (retryable) {
showRetryButton();
}
});
Next Steps
- Render structured errors inside the React UI: 10. React UI Components