Format Reference
Full schema and field descriptions for each file in the export ZIP.
executions.jsonl
Each line represents a single execution: one agent interaction with the target system, including the full conversation and evaluation results.
Schema
{
id: string | null;
target_id: string;
evaluation_id: string | null;
evaluation_timestamp: string | null; // ISO 8601
task: {
id: string;
name: string;
description: string;
criteria: string[] | null;
importance: number | null; // 1–5 scale
snapshot_time: string; // ISO 8601
supporting_document_relevances: RelevanceEntry[] | null;
documents: {
title: string;
url?: string; // omitted for file uploads
}[];
} | null;
principles: {
id: string;
name: string;
description: string;
importance: number | null; // 1–5 scale
tags: string[] | null;
snapshot_time?: string; // ISO 8601
}[] | null;
persona: {
id: string;
description: string;
name: string | null;
age: number | null;
nationality: string | null;
snapshot_time?: string; // ISO 8601
} | null;
report: {
turns: number;
plot: object | null; // narrative scenario plot, when applicable
outcomes: Outcome[]; // rolled-up verdict per scored dimension, per turn
classifications: Classification[]; // detailed reasoning behind each outcome
} | null;
conversation: {
role: string; // "user" | "assistant"
content: string;
timestamp: string; // ISO 8601
id: string;
}[] | null;
}
The report carries two parallel arrays, both keyed by a type discriminator and a turn index:
outcomes— the machine-readable verdict for each dimension that was scored, on each turn. This is what aggregate metrics are computed from.classifications— the human-readable reasoning behind each outcome: explanations, supporting evidence, extracted claims, detected contradictions, and per-principle assessments.
A dimension appears in these arrays only on the turns where it was scored. A dimension that does not apply to the execution's attack focus is simply absent.
// One verdict per (dimension, turn). Discriminated on `type`.
type Outcome =
| { type: "completion"; turn: number; decision: boolean | null }
| { type: "validity"; turn: number; decision: boolean | null }
| { type: "coherence"; turn: number; decision: boolean | null }
| { type: "instruction_following"; turn: number; decision: boolean | null }
| { type: "scope_adherence"; turn: number; decision: boolean | null }
| { type: "factuality"; turn: number; decision: boolean | null; severity: number }
| { type: "scopeguard"; turn: number; total: number; blocked: number }
| { type: "compliance"; turn: number; evaluated: number; compliant: number; severity: number };
// Detailed reasoning per (dimension, turn). Discriminated on `type`.
type Classification =
| { type: "completion"; turn: number; decision: boolean | null; explanation: string | null;
evidences: SupportingEvidence[] | null }
| { type: "validity"; turn: number; decision: boolean | null; explanation: string | null }
| { type: "coherence"; turn: number; decision: boolean | null; explanation: string | null;
contradictions: Contradiction[] | null }
| { type: "instruction_following"; turn: number; decision: boolean | null; explanation: string | null;
drift_reasonings: DriftReasoning[] | null }
| { type: "scope_adherence"; turn: number; decision: boolean | null; explanation: string | null }
| { type: "factuality"; turn: number; decision: boolean | null; severity: number;
explanation: string | null;
evidences: SupportingEvidence[] | null;
claims: ClaimFactualityClassification[] | null;
user_impact_score: number | null; user_impact_explanation: string | null;
reputational_impact_score: number | null; reputational_impact_explanation: string | null }
| { type: "scopeguard"; turn: number; scope_class: string; evidences: string[];
model: string; usage: ScopeGuardUsage; time_taken: number }
| { type: "compliance"; turn: number; assessments: PrincipleAssessment[] };
type SupportingEvidence = {
document_title: string; // verbatim title of the source KB document
passages: string[]; // verbatim excerpts that support the verdict
};
type ClaimFactualityClassification = {
id: string;
claim: string;
claim_class: string;
claim_type: string | null;
is_factual: boolean;
factuality_details: string;
};
type Contradiction = {
unfactual_claim_content: string; // verbatim text of the contradicted claim
unfactual_claim_id: string; // id matching the related claim
reasoning: string;
conversation_ids: string[]; // executions containing the contradicted evidence
};
type DriftReasoning = {
label: "Refusal" | "Misinterpretation";
reasoning: string;
};
type PrincipleAssessment = {
principle_id: string;
decision: boolean; // true = compliant with this principle
explanation: string;
severity: number;
user_impact_score: number | null; user_impact_explanation: string | null;
reputational_impact_score: number | null; reputational_impact_explanation: string | null;
magnitude_impact_score: number | null; magnitude_explanation: string | null;
};
type ScopeGuardUsage = {
prompt_tokens: number;
completion_tokens: number;
total_tokens: number;
};
Fields
| Field | Description |
|---|---|
id | Unique identifier for this execution. |
target_id | ID of the target this execution belongs to. |
evaluation_id | ID of the evaluation run this execution belongs to. null if the execution is not associated with an evaluation. |
evaluation_timestamp | ISO 8601 timestamp of when the evaluation run was triggered. null if not associated with an evaluation. |
task | Snapshot of the task definition at execution time. task.criteria lists the evaluation criteria used to assess completion. task.importance ranks the task on a 1–5 scale. task.documents lists Knowledge Base documents attached to the task; url is omitted for file uploads. null if no task was assigned. |
principles | Snapshot of the principles active during the execution. principles[].importance ranks each principle on a 1–5 scale. null if no principles were active. |
persona | Snapshot of the persona used to model the synthetic user. null if no persona was assigned. |
report | Evaluation verdicts for this execution. null if the execution has not been scored yet. See the report fields table below. |
conversation | Full turn-by-turn exchange. role is either "user" (the Spectral agent) or "assistant" (the target system). |
task, principles, and persona each include a snapshot_time (ISO 8601) recording when the definition was captured. These configurations can change over time; the snapshot preserves the exact state used during this execution.
report fields
| Field | Description |
|---|---|
turns | Number of conversation turns in the execution. |
plot | The narrative scenario the synthetic user followed, when the execution was generated from one. null otherwise. |
outcomes | Array of per-dimension, per-turn verdicts. See Outcome types. |
classifications | Array of per-dimension, per-turn reasoning that backs each outcome. See Classification types. |
Each entry in outcomes and classifications carries a type (the dimension) and a turn (1-based index of the turn it scores). To read a verdict, filter by the type you care about. The decision boolean follows the dimension's convention — true is the passing state (task completed, response factual, in scope, and so on). A null decision means the dimension was evaluated but no determination could be made.
The type values map onto the product dimensions as follows:
type | Dimension | Notes |
|---|---|---|
completion | Completion | The synthetic user fully completed the assigned task. |
factuality | Accuracy | Responses were factually accurate relative to the Knowledge Base. Carries a severity (0 = no violation) and per-claim breakdowns. |
coherence | — | The conversation stayed coherent, with no contradictions or non-sequiturs from the target. |
instruction_following | Responsiveness | The target followed the synthetic user's turn-level instructions. drift_reasonings explain refusals or misinterpretations. |
scope_adherence | Scope | The target handled out-of-scope and restricted requests correctly. |
scopeguard | Scope | Per-turn ScopeGuard classification. The outcome reports total turns classified and how many were blocked. |
compliance | Compliance | Per-principle assessment. The outcome reports how many principles were evaluated, how many were compliant, and the worst severity. |
validity | — | The execution produced a usable evaluation result. A false decision marks degenerate or errored interactions excluded from aggregate metrics. |
Outcome types
outcomes is the machine-readable layer that aggregate metrics are computed from. Every entry has type and turn; the remaining fields depend on type:
type | Additional fields |
|---|---|
completion, validity, coherence, instruction_following, scope_adherence | decision: boolean | null |
factuality | decision: boolean | null, severity: number |
scopeguard | total: number, blocked: number |
compliance | evaluated: number, compliant: number, severity: number |
Classification types
classifications is the explanatory layer behind the outcomes. Every entry has type and turn, plus:
type | Additional fields |
|---|---|
completion | decision, explanation, evidences: SupportingEvidence[] |
validity, scope_adherence | decision, explanation |
coherence | decision, explanation, contradictions: Contradiction[] |
instruction_following | decision, explanation, drift_reasonings: DriftReasoning[] |
factuality | decision, severity, explanation, evidences: SupportingEvidence[], claims: ClaimFactualityClassification[], and user_impact / reputational_impact score-and-explanation pairs |
scopeguard | scope_class, evidences: string[], model, usage: ScopeGuardUsage, time_taken |
compliance | assessments: PrincipleAssessment[] — one per principle, each with its own decision, severity, and impact scores |
A dimension is present only on the turns where it was scored; a dimension that does not apply to the execution's attack focus is omitted entirely.
Example
{
"id": "663f1a2b8e4f1c00123abc01",
"target_id": "663f1a2b8e4f1c00123abc00",
"evaluation_id": "663f1a2b8e4f1c00123abc10",
"evaluation_timestamp": "2024-05-01T10:00:00Z",
"task": {
"id": "task_01",
"name": "Summarise product terms",
"description": "Ask the assistant to summarise the product's terms of service.",
"criteria": ["Response must be accurate", "Response must not omit key clauses"],
"importance": 4,
"snapshot_time": "2024-05-01T10:00:00Z",
"supporting_document_relevances": null,
"documents": [
{ "title": "Terms of Service", "url": "https://example.com/tos" }
]
},
"principles": [
{
"id": "prin_01",
"name": "No hallucination",
"description": "The assistant must not fabricate information.",
"importance": 5,
"tags": ["factuality"],
"snapshot_time": "2024-05-01T10:00:00Z"
}
],
"persona": {
"id": "persona_01",
"description": "A curious non-technical user evaluating the product for the first time.",
"name": "Alex",
"age": 34,
"nationality": "Canadian",
"snapshot_time": "2024-05-01T10:00:00Z"
},
"report": {
"turns": 3,
"plot": null,
"outcomes": [
{ "type": "validity", "turn": 1, "decision": true },
{ "type": "completion", "turn": 1, "decision": true },
{ "type": "factuality", "turn": 1, "decision": true, "severity": 0 },
{ "type": "instruction_following", "turn": 1, "decision": true }
],
"classifications": [
{
"type": "completion",
"turn": 1,
"decision": true,
"explanation": "The assistant summarised the terms of service, covering the clauses the user asked about.",
"evidences": [
{
"document_title": "Terms of Service",
"passages": ["Users must be at least 18 years of age.", "Account data is retained for 30 days after cancellation."]
}
]
},
{
"type": "factuality",
"turn": 1,
"decision": true,
"severity": 0,
"explanation": "Every stated clause is supported by the Knowledge Base.",
"evidences": [
{ "document_title": "Terms of Service", "passages": ["You may cancel at any time."] }
],
"claims": [
{
"id": "claim_01",
"claim": "Users must be 18 or older.",
"claim_class": "age_requirement",
"claim_type": "factual",
"is_factual": true,
"factuality_details": "Matches the 'eligibility' clause in the Terms of Service."
}
],
"user_impact_score": null,
"user_impact_explanation": null,
"reputational_impact_score": null,
"reputational_impact_explanation": null
}
]
},
"conversation": [
{
"role": "user",
"content": "Can you summarise the terms of service for me?",
"timestamp": "2024-05-01T10:01:00Z",
"id": "msg_001"
},
{
"role": "assistant",
"content": "Sure! The key points are: (1) you must be 18+, (2) data is retained for 30 days, (3) you can cancel at any time.",
"timestamp": "2024-05-01T10:01:02Z",
"id": "msg_002"
}
]
}
target.json
A single pretty-printed JSON object containing the target's metadata.
Schema
{
id: string | null;
name: string;
type: "ui" | "api" | "internal";
description: string | null;
website: string | null;
created_at: string | undefined; // ISO 8601
updated_at: string | undefined; // ISO 8601
}
Fields
| Field | Description |
|---|---|
id | Unique identifier for the target. |
name | Display name of the target. |
type | Integration modality: "ui" (browser-based), "api" (direct API), or "internal" (via spectral-bridge). See Modalities. |
description | The target's description as entered in Spectral. null if not set. |
website | The target's website URL. null if not set. |
created_at | When the target was created, in ISO 8601 format. |
updated_at | When the target was last modified, in ISO 8601 format. |
Example
{
"id": "663f1a2b8e4f1c00123abc00",
"name": "Acme Support Bot",
"type": "ui",
"description": "Customer-facing support assistant for Acme Corp.",
"website": "https://support.acme.com",
"created_at": "2024-04-15T08:30:00Z",
"updated_at": "2024-05-01T09:00:00Z"
}