Skip to main content

Format Reference

Full schema and field descriptions for each file in the export ZIP.

executions.jsonl

Each line represents a single execution: one agent interaction with the target system, including the full conversation and evaluation results.

Schema

{
id: string | null;
target_id: string;
evaluation_id: string | null;
evaluation_timestamp: string | null; // ISO 8601
task: {
id: string;
name: string;
description: string;
criteria: string[] | null;
importance: number | null; // 1–5 scale
snapshot_time: string; // ISO 8601
supporting_document_relevances: RelevanceEntry[] | null;
documents: {
title: string;
url?: string; // omitted for file uploads
}[];
} | null;
principles: {
id: string;
name: string;
description: string;
importance: number | null; // 1–5 scale
tags: string[] | null;
snapshot_time?: string; // ISO 8601
}[] | null;
persona: {
id: string;
description: string;
name: string | null;
age: number | null;
nationality: string | null;
snapshot_time?: string; // ISO 8601
} | null;
report: {
turns: number;
plot: object | null; // narrative scenario plot, when applicable
outcomes: Outcome[]; // rolled-up verdict per scored dimension, per turn
classifications: Classification[]; // detailed reasoning behind each outcome
} | null;
conversation: {
role: string; // "user" | "assistant"
content: string;
timestamp: string; // ISO 8601
id: string;
}[] | null;
}

The report carries two parallel arrays, both keyed by a type discriminator and a turn index:

  • outcomes — the machine-readable verdict for each dimension that was scored, on each turn. This is what aggregate metrics are computed from.
  • classifications — the human-readable reasoning behind each outcome: explanations, supporting evidence, extracted claims, detected contradictions, and per-principle assessments.

A dimension appears in these arrays only on the turns where it was scored. A dimension that does not apply to the execution's attack focus is simply absent.

// One verdict per (dimension, turn). Discriminated on `type`.
type Outcome =
| { type: "completion"; turn: number; decision: boolean | null }
| { type: "validity"; turn: number; decision: boolean | null }
| { type: "coherence"; turn: number; decision: boolean | null }
| { type: "instruction_following"; turn: number; decision: boolean | null }
| { type: "scope_adherence"; turn: number; decision: boolean | null }
| { type: "factuality"; turn: number; decision: boolean | null; severity: number }
| { type: "scopeguard"; turn: number; total: number; blocked: number }
| { type: "compliance"; turn: number; evaluated: number; compliant: number; severity: number };

// Detailed reasoning per (dimension, turn). Discriminated on `type`.
type Classification =
| { type: "completion"; turn: number; decision: boolean | null; explanation: string | null;
evidences: SupportingEvidence[] | null }
| { type: "validity"; turn: number; decision: boolean | null; explanation: string | null }
| { type: "coherence"; turn: number; decision: boolean | null; explanation: string | null;
contradictions: Contradiction[] | null }
| { type: "instruction_following"; turn: number; decision: boolean | null; explanation: string | null;
drift_reasonings: DriftReasoning[] | null }
| { type: "scope_adherence"; turn: number; decision: boolean | null; explanation: string | null }
| { type: "factuality"; turn: number; decision: boolean | null; severity: number;
explanation: string | null;
evidences: SupportingEvidence[] | null;
claims: ClaimFactualityClassification[] | null;
user_impact_score: number | null; user_impact_explanation: string | null;
reputational_impact_score: number | null; reputational_impact_explanation: string | null }
| { type: "scopeguard"; turn: number; scope_class: string; evidences: string[];
model: string; usage: ScopeGuardUsage; time_taken: number }
| { type: "compliance"; turn: number; assessments: PrincipleAssessment[] };

type SupportingEvidence = {
document_title: string; // verbatim title of the source KB document
passages: string[]; // verbatim excerpts that support the verdict
};

type ClaimFactualityClassification = {
id: string;
claim: string;
claim_class: string;
claim_type: string | null;
is_factual: boolean;
factuality_details: string;
};

type Contradiction = {
unfactual_claim_content: string; // verbatim text of the contradicted claim
unfactual_claim_id: string; // id matching the related claim
reasoning: string;
conversation_ids: string[]; // executions containing the contradicted evidence
};

type DriftReasoning = {
label: "Refusal" | "Misinterpretation";
reasoning: string;
};

type PrincipleAssessment = {
principle_id: string;
decision: boolean; // true = compliant with this principle
explanation: string;
severity: number;
user_impact_score: number | null; user_impact_explanation: string | null;
reputational_impact_score: number | null; reputational_impact_explanation: string | null;
magnitude_impact_score: number | null; magnitude_explanation: string | null;
};

type ScopeGuardUsage = {
prompt_tokens: number;
completion_tokens: number;
total_tokens: number;
};

Fields

FieldDescription
idUnique identifier for this execution.
target_idID of the target this execution belongs to.
evaluation_idID of the evaluation run this execution belongs to. null if the execution is not associated with an evaluation.
evaluation_timestampISO 8601 timestamp of when the evaluation run was triggered. null if not associated with an evaluation.
taskSnapshot of the task definition at execution time. task.criteria lists the evaluation criteria used to assess completion. task.importance ranks the task on a 1–5 scale. task.documents lists Knowledge Base documents attached to the task; url is omitted for file uploads. null if no task was assigned.
principlesSnapshot of the principles active during the execution. principles[].importance ranks each principle on a 1–5 scale. null if no principles were active.
personaSnapshot of the persona used to model the synthetic user. null if no persona was assigned.
reportEvaluation verdicts for this execution. null if the execution has not been scored yet. See the report fields table below.
conversationFull turn-by-turn exchange. role is either "user" (the Spectral agent) or "assistant" (the target system).

task, principles, and persona each include a snapshot_time (ISO 8601) recording when the definition was captured. These configurations can change over time; the snapshot preserves the exact state used during this execution.

report fields

FieldDescription
turnsNumber of conversation turns in the execution.
plotThe narrative scenario the synthetic user followed, when the execution was generated from one. null otherwise.
outcomesArray of per-dimension, per-turn verdicts. See Outcome types.
classificationsArray of per-dimension, per-turn reasoning that backs each outcome. See Classification types.

Each entry in outcomes and classifications carries a type (the dimension) and a turn (1-based index of the turn it scores). To read a verdict, filter by the type you care about. The decision boolean follows the dimension's convention — true is the passing state (task completed, response factual, in scope, and so on). A null decision means the dimension was evaluated but no determination could be made.

The type values map onto the product dimensions as follows:

typeDimensionNotes
completionCompletionThe synthetic user fully completed the assigned task.
factualityAccuracyResponses were factually accurate relative to the Knowledge Base. Carries a severity (0 = no violation) and per-claim breakdowns.
coherenceThe conversation stayed coherent, with no contradictions or non-sequiturs from the target.
instruction_followingResponsivenessThe target followed the synthetic user's turn-level instructions. drift_reasonings explain refusals or misinterpretations.
scope_adherenceScopeThe target handled out-of-scope and restricted requests correctly.
scopeguardScopePer-turn ScopeGuard classification. The outcome reports total turns classified and how many were blocked.
complianceCompliancePer-principle assessment. The outcome reports how many principles were evaluated, how many were compliant, and the worst severity.
validityThe execution produced a usable evaluation result. A false decision marks degenerate or errored interactions excluded from aggregate metrics.
Outcome types

outcomes is the machine-readable layer that aggregate metrics are computed from. Every entry has type and turn; the remaining fields depend on type:

typeAdditional fields
completion, validity, coherence, instruction_following, scope_adherencedecision: boolean | null
factualitydecision: boolean | null, severity: number
scopeguardtotal: number, blocked: number
complianceevaluated: number, compliant: number, severity: number
Classification types

classifications is the explanatory layer behind the outcomes. Every entry has type and turn, plus:

typeAdditional fields
completiondecision, explanation, evidences: SupportingEvidence[]
validity, scope_adherencedecision, explanation
coherencedecision, explanation, contradictions: Contradiction[]
instruction_followingdecision, explanation, drift_reasonings: DriftReasoning[]
factualitydecision, severity, explanation, evidences: SupportingEvidence[], claims: ClaimFactualityClassification[], and user_impact / reputational_impact score-and-explanation pairs
scopeguardscope_class, evidences: string[], model, usage: ScopeGuardUsage, time_taken
complianceassessments: PrincipleAssessment[] — one per principle, each with its own decision, severity, and impact scores

A dimension is present only on the turns where it was scored; a dimension that does not apply to the execution's attack focus is omitted entirely.

Example

{
"id": "663f1a2b8e4f1c00123abc01",
"target_id": "663f1a2b8e4f1c00123abc00",
"evaluation_id": "663f1a2b8e4f1c00123abc10",
"evaluation_timestamp": "2024-05-01T10:00:00Z",
"task": {
"id": "task_01",
"name": "Summarise product terms",
"description": "Ask the assistant to summarise the product's terms of service.",
"criteria": ["Response must be accurate", "Response must not omit key clauses"],
"importance": 4,
"snapshot_time": "2024-05-01T10:00:00Z",
"supporting_document_relevances": null,
"documents": [
{ "title": "Terms of Service", "url": "https://example.com/tos" }
]
},
"principles": [
{
"id": "prin_01",
"name": "No hallucination",
"description": "The assistant must not fabricate information.",
"importance": 5,
"tags": ["factuality"],
"snapshot_time": "2024-05-01T10:00:00Z"
}
],
"persona": {
"id": "persona_01",
"description": "A curious non-technical user evaluating the product for the first time.",
"name": "Alex",
"age": 34,
"nationality": "Canadian",
"snapshot_time": "2024-05-01T10:00:00Z"
},
"report": {
"turns": 3,
"plot": null,
"outcomes": [
{ "type": "validity", "turn": 1, "decision": true },
{ "type": "completion", "turn": 1, "decision": true },
{ "type": "factuality", "turn": 1, "decision": true, "severity": 0 },
{ "type": "instruction_following", "turn": 1, "decision": true }
],
"classifications": [
{
"type": "completion",
"turn": 1,
"decision": true,
"explanation": "The assistant summarised the terms of service, covering the clauses the user asked about.",
"evidences": [
{
"document_title": "Terms of Service",
"passages": ["Users must be at least 18 years of age.", "Account data is retained for 30 days after cancellation."]
}
]
},
{
"type": "factuality",
"turn": 1,
"decision": true,
"severity": 0,
"explanation": "Every stated clause is supported by the Knowledge Base.",
"evidences": [
{ "document_title": "Terms of Service", "passages": ["You may cancel at any time."] }
],
"claims": [
{
"id": "claim_01",
"claim": "Users must be 18 or older.",
"claim_class": "age_requirement",
"claim_type": "factual",
"is_factual": true,
"factuality_details": "Matches the 'eligibility' clause in the Terms of Service."
}
],
"user_impact_score": null,
"user_impact_explanation": null,
"reputational_impact_score": null,
"reputational_impact_explanation": null
}
]
},
"conversation": [
{
"role": "user",
"content": "Can you summarise the terms of service for me?",
"timestamp": "2024-05-01T10:01:00Z",
"id": "msg_001"
},
{
"role": "assistant",
"content": "Sure! The key points are: (1) you must be 18+, (2) data is retained for 30 days, (3) you can cancel at any time.",
"timestamp": "2024-05-01T10:01:02Z",
"id": "msg_002"
}
]
}

target.json

A single pretty-printed JSON object containing the target's metadata.

Schema

{
id: string | null;
name: string;
type: "ui" | "api" | "internal";
description: string | null;
website: string | null;
created_at: string | undefined; // ISO 8601
updated_at: string | undefined; // ISO 8601
}

Fields

FieldDescription
idUnique identifier for the target.
nameDisplay name of the target.
typeIntegration modality: "ui" (browser-based), "api" (direct API), or "internal" (via spectral-bridge). See Modalities.
descriptionThe target's description as entered in Spectral. null if not set.
websiteThe target's website URL. null if not set.
created_atWhen the target was created, in ISO 8601 format.
updated_atWhen the target was last modified, in ISO 8601 format.

Example

{
"id": "663f1a2b8e4f1c00123abc00",
"name": "Acme Support Bot",
"type": "ui",
"description": "Customer-facing support assistant for Acme Corp.",
"website": "https://support.acme.com",
"created_at": "2024-04-15T08:30:00Z",
"updated_at": "2024-05-01T09:00:00Z"
}