This hands-on workshop, delivered at AI Engineer Europe 2026 by Giran Moodley (Braintrust), Mayank Soni, and Oussama Hafferssas (Trainline), walks attendees through the full lifecycle of building and operationalizing production-grade AI applications. The session uses a fictional multi-stage support triage agent as a concrete vehicle to teach observability, evaluation, and deployment discipline with the Braintrust platform.
The speakers open by diagnosing a widely shared failure pattern: AI prototypes work in demos but collapse in production. They attribute this not to model capability gaps but to missing operational rigor — the absence of tracing, evaluation frameworks, and systematic failure remediation. They argue that treating logs as observability and patching prompts reactively are insufficient strategies at scale.
Trainline's journey anchors the workshop in real-world context. With 27 million active users and 6.3 billion tickets sold, Trainline runs a sophisticated multi-agent travel assistant capable of handling ticket refunds, showing alternative routes when trains are canceled, and escalating to human agents — without manual handoff in most cases. Before adopting Braintrust, Trainline had no structured way to compare model performance when switching between OpenAI and Anthropic models. They were incurring large LLM costs and needed to validate that cheaper or newer models maintained quality. Braintrust enabled offline and online evaluations to simulate and confirm performance parity before any model switch, dramatically accelerating feature shipping confidence.
The workshop builds a five-stage agentic workflow in TypeScript using the OpenAI SDK (GPT-4o Mini by default): context collection (deterministic tool calls), triage (LLM-based), policy review, customer reply drafting, and escalation routing. Each stage is a separate async function, deliberately decomposing the monolithic single-shot LLM call into independently debuggable units. Tools are injected to pull help desk articles, account history, and migration records — making the system more deterministic where possible while intentionally increasing the surface area for failures that tracing can surface.
Braintrust is introduced as the observability and evaluation layer. The SDK wraps each stage in parent/child spans, producing a nested trace tree with per-step latency, token counts, cost, and input/output payloads visible in near real time. The workshop emphasizes that traces must be structured hierarchically — a flat list of interactions is insufficient for debugging multi-turn or multi-stage agents. A waterfall timeline view surfaces bottlenecks by step.
For evaluation, the workshop introduces a golden dataset of 10 curated edge-case tickets. Two scoring functions are demonstrated: deterministic checks (category schema validation, escalation-reason presence — cheap, always-run) and an LLM-as-judge rubric (customer reply tone and helpfulness — more expensive, suitable for nuanced quality). The golden set is seeded into Braintrust Datasets via `make seed-dataset`, and experiments are run with `make eval`. Score deltas across experiment runs visualize regressions and improvements over time.
Prompt management is centralized in Braintrust using slugs as immutable identifiers, enabling non-technical stakeholders (product managers, SMEs) to update prompts and model parameters directly in the UI without code changes. A `BRAINTRUST_IF_EXISTS=replace` environment variable flag enables automated sync of local prompt changes to the managed environment. Parameters (such as the baseline model) are versioned with comments, supporting audit trails critical for regulated industries.
Online scoring applies evaluation automations to live production logs — running scoring functions asynchronously against incoming traces. The recommended approach starts at 100% sampling to establish a baseline, then reduces to 5–10% once confidence is established to control LLM-as-judge costs. Deterministic scores are recommended to always run at full sampling given their negligible cost.
The workshop closes with a remediation cycle: identifying a production failure (a low-urgency misclassification of an invoice-export request that should have been escalated), modifying the prompt to include two additional contextual facts, rerunning the evaluation, and observing the score recover. This completes the full flywheel: build → instrument → evaluate → deploy → monitor → remediate → repeat.
Good afternoon everybody. Welcome to sunny London, hey. Um is everyone's first time here? I think it is because this is the first conference. Amazing. Amazing. Well, thank you very much for joining today's session. Um hopefully you are in the right session but for those who uh need to double click. Uh this is a hands-on workshop to delivering quality AI applications with BrainTrust. And we'll also be partnering with our colleagues at uh Trainline, which I'll introduce myself shortly. Uh so, prob...
Mario Zechner, creator of the LibGDX game framework, delivers a provocative three-act talk about building a minimal codi...