This recording captures the full Day 2 programme of AI Engineer Singapore 2026, held May 17, 2026, spanning roughly nine hours of technical sessions across agent engineering, document AI, robotics, design, and organisational knowledge systems, with approximately 25 speakers.
SallyAnn DeLucia (Arize AI) opened with a post-mortem on building Alyx. She outlined four lessons: planning must live in explicit tool states (pending/in-progress/blocked/completed) outside conversation history to prevent truncation; large-JSON abstractions compress string values but preserve full schema structure so agents can navigate 100,000-token experiment data; production traces—not hand-authored golden sets—are the best eval ground truth; and Arize Skills now feed traces directly into coding agents like Claude Code to close issue-to-fix cycles in minutes.
Abhishek Kankani (Cloudflare) presented Code Mode: converting tool definitions to TypeScript type declarations drops context from 1.7 million tokens (for 2,500 Cloudflare APIs as standard MCP tools) to roughly 1,000—a 99.9% reduction. He also covered V8 isolate execution via Cloudflare Workers, which provides zero-cold-start secure sandboxes for model-generated code, eliminating container cold-start latency.
Tejas Kumar (IBM) delivered a live-coded harness tutorial, building from a failing GPT-3.5 Turbo browser agent to a working one by adding max-iteration guardrails, context trimming, a harness-level verify step that reads the message event trace, and an auto-login handler injecting credentials at the harness layer without prompt changes. His central point: a good harness allows a weak model with a poor prompt to succeed reliably. IBM's OpenRAG was cited as an open-source enterprise harness reference.
JJ Geewax (Google DeepMind) argued against monolithic LLM routers in production. His pattern decomposes pipelines into: an LLM-as-classifier routing layer (multiple-choice, not open-ended); a deterministic JSON-to-JSON transform layer; a generation layer; and a small-model safety classifier. Key warnings: temperature-zero is not deterministic (subtle text differences cause large output variance); RAG pipelines can propagate policy-breaking training artefacts (e.g., a $1 car sale in a test doc leaking into live responses). On-device fast models combined with slower cloud models solve real-time streaming video classification problems.
Geoff Huntley argued software development now costs less than minimum wage, presenting data on a New Zealand company that cut from 60 to 20 people by not backfilling and achieved higher velocity. He characterised AI-native startups as lean apex predators versus large incumbents undergoing a 3–4 year J-curve transformation. He introduced the Ralph Loop—a context-engineering memory technique that wraps tool calls in a loop—now embedded in many agent tools.
Vincent Koc (OpenClaw Foundation) reported 1 million+ npm downloads per week, 50,000 commits, 1,600 contributors. He described transitioning to a plugin architecture (breaking internal/external boundaries), introduced Clownfish—GitHub Actions harnesses that reduced 10,000 open PRs to 3,000 in two days—and git-crawl/Crabbox tooling for ephemeral sandboxing and hourly-fresh distributed issue clustering.
Sara Hooker (Adaption Labs) challenged the scaling consensus: small models outperform much larger ones; 95% of weights can be pruned post-training; doubling model size only captures a long tail of rare artefacts at high cost. She announced AutoScientist, which self-improves training co-optimisation across 30+ models on Together AI and can train a frontier model in two days, outperforming their own human research staff on cross-architecture tasks. Adaption covers 242 languages and has processed 27 million data points in four weeks.
Pierre-Loic Doulcet (LlamaIndex) reported parsing over one billion documents via agentic loops and catalogued failure modes: whitespace loops (Anthropic Sonnet class is particularly susceptible; spaces cannot be stop tokens due to tokeniser multi-space-token design), repeated-output loops affecting ~0.5–1% of queries, and structural PDF failures. He built PassBench (open-source leaderboard on Kaggle) and LightPass (non-LLM fallback, 500 pages/second on CPU) as essential production safeguards.
Jun Yu Tan (Tusk) introduced Fence, a deterministic OS-level execution boundary. He calculated that a 99% reliable probabilistic classifier yields only ~30% zero-error probability over a 120-tool-call session and essentially zero over 1,000 calls. Fence enforces file-system, network, and command policies at kernel level with no container runtime, providing the middle layer in a classifier → OS policy → container isolation stack.
Conor Brennan-Burke (Hyperspell, $6.7M raised) argued that RAG connectors provide access but not understanding. He proposed context graphs—file-system representations of company knowledge combining Slack threads, emails, meeting notes, and agent execution traces, deduplicated and confidence-scored—as the missing enterprise deployment layer. Vincent Wu (MiniMax) outlined inference exchanges where autonomous agents pre-declare their token profiles (cache-hit ratio, output distribution) to enable GPU fleet utilisation optimisation and off-peak cost reduction.
Waves [music] crashing [singing] night waves crashing ocean know it. You need Hey, hey, hey. Hey, hey, hey, hey, hey. >> [music] [music] >> Hey, hey, hey. Hey, hey, hey. Hey, hey, hey. Hey, hey, hey. >> [music] >> Hey, hey, hey. Hey, hey, hey. Hey, hey, hey. Hey, hey, hey. Heat. [music] Heat. Hey, hey, hey. of this event, co-founder 65 Labs, and thank you so much for showing up. I know it's day three, Sunday morning, and all of you here in this room have chosen sleep deprivation over missing a s...
Mario Zechner, creator of the LibGDX game framework, delivers a provocative three-act talk about building a minimal codi...