Day 2 of AI Engineer Europe 2026 (April 10, London) covered MCP architecture, coding agent design philosophy, vertical AI, and non-coding agent workflows.
Omar Sanseviero (Google DeepMind) argued open models like Gemma 4 can now run highly agentic tasks entirely on-device and predicted customizable on-device models within 6 to 12 months.
David Soria Parra (Anthropic), creator of MCP, reported 110 million monthly downloads — reached twice as fast as React. He introduced two key 2026 patterns: progressive discovery (defer tool loading via a tool-search primitive instead of injecting all tools at once, which measurably cut Claude Code's context footprint) and programmatic tool calling (model writes a script composing tool calls rather than chaining sequential agentic calls, reducing latency and tokens). June spec additions: stateless HTTP transport, async agent-to-agent task primitives, server discovery via well-known URLs, cross-app enterprise SSO, and skills-over-MCP.
Ido Salomon (MCP Apps) demonstrated AgentCraft, a multi-agent orchestrator with RTS game mechanics — the file system renders as a map, agents are visible units, and a collision heatmap surfaces concurrent write conflicts.
Mario Zechner (Pi) built a minimal four-tool self-modifying harness after finding Claude Code injected system reminders mid-context and obscured tool definition changes. Terminal Bench (December 2025) shows a bare keystroke-to-tmux harness outperforms native harnesses across most model families. His warning: agents compound errors with zero learning, producing enterprise-grade complexity in days; silent local-failure recovery creates brittle systems.
Armin Ronacher and Cristina Poncela Cubeiro (Earendil) argued codebases are infrastructure requiring agent-legible design: enforce unique function names, erasable-syntax-only TypeScript, single SQL query interface, no bare catch blocks, one UI primitives library. Their Pi extension routes auto-fixable bugs to the agent and flags migrations and permission changes for human review.
David Gomes (Cursor) replaced 15,000 lines of code for Cursor's worktrees feature with ~40 lines of markdown. The /worktree and /best-event commands are backend-served for prompt iteration without a new release. Best-event spawns one sub-agent per model (Gemini, Grok, Composer, GPT, Opus) in isolated worktrees; the parent produces a comparative table.
Lawrence Jones (Incident.io) built eval-tool — a read/edit/add CLI for YAML eval suites — so coding agents can modify prompts without context overflow. His team exports agent trace UIs as file-system packages for Claude Code debugging and runs 25 parallel sub-agents over daily back-tests, then clusters results to produce pattern-level RCA reports.
Ben Burtenshaw (Hugging Face) demonstrated coding agents writing CUDA kernels via the kernels library (TOML hardware-tagged hub repos), achieving a 94% speedup on Qwen 38B for H100. He also showed a four-role multi-agent auto-research lab (researcher, planner, workers, reporter) iterating ML training scripts and tracking results in a Draxio Parquet dashboard.
Liam Hampton (Microsoft) showed VS Code as a single entry point for three simultaneous agent tiers: local (human-in-loop, tests), background in git worktrees via GitHub Copilot CLI (UI building), and cloud agents in GitHub Actions (GitHub MCP + Playwright MCP, firewalled, no main-branch access).
Tuomas Artman (Linear) shared that 10% of bugs are now auto-resolved and merged without engineer involvement, expected to approach 100%. All features still require design and customer research; taste is Linear's moat.
Jacob Lauritzen (Legora, 1,000+ law-firm customers, 50+ markets) applied the verifiers rule to vertical agents: easy-to-verify tasks (formatting, definition-checking) are agent-solvable; litigation strategy is unverifiable. He advocates durable artifact UIs — clause-level document editors and tabular review views — over chat.
Peter Gostev (Arena.ai) introduced the Bullshit Benchmark — 155 nonsense questions graded by LLM-as-judge across 700+ tracked models. Claude models showed the highest clear-pushback rate; GPT and Gemini accepted nonsense roughly 50% of the time.
swyx closed by describing how his nine-person team runs a $9M+ conference through agents (Figma-to-website, schedule management, ETL, procurement), calling agents-for-non-coding-work a top 2026 trend.
Heat. Heat. Heat. It doesn't knock. It doesn't name itself. It calls you. No face. It needs no crown. No single hand to strike you down. It moves through mouths that claim they know what's right. What must be so? It speaks in care. It speaks in good. Cuts everywhere. No god, no code, no line to cross. Just necessary justify the fracture. Every truth becomes a weapon. Everyone >> wor it isn't flesh or bone. It thinks through what you take your body. It takes the hijacks choice rewrites in the cha...
Mario Zechner, creator of the LibGDX game framework, delivers a provocative three-act talk about building a minimal codi...